19:50
2026-09-16
discuss.huggingface.co
large-language-models
Quick recap of ByteLex's Cross-Tokenization and the faults leading to successes
ByteLex reported that a tokenizer-free coordinate map built from 11 tokenizer vocabularies over 3-byte windows predicted its 237M-parameter byte-level language model's per-word errors with a Spearman …