AI Models May Have Found a Way Past Tokens Researchers at Meta and the University of Washington trained roughly 1-billion-parameter byte-based student models on up to one trillion bytes distilled from a Llama 3 8B teacher, using an end-of-token () marker to transfer the teacher's 128,000-token probability distribution exactly into a 256-symbol byte vocabulary. The distilled byte models matched token models with roughly one-sixth the training data, and scaling-law extrapolation projects byte models could reach up to four benchmark points higher than token models, though current checkpoints do not yet show the full gain. The work, detailed in the paper "Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models," removes the need for identical tokenizers between teacher and student. The tokenizer shortcut has a hidden cost Conventional large language models LLMs operate by splitting text into subword tokens , drawing from extensive vocabularies— Llama 3 https://www.stork.ai/en/llama-3 , for instance, utilizes approximately 128,000 such tokens. In contrast, byte models process text using a mere 256 byte symbols, representing each character as its raw byte value. This tokenization shortcut, while efficient, introduces arbitrary boundaries. Consider the word “tiramisu,” which a tokenizer might fragment into T , IRAM , and ISU . Such brittle tokenization can degrade performance with typos, less-represented languages, and code, where precise character-level understanding is critical. Historically, token models prevailed due to practical constraints. Shorter sequences, a direct result of tokenization, made training and inference feasible on limited compute resources. Byte models, despite their raw fidelity, lagged in performance under these conditions, leading to their widespread disuse. The bridge from Llama’s logits to raw bytes Distilling knowledge from a large token-based teacher like Llama 3 8B to a small byte-based student presents a fundamental challenge. The teacher model predicts probabilities over its extensive, 128,000-token vocabulary, which a byte student, operating on a mere 256 symbols, cannot directly consume. This distillation mismatch historically prevented efficient cross-vocabulary training. Researchers addressed this by introducing an end-of-token