# AI Models May Have Found a Way Past Tokens

> Source: <https://www.stork.ai/blog/ai-models-may-have-found-a-way-past-tokens>
> Published: 2026-10-02 20:36:32+00:00

## The tokenizer shortcut has a hidden cost

Conventional large language models (LLMs) operate by splitting text into **subword tokens**, drawing from extensive vocabularies—[Llama 3](https://www.stork.ai/en/llama-3), for instance, utilizes approximately 128,000 such tokens. In contrast, **byte models** process text using a mere 256 byte symbols, representing each character as its raw byte value.

This tokenization shortcut, while efficient, introduces arbitrary boundaries. Consider the word “tiramisu,” which a tokenizer might fragment into `T`, `IRAM`, and `ISU`. Such brittle tokenization can degrade performance with typos, less-represented languages, and code, where precise character-level understanding is critical.

Historically, token models prevailed due to practical constraints. Shorter sequences, a direct result of tokenization, made training and inference feasible on limited compute resources. Byte models, despite their raw fidelity, lagged in performance under these conditions, leading to their widespread disuse.

## The bridge from Llama’s logits to raw bytes

Distilling knowledge from a large token-based teacher like Llama 3 8B to a small byte-based student presents a fundamental challenge. The teacher model predicts probabilities over its extensive, 128,000-token vocabulary, which a byte student, operating on a mere 256 symbols, cannot directly consume. This **distillation mismatch** historically prevented efficient cross-vocabulary training.

Researchers addressed this by introducing an **end-of-token (<eot>)** marker. This special symbol captures any probability mass that would otherwise be lost when translating a token prediction into a sequence of byte predictions. For instance, if a token like "tiramisu" is broken into `t`, `iram`, `isu` bytes, the <eot> marker ensures that the full probability of the original token is conserved across its byte representation.

This innovative mechanism enables an **exact probability transfer** in a single pass. By appending the <eot> symbol, the researchers guaranteed that the teacher’s complete probability distribution could be precisely mapped to the byte student’s vocabulary without approximation.

This breakthrough allows small byte models to learn directly and effectively from powerful, existing token models such as Llama 3 8B. It eliminates the need for identical tokenizers between teacher and student, unlocking a new pathway for developing more robust and efficient language models.

## Why bytes may win when training goes long

Researchers at Meta and the University of Washington conducted a critical experiment, training roughly **1-billion-parameter** student models with data up to one trillion bytes. They distilled knowledge from a Llama 3 8B teacher, meticulously converting its token-based predictions into byte-level probabilities.

Early in training, token models consistently outperformed their byte counterparts, a common pattern due to their ability to process more semantic information per step. However, as training extended, byte models demonstrated a steeper, more sustained improvement curve, eventually surpassing the token models.

This extended training revealed a striking efficiency: distilled byte models achieved performance matching token models with roughly one-sixth the training data. For further details on the methodology, consult the paper [Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models](https://arxiv.org/abs/2609.12303).

The long-term advantage of byte models is significant. While current checkpoints do not yet show the full gain, **scaling-law extrapolation** projects byte models could achieve up to four benchmark points higher than token models. This projection, not an already realized score, highlights their potential in compute-rich training environments.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

## Cheaper to teach, slower to answer

Byte-level distributions offer a compelling storage advantage. Instead of relying on lossy top-k truncation, researchers preserved all predictions, reducing **logit storage** to approximately one-fifth of the space required for conventional token distributions. This allows for complete fidelity in the distillation process, a critical factor for capturing nuanced teacher model behavior.

This efficiency, however, introduces a trade-off during inference. Byte sequences are typically four to five times longer than their tokenized counterparts. This extended length directly translates to increased **inference time** and potentially higher **KV-cache memory requirements**, posing a challenge for real-time applications and resource-constrained deployments.

Byte models show promise for scenarios prioritizing compute-rich training environments and those vulnerable to tokenizer-induced fragility, such as handling novel languages or complex character sets. The projected four-point benchmark edge and reduced training data requirements are significant. However, their practical speed, deployment costs, and the ultimate validation of this projected performance gain remain areas for further investigation.

## Frequently Asked Questions

### What is a byte model?

A byte model reads text as sequences of bytes rather than splitting it into subword tokens from a learned vocabulary.

### How does Meta distill a token model into a byte model?

The method maps the teacher’s token probabilities to bytes and uses an end-of-token marker to preserve the full probability distribution.

### Do byte models already outperform token models?

The reported results show byte models improving with more training, but the projected four-point benchmark gain is an extrapolation, not a measured final result.

### What is the main drawback of byte models?

Their sequences are typically four to five times longer than token sequences, which can make inference slower and increase memory use.
