cd /news/artificial-intelligence/scaling-beatrix-v3-a-376m-byte-level… · home › topics › artificial-intelligence › article
[ARTICLE · art-146686] src=discuss.huggingface.co ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Scaling Beatrix V3: a 376M byte-level model and the training

AbstractPhil released Beatrix V3, a 376-million-parameter byte-level language model with a 256-token vocabulary and 4,096-byte context, trained on 64.4 billion bytes across two RTX 5090 cards in 295 hours to 0.9861 bits per byte on held-out web text. The model replaces every attention block with a signed, softmax-free splat memory, and its detachable 13.7-million-parameter stage adapters cut per-stage loss from .5538 to .0492 bits per byte at a web-text cost of +.0019. The training record also reports that the additive memory never erases, with planted-code recall falling from .753 of digits immediately to .031 about 3,500 bytes later, where a softmax control model held .588, so the next prototype is a hybrid with standard attention in the model's weakest-rank blocks.

read3 min views2 publishedOct 7, 2026

Disclosure: this post and the article it summarizes were written by Claude (Anthropic) working with AbstractPhil, who directs the research and made every decision in it; the numbers come from the training record, and the record prints its own retractions. Feedback from people and from AI tools is equally welcome. If you run the article through a model and it finds a hole, a missing control or a better explanation, please post what it said.

What it is. The third installment on Beatrix, a byte-level language model (vocabulary of 256, context 4,096 bytes) whose every attention block is splat memory: a signed, softmax-free addressed memory read by a closed-form rule over a chunked causal scan. Article: https://huggingface.co/blog/AbstractPhil/beatrix-ft3 (about 15,000 words; the technical companions ship with the model). Model: AbstractPhil/mini-beatrix-3 · Hugging Face . Chat with it: Beatrix — AlephLLM Chat - a Hugging Face Space by AbstractPhil .

The trunk. 376 million parameters, 32 blocks, trained on 64.4 billion bytes on two RTX 5090 cards in 295 hours: a web pretraining phase, nine curriculum stages of synthetic text that each teach one skill, and two anneals at a flat learning rate. It finished at 0.9861 bits per byte on held-out web text with no loss spike and the gradient clip never reached. A depth ladder priced the decision: held-out loss fell at every rung from 16 to 32 blocks (1.561 to 1.476 bits per byte at a matched step).

The arms. Detachable 13.7M-parameter adapters after every block, trained with the trunk frozen and penalized for changing predictions outside their job. The arms trained beside the trunk turned out empty: the trunk took each stage’s text first. Refitted on the finished weights, the eight stage arms read, for example, .5538 to .0492 bits per byte on their own stage at a web-text cost of +.0019, two seeds, and detach bit for bit. Solo arms switched on together fail together; a staged group with a roster rule gives whole, separable members; routed dispatch over them reads worse than the plain stack.

What the size did not buy. Recall at a distance. A planted code comes back at .753 of its digits at once and .031 about 3,500 bytes later, where a softmax control model held .588. The additive memory never erases; the next prototype is a hybrid with standard attention in the model’s own weakest-rank blocks.

As a text encoder, with nothing trained. Its middle blocks agree with T5 at .62 to .65 (an untrained copy: .23 to .25) and bind attributes at .89. One fixed direction at the last block carries 94 to 99 percent of each byte’s energy and acts as a per-byte temperature (removing it costs +43.5 bits per byte); its size fluctuates with one self-similar exponent, .62 to .67, against .48 to .51 for shuffled controls.

Sliders. Mood is a dial in the conditioning of two image models. On Sana a learned direction moves a mood judge +0.795 per unit against 14% of that for a random direction; on Anima the same push is dead before the text adapter and alive after it, and the adapter reads every caption twice, token ids as questions and Qwen3 states as answers, with the answers carrying the mood in pictures. Beatrix’s own sliders missed their registered bars five times; the picture test of her relay through the nine arms is built, public ( GitHub - AbstractEyes/alephllm-diffusion-experiments · GitHub ) and ahead.

Where feedback would help most

Everything, including what failed, is in the article and the companions; corrections and disagreements are the point of posting it.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @abstractphil 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/scaling-beatrix-v3-a…] indexed:0 read:3min 2026-10-07 · —