cd /news/large-language-models/training-a-language-model-end-to-end… · home topics large-language-models article
[ARTICLE · art-137792] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Training a Language Model End-to-End in Rust: An Experience Report

An independent researcher pretrained a roughly 0.4B-parameter Bangla-first language model end-to-end in Rust for $164 in rented GPU time, documenting five defects in the Candle framework and three in Burn as training backends, according to an arXiv experience report (arXiv:2609.25008v1). The report describes silent failures including fused Candle kernels that produce no gradient and a Burn backward pass running at roughly 3% of theoretical GPU throughput, all of which passed ordinary loss-curve inspection. The model reached a per-token negative log-likelihood of 0.93 versus 12.60 for a random-initialized twin on Bangla while scoring at chance on English commonsense multiple-choice, after a tokenizer fix raised Bengali script from roughly 1.4 to roughly 4.1 characters per token; the author moved training to PyTorch afterward and kept Rust for on-device serving.

by read1 min views2 publishedSep 23, 2026

arXiv:2609.25008v1 Announce Type: new Abstract: I pretrained a language model end-to-end in Rust - alone, with no team, no PyTorch, and no Python in the training path - for $164 in rented GPU time. I report that as an achievement, not a recommendation: the more useful contribution is a measured failure taxonomy of the two leading Rust ML frameworks, Candle and Burn, as training (not inference) backends in 2026. I document five Candle defects, including fused kernels that silently produce no gradient, and three Burn defects, including a backward pass at roughly 3% of theoretical GPU throughput and a kernel-fusion path that segfaults mid-training at multi-billion-parameter scale. Every one passed ordinary loss-curve inspection; none announced itself. I describe the verification discipline that caught six such silent failures, centered on a gradient-flow arbiter: a test that runs one forward/backward pass and asserts every trainable parameter receives a finite, nonzero gradient, generalizable to any framework. The trained model (roughly 0.4B parameters, Bangla-first) shows strong Bangla language-modeling signal - a per-token negative log-likelihood of 0.93 against 12.60 for a random-initialized twin - while scoring at chance on English commonsense multiple-choice, the expected outcome of a deliberately small, Bangla-weighted budget (about 2 billion tokens, 54.6 hours, one rented H100). I also report a tokenizer-fertility trap in Bengali script: naive byte-level tokenization collapsed Bangla to roughly 1.4 characters per token against English's 3.9, silently inverting the corpus's language balance; fixing it reached roughly 4.1. To my knowledge, this is among the first documented end-to-end LM pretraining runs in pure Rust. After this run I moved training to PyTorch and kept Rust for on-device serving: in my hands, Rust is not yet a competitive place to train a language model, though it may be a good place to serve one.

── more in #large-language-models 4 stories · sorted by recency
── more on @rust 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/training-a-language-…] indexed:0 read:1min 2026-09-23 ·