PSSA: A non-transformer language model written from scratch in Rust A from-scratch Rust language model called PSSA, which replaces transformer attention with a recurrent state-space core and a 512-slot episodic memory bank, reached 3.98 training cross-entropy versus 4.43 for a matched transformer over 12.7M tokens of cleaned WikiText-103, a 0.45-nat gap. On a 198,939-token held-out slice, PSSA scored 3.997 cross-entropy and 54.4 perplexity against the transformer's 4.429 and 83.8, with next-token accuracy of 24.1% versus 18.0%, and generated 200 tokens in 226 ms versus 2,735 ms on the same CPU. The author cautions both are 1.5M-parameter research prototypes with poor text quality at this scale, so the result concerns learning efficiency rather than fluency. PSSA is a small language model that is not a transformer. It reads text one token at a time through a recurrent state-space layer, keeps a bank of episodic memories it can look things up in, and rewrites part of its own weights while it runs. It is written in Rust from scratch, with no PyTorch, no TensorFlow, and no ML framework of any kind underneath it. At matched parameters and on the same corpus, it learns faster than a transformer and generates text about twelve times quicker on the same CPU. A transformer scores every pair of tokens in the context, so its cost per step grows with the square of the sequence length and the whole context is re-read at every step. PSSA carries one fixed-size state along the sequence in a single left-to-right pass, and looks things up in a memory bank instead of re-reading the context, so cost grows linearly with length. Two models, same corpus, same tokenizer, same optimizer schedule, same seed, same number of parameters. One is PSSA, one is a standard transformer. Over 12.7M tokens of cleaned WikiText-103: PSSA finished at 3.98 training cross-entropy, the transformer at 4.43 . That is a gap of 0.45 nats , perplexity 53.7 against 83.7. The transformer spent its entire 12.7M-token budget to reach a loss PSSA had already passed around 2M tokens in. The two curves never cross, and they never touch: Training loss only says a model fit the stream it was fed. So both checkpoints were scored on a 198,939-token slice cut from a part of the corpus neither run ever touched: Every checkpoint of both runs, 64 PSSA links and 43 transformer links, scored on a bounded 9,934-token window of that unseen slice. The curves never cross: PSSA is ahead from the first link and finishes 0.51 nats lower. The table below is the final checkpoint of each run on the full slice. | Held-out slice, 198,939 unseen tokens | PSSA | Transformer | |---|---|---| | Cross-entropy | 3.997 | 4.429 | | Perplexity | 54.4 | 83.8 | | Next-token accuracy | 24.1% | 18.0% | The held-out gap, 0.43 nats, is essentially the training gap. PSSA is not memorizing harder, it is generalizing better. Generating 200 tokens on the same CPU, same prompt, same sampler: | | PSSA | Transformer | |---|---|---| | 200 tokens | 226 ms | 2,735 ms | | Relative | 12x faster | baseline | A recurrent model carries a fixed-size state, so the cost of each new token does not grow with the length of what came before. A transformer re-reads its whole context every step. - A recurrent state-space core. Learned continuous state matrices carry information forward in a fixed-size state, instead of attention over the full context window. - An episodic memory bank. 512 slots with hyperbolic Poincare-style retrieval and bounded top-4 search, written to and read from during the run. - Plastic weights. Fast updates reinforce what works, novelty drives growth, and a refractory gate rate-limits overwrites so repeated contradictory input does less damage. - Closed-form consolidation. A ridge-regression step folds the fast plastic updates back into the base transition matrix, the way sleep consolidates a day's learning. - No framework. Hand-written linear algebra in Rust, with a CUDA path for training and a scalar CPU reference that every gradient is checked against max gradient difference 2.98e-8 . Being straight about the scale, because the numbers above are easy to over-read: - These are 1.5M-parameter models on 12.7M tokens. That is a research prototype, not a competitor to anything you have heard of. - Text quality at this scale is poor for both models. PSSA emits "a barget of the Prian Academy", the transformer "a material circulation of the United States". The comparison is about learning efficiency, not fluency. - The speed comparison is CPU-to-CPU, which is fair. The training throughput numbers further down are not hardware-matched and should not be read as an architecture result. - Two experiments are still unmeasured: retention of earlier skills after a corpus switch, and whether ablating the memory bank changes the loss. git clone https://github.com/Sparticle62ops/pssa.git cd pssa cargo build --release ./target/release/oxide ai pssa Running it with no arguments gives you a home screen listing every command plus any checkpoint and corpus it finds in the working directory. The whole result above was trained on a free hosted notebook with a single entry-level GPU, in 200,000-token links, because a session gets cut after a few hours. Every interesting question left, whether the gap holds at 10x or 100x these parameters, whether the memory bank matters at scale, how it does against a modern recurrent baseline, needs one thing: a GPU with real VRAM and allocations measured in days instead of hours. Anything meaningfully above the entry-level card this ran on changes what can be asked. If you have compute to grant, or you work somewhere that does, that is the single highest-leverage thing anyone can offer this project. Sponsorship funds compute and nothing else. In return you get named here and in the write-up of any result your hardware made possible. Get in touch before sending anything so the details can be agreed. Issues and pull requests are welcome. The parts most in need of hands: kernel performance, a modern recurrent baseline to compare against, and evaluation beyond next-token loss. Validate any branch with cargo test --release before opening a PR. Solana: 4XPZ9uAa2BMoth6msoHRxTWL4mUrMfq3LGrxbAGja96h Everything below is for running, training, and working on the project. - Rust toolchain with Edition 2024 support, including Cargo. - Network access only when using an HTTP/HTTPS dataset or a Hugging Face dataset. - Enough memory and disk for larger corpora and serialized models. - Optional: a CUDA device for the GPU training path. The CPU path is the reference and always available. Direct runtime dependencies are ureq https://crates.io/crates/ureq for dataset downloads and tokenizers https://crates.io/crates/tokenizers for byte-level BPE. Both chains ran 64 links of 200,000 encoded tokens, each link resuming from the previous checkpoint, so the learning-rate schedule and optimizer state continue across the whole run instead of restarting per link. - Identical corpus: one clean-wikitext pass over WikiText-103, reused byte for byte. - Identical token IDs: the baseline pins --tokenizer-from to the PSSA chain's own checkpoint, so neither model sees a different vocabulary. - Identical optimization: 30,000-update cosine horizon, no warm-up restart, 512 supervised target tokens per update, seed 42. - PSSA: latent 256, recurrent state 16, 512 memory slots, key width 32, vocab 2,048. - Baseline: 1,541,120 parameters, 1 layer, width 256, 4 heads, FFN 448, vocab 2,048. End-of-link training cross-entropy: | Link | Tokens seen | PSSA | Transformer | |---|---|---|---| | ck01 | 200,000 | 5.733 | 6.461 | | ck05 | 1,000,000 | 4.617 | 5.467 | | ck10 | 2,000,000 | 4.447 | 5.082 | | ck15 | 3,000,000 | 4.292 | 4.858 | | ck20 | 4,000,000 | 4.185 | 4.704 | | ck25 | 5,000,000 | 4.221 | 4.704 | | ck30 | 6,000,000 | 4.070 | 4.561 | | ck35 | 7,000,000 | 4.039 | 4.523 | | ck37 | 7,400,000 | 3.960 | 4.465 | | ck44 | 8,800,000 | 4.004 | 4.480 | | ck48 | 9,600,000 | 3.937 | 4.415 | | ck52 | 10,400,000 | 3.846 | 4.344 | | ck56 | 11,200,000 | 3.887 | 4.375 | | ck60 | 12,000,000 | 3.972 | 4.418 | | ck64 | 12,800,000 | 3.982 | 4.428 | The baseline's first session was cut at link 43 by the notebook session limit and its loss CSV did not survive, so links 1 to 43 are read back from that session's own run log instead. The chain resumed from ck43 in a second session and finished all 64 links, and both curves above now cover the full run. PSSA trained on a Kaggle T4 at roughly 900 tokens/second. The baseline is CPU-only, because train-transformer has no GPU path, and held 212 tokens/second. Those two numbers say nothing about the architectures. On the same CPU-only Kaggle hardware the batched PSSA path measures 375 tokens/second against the baseline's 212, and the loss comparison above is unaffected either way, since it is matched on tokens and updates rather than on time. The losses are end-of-link training cross-entropy on the stream being fit, not held-out evaluation. For a held-out comparison on an unseen slice, use the compare command described in docs/COMPARISON.md https://github.com/Sparticle62ops/pssa/blob/main/docs/COMPARISON.md . Generation quality at this scale is poor for both models: PSSA emits "a barget of the Prian Academy", the baseline "a material circulation of the United States". Two experiments are not yet measured: retention of earlier skills after a corpus switch, and whether ablating the 512 memory slots changes loss. bash kaggle/kaggle continue.sh the PSSA chain bash kaggle/kaggle transformer baseline.sh the parameter-matched baseline Both read TOTAL , WINDOW and FRESH from the environment and write --loss-csv , so the curve survives a cut session. General form: oxide ai pssa