Tupoi: An attention-free LLM with strictly O(1) memory and 6 KB state NARE LABS, built by a 15-year-old, released TUPOI, an attention-free language model that replaces dense attention with a Symplectic Hamiltonian Integrator, achieving strictly O(1) memory with a fixed 6 KB state. In benchmarks, TUPOI-300M achieved a validation perplexity of 46.39 versus 50.80 for a 355M-parameter Transformer baseline, using 14% fewer parameters, and maintained constant memory usage across generation lengths up to 65,536 tokens, where the Transformer's KV-cache exceeded 6 GB and caused CUDA OOM. A sub-quadratic, attention-free sequence modeling architecture that replaces dense attention matrices with a Symplectic Hamiltonian Integrator Velocity-Verlet . Author: NARE LABS Built by 15 y.o. This repository contains the reference implementation, pre-trained weights, and empirical benchmark suite for the TUPOI language model, a post-transformer architecture that completely eliminates the Attention mechanism and the persistent KV-cache in favor of continuous Hamiltonian phase-space dynamics. The primary objective of this research is to investigate whether modeling sequence context as particles moving through a symplectic phase space $O 1 $ memory consumption and zero state dissipation TUPOI replaces The hidden state is decomposed into canonical coordinates Velocity-Verlet Leapfrog dynamical scheme: Because the Jacobian determinant of this transformation satisfies Information is neither created nor destroyed during the forward pass. To stabilize language representations over infinite generation horizons: - IgnoranceGate: A learned smooth sigmoid filter$x' = x \odot \sigma W g x $ that attenuates token noise before perturbing phase space. - OpinionAnchor: A running exponential moving average$a t = 1-\eta a {t-1} + \eta \cdot \bar{x}'$ acting as a canonical phase attractor. Our evaluations focus on rigorous, apples-to-apples comparisons against a standard GPT-style Transformer BaselineLM under identical training constraints seed, optimizer, token budget, parameter scale . Setup: GPT-2 BPE tokenizer vocab: 50,257 , seq len=512, trained for 10,000 steps on Tesla T4. | Model Tier | Active Parameters | Final Val Loss | Perplexity PPL | Complexity | Status | |---|---|---|---|---|---| | TUPOI-17M S | 16.9 M | 4.1516 | 63.53 | Verified | | | TUPOI-47M M | 47.0 M | 4.1078 | 60.81 | Verified | | BaselineLM Transformer | 355.0 M | 3.9280 | 50.80 | Baseline | | TUPOI-300M L | 304.2 M | 3.8371 | 46.39 | Winner 🏆 | Analysis: TUPOI converges reliably across 10,000 steps and achieves superior validation perplexity 46.39 vs 50.80 while utilizing 14% fewer parameters than the Transformer baseline. The power-law fit confirms that TUPOI scales predictably without hitting capacity saturation. Setup: Context memory footprint during sequential autoregressive token-by-token generation FP16 . | Generated Tokens | Transformer KV-Cache 24L, 16H | TUPOI Phase State | Memory Advantage | |---|---|---|---| 512 | 48.00 MB | 6.00 KB | | 2,048 | 192.00 MB | 6.00 KB | | 8,192 | 768.00 MB | 6.00 KB | | 32,768 | 3.00 GB | 6.00 KB | | 65,536 | 6.00 GB CUDA OOM | 6.00 KB | Analysis: Standard Transformers accumulate an TUPOI maintains a strictly invariant 6.00 KB state vector , allowing infinite streaming generation on consumer hardware. Setup: Forward-pass benchmark evaluating peak VRAM allocated MB measured directly via torch.cuda.max memory allocated . | Context Length | Baseline Transformer MB | TUPOI-300M VRAM MB | VRAM Saved | Status | |---|---|---|---|---| 512 tokens | 1,695.0 MB | 1,273.3 MB | -24.9% -421.7 MB | ✅ Verified | 1,024 tokens | 1,743.3 MB | 1,325.3 MB | -24.0% -418 MB | ✅ Verified | 2,048 tokens | 1,791.4 MB | 1,377.5 MB | -23.1% -414 MB | ✅ Verified | 4,096 tokens | 1,889.6 MB | 1,483.6 MB | -21.5% -406 MB | ✅ Verified | 8,192 tokens | 2,272.9 MB | 1,805.4 MB | -20.6% -467 MB | ✅ Verified | Analysis: TUPOI operates with 400 MB to 1.2 GB lower peak memory allocation across all sequence lengths. Across an 8,192 token context, TUPOI's total activation delta is merely ~500 MB above the static weights base ~1.21 GB . To maintain scientific rigor, we note the following current limitations: - Context Window Scale: The 300M model has been validated up to 8,192 token hardware contexts. Extreme horizons $ 128\text{k}$ require dedicated Flash-Verlet CUDA kernels. - Downstream Task Benchmarks: Supervised fine-tuning on reasoning benchmarks GSM8k, MetaMathQA is currently in progress. - Multi-Billion Scaling: Validating power-law behavior beyond 1B+ parameter regimes remains the primary objective for next iterations. This repository contains all necessary scripts to reproduce the findings reported above. git clone https://github.com/narelabs/TUPOI.git cd TUPOI pip install -r requirements.txt python scripts/generate cli.py python benchmarks/benchmark vram.py python scripts/plot benchmarks.py python scripts/generate dual axis charts.py @article{tupoi2026, title={TUPOI: Symplectic Post-Transformer Language Model with O 1 Memory}, author={NARE LABS Built by 15 y.o. }, year={2026}, journal={arXiv preprint}, url={https://github.com/narelabs/TUPOI} } Distributed under the MIT License .