cd /news/large-language-models/tupoi-an-attention-free-llm-with-str… Β· home β€Ί topics β€Ί large-language-models β€Ί article
[ARTICLE Β· art-98172] src=github.com β†— pub= topic=large-language-models verified=true sentiment=↑ positive

Tupoi: An attention-free LLM with strictly O(1) memory and 6 KB state

NARE LABS, built by a 15-year-old, released TUPOI, an attention-free language model that replaces dense attention with a Symplectic Hamiltonian Integrator, achieving strictly O(1) memory with a fixed 6 KB state. In benchmarks, TUPOI-300M achieved a validation perplexity of 46.39 versus 50.80 for a 355M-parameter Transformer baseline, using 14% fewer parameters, and maintained constant memory usage across generation lengths up to 65,536 tokens, where the Transformer's KV-cache exceeded 6 GB and caused CUDA OOM.

read3 min views1 publishedAug 15, 2026
Tupoi: An attention-free LLM with strictly O(1) memory and 6 KB state
Image: Michielbdejong (auto-discovered)

A sub-quadratic, attention-free sequence modeling architecture that replaces dense attention matrices with a Symplectic Hamiltonian Integrator (Velocity-Verlet).

Author: NARE LABS (Built by 15 y.o.)

This repository contains the reference implementation, pre-trained weights, and empirical benchmark suite for the TUPOI language model, a post-transformer architecture that completely eliminates the Attention mechanism and the persistent KV-cache in favor of continuous Hamiltonian phase-space dynamics.

The primary objective of this research is to investigate whether modeling sequence context as particles moving through a symplectic phase space ** $O(1)$ memory consumption** and zero state dissipation (

TUPOI replaces

The hidden state is decomposed into canonical coordinates Velocity-Verlet Leapfrog dynamical scheme:

Because the Jacobian determinant of this transformation satisfies Information is neither created nor destroyed during the forward pass.

To stabilize language representations over infinite generation horizons:

IgnoranceGate: A learned smooth sigmoid filter$x' = x \odot \sigma(W_g x)$ that attenuates token noise before perturbing phase space. - OpinionAnchor: A running exponential moving average$a_t = (1-\eta)a_{t-1} + \eta \cdot \bar{x}'$ acting as a canonical phase attractor.

Our evaluations focus on rigorous, apples-to-apples comparisons against a standard GPT-style Transformer (BaselineLM) under identical training constraints (seed, optimizer, token budget, parameter scale).

Setup: GPT-2 BPE tokenizer (vocab: 50,257), seq_len=512, trained for 10,000 steps on Tesla T4.

Model Tier Active Parameters Final Val Loss Perplexity (PPL) Complexity Status
TUPOI-17M (S) 16.9 M 4.1516 63.53 Verified
TUPOI-47M (M) 47.0 M 4.1078 60.81 Verified
BaselineLM (Transformer)
355.0 M
3.9280 50.80 Baseline
TUPOI-300M (L)
304.2 M
3.8371
46.39
Winner πŸ†

Analysis: TUPOI converges reliably across 10,000 steps and achieves superior validation perplexity (46.39

vs 50.80

) while utilizing 14% fewer parameters than the Transformer baseline. The power-law fit confirms that TUPOI scales predictably without hitting capacity saturation.

Setup: Context memory footprint during sequential autoregressive token-by-token generation (FP16).

| Generated Tokens ( | Transformer KV-Cache (24L, 16H) | TUPOI Phase State | Memory Advantage | |---|---|---|---| 512 | 48.00 MB | 6.00 KB | | 2,048 | 192.00 MB | 6.00 KB | | 8,192 | 768.00 MB | 6.00 KB | | 32,768 | 3.00 GB | 6.00 KB | | 65,536 | 6.00 GB (CUDA OOM) | 6.00 KB |

Analysis: Standard Transformers accumulate an TUPOI maintains a strictly invariant 6.00 KB state vector, allowing infinite streaming generation on consumer hardware.

Setup: Forward-pass benchmark evaluating peak VRAM allocated (MB) measured directly via torch.cuda.max_memory_allocated().

Context Length Baseline Transformer (MB) TUPOI-300M VRAM (MB) VRAM Saved Status
512 tokens
1,695.0 MB 1,273.3 MB
-24.9% (-421.7 MB)
βœ… Verified
1,024 tokens
1,743.3 MB 1,325.3 MB
-24.0% (-418 MB)
βœ… Verified
2,048 tokens
1,791.4 MB 1,377.5 MB
-23.1% (-414 MB)
βœ… Verified
4,096 tokens
1,889.6 MB 1,483.6 MB
-21.5% (-406 MB)
βœ… Verified
8,192 tokens
2,272.9 MB 1,805.4 MB
-20.6% (-467 MB)
βœ… Verified

Analysis: TUPOI operates with 400 MB to 1.2 GB lower peak memory allocation across all sequence lengths. Across an 8,192 token context, TUPOI's total activation delta is merely ~500 MB above the static weights base (~1.21 GB).

To maintain scientific rigor, we note the following current limitations:

Context Window Scale: The 300M model has been validated up to 8,192 token hardware contexts. Extreme horizons ($>128\text{k}$ ) require dedicated Flash-Verlet CUDA kernels. - Downstream Task Benchmarks: Supervised fine-tuning on reasoning benchmarks (GSM8k, MetaMathQA) is currently in progress. - Multi-Billion Scaling: Validating power-law behavior beyond 1B+ parameter regimes remains the primary objective for next iterations.

This repository contains all necessary scripts to reproduce the findings reported above.

git clone https://github.com/narelabs/TUPOI.git
cd TUPOI
pip install -r requirements.txt
python scripts/generate_cli.py
python benchmarks/benchmark_vram.py
python scripts/plot_benchmarks.py
python scripts/generate_dual_axis_charts.py
@article{tupoi2026,
  title={TUPOI: Symplectic Post-Transformer Language Model with O(1) Memory},
  author={NARE LABS (Built by 15 y.o.)},
  year={2026},
  journal={arXiv preprint},
  url={https://github.com/narelabs/TUPOI}
}

Distributed under the MIT License.

── more in #large-language-models 4 stories Β· sorted by recency
── more on @nare labs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/tupoi-an-attention-f…] indexed:0 read:3min 2026-08-15 Β· β€”