Show HN: CellularFlow – Continual-learning LLM using associative memory CellularFlow, a continual-learning LLM architecture from developer celcilin, replaces dense feed-forward networks with Multi-Head Associative DNA Memory Banks and an Episodic Memory Slot Buffer, achieving 83.9% retention across sequential domains versus 61.8% for standard Transformers, and enabling zero-backprop streaming learning during inference. The open-source project, available on GitHub, supports three modes: live learning via EMA, selective fine-tuning that freezes ~85% of the backbone, and episodic fact injection with temporal decay. CellularFlow is a memory-augmented neural architecture designed as a continual-learning alternative to standard Transformers. By replacing dense Feed-Forward Networks FFN/MLP with Multi-Head Associative DNA Memory Banks and an Episodic Memory Slot Buffer , CellularFlow decouples factual knowledge storage from sequence reasoning. It achieves state-of-the-art catastrophic forgetting mitigation 83.9% retention across sequential domains and enables zero-backprop streaming learning during inference. | Feature | Standard Transformer LLaMA/GPT | CellularFlow v4 | |---|---|---| | Parametric Architecture | Dense FFN / SwiGLU | Multi-Head Associative DNA Memory CMCLayer | | Sequential Adaptation | Severe catastrophic forgetting 61.8% retention | 83.9% retention via Selective Fine-Tuning Mode 2 | | Real-time Live Learning | ❌ Impossible without retraining | ✅ Mode 1: EMA streaming forward update 0 backprop | | Instant Fact Injection | ❌ Requires finetuning or external RAG | ✅ Mode 3: Episodic slot buffer with decay & consolidation | | Sequence Attention | | | | Inference Efficiency | Full recompute or dense KV cache | Decoupled memory lookup + incremental KV-cache | | Knowledge Inspectability | Diffuse, entangled weights | Discrete, addressable, and prunable memory slots | CellularFlow fuses two computational pathways into a unified Hybrid CMC Layer : Input Sequence: X B, T, d │ ┌──────────────┴──────────────┐ ▼ ▼ ┌─────────────────────────┐ ┌─────────────────────────┐ │ Multi-Head DNA Memory │ │ Episodic Memory Slot │ │ Associative Banks │ │ Buffer Fast-Write │ └────────────┬────────────┘ └────────────┬────────────┘ │ │ └──────────────┬──────────────┘ │ Gated Memory Enrichment ▼ ┌───────────────────────────────────────────────────────┐ │ Causal Multi-Head Self-Attention with RoPE FlashAttn │ └───────────────────────────┬───────────────────────────┘ │ ▼ Output Sequence: Y B, T, d Each head - Specialized Subspaces: Heads specialize independently across syntax, semantics, and domain knowledge. - Exploration Noise: Gaussian perturbation prevents dead memory slots during Top-K sparse routing. - Mode 1 — Live Learning inference write=True : Updates DNA memory values on the fly during inference via Exponential Moving Average EMA with zero backward pass. Protected by Spherical Anisotropy Regularization to prevent vector collapse. - Mode 2 — Selective Fine-Tuning set mode "selective" : Freezes ~85% of the backbone projections, embeddings, LayerNorms and trains only the DNA banks. Retains foundational knowledge while rapidly absorbing new domains. - Mode 3 — Episodic Fact Injection inject fact : Writes facts into slot-based episodic memory with temporal age decay exp -0.005 age and consolidates top facts into DNA banks post-epoch. Requires Python $\ge$ 3.11 and PyTorch$\ge$ 2.4.0 . Clone the repository git clone https://github.com/celcilin/cellularflow.git cd cellularflow Install dependencies using UV recommended or pip pip install -e . For GPU acceleration CUDA 12.4+ : pip install torch --index-url https://download.pytorch.org/whl/cu124 python import torch from cellularflow import CellularFlowLM, CellularFlowTrainer, BPEDataset 1. Initialize tokenizer & dataset dataset = BPEDataset "Alice was beginning to get very tired of sitting by her sister...", context len=256 2. Instantiate CellularFlow LM model = CellularFlowLM vocab size = dataset.vocab, dim = 512, n layers = 6, n heads = 8, n entries = 128, context len = 256, use episodic = True 3. Pretraining trainer = CellularFlowTrainer model, dataset, device="cuda" if torch.cuda.is available else "cpu" trainer.pretrain epochs=100, seed dna=True 4. Fast Generation with KV-Cache prompt = "The journey into" print trainer.generate prompt, max new=100, temperature=0.8 5. Continual Learning: Mode 3 Fact Injection trainer.inject fact "The hidden archives are kept inside Vault 42." 6. Continual Learning: Mode 2 Selective Fine-Tuning Backbone Frozen trainer.selective finetune "Technical medical notes on neurology...", epochs=10 7. Continual Learning: Mode 1 Live Streaming Learning 0 Backprop trainer.live learn "Streaming log telemetry received in real time..." CellularFlow includes an interactive glassmorphic web dashboard for real-time inference, fact injection, and memory inspection: Start the FastAPI server uvicorn server.app:app --host 0.0.0.0 --port 8000 Open http://localhost:8000 in your browser to interactively generate text, inspect layer-wise episodic slot utilization, and test live fact injections. Interactive CLI Playground python analysis.py --checkpoint checkpoint/CMC BaseModel.pt --interactive Evaluated on a standardized 62KB multi-domain corpus: | Architecture | Parameters | Perplexity | Accuracy | |---|---|---|---| | GPT-mini Vanilla Transformer | 810K | 8.51 | 36.4% | | CellularFlow v4 Hybrid CMC | 379K 2.1× fewer | 2.54 −70.3% | 73.7% | Trained sequentially across Literature , Science , History , Technical , and Poetry : | Fine-Tuning Strategy | Overall Domain Retention | |---|---| | Full Fine-Tuning All Weights | 61.8% | | Mode 2: Selective DNA Fine-Tuning | 83.9% +22.1 pp | cellularflow/ ├── cellularflow/ │ ├── core.py CMCLayer, HybridCMCLayer, EpisodicMemory, CellularFlowLM │ ├── trainer.py Pretraining, selective fine-tuning, live learning, mixed precision │ ├── extensions.py Blockwise Attention, Compressed KV CKV , MTP, Beaconing │ ├── corpus.py Multi-domain benchmark corpora │ └── swarm.py DNASwarm evolutionary optimizer ├── benchmarks/ │ └── evaluate checkpoint.py Evaluation harness for perplexity, accuracy, and memory norms ├── server/ │ └── app.py FastAPI server + WebSocket endpoint ├── dashboard/ │ ├── index.html Web dashboard UI │ ├── app.js Frontend WebSocket and API client │ └── styles.css Dark glassmorphic design system ├── sft/ │ ├── sft dataset.py ChatML templates and target loss masking │ └── sft trainer.py Supervised fine-tuning curriculum engine ├── scripts/ │ ├── train tokenizer.py ByteLevelBPE tokenizer builder │ └── test extensions.py Architecture extension verification ├── analysis.py CLI exploration & interactive REPL ├── pyproject.toml Project build & dependency definitions └── CONTRIBUTING.md Contribution guidelines & developer standards We welcome contributions from researchers, engineers, and developers worldwide Please review CONTRIBUTING.md /celcilin/cellularflow/blob/main/CONTRIBUTING.md for instructions on setting up your environment, adhering to XLA/TPU graph rules, and submitting pull requests. Celcilin C S - GitHub: @celcilin https://github.com/celcilin - Email: celcilin204@gmail.com mailto:celcilin204@gmail.com This project is licensed under the MIT License — see the LICENSE /celcilin/cellularflow/blob/main/LICENSE file for details.