MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models Researchers propose MIITA, a Memory-Induced Inference-Time Adaptation framework that enables small language models to perform continual learning under constrained storage without catastrophic forgetting. MIITA stores supervised experiences as compact correction-direction prototypes with semantic anchors and retrieves them at inference time using semantic and uncertainty-based cues, applying gated temporary hidden-state adaptation. Experiments across diverse supervised continual learning settings show MIITA consistently improves final performance and mitigates forgetting under fixed memory budgets. arXiv:2607.22556v1 Announce Type: new Abstract: Continual learning CL is essential for small language models SLMs to adapt to evolving real-world needs in resource-constrained deployments. However, directly updating their limited parameter space causes catastrophic forgetting. While memory-based methods naturally address this by decoupling knowledge retention from parameters, existing approaches designed for large language models LLMs rely on abundant storage and strong in-context reasoning that SLMs lack. To address these challenges, we propose MIITA, a Memory-Induced Inference-Time Adaptation framework for supervised CL under constrained storage. MIITA stores supervised experiences as compact correction-direction prototypes with semantic anchors, and retrieves them at inference time using semantic and uncertainty-based cues. The retrieved directions are applied through gated temporary hidden-state adaptation, enabling non-destructive reuse of past supervision without backbone updates, prompt extensions, or test-time backpropagation. A local theoretical analysis links this design to first-order loss reduction, uncertainty-guided retrieval, and directional coverage for retaining old-stage knowledge. Extensive experiments across diverse supervised CL settings show that MIITA consistently improves final performance and mitigates forgetting under fixed memory budgets.