{"slug": "opt-gear-technical-report", "title": "Opt.Gear Technical Report", "summary": "Opt.Gear, a new foundation model family from an unnamed research team, achieves up to 4.9x faster prefill and decoding speeds on NPUs compared to similar-scale models, with dense models of 1M, 270M, and 1B parameters supporting a 64K context length. The models are trained on a curated 0.5T token subset from a 2T token corpus, and all weights and deployment binaries are open-sourced for ONNX, Qualcomm NPU, and Apple ANE. The smallest variant, Opt.Gear-1M, is the first generative language model to reach 20 tokens per second with W4A32 quantization on the ARM Cortex-M7 CPU of the STM32H747I-DISCO.", "body_md": "arXiv:2608.01034v1 Announce Type: new\nAbstract: We introduce Opt.Gear, a foundation model designed for efficient on-device deployment, real-tim inference, and strong task capability. It includes a dense model (1M, 270M, and 1B) with a context length of 64K. We designed a new hybrid architecture that combines a convolutional key-value gated mixer with local-global attention to reduce the KV-cache memory that tends to increase exponentially with long context. This architecture delivers up to X4.9 faster prefill and decoding speeds on the NPUs compared to models of a similar scale models. From a 2T tokens candidate corpus, Opt.Gear is trained on a curated 0.5T tokens subset without knowledge distillation. This is the most data-efficient of the existing foundation models. All models are released with open weights and deployment binaries for ONNX, Qualcomm NPU, and Apple ANE making Opt.Gear a practical base for edge applications that need fast, memory-efficient inference and strong task capabilities. Furthermore, to expand the ecosystem of on-device generative language models, we are introducing the Opt.Gear-1M that can be deployed on Micro-Controller Units (MCUs), a Tiny Language Model (TLM). Opt.Gear-1M is the first generative language model to achieve 20 TPS with W4A32 quantization on the ARM Cortex-M7 CPU of the STM32H747I-DISCO.", "url": "https://wpnews.pro/news/opt-gear-technical-report", "canonical_source": "https://www.machinebrief.com/news/optgear-technical-report-ujs6", "published_at": "2026-08-04 04:00:00+00:00", "updated_at": "2026-08-04 06:33:03.397231+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-research"], "entities": ["Opt.Gear", "Qualcomm", "Apple", "ONNX", "ARM Cortex-M7", "STM32H747I-DISCO"], "alternates": {"html": "https://wpnews.pro/news/opt-gear-technical-report", "markdown": "https://wpnews.pro/news/opt-gear-technical-report.md", "text": "https://wpnews.pro/news/opt-gear-technical-report.txt", "jsonld": "https://wpnews.pro/news/opt-gear-technical-report.jsonld"}}