{"slug": "i-pretrained-the-best-small-model-under-100m-parameters", "title": "I pretrained the best small model under 100M parameters", "summary": "BarunLM, a 35-million-parameter language model introduced by developer Barun, outperforms models over 6x larger, including LiquidAI's lfm2.5-230m, while training on a single H200 GPU. The model's architecture, inspired by DeepSeek and Kimi, combines hybrid attention, grouped query attention, partial RoPE, gated attention, bounded SwiGLU, and residual selectors, and is further refined with reinforcement learning using verifiable rewards.", "body_md": "introducing BarunLM (35m) 🤗\ni pretrained the world's best small model under 100m parameters.\na 35m parameter language model that outperforms models over 6x larger, including LiquidAI's lfm2.5-230m, while training on a single H200 gpu.\nthe gains come from a better architecture. we took inspiration from deepseek, kimi, and other frontier open source models.\nBarunLM combines hybrid local and global attention, grouped query attention to reduce memory, partial rope, gated attention, bounded swiglu, and residual selectors to improve how information flows through the network.\nafter pretraining, we further improved the model with rlvr (reinforcement learning with verifiable rewards) using deterministic verifiers and rloo, allowing the model to reinforce correct reasoning without relying on human preference labels.", "url": "https://wpnews.pro/news/i-pretrained-the-best-small-model-under-100m-parameters", "canonical_source": "https://twitter.com/HarshalsinghCN/status/2083223503554449745", "published_at": "2026-07-31 19:00:05+00:00", "updated_at": "2026-07-31 19:22:23.426112+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "ai-research"], "entities": ["BarunLM", "LiquidAI", "lfm2.5-230m", "DeepSeek", "Kimi", "H200 GPU"], "alternates": {"html": "https://wpnews.pro/news/i-pretrained-the-best-small-model-under-100m-parameters", "markdown": "https://wpnews.pro/news/i-pretrained-the-best-small-model-under-100m-parameters.md", "text": "https://wpnews.pro/news/i-pretrained-the-best-small-model-under-100m-parameters.txt", "jsonld": "https://wpnews.pro/news/i-pretrained-the-best-small-model-under-100m-parameters.jsonld"}}