{"slug": "thinking-machines-bets-on-efficiency-over-size-with-its-second-model-inkling", "title": "Thinking Machines bets on efficiency over size with its second model, Inkling Small", "summary": "Thinking Machines, the AI lab founded by former OpenAI CTO Mira Murati, released Inkling Small, an open-weights reasoning model that scores 40 on the Artificial Analysis Intelligence Index, one point below its larger sibling Inkling (41) but with less than a third of the parameters (276 billion total, 12 billion active). According to Artificial Analysis, no open model of equal or smaller size scores higher, and Inkling Small outperforms Inkling on several benchmarks, including Humanity's Last Exam (32% vs. 30%) and GPQA Diamond (89% vs. 87%), while being more token-efficient, averaging 24K output tokens per task compared to 45K for Deepseek V4 Flash and 78K for GPT-5.4 mini. The model handles text, image, and speech inputs, has a 256K-token context window, and is available under Apache 2.0 on Hugging Face.", "body_md": "# Thinking Machines bets on efficiency over size with its second model, Inkling Small\n\n**Thinking Machines, the AI lab from former OpenAI CTO Mira Murati, has released Inkling Small.** [According to Artificial Analysis](https://artificialanalysis.ai/articles/inkling-small-lands-within-a-point-of-inkling-on-the-artificial-analysis-intelligence-index-with-less-than-a-third-of-the-parameters), the open-weights reasoning model scores 40 on the Intelligence Index, one point below Inkling (41), with less than a third of the parameters (276 billion total, 12 billion active). AA says no open model of equal or smaller size scores higher.\n\nInkling Small beats its bigger sibling on several coding and reasoning tests, including Humanity's Last Exam (32% vs. 30%) and GPQA Diamond (89% vs. 87%). It falls behind on agent-based tasks and factual knowledge but is far more token-efficient, averaging 24K output tokens per task compared to 45K for [Deepseek V4 Flash](https://the-decoder.com/new-deepseek-flash-model-matches-openais-gpt-5-6-luna-at-roughly-60-percent-lower-cost/) and 78K for GPT-5.4 mini.\n\nThe model handles text, image, and speech inputs, has a 256K-token context window, and ships under Apache 2.0. Weights are on [Hugging Face](https://huggingface.co/thinkingmachines/Inkling-Small), and users can fine-tune it in the browser [via Tinker Playground](https://tinker.thinkingmachines.ai/playground?utm_source=blog&utm_campaign=inkling_small_model_release). Thinking Machines positions its models as a foundation for [fine-tuning with users' own data](https://the-decoder.com/thinking-machines-lab-ships-its-first-model-and-argues-interactivity-is-what-openai-gets-wrong-about-voice/). Some see this as [the next frontier in AI](https://the-decoder.com/ex-openai-researcher-bets-100-billion-will-flow-into-training-data-because-scaling-alone-wont-cut-it/).\n\n```\nAI News Without the Hype – Curated by Humans\n\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive \"AI Radar\" frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t\n\n\t\t\t\t\tSubscribe now\n```\n\n[Thinking Machines](https://thinkingmachines.ai/news/inkling-small/)|\n\n[Artificial Analysis](https://artificialanalysis.ai/articles/inkling-small-lands-within-a-point-of-inkling-on-the-artificial-analysis-intelligence-index-with-less-than-a-third-of-the-parameters)", "url": "https://wpnews.pro/news/thinking-machines-bets-on-efficiency-over-size-with-its-second-model-inkling", "canonical_source": "https://the-decoder.com/thinking-machines-bets-on-efficiency-over-size-with-its-second-model-inkling-small/", "published_at": "2026-07-31 17:41:50+00:00", "updated_at": "2026-07-31 18:09:51.172481+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products"], "entities": ["Thinking Machines", "Mira Murati", "Inkling Small", "Inkling", "Artificial Analysis", "Deepseek V4 Flash", "GPT-5.4 mini", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/thinking-machines-bets-on-efficiency-over-size-with-its-second-model-inkling", "markdown": "https://wpnews.pro/news/thinking-machines-bets-on-efficiency-over-size-with-its-second-model-inkling.md", "text": "https://wpnews.pro/news/thinking-machines-bets-on-efficiency-over-size-with-its-second-model-inkling.txt", "jsonld": "https://wpnews.pro/news/thinking-machines-bets-on-efficiency-over-size-with-its-second-model-inkling.jsonld"}}