{"slug": "meet-microsofts-mai-thinking-1-what-it-is-and-how-it-works", "title": "Meet Microsoft’s MAI-Thinking-1: What It Is and How It Works", "summary": "Microsoft introduced MAI-Thinking-1, a 35B-active, ~1T-total parameter sparse Mixture of Experts reasoning model that matches leading models on software engineering benchmarks and is preferred to Sonnet 4.6 in blind human evaluations. The model, trained without distillation from third-party models on clean, traceable data, achieves 97.0% on AIME 2025 and 94.5% on AIME 2026, and is part of Microsoft's broader effort toward Humanist Superintelligence.", "body_md": "Today Microsoft is introducing MAI-Thinking-1, Microsoft AI’s reasoning model. It is a medium-sized model that stands among the strongest models in its weight class. It matches leading models on key software engineering benchmarks, demonstrates advanced mathematical reasoning capabilities, and is preferred to Sonnet 4.6 in our blind human side-by-side evaluations. We don’t distill from other labs and we don’t rely on opaque data. Their datasets are clean, traceable, and enterprise-grade.\n\nMAI-Thinking-1 is a step in our broader work to build towards Humanist Superintelligence: advanced AI capabilities designed to serve people and organizations, not to replace them. The model matters on both axes: what it can do, and how it was built.\n\nMore than a single model, Microsoft introduced their Hill-Climbing Machine: a co-designed pipeline built to make every component of model development climbable, so capabilities improve continually and reliably over time. The aim is a repeatable system that can absorb better data, stronger rewards, more capable environments, and more compute.\n\n**Three main pillars guide of MAI’s philosophy.**\n\n**First, capabilities should be learned, not inherited.** Although faster to acquire, inherited intelligence lacks the steerability essential for real world usage: an imitator is fundamentally tied to the design choices of its teacher and struggles to adapt to new situations. MAI-Thinking-1 was trained without distillation from third party models, forcing our model to truly learn the tasks at hand.\n\n**Second, clean data.** They trained it from the ground up on clean, traceable and enterprise-grade data, without distillation from third-party models. This matters for quality, provenance, and control. Microsoft says “If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it”.\n\n**Third, self-sufficiency across the entire stack.** All the way from co-design of MAI’s models with MSFT’s own accelerators through to their reinforcement learning framework, they have focused efforts on in-house training infrastructure. This is a crucial part of building the hill-climbing machine, to ensure they can fully optimize and shape MAI’s systems end-to-end to best serve the needs.\n\nMAI-Thinking-1 is a 35B-active, ~1T-total parameters, sparse Mixture of Experts model, a smaller inference footprint than much larger models. Despite this, this model is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro. That matters for developers and enterprises because model size determines where advanced coding assistance can be deployed, how often it can be used, and whether it can move from exceptional tasks into daily workflows.\n\nMicrosoft have invested heavily in the training environments needed for agentic coding. Each verified environment is deterministic, executable, and graded by real test suites. This gives the model practice on the kind of multi-step work developers actually do: reading code, editing files, running tests, observing failures, and recovering from intermediate mistakes.\n\nMAI-Thinking-1 reaches 97.0% on AIME 2025, and 94.5% on AIME 2026, showing strong mathematical and scientific reasoning for its weight class. Strong performance here gives them confidence that MAI’s training loop can create real reasoning gains — climbing all the way from the ground up — from their own data, rewards, and evaluation process, enabling this intelligence to generalize to other domains over time.\n\nPeople care about whether a model understands the task, follows instructions, uses the right level of detail, writes clearly, and respects their time.\n\nMicrosoft have built a blind side by side human evaluation with one of their partners, Surge, using their pool of professional raters to measure various models on these traits. The evaluation spanned 1,276 tasks across a wide variety of use cases in both single-turn and multi-turn conversations, with a focus on measuring how helpful each response is and whether it actually advances the user’s goals. In these evaluations, users preferred MAI-Thinking-1 over Claude Sonnet 4.6.\n\nThis has been a core focus of post-training. Microsoft want the model to be capable without being brittle, concise without being incomplete, and helpful without overreaching. Human preference data gives us a direct signal on whether benchmark improvements translate into better experiences for users.\n\nMAI-Thinking-1 is built with enterprise readiness in mind. It supports long context with a 256k token window (enough to fit a 600 page document), function calling, and the flexibility to add developer instructions. We trained the model to follow multiple layers of instructions and aligned its default style to enterprise needs. It’s compatible with the widely used Chat Completions API. All MAI models come with enterprise-grade security and compliance through Microsoft Foundry.\n\nMicrosoft have report results in two views: post-trained MAI-Thinking-1 evaluations, and pre-training metrics for their base model.\n\n**Table 1. MAI-Thinking-1 metrics**\n\nPost-trained model evaluation results on public STEM and agentic coding benchmarks. Other model numbers are taken from respective official model cards. Scores are percentages unless otherwise noted; dashes indicate unavailable model values.\n\n**Table 2. Pre-training metrics**\n\nMicrosoft is building towards Humanist Superintelligence: advanced AI capabilities designed to serve people and organizations, not replace them. Their models must remain subordinate technologies under human control with the goal of upholding human autonomy and being helpful. That means MAI models must not refuse legitimate requests under the guise of safety and compliance as then they are not truly serving humans.\n\nStriking the delicate balance between being helpful and safe is not easy. For MAI-Thinking-1, Microsoft has aimed to achieve this balance by treating unsafe compliance and unnecessary refusal as defects in the same reward construction where aggregation is based on severity of potential of harm. Safety is trained with the same reinforcement learning infrastructure used for capability, so safety rewards are part of the same hill-climbing loop ensuring safety is always aligned to the capabilities and not incidental.\n\nAs a result, Microsoft see that their model can balance ensuring a safety bar on sensitive unsafe requests while also being helpful on non-sensitive content.\n\nMAI-Thinking-1 is now available in public preview. Try it now in [Microsoft Foundry](https://aka.ms/mai-thinking-1-foundrycard)\n\nTo Summarize, with cost-efficient reasoning for a wide-range of intensive enterprise tasks, MAI achieves SOTA performance on maths, knowledge and coding for its weight class.\n\nThe model provides clean, traceable and enterprise-grade data. It uses Microsoft Foundry’s integrated evaluation, observability, safety and deployment capabilities, making it a perfect fit for enterprise use cases needing quality, provenance, control and cost efficiency.\n\n[Meet Microsoft’s MAI-Thinking-1: What It Is and How It Works](https://pub.towardsai.net/meet-microsofts-mai-thinking-1-what-it-is-and-how-it-works-3e84374d0d89) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/meet-microsofts-mai-thinking-1-what-it-is-and-how-it-works", "canonical_source": "https://pub.towardsai.net/meet-microsofts-mai-thinking-1-what-it-is-and-how-it-works-3e84374d0d89?source=rss----98111c9905da---4", "published_at": "2026-08-13 22:01:02+00:00", "updated_at": "2026-08-13 22:17:31.435819+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products"], "entities": ["Microsoft", "MAI-Thinking-1", "Sonnet 4.6", "Claude Opus 4.6", "Surge", "AIME 2025", "AIME 2026"], "alternates": {"html": "https://wpnews.pro/news/meet-microsofts-mai-thinking-1-what-it-is-and-how-it-works", "markdown": "https://wpnews.pro/news/meet-microsofts-mai-thinking-1-what-it-is-and-how-it-works.md", "text": "https://wpnews.pro/news/meet-microsofts-mai-thinking-1-what-it-is-and-how-it-works.txt", "jsonld": "https://wpnews.pro/news/meet-microsofts-mai-thinking-1-what-it-is-and-how-it-works.jsonld"}}