cd /news/large-language-models/minimax-quietly-ships-a-coding-only-… · home › topics › large-language-models › article
[ARTICLE · art-140396] src=startupfortune.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

MiniMax quietly ships a coding-only model as China's AI models flood the market

MiniMax confirmed on September 27 that M3.1-Flash-Preview, a coding-only model, had gone live inside its MiniMax Code assistant with no model card, benchmarks or published price, a day after developers spotted it in the product's model picker. The release contrasts with MiniMax's June launch of M3, a 428-billion-parameter mixture-of-experts model with roughly 23 billion active parameters per token, a million-token context window and native image and video input, which shipped with architecture details, SWE-bench Verified and SWE-Bench Pro scores of 80.5% and 59.0%, and API pricing of about $0.30 per million input tokens and $1.20 per million output tokens. MiniMax is shipping the Flash tier into a crowded Chinese market that already includes Alibaba's Qwen3.8-Max, a 2.4-trillion-parameter flagship released August 3 at $2 per million input tokens, and Zhipu's GLM-5.3, released August 14.

by read5 min views1 publishedSep 27, 2026
MiniMax quietly ships a coding-only model as China's AI models flood the market
Image: Startupfortune (auto-discovered)

MiniMax slipped a new coding model into its developer tool with no announcement, no benchmarks and no price. That quiet rollout says more about the state of the China AI race than a launch event would have.

On September 27, MiniMax's official account confirmed what developers had already spotted in the product: M3.1-Flash-Preview had gone live inside MiniMax Code, the company's coding assistant. The company's post called it fast, reliable and "ready for real work, from quick bug fixes to full features." That's it. No model card. No benchmark table. No price per million tokens. Chinese AI watchers on X had noticed the model appear in MiniMax Code's picker a full day before MiniMax said anything at all, according to posts from accounts tracking the rollout on the platform.

That's not how MiniMax shipped its last flagship. When M3 launched in June, the company published architecture details, SWE-bench scores and a pricing sheet that undercut rivals by 90%. This time, the API endpoint for M3.1-Flash-Preview is gated, and the only way to try it is inside MiniMax's own product, with reasoning levels running from low up through a new "max" tier that sits alongside the older M3, M2.7 and M2.7-highspeed options.

You'd think a company competing for developer mindshare would want the noise. Instead, MiniMax appears to be A/B testing a product feature and letting the internet find out. That's either confidence that the model will speak for itself once people use it, or a sign that this release isn't the big one, just a faster, cheaper sibling of M3 built to eat routine coding tasks while the real flagship work happens elsewhere.

The base M3 model gives a sense of what MiniMax is capable of shipping when it does want attention. It's a 428-billion-parameter mixture-of-experts model with roughly 23 billion active per token, a million-token context window, and native image and video input. On SWE-bench Verified, MiniMax reported 80.5%. On the harder SWE-Bench Pro, it scored 59.0%, which the company says puts it ahead of GPT-5.5 and Gemini 3.1 Pro on that specific test. Those are MiniMax's own numbers, run on its own infrastructure, so treat them as a claim rather than an independent verdict. The API price for M3 tells its own story regardless: about $0.30 per million input tokens and $1.20 per million output through MiniMax directly, or as low as $0.23 and $0.96 through OpenRouter. That's roughly 5% to 10% of what a comparable proprietary model charges.

MiniMax's revenue surge shows China's AI labs are learning to sell MiniMax has more than doubled sales as it prepares a new model launch, showing that Chinese AI labs are starting to turn technical momentum into real revenue. - how to monetize AI models at scale - Chinese AI startups revenue growth strategy

Flash is the version built to make that speed and price advantage sharper for everyday coding, the quick bug fix and small feature work that a developer runs through dozens of times a day rather than the hard agentic task you save the big model for. Splitting a model line this way, a large frontier version and a fast cheap version, isn't new. OpenAI and Google both do it. What's notable is that MiniMax is doing it in open weight and doing it for essentially free to try, at a moment when every major Chinese lab is racing to own the entry-level coding workload.

The bigger story is the whole shelf, not one model #

MiniMax isn't shipping into a quiet market. Alibaba put out Qwen3.8-Max on August 3, a 2.4-trillion-parameter flagship priced at $2 per million input tokens, then followed it on August 12 with the first open-weight version of a Max-tier Qwen model. Zhipu's GLM-5.3 landed August 14 with a million-token context window of its own. DeepSeek and Kimi are running the same playbook: ship often, price low, publish weights. MiniMax's Flash drop, quiet as it was, fits the same pattern of near-constant releases designed to keep developers evaluating Chinese models instead of settling on one Western default.

And developers are noticing. According to CNBC's reporting on OpenRouter data, Chinese models accounted for 57% to 67% of total token usage on the platform for the week that included September 14, up from just 6% to 13% back in February. Vercel saw a similar jump, with Chinese models' share of usage rising to 55% in August from 11% in January. The reason is not mysterious: for the agentic coding and customer-service workloads that now eat most enterprise AI budgets, Chinese open models are running 60% to 90% cheaper than the leading U.S. alternatives, per CNBC's analysis, and that gap is large enough that price-sensitive teams stop treating it as a rounding error.

Washington is watching that shift with more alarm than enthusiasm. Daniel Remler, a senior fellow at the Center for a New American Security, told CNBC that Chinese AI adoption represents

Also read: Why Developers Are Running AI Models on Mac Minis Instead of Nvidia GPUs • Tesla Finally Puts the Semi Into Volume Production, Seven Years Late • Nubank is in early talks to buy Monzo at up to £10 billion

This article is posted in Entrepreneurship News, check it out for more related stories.

CZI bets on an open AI model for drug discovery CZI's new AI world model for drug discovery arrives as the sector turns into a capital-heavy race between open scientific infrastructure and proprietary platforms. - open source AI models for drug discovery - protein biology AI platform for academic research labs

Join the discussion #

Open in the community → Almost there. Sign in and your reply posts straight away.

── more in #large-language-models 4 stories · sorted by recency
── more on @minimax 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/minimax-quietly-ship…] indexed:0 read:5min 2026-09-27 · —