{"slug": "ai-model-radar-1", "title": "AI Model Radar #1", "summary": "Moonshot's Kimi K3, a 2.8 trillion parameter open-weight model with weights promised for July 27, leads the July 2026 AI Model Radar, which also includes Thinking Machines Inkling (975B), Tencent Hy3 (295B MoE), NVIDIA Nemotron-Labs-Audex-2B (2B), and Prism ML Ternary Bonsai 27B (27B ternary). The radar independently verifies claims and provides actionable verdicts, noting that K3's API is already live and has scored number one in frontend and agentic coding across six of seven Arena sub-domains, while Inkling and Audex-2B are downloadable now.", "body_md": "AI\n\n# AI Model Radar #1\n\nKimi K3 leads July 2026's radar, with Thinking Machines Inkling, Tencent Hy3, NVIDIA Audex-2B, and Ternary Bonsai 27B. What is verified and what to try first.\n\n## On the radar: five new AI models for July 2026[#](#on-the-radar-five-new-ai-models-for-july-2026)\n\nOne industry tracker counts [five major open-weight moves between July 4 and July 17](https://www.digitalapplied.com/blog/open-weight-model-wave-july-2026-momentum-tracker) alone, and that pace is exactly why a triage exists. This first issue of AI Model Radar picks, from all of June and July 2026, the five new AI models most worth your time - not to repeat the news, but to sort what is independently verified from what is still a vendor claim sheet, and to say what you should actually do about each one. That is the series premise: a numbered digest of new model releases, published whenever enough of them clear the bar rather than on a calendar.\n\nThe bar for this issue: a verdict you can act on this week. For four of the five entries that means weights you can download today. The fifth, Moonshot’s Kimi K3, gets the lead slot without downloadable weights - its API is live, independent leaderboards have already scored it, and the promised July 27 weights drop would redraw what “open-weight frontier” means. Waiting for the weights would mean covering it two weeks after everyone else - serving the calendar, not the reader.\n\n**This issue:** Moonshot’s Kimi K3 (a 2.8T frontier claim with weights promised July 27), Thinking Machines’ Inkling (a 975B open flagship positioned for fine-tuning), Tencent’s Hy3 (295B MoE with bold but unverified claims), NVIDIA’s Nemotron-Labs-Audex-2B (the whole voice loop in one small model), and Prism ML’s Ternary Bonsai 27B (27B-class reasoning in 7.2GB).\n\n## The radar at a glance[#](#the-radar-at-a-glance)\n\n| Model | Params (total / active) | Context | License | Serving floor |\n|---|---|---|---|---|\n| Moonshot Kimi K3 | 2.8T / undisclosed (MoE) | 1M | Modified MIT (expected) | API only (weights promised July 27) |\n| Thinking Machines Inkling | 975B / 41B (MoE) | 1M | Apache 2.0 | Multi-GPU cluster, ~520GB quantized |\n| Tencent Hy3 | 295B / 21B (MoE) | 256K | Apache 2.0 | 8-GPU node, ~300GB at FP8 |\n| NVIDIA Nemotron-Labs-Audex-2B | 2B (dense) | 128K | NVIDIA noncommercial | Single GPU, ~4GB |\n| Prism ML Ternary Bonsai 27B | 27B (dense, ternary) | 262K | Apache 2.0 | Laptop, 7.2GB |\n\nThe serving-floor column is the issue’s real axis. Three of these models need datacenter hardware no matter how clever their sparsity is; two fit on hardware you already own. The footprints for Inkling and Audex-2B are arithmetic from their published checkpoints, K3’s is arithmetic for a 4-bit checkpoint that does not exist yet, and Hy3 and Bonsai are vendor-published sizes - and the gap between the groups is not a spectrum, it is a cliff.\n\n## Moonshot Kimi K3: a 2.8T frontier claim you can verify before the weights land[#](#moonshot-kimi-k3-a-28t-frontier-claim-you-can-verify-before-the-weights-land)\n\n**Use cases:**\n\n**Frontend and agentic coding**- the one domain with independent number-one evidence, across 6 of 7 Arena sub-domains** Long-horizon work over huge repos and documents**- the 1M-token context with reasoning on by default** Multimodal pipelines over screenshots, mockups, and video**- native image and video input** General chat and knowledge work**- very good, but ninth on the Text Arena at frontier prices\n\nThe biggest open-weights promise ever made is five days old, and you do not have to take anyone’s word for it. Moonshot [announced Kimi K3 on July 16](https://simonwillison.net/2026/Jul/16/kimi-k3/) - 2.8 trillion parameters, which would make it the largest open-weight model ever released if the weights land on July 27 as promised - and because the API went live the same day, third parties have been scoring it all week.\n\nThe scale is managed by aggressive sparsity: a mixture-of-experts design routes each token to 16 of 896 experts, under 2% of the pool. Moonshot has not disclosed the active-parameter count - the community shorthand circulating is 2.8T-A50B, roughly 50 billion active, but treat that as directional. The rest of the sheet: a 1,048,576-token context window, text, image, and video input, reasoning on by default, and two architecture changes Moonshot credits for the efficiency (Kimi Delta Attention and Attention Residuals). API pricing sits at $3 per million input tokens and $15 per million output.\n\nK3 inverts this issue’s verification problem. Hy3 has weights you can download and numbers you cannot check yet; K3 has numbers you can check today and weights you have to wait for. [Artificial Analysis](https://artificialanalysis.ai/models/kimi-k3) scores it 57 on its Intelligence Index, fourth among 189 models, and Arena’s blind developer testing puts it [first on the Frontend Code Arena at 1,679 points](https://x.com/arena/status/2077824029126504525) - a 17-place jump from K2.6’s 18th, ahead of the closed frontier models it is priced against. The profile is spiky, though: the same model sits ninth on the general Text Arena, and early hands-on reviews are more measured than the leaderboards - [one of the first detailed ones](https://www.buildfastwithai.com/blogs/kimi-k3-review) calls the upgrade from K2.7 worth it for multimodal and long-context work but a waste for high-volume coding, where the cheaper K2.7 Code remains the better per-dollar tool. Frontier-class at frontend and agentic work, merely very good at everything else, is a fair reading of the evidence so far.\n\nTwo things remain promises rather than facts. The license is expected to be the Modified MIT family Moonshot used for the whole K2 line - [plain MIT terms until a product passes 100 million monthly active users or $20 million in monthly revenue](https://github.com/moonshotai/Kimi-K2/blob/main/LICENSE), after which the UI must credit the model - but K3’s actual text does not exist publicly until the repository lands. And the serving floor, once weights arrive, is not a self-hosting story at all: 2.8T parameters is roughly 1.4TB at 4-bit, arithmetic that puts K3 in multi-node territory and makes the open weights matter for inference providers, researchers, distillation, and sovereignty requirements rather than anyone’s homelab.\n\nPoint your hardest frontend or agentic workload at the API this week - that is where the independent numbers are strongest and where a real advantage would show up in your own harness. Then, on July 27, read the license text before you build anything on it. If the weights and license land as promised, K3 resets the open-weight frontier; if they slip, it is a very good API model wearing an open-weight story.\n\n## Thinking Machines Inkling: a 975B flagship built to be fine-tuned, not crowned[#](#thinking-machines-inkling-a-975b-flagship-built-to-be-fine-tuned-not-crowned)\n\n**Use cases:**\n\n**Fine-tuning a custom domain model**- the release is built around Tinker, and Apache 2.0 removes downstream friction** Assistants that must follow instructions and not embellish**- IFBench and SimpleQA are where it independently wins** Multimodal RAG and agent backbones**- text, image, audio, and video in one open base with 1M context** Off-the-shelf coding assistant**- OK at best; it independently trails GLM-5.2 on SWE-Bench Pro\n\nThe most useful sentence in Inkling’s launch material was written by the vendor about its own model. Thinking Machines states, in its own [announcement](https://thinkingmachines.ai/news/introducing-inkling/), that “Inkling is not the strongest overall model available today, open or closed” - an admission that reframes what the other 975 billion parameters are for. This is Mira Murati’s lab shipping its first from-scratch model, and the launch on July 15 was less a leaderboard bid than a statement about who should be shaping model behavior: the developer fine-tuning it, not the lab that trained it.\n\nThe shape of the thing, from the model card and announcement:\n\n- 975B total / 41B active\n- 6 of 256 experts + 2 shared\n- Text, image, audio, video in\n- 1M-token context\n- 45T training tokens\n- Apache 2.0\n\nThe architecture is genuinely non-standard, and the best independent read so far is [Sebastian Raschka’s analysis](https://sebastianraschka.com/blog/2026/inkling-architecture-benchmark-notes.html): short kernel-4 convolutions after the key/value projections for cheap local token mixing, a learned relative-position bias in place of RoPE, and 55 of 66 layers running local 512-token attention windows with only 11 global layers. His benchmark read matches the positioning. Against GLM-5.2, Inkling wins on instruction following (IFBench 79.8 vs 73.3) and factual reliability (SimpleQA Verified 43.9 vs 38.1) while losing on hard reasoning (HLE 29.7 vs 40.1) and agentic coding (SWE-Bench Pro 54.3 vs 62.1). He calls it an honest all-rounder, and that reads right: a broad base that follows instructions and hallucinates less, rather than a benchmark specialist.\n\nTreat the headline numbers on the model card - 97.1% on AIME 2026, 77.6% on SWE-Bench Verified - as what they are: results from Thinking Machines’ own evaluation harness at maximum thinking effort, with no independent reproduction yet. The business logic is easier to verify. Inkling is the demo for Tinker, the company’s fine-tuning platform, where the model is available for customization at a limited-time 50% discount, and hosted inference runs through Together AI, Fireworks, Modal, Databricks, and Baseten. The model is the funnel; the platform is the product.\n\nIf your team fine-tunes models, this is the most interesting open base of the year - multimodal, 1M context, permissively licensed, and built by a lab whose entire thesis is customization. If you just want the strongest general model behind an API, the vendor has already told you to look elsewhere. Believe them.\n\n## Tencent Hy3: flagship claims on a 21B active-parameter budget[#](#tencent-hy3-flagship-claims-on-a-21b-active-parameter-budget)\n\n**Use cases:**\n\n**Self-hosted coding agents at serving scale**- 21B active parameters is the cost story, and frontend work is its top claimed strength** Data-pipeline and CI/CD automation**- the categories where Tencent’s blind expert eval showed the largest advantage** Long-context internal document work**- a 256K window with productivity-focused post-training** Anything customer-facing**- OK only after your own evaluation harness confirms the vendor numbers\n\nTencent’s claim sheet for Hy3 is bold: a 295B mixture-of-experts model that activates only 21B parameters per token, matches open flagships with 2 to 5 times its parameter count, and ships under Apache 2.0 with a 256K context window. The catch is a single sentence long: every number on that sheet was produced by Tencent.\n\nThe engineering is real and inspectable on the [model card](https://huggingface.co/tencent/Hy3): 80 layers routing top-8 of 192 experts, plus a 3.8B-parameter multi-token-prediction layer that vLLM and SGLang use for speculative decoding. The BF16 weights come to roughly 600GB and the FP8 build to about half that, which puts Hy3 in 8-GPU-node territory - Tencent recommends H20-3e-class hardware. This is the follow-up to April’s preview release, with what Tencent describes as scaled-up post-training on higher quality data.\n\nThe evidence, so far, is all in-house. A blind evaluation by 270 experts scored Hy3 at 2.67 out of 4 against GLM-5.1’s 2.51, with the largest gaps in frontend development, data work, and CI/CD tasks. Internal metrics report the hallucination rate falling from 12.5% to 5.4% and multi-turn failure rates from 17.4% to 7.9% between versions. [Developers Digest noted](https://www.developersdigest.tech/blog/tencent-hy3-open-source-moe-model) that the launch-day agentic numbers - 75.8% on SWE-Bench Multilingual, 71.7% on Terminal Bench - did not appear on the Hugging Face card at publish time, which is exactly the kind of gap that should keep a claim in the “directional” bucket. And a hallucination rate you measured on your own benchmark is a claim about the benchmark as much as the model - I wrote about [why self-reported reliability numbers need an abstention lens](/blog/ai-abstention-matters-more-than-accuracy/) and the argument applies verbatim to launch cards.\n\nWhat makes Hy3 worth tracking despite all that is the economics. Active parameters drive the compute bill, and 21B active with flagship-adjacent output would be a materially cheaper serving profile than the 40B-active class it claims parity with. If the claims survive independent runs, this becomes the value pick for self-hosted agentic workloads.\n\nIf you serve models at scale and run your own evaluation harness, Hy3 earns a slot in it this week - your repo and your failure modes will tell you more than any launch table. Everyone else should wait for the first wave of independent numbers. Nobody should re-platform on a claim sheet.\n\n## NVIDIA Nemotron-Labs-Audex-2B: the entire voice loop in one 2B model[#](#nvidia-nemotron-labs-audex-2b-the-entire-voice-loop-in-one-2b-model)\n\n**Use cases:**\n\n**Prototyping native voice agents**- the single-model speech-to-speech loop is the whole point, and one GPU suffices** Research on unified audio-text models**- the training recipe preserves text intelligence, per NVIDIA’s report** Stress-testing ASR on accented and noisy speech**- cheap to stand up next to your existing cascade** Production use of any kind**- not a use case at all until a commercial license exists\n\nEvery production voice agent today is three systems in a trench coat: speech recognition feeding a language model feeding a speech synthesizer, with latency, errors, and lost nuance compounding at each seam. Audex-2B, released June 8, is NVIDIA’s argument that the whole loop belongs in one model - speech recognition, translation, audio understanding, text-to-speech, audio generation, and direct speech-to-speech, in a single 2B-parameter checkpoint that keeps the 128K context and reasoning of its text backbone, with a thinking mode and an instruct mode. At roughly 4GB in BF16, it runs on one GPU.\n\nThe interesting claim is not the task list but what unified audio models usually sacrifice to get it. Adding audio output typically costs text intelligence - the “text tax” that has kept unified models a research curiosity. For the larger Audex-30B-A3B sibling released in early July, NVIDIA’s [technical report puts the tax at roughly zero](https://arxiv.org/abs/2607.05196): 86.4 on MMLU-Redux against the text backbone’s 86.3, and 81.1 on IMO AnswerBench against 79.3. Those numbers belong to the 30B model, not this one - the 2B was trained with the same recipe, and NVIDIA claims marginal-to-no regression for it, but nobody outside NVIDIA has published numbers for either, and the [2B model card](https://huggingface.co/nvidia/Nemotron-Labs-Audex-2B) leans on charts rather than tables.\n\nThen comes the sentence that decides everything: the model ships under the NVIDIA OneWay Noncommercial License. Not Apache, not MIT - noncommercial, full stop. The license, not the benchmark table, is the spec that determines whether Audex-2B can ship in your product, and for anything revenue-adjacent the answer is no.\n\nThis is the most instructive weekend project on the radar. Stand it up on vLLM, run your own accented and noisy audio through it, and measure the end-to-end latency of a native voice loop against your ASR-LLM-TTS cascade - that comparison will teach you where voice agents are heading regardless of whose model you eventually ship. And watch for a commercially licensed sibling: NVIDIA has walked this road in audio before, shipping its original Canary ASR model noncommercial and following with [Canary-1B-v2 under commercial-friendly CC-BY-4.0](https://huggingface.co/nvidia/canary-1b-v2). Build the prototype; do not build the product.\n\n## Prism ML Ternary Bonsai 27B: 95% of a 27B reasoner in 7.2GB[#](#prism-ml-ternary-bonsai-27b-95-of-a-27b-reasoner-in-72gb)\n\n**Use cases:**\n\n**Private, offline reasoning on a laptop**- math and code stay within two to three points of full precision** On-device analysis of long documents**- 262K context with a 4-bit KV cache fits 16GB machines** Phone-class deployment experiments**- the 3.9GB 1-bit sibling targets current flagship phones** Agentic tool use**- the weakest category by the vendor’s own numbers, and it says so\n\nThe rule of thumb for years was that reasoning models collapse below 4 bits per weight. Prism ML’s Ternary Bonsai 27B is the strongest public counterexample yet: a 27B model built on Qwen3.6-27B with every backbone weight stored as `{-1, 0, +1}`\n\n- about 1.71 bits per weight with FP16 group scales every 128 weights - that shrinks a 54GB FP16 footprint to 7.2GB deployed and, per [Prism ML’s published suite](https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf), keeps 94.6% of the full-precision benchmark average: 80.49 against 85.07 across 15 benchmarks.\n\nThe category breakdown is where the useful signal lives. Math holds within two points of full precision (93.40 vs 95.33) and coding within three (85.96 vs 88.74) - exactly the demanding tasks where conventional sub-4-bit quantizations fall apart. Agentic tool use takes the largest hit (74.01 vs 80.00), and Prism ML itself flags agentic coding as not yet a strong target, which is the kind of self-reported weakness that makes the rest of a claim sheet more believable, not less. There is also a 1-bit sibling at 3.9GB that retains 89.5% - the quality-size ladder is remarkably gradual for this regime.\n\nSmall only matters if it is also usable, and the deployment numbers say it is: 262K tokens of context on-device, kept affordable by the Qwen3.6 hybrid-attention backbone (about 75% linear attention) and a 4-bit KV cache. Prism ML reports 26.2 tokens per second decoding on an Apple M5 Pro and 44.0 on an M5 Max, and a bundled speculative drafter delivers a lossless 1.34x speedup on the CUDA path (98 to 131.8 tokens per second on an H100). Apache 2.0, with a hosted endpoint on Together AI if you want to taste it before downloading. The one real friction point: the ternary format needs Prism ML’s fork of llama.cpp for its custom 2-bit kernels, not mainline - an extra build step and an extra trust surface until upstream support lands.\n\nIf you have a 16GB M-series laptop, this is the model to spend an evening with. Every retention number above is self-reported - but the weights are public, the download is 7.2GB, and your own prompts are the benchmark that matters, which makes Bonsai the cheapest claim on this radar to verify yourself.\n\n## So what?[#](#so-what)\n\nSpend your evening on Ternary Bonsai 27B: pull the 7.2GB GGUF, build Prism ML’s llama.cpp fork, and run your actual workload - expect math and code near full precision and agentic tool use noticeably weaker. If you serve models at scale, put Hy3 into your own evaluation harness this week instead of waiting for the discourse to settle; at 21B active parameters, the serving economics justify the test even if the flagship claims deflate. And put July 27 on the calendar: that is the day Kimi K3’s open-weight promise becomes either the largest weights release in history or a missed deadline worth knowing about.\n\n## Related reading[#](#related-reading)\n\n[Why AI Agents Forget by Design](/blog/why-ai-agents-forget/)- context windows and memory limits behind the agentic claims every model card now makes.[10 Machine Learning Algorithms You’ll Actually Use in Production](/blog/machine-learning-algorithms-production/)- the evergreen counterpoint: most production ML still is not an LLM.", "url": "https://wpnews.pro/news/ai-model-radar-1", "canonical_source": "https://stantyan.com/blog/ai-model-radar-1/", "published_at": "2026-07-20 19:00:00+00:00", "updated_at": "2026-07-20 22:29:15.536022+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-research"], "entities": ["Moonshot", "Kimi K3", "Thinking Machines", "Inkling", "Tencent", "Hy3", "NVIDIA", "Nemotron-Labs-Audex-2B"], "alternates": {"html": "https://wpnews.pro/news/ai-model-radar-1", "markdown": "https://wpnews.pro/news/ai-model-radar-1.md", "text": "https://wpnews.pro/news/ai-model-radar-1.txt", "jsonld": "https://wpnews.pro/news/ai-model-radar-1.jsonld"}}