🤖 AI Agents Weekly: Kimi K3, DeepSeek-V4-Flash API, GPT-5.6 Price Cuts, Inkling-Small, YC's QM Harness, Gemini Robotics 2, Codex Security CLI, and More Moonshot AI released Kimi K3, a 2.8T-parameter open-weight mixture-of-experts model with native vision, a 1-million-token context window, and 104B active parameters per token, achieving frontier-level performance on long-horizon coding, agentic, reasoning, and vision tasks, trailing only Claude Fable 5 and GPT-5.6 Sol among models evaluated. The model uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE, with million-token agentic reinforcement learning and multiple reasoning-effort levels. 🤖 AI Agents Weekly: Kimi K3, DeepSeek-V4-Flash API, GPT-5.6 Price Cuts, Inkling-Small, YC's QM Harness, Gemini Robotics 2, Codex Security CLI, and More Kimi K3, DeepSeek-V4-Flash API, GPT-5.6 Price Cuts, Inkling-Small, YC's QM Harness, Gemini Robotics 2, Codex Security CLI, and More In today's issue: Moonshot open-sources Kimi K3 DeepSeek ships V4-Flash agent API OpenAI cuts GPT-5.6 prices 80% Thinking Machines drops Inkling-Small Google launches Gemini Robotics 2 OpenAI open-sources Codex Security CLI YC open-sources its QM agent harness Microsoft Foundry adds tool search Moonshot releases agent RL infra Nous adds wake word to Hermes Cursor lands on iPad ResearchArena probes AI R&D sabotage HANDBOOK.md tests long policy files Study exposes coding agent harness effects SlopCodeBench stress-tests Opus 5 And all the top AI dev news, papers, and tools. Top Stories Moonshot Open-Sources Kimi K3 Moonshot AI released Kimi K3, a 2.8T-parameter open-weight MoE model with native vision that lands closer to the closed frontier than any prior open release. Architecture: Combines Kimi Delta Attention, Attention Residuals, and Stable LatentMoE, activating 16 of 896 routed experts and 104B parameters per token. Scale and context: Ships a 1-million-token context window and roughly 2.5x better scaling efficiency than Kimi K2. Agentic post-training: Uses million-token agentic RL with persistent rollout and sandbox state, plus multiple reasoning-effort levels for long-horizon execution. Where it lands: Frontier-level on long-horizon coding, agentic, reasoning, and vision tasks, trailing only Claude Fable 5 and GPT-5.6 Sol among models evaluated.