The Best AI Coding Agent in the World Ships 35% of Your Feature Tickets Cognition released SWE-2, a coding model scoring 50% on FrontierCode 1.1 Main — within one point of Fable 5.1 — at 64% lower cost, built on a 2.8-trillion-parameter Kimi K3 base with RL post-training. Separately, the Ruby on Rails core team's Agents on Rails Stage 2 benchmark found the best agent, GPT-6 Astra, completed only 35% of 60 real feature tickets, with GPT-5.6 Luna finishing zero, as agents systematically skipped edge cases and unstated requirements. AI coding agents are shipping faster than anyone can benchmark them. This week, three separate tests landed that together paint a more grounded picture than any single press release: real completion rates, real cost data, and a new model that quietly upends the price-performance calculus for everyone building on top of these systems. 1. Cognition's SWE-2 Changes the Cost/Performance Equation — and That Is the Actual Story Cognition released SWE-2 https://cognition.com/blog/swe-2 this week, their new coding model. It scores 50% on FrontierCode 1.1 Main — within one point of Fable 5.1 — at 64% lower cost. The base model is Kimi K3 at 2.8 trillion parameters, with reinforcement learning post-training adding five to six points across benchmarks. Notably, their RL algorithm trains all reasoning-effort levels in a single run rather than separate passes per capability tier. SWE-2 is live now in Devin Desktop, CLI, and Web. The capability numbers are competitive. But the more interesting signal is the cost trajectory. Matching frontier performance at a fraction of the price is not a minor efficiency win — it's the mechanism by which this market compresses. Models that were cost-prohibitive at scale six months ago become the default choice when something 64% cheaper delivers equivalent output. The frontier keeps moving, but so does the floor, and the floor is what changes actual adoption curves. Why it matters: For ICs: Before defaulting to whichever frontier model had the best launch event, SWE-2's price/performance ratio is worth a test run on your actual workflow. For leaders: The cost/performance compression means your AI coding budget buys roughly twice the capability it did six months ago. That pace is not slowing. For founders: The new reference point for "frontier coding performance" is now 64% cheaper than it was. Reprice your cost assumptions accordingly. The RL post-training is the real moat here, not the base model. Whoever gets RL fine-tuning right at this scale has a compounding advantage the base model race alone won't close. 2. The Rails Team Asked AI Agents to Ship Features. The Best One Made It 35% of the Time. The Ruby on Rails core team released Stage 2 of their Agents on Rails benchmark https://rubyonrails.org/2026/9/9/agents-on-rails-stage-2 , testing 10 models against 20 real feature tickets for Fizzy, a kanban app. Stage 1 tested Rails knowledge; Stage 2 asks whether a model can take a feature description and deliver working software. The difference between the two is the difference between a coding assistant and a coding agent. Results: GPT-6 Astra leads at 35% 21 of 60 tasks completed , Claude Fable 5.1 at 30%, Gemini 3.8 Flash at 28%. GPT-5.6 Luna — which performed well on Stage 1's knowledge tasks — completed zero of sixty feature tickets. The dominant failure mode: shipping the happy path and stopping. Edge cases, unstated requirements, error states — systematically skipped. Agents solved the problem the ticket described, not the problem the product actually had. This matters because real feature tickets are underspecified by design. A competent engineer fills the gaps from context and domain knowledge. Current agents don't — they execute the prompt as written. 35% completion on a well-defined kanban app, not a legacy monolith or microservices maze, is where the frontier sits today. That's the honest baseline. Why it matters: For ICs: Use agents for the defined-subtask layer — specific implementations, boilerplate, tests for known paths. Interpreting unstated intent is still your job. For leaders: If you're benchmarking agent performance internally, test against real tickets, not knowledge questions or toy problems. The gap between Stage 1 and Stage 2 results is substantial. For founders: Products built on coding agents need human review not just for correctness but for completeness. The missing edge case is the failure mode, and it won't announce itself. 3. An AI Coding Cost Tool Promised 60–90% Savings. Independent Testing Found 5%. RTK Rust Token Killer has been circulating as a way to dramatically cut AI coding costs by compressing terminal output before agents process it. Claimed savings: 60–90% token reduction. An independent benchmark from Quesma https://quesma.com/blog/does-rtk-make-ai-coding-cheaper/ ran 1,740 task attempts at over $1,500 in tokens and found something different. With Claude Code and Fable 5.0, costs dropped 5%. With OpenCode and DeepSeek V4 Pro, costs rose 5%. DeepSeek's average per-task cost increased 17% with RTK enabled. Pass rates fell slightly across configurations. The core problem: RTK's own metric measures raw bytes removed, not actual token cost. Terminal output represents 11% of Fable's total input tokens and 40% of DeepSeek's — the rest is system prompt, code context, and conversation history. Compressing 11% doesn't move the needle, and if extra agent turns get triggered by compressed output, you've made things worse. This is a recurring pattern in the AI tooling space: a metric that sounds like cost savings turns out to measure something adjacent to cost. Token compression is real. Whether it translates to bill compression depends on what fraction of your actual bill is the thing being compressed. Why it matters: For ICs: When evaluating AI coding cost tools, require task-level cost data with pass rates included — not just bytes saved. A tool that saves tokens but triggers more turns is a net loss. For leaders: "We reduce token usage by 80%" is not the same as "your bill goes down 80%." Ask for the methodology and the pass rate delta. For founders: The real AI coding cost lever is model selection and task routing, not compression at the edges. Optimize the 89%, not the 11%. The Verdict: Real or Hype? SWE-2's cost/performance compression → Real. Matching frontier capability at 64% lower cost is the mechanism by which the whole market reprices — and it keeps happening faster than anyone budgets for. 35% real-world feature completion → Real, and worth using as a baseline. This is calibration, not failure. Know what you're buying before you build a workflow around it. AI coding cost optimization tooling → Hype, mostly. The savings exist on paper. Whether they show up on your invoice depends entirely on what's actually driving your costs — and that is almost never the thing being compressed.