The Hardest Software Is Falling to Agents. The Easiest Isn't. A single developer, working with an AI collaborator on a roughly $100/month subscription, produced a formal C17 compiler written entirely in ARM64 assembly that self-hosts to a byte-identical binary, cross-compiles to x86-64, builds Lua, SQLite and DOOM, and boots a Linux kernel. A separate project generating a web rendering engine from specifications admitted code only when its layout matched Chromium, Firefox and WebKit, agreeing on 699 of 705 browser-pair comparisons across 235 documents, though only 0.6% of the engine surface is generated and 93% remains unreached. Both efforts point to automated verification — differential testing, fixpoint checks and multi-engine voting — as the binding constraint on what agents can accomplish. Three projects landed this week on opposite sides of the same line, and the line is not where anyone would have drawn it. 1. One Developer Shipped a C Compiler That Boots Linux. The Interesting Part Is How We Know It Works. A formal C17 compiler written entirely in ARM64 assembly https://github.com/LiterateDrivenDevelopment/kcc now compiles itself to a byte-identical binary, cross-compiles to x86-64, builds and runs Lua, SQLite and DOOM, and boots a Linux kernel. It was built by one person with an AI collaborator on a hundred-dollar-a-month subscription. If you have ever staffed a compiler team, that should stop you. Read the build targets and the trick becomes obvious. A differential test compiles a corpus with both this compiler and clang and compares the output. A diagnostic test confirms invalid programs are rejected and warnings fire. A self-host fixpoint recompiles a module, relinks, and demands the binary come out identical. Booting a kernel is not a demo, it is the most unforgiving integration test available. At no point does a human read generated assembly and form an opinion about it. This did not win because the model is brilliant at register allocation. It won because a wrong answer is caught automatically, in seconds, every time — so the loop runs thousands of times without a person in it. The benchmark world has formalized this. KernelBench https://github.com/ScalingIntelligence/KernelBench counts a GPU kernel task solved only when the result is verified correct against the reference operator on randomized inputs and faster than the PyTorch baseline by a chosen margin. Correctness and speed, both mechanized, no judge model anywhere. Where you can write that function, agents compound. Where you cannot, they wander. Why it matters: - For ICs: Before you point an agent at a hard problem, ask what will tell it that it failed. If the answer is "I will review the diff," you are the bottleneck and the loop runs at human speed. - For leaders: Your agent's ceiling is your verification infrastructure, not your model subscription. Differential testers, golden corpora and fixpoint checks are now capacity investments, not hygiene. - For founders: Domains with a free, brutal oracle — compilers, codecs, query planners, protocol implementations — just got dramatically cheaper to enter. Incumbents there were protected by headcount, and headcount was the moat. 2. Someone Finally Did the Arithmetic on Generating a Browser Engine A project generating a web rendering engine from specifications https://tangled.org/burrito.space/bez is attacking the canonical "only three companies can afford this" problem, and the oracle is the clever part: candidate code is admitted only when its layout matches Chromium, Firefox and WebKit, with the third engine naming the odd one out. Across 235 documents, 699 of 705 browser-pair comparisons agreed, and the vote resolved every disagreement. It scales: a stable three-engine majority covers 94.8% of Web Platform Tests keys. The honesty makes it useful. Overall, 0.6% of the engine surface is generated and 93% is unreached. Nine CSS 2.1 layout rules are in place; eight were model-written and admitted by the vote, and block height still uses the hand-written rule because no candidate beat it. Then the number that reframes the exercise: roughly 55–60% of engine-relevant platform entries have both a usable automated oracle and generatable spec prose, and 8–18% have neither. That is not a progress report, it is a map of where the method can ever reach. The oracle pays off even when generation fails. The three-browser check surfaced a genuine compatibility bug: Firefox rounds lengths to 1/60 of a pixel where Blink and WebKit use 1/64, making a flex item wrap in one browser and not the others, with reproductions matching breakage already diagnosed on real sites. Machinery built to grade agents turns out to be a fine bug-finding instrument pointed at humans. Why it matters: - For ICs: Differential testing against several independent implementations beats any single reference, and it tells you which one is wrong. Use it wherever competing implementations exist. - For leaders: Ask for oracle coverage as a planning number before approving an agent-heavy project. "What fraction of this surface can a machine grade?" is the forecast; model quality is not. - For founders: The expensive part of cloning a hard product is shifting from writing it to checking it. If your defensibility is accumulated implementation labor, price that in. 3. Documentation Looked Like the Easy Win. It Scores Worse Than Compilers Do. Now the other side of the line. A new benchmark for user-facing documentation https://dogbench.ai/ draws 292 tasks from real projects, and the best combined score across seven agent lanes is 47.3 out of 100, with the best no-critical-defect delivery rate at 39.0%. In a wider audit of 1,267 submissions, 45.5% had a task-completion gap — a missing prerequisite, step, or recovery path — 36.6% contained technical inaccuracies, and 6.1% invented classes, flags or endpoints outright. Note what is hard. Of the 292 tasks, 87 require changing nothing at all; agents document internal refactors that need no user guidance, then miss necessary updates because no page exists yet. The benchmark scores the update decision separately from patch quality and combines them with a harmonic mean, precisely so fluent prose cannot paper over bad judgment. And note the cost of hand-building this oracle: 3,273 rubric criteria validated with maintainers, 798 of them blocking. Compilers get their grader for free. Documentation needed a research project, and the result still only measures progress toward a standard rather than replacing the reviewer. Why it matters: - For ICs: Treat an agent's doc patch as a draft with a 45% chance of a hole in the procedure. Walk the steps yourself; polished prose is not evidence of a working path. - For leaders: Do not reassign technical writers on the theory that this is solved. Judging whether users need guidance is the job, and it is the part scoring worst. - For founders: Selling agents into judgment-shaped work means you ship the evaluation rubric too. That is the real product, and it is expensive. - The general rule: fluency is now free and correctness is not. Wherever those two came bundled, your quality instincts are miscalibrated. The Verdict: Real or Hype? Agent-built systems software behind a mechanical oracle → Real. A self-hosting compiler that boots a kernel, from one developer, settles the capability-ceiling argument. Generating the web platform from specs → Real but early. 0.6% generated with 93% unreached, but the oracle-coverage arithmetic tells you exactly how far it can go. Agents owning user-facing documentation → Hype. The task everyone assumed was the easiest AI win scores 47 out of 100, because no machine can tell the agent it left out a step.