I graded 18 dated AI predictions from the last two years against what actually happened. The scoreboard is brutal, and the one guy who scored well got nothing for it.
In March 2025, Dario Amodei told the Council on Foreign Relations that AI would be writing 90% of code within three to six months, and essentially all of it within twelve. That twelve-month deadline passed in March. Almost nobody went back to check.
That is the strange etiquette of AI forecasting. The predictions are loud, dated, and delivered with total confidence, and then the date arrives and everyone has moved on to the next one. There is no scoreboard. So I built one.
In February 2025, I scoped agent workflows into Whizi on the assumption Altman was right about 2025. I killed the project six weeks later, after doing the math on what a single multi-step agent run costs against a flat monthly subscription. Six weeks of my own roadmap, spent on someone else's forecast. I pulled every dated 2024 and 2025 prediction I could find with a primary source and a deadline that has already passed. Eighteen made the cut.
The grading rule, so you can argue with it: a prediction is graded against what its own words promised, on its own deadline. If you said 90% by September and September brought 30%, "well, eventually" is not a grade. It is an excuse.
Sam Altman, January 2025: "We believe that, in 2025, we may see the first AI agents 'join the workforce' and materially change the output of companies."
MIT's NANDA project found in August 2025 that 95% of enterprise generative AI pilots produced no measurable P&L impact. Carnegie Mellon staffed a fake company entirely with AI agents and measured the office work they finished: the best model completed 30.3% of tasks. A 30% employee does not join your workforce. They exit your probation period.
One carve-out saves this from an F: inside software companies, coding agents genuinely did change output. Grade: C-.
Marc Benioff, September 2024: "Our vision is bold: to empower one billion agents with Agentforce by the end of 2025."
Salesforce reported roughly 29,000 cumulative Agentforce deals by early 2026 and crossed $1B in annual recurring revenue. A real business, and about 34,000 times short of the promise, counting deals rather than agents. My favorite detail: Salesforce now reports "agentic work units served" (3.8 billion of them) instead of a count of agents. When you cannot hit the number, change the unit. Grade: F.
Klarna, February 2024: its AI assistant "is doing the equivalent work of 700 full-time agents."
Then in May 2025 CEO Sebastian Siemiatkowski reversed course and started rehiring humans: "We focused too much on efficiency and cost. The result was lower quality." The AI kept the routine tickets. The humans came back for everything that mattered. Credit for running the experiment in public and admitting the result. Grade: C.
Amodei's 90%. The direction was right and the denominator was wrong. Google says 75% of its new code is AI-generated as of mid-2026, up from 25% in late 2024, though that counts every accepted autocomplete and humans still review before deploy. Industry-wide estimates in early 2026 clustered around 25-41%.
So: 90% of all code in six months, no. Essentially all code in twelve, not close. A historic shift in how code gets written, absolutely yes. The prediction described a real revolution and got every number wrong. Grade: C+ for the 90%, F for "essentially all." Eric Schmidt, April 2025: "in the next one year, the vast majority of programmers will be replaced by AI programmers."
The year is up. Programmers remain employed almost everywhere they were employed before. The real labor damage is narrower and crueler than the prediction: Stanford and ADP payroll data show employment for 22-to-25-year-olds in the most AI-exposed jobs down roughly 16% since late 2022 and still shrinking. Entry-level took the hit; the vast majority kept their jobs and got Copilot licenses. Grade: F.
And the number nobody predicted at all: Cursor went from $100M to $4B in annual recurring revenue in about seventeen months, and on August 14 SpaceX closed its $60B acquisition of the company, the largest startup acquisition ever recorded. Not one 2024 prediction list I checked has "a code editor becomes the biggest startup exit in history" on it.
Mark Zuckerberg, January 2025, on Joe Rogan: "Probably in 2025, we at Meta... are going to have an AI that can effectively be a sort of midlevel engineer that you have at your company that can write code." On the earnings call weeks later he added two more for 2025: Meta AI as the leading billion-user assistant, and Llama as the most advanced and widely used model.
None of the three happened. Llama 4 landed with a thud while Gemini 3 took the frontier crown in November 2025. And Meta spent the year making the opposite bet with its wallet: $14.3B for half of Scale AI, researcher pay packages reported between $100M and $450M, then 600 people cut from the superintelligence lab in October and roughly 8,000 layoffs in May 2026.
Meta predicted an AI that codes like a mid-level engineer, and instead spent eighteen months setting the world-record market price for human ones. Grade: F.
Elon Musk, April 2024: AI "smarter than any one human probably around the end of next year."
By end of 2025, the flagship xAI model was Grok 4, launched as "the smartest AI in the world," scoring 25.4% on a benchmark literally named Humanity's Last Exam. The claim did not survive contact with the exam. Grade: F.
Mira Murati, June 2024: PhD-level intelligence "for specific tasks" in about eighteen months.
For math it basically happened: Google DeepMind's Deep Think took an officially certified gold at the 2025 International Math Olympiad, years ahead of most 2024 forecasts. The qualifier she used, "for specific tasks," is doing heavy lifting, and it is exactly the qualifier her louder peers refused to use. The hedged version came true. The unqualified one, shipped as GPT-5's "PhD-level" branding in August 2025, went badly enough that Altman admitted OpenAI "totally screwed up" the rollout. Grade: B-. Gary Marcus published 25 dated predictions on January 1, 2025. The industry mostly rolls its eyes at him. Grading his gradeable claims:
Grade: A-. The docked minus is for what the list does not contain: nothing in it predicted a $4B coding-tool category, an IMO gold, or agents becoming genuinely useful in the one domain where output is checkable. The best forecaster of 2025 called every failure and missed every success.
And the uncomfortable second-order fact: being right earned him approximately nothing. The investors who bet against his entire worldview are the ones who priced Cursor at $60B. Forecast accuracy and financial returns run on separate scoreboards, and only one of them compounds.
Sort the 18 predictions into two piles and the noise disappears.
Predictions about capability, what models would be able to do, aged well and sometimes came in early. The length of tasks agents can finish has been doubling roughly every four to seven months, per METR, and the trend held.
Predictions about deployment, what organizations would let AI actually do, failed almost uniformly. Workforce agents, billion-agent platforms, replaced programmers, AI employees: all of it hit the same wall of quality bars, liability, integration cost, and the stubborn fact that a system completing 30% of tasks is a demo, not a hire.
Altman said agents would join the workforce. They joined the IDE, because the IDE is the one workplace where a wrong answer gets caught by a compiler instead of a customer.
So here is the decision rule I use now, with my own money on the line: when you hear a capability prediction, take it seriously, even the wild ones, because the trend lines keep winning. When you hear a deployment prediction with a date attached, double the timeline, then check what the person saying it is selling.
Every quarter from here I grade whatever came due, and the next edition opens with my own dated predictions, so you can run me through the same rubric. Grading is free. Being graded is the price.
If you ever wonder what I do when I'm not working: I run Whizi. One subscription, a few hundred models, GPT, Claude, Kimi and more. Come break it and tell me what happened.