cd /news/artificial-intelligence/show-hn-i-audited-my-ai-leaderboard-… · home topics artificial-intelligence article
[ARTICLE · art-80319] src=agiranker.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Show HN: I audited my AI leaderboard scale – every score dropped 6-15 points

AGI Ranker, an independent AI leaderboard, released v2.0.0 of its AGI Score, causing every model's score to drop by 6 to 15 points due to a correction in how human-parity ceilings are applied. The revision fixes an inaccuracy where most benchmarks were previously described as using best-human ceilings when they actually used benchmark maximums, with only GPQA Diamond (0.81) retaining a measured human ceiling. The AGI Score now uses a mixed scale, and the change primarily accounts for the score drops across all models.

read47 min views1 publishedJul 30, 2026
Show HN: I audited my AI leaderboard scale – every score dropped 6-15 points
Image: source

#

ONE SCORE.

INFINITE CLARITY.

AGI Ranker measures how close each frontier AI is to AGI. One transparent score (0-100) per model, distilled from 10 public benchmarks. Score 100 marks the AGI threshold.

Independent verification preferred over lab self-reports. Scores we can't verify are flagged, not invented. Every correction is logged publicly.

The AGI Ranker Leaderboard #

estimatefrom list prices, not a measured per-task bill. Capability is the published score for the selected area - identical to that area's tab (Overall = the AGI Score). The 💎 best-value pick is the model giving the most extra capability per dollar;

see the methodologyfor exactly how value for money is derived. Cost is an

estimatefrom published list prices at a typical token mix for that workload (cache-aware), not a measured per-task bill.

| # | Model | AGIAGI Score | Arena Elo | Coverage | Price $/M tok | ·Details |

|---|

[See full methodology →](/methodology)

Interactive Model Explorer #

Adjust domain weights and instantly see how the AGI Score changes. Transparency at its core.

How the AGI Score is Calculated #

A transparent, reproducible composite index designed for maximum signal and minimum noise.

Source quality is tiered: independent benchmark leaderboards (Tier 1) count fully at 1.00×; third-party evaluators with provider involvement (Tier 2) at 0.85×; self-reports from labs with verified track records (Tier 3) at 0.75×; non-verified labs and commentary content (Tier 4) are excluded from scoring entirely.

Note on the AA Intelligence Index (removed June 2026): We previously carried Artificial Analysis's composite Intelligence Index as a Language signal. We have removed it from the AGI Score. Their v4.1 revision turned it into an explicitly agentic composite (GDPval-AA, Terminal-Bench 2.1, τ³-Banking, SciCode, HLE, GPQA and more) that re-bundles benchmarks we already score directly - keeping it would double-count those signals across components and mislabel agency as language. Language now rests on LiveBench; a dedicated language/writing benchmark is on our roadmap to restore a second independent source. We continue to track the AA Index as an external reference.

Revised in v2.0.0: each raw benchmark score is rescaled by a ceiling. Where a published human study exists under a protocol comparable to the models’, that measurement is the ceiling and

100 means human parity. Where none exists, the ceiling is the benchmark maximum and

100 means a perfect score, with no human claim attached. Today exactly one scored benchmark carries a measured human ceiling, GPQA Diamond at 0.81, and the other nine are scored against the benchmark maximum, so the AGI Score is currently a mixed scale rather than a pure human comparison. Three further benchmarks passed the same ceiling audit and none of them is on the board: OSWorld (0.72) was retired in v2.0.0, and FrontierMath (0.35) and SimpleBench (0.837) have no harvested coverage yet. We previously described every ceiling as best-human; that was not accurate for most of them, and correcting it is the main reason scores moved in v2.0.0. AGI Score = 100 remains our marker for the

genesis of AGI, per our working definition in full: “AGI is an artificial intelligence that surpasses the best human on every purely brain-based intellectual task with no involvement of a physical body.” Scores climb past 100 as the AI grows from genesis into super-human (ASI-direction) territory; we preserve those scores rather than clamping, so the index stays informative in an ASI / post-AGI world.

**Agency**(35%),

**Fluid Reasoning**(29%),

**World Knowledge**(15%),

**Visual Reasoning**(11%),

Language Production(10%). Within each component, the score is the weighted mean of all the model's contributing benchmarks; each benchmark's effective weight is

sub_weight × source_tier_multiplier

. Sub-weights are fixed per benchmark and reflect each benchmark's relative importance within a component (e.g., ARC-AGI-2 carries 25% of Reasoning, GPQA carries 20%). The

saturation rule kicks in dynamically: if the IQR of normalized scores among the top-10 ranked models on a benchmark drops below 5 percentage points, that benchmark's sub-weight is halved and the freed weight is redistributed to non-saturated benchmarks in the same component - keeping the component score responsive to whichever benchmarks still discriminate frontier models.

Source tiering applies a multiplier per cell based on source independence (T1 1.00×, T2 0.85×, T3-Verified-Lab 0.75×; T4 excluded from scoring entirely). Component weights are user-adjustable in the Explorer below.

asymmetric pull-down shrinkage: if a model's coverage in a component is below 60% AND the raw score is above the population median, we pull it toward the median proportionally to coverage shortfall. Low scores with thin coverage stay low - no reward for hiding weaknesses. High scores with thin coverage get discounted toward "typical." Models need ≥3 components and ≥5 benchmarks (with Reasoning + Agency required) to be ranked.

Cross-component coverage floor (v1.4): a model with 2+ components below their coverage thresholds gets demoted from RANKED to PROVISIONAL - even if its present cells are exceptional. Thresholds: 30% sub-weight for World Knowledge / Visual Reasoning / Language;

40% for the core components(Fluid Reasoning and Agency, the load-bearing AGI dimensions per the locked definition). This prevents selection-bias rankings: a model can only claim near-AGI position if its evidence is reasonably broad across capability dimensions. The

saturation rule separately protects against dead benchmarks: any benchmark whose IQR among the top-10 ranked models drops below 5pp gets its sub-weight halved and redistributed to non-saturated benchmarks in the same component.

AGI Score and task-focused views. Each

specialty(currently Coding and Knowledge) re-ranks models using only the benchmarks that test that capability, with within-set sub-weights renormalized to 100%. Specialty scores apply the same source-tier weighting and asymmetric pull-down shrinkage as the main score - so a model with thin coverage in a specialty is pulled toward the population median for that specialty, never the other way.

Specialties are shown on a separate 0 to 10 index, not on the AGI scale, because they are a different kind of claim:

10 means a perfect score on every benchmark in that set, and it is

nota statement about human performance. Most of these benchmarks have no published human study, so there is no human bar to measure distance to.

A specialty index is never AGI- AGI requires the full battery, which only the canonical AGI Score measures. The

Custom tab (visible when you adjust the Explorer's weight sliders) shows the AI Score under your settings; the AGI tab always reverts to default weights, keeping the canonical view unambiguous.

Constant Value Used in
Asymmetric shrinkage coverage threshold (main AGI Score & specialties with 3+ benchmarks) 60% Step 4
Asymmetric shrinkage coverage threshold (specialties with 2 benchmarks, e.g. Reasoning) 85% Step 5
Saturation IQR threshold (top-10 ranked models) 5 pp Step 4
Coverage floor - Fluid Reasoning & Agency (core) 40% Step 4
Coverage floor - Knowledge / Visual Reasoning / Language 30% Step 4
Max thin components allowed in the RANKED tier 1 Step 4
Eval-date grace window vs. model release 30 days below

2. Reasoning previously ran at 1, on the grounds that it had only two benchmarks; that exception is what let it publish a one-benchmark ranking under a two-benchmark label, and it was withdrawn in v2.0.0 rather than kept. Models below the minimum but with at least one relevant cell appear in a separate "insufficient evidence" section below the ranking.

Note on Coding methodology (v2.0): Real-world coding in 2026 is agentic - it happens via tools like Claude Code, Codex, Cursor, not bare LLM completion, and the composition reflects that. SWE-bench Verified and Terminal-Bench 2.1 are both run under standardized agentic harnesses, and LiveBench Agentic Coding joins as a third independent evaluator. SWE-bench Pro was retired in v2.0.0: none of our cells for it came from a clean principal source, and the owner leaderboard has not been refreshed in months. We would rather drop a benchmark than keep one we cannot stand behind.

Contamination note for SWE-bench Verified: Public-test-set SWE-bench scores can be inflated by training-data leakage. Private-test-set evaluators (e.g., vals.ai) tend to report 5-15pp lower numbers than public-leaderboard aggregators for the same model. We prefer T1 private-test sources where available; see the Corrections Log for the DeepSeek V4 Pro GPQA correction (0.901 self-report -> 0.729 independent T2) as a concrete example of the contamination correction in action.

Pre-release builds: preview, beta and release-candidate models are

ranked normally and flagged. They meet the same evidence requirements as any other entry and are not given easier treatment. Holding them off the board would mean omitting models people are already using, and parking them in Provisional would misuse a tier that means

thin evidencerather than

unsettled product. What does deserve saying is that the weights behind a preview can change before general availability, so the score describes the build that was measured, not a finished product. Those entries carry a Pre-release tag. If a lab ships a materially different build under the same name, we treat that as a new model rather than silently updating the old row.

Source-choice sensitivity, GPQA Diamond and MMMU-Pro: both moved from Artificial Analysis to vals.ai on 2026-07-26, to break a single-evaluator dependency: Artificial Analysis had been supplying 97% of Knowledge and 100% of Visual Reasoning. As with Terminal-Bench we measured the gap before switching rather than after. On

GPQA Diamond the two boards agree closely, with a −0.30pp offset and 1.38pp scatter across all 20 models, and vals additionally publishes a standard error per row. On

MMMU-Pro vals runs

+5.71pp higher, with 1.67pp scatter and vals higher on all eleven comparable models, so scores on that benchmark rise by design rather than by accident.

One row is held rather than used: vals reports Grok 4.5 at 61.8 on MMMU-Pro, below Grok 4.3 (83.1), below Grok 4 (76.3) and below xAI’s own non-reasoning variants. A flagship scoring beneath its predecessor and beneath the non-reasoning builds is not a capability measurement, so we cannot say what that row measured and do not use it. Grok 4.5 therefore has no Visual Reasoning cell. We do not mix evaluators inside one benchmark column: taking vals for eleven models and Artificial Analysis for the twelfth would put two measurement stacks in one column and call the difference capability.

Source-choice sensitivity, Terminal-Bench 2.1: we report this separately rather than folding it into the uncertainty band, because it is a different kind of doubt. Terminal-Bench 2.1 is published by two independent evaluators, and they disagree systematically: across the 19 models present on both boards, vals.ai scores

7.77 points lower on average, and lower on 18 of the 19. Matching the declared effort level does not close it; Claude Opus 4.8 is max-against-max and still 12.7 points apart. We source from vals.ai, so our Terminal-Bench column sits about 7.8 points below where it would sit had we chosen the other board.

What that is worth on the AGI Score: 0.71 points on average, 1.01 at most, because Terminal-Bench is 25% of Agency and Agency is 35% of the score. Had we chosen the other evaluator, only 2 of 20 models would change rank, and those two are 0.01 apart in any case. The published interval already spans roughly ±9.5 points, so this sits comfortably inside it; folding it in would count the same doubt twice. The remaining per-model disagreement after removing that systematic gap, 4.54 points,

isinside the interval, where it belongs.

Harness disclosure for Terminal-Bench 2.1: Agentic-benchmark scores depend on the scaffold used to execute tasks. Every Terminal-Bench 2.1 cell comes from vals.ai using the Terminus 2 harness, pass@1 over 89 tasks. We deliberately use one evaluator rather than blending several, because two evaluators running the same 89 tasks do not agree: across the 19 models present on both the vals.ai and Artificial Analysis boards, vals.ai sits an average of

7.8 points lower, and it is lower on 18 of those 19. Matching the declared effort level does not close the gap. Blending the two would produce a number neither evaluator published, so we pick one, name it, and carry the disagreement into the uncertainty band instead. Native-agent rows (Claude Code, Codex CLI, Cursor CLI) are excluded entirely: they measure a product, not a model. Real users running each model with its own agentic tool may see different relative performance than these scores predict.

Specialty Benchmarks & within-set sub-weights
Coding SWE-bench Verified 40 · Terminal-Bench 2.1 35 · LiveBench Agentic Coding 25
Reasoning ARC-AGI-2 70 · AIME 2025 30
Knowledge HLE 40 · GPQA Diamond 35 · MMMU-Pro 25
Tool Use withdrawn in v2.0.0 — OSWorld, BrowseComp and the two Tau-bench domains all left the board in the v2 evidence review: OSWorld had 2 clean cells out of 22, BrowseComp had no uniform browsing harness across models, and the Tau-bench retail and airline pair is superseded by τ³-Banking. That leaves a single clean non-coding tool-use benchmark, and a one-benchmark ranking is exactly what our two-benchmark minimum exists to prevent. The tab returns when a second one exists.

Estimated cost is a blended price per 1M tokens at a typical input:output mix for that workload (coding is input-heavy over a large cached context; reasoning is output-heavy; etc.), using each lab's published list prices and its cached-input rate where available (otherwise about 10% of the input rate for the cached share). It is an estimate of typical cost, not a measured per-task bill.

Value for money is the extra capability a model delivers over the weakest option shown, per estimated dollar: (score − lowest score in view) / cost. We subtract that floor on purpose - a capability score has no true zero (a coding score of 0 is not "no value"), so raw score-per-dollar would over-reward the cheapest model no matter how capable it is. Subtracting the floor measures the extra capability you actually buy.

**Top**(highest capability),

**Budget**(cheapest to run), and

Best value(the most extra capability per dollar). The ranking shows the picks rather than a raw value number, which is ambiguous to read in isolation.

Pro, Anthropic

thinking, etc.) must appear in the source label. Effort-knob suffixes are

not variants.

eval_date

predates the model's release_date

by more than 30 days is rejected as a mis-attribution. The 30-day grace window covers pre-release lab evaluations; anything older almost certainly tested a preceding model that happens to share part of the name.Our social preview card was still showing pre-v2 numbers, weeks after the numbers changed. Anyone sharing a link to this site saw a card reading

14 live benchmarks,

20 models trackedand “multimodal”. The truth is 10, 21 and Visual Reasoning. The image file on our server was in fact correct and had been regenerated with the rest of the v2 release; what was wrong was our assumption that replacing a file replaces what the world sees. Social platforms cache preview images against the URL, so re-rendering the same filename changes nothing for anyone who has already shared the link, and we had verified the file rather than the card.

Every preview image now lives at a versioned URL, the homepage card and all 21 model cards, which forces every platform to fetch it again. The previous files stay where they are, so an existing embed shows an old image rather than a broken one. The card generator carries the version token now, so this is a one-line change next time rather than a rediscovery.

This is the same failure as the human-ceiling counts two releases ago: a rendered artifact that nobody re-derived after the thing underneath it changed. The site text was clean, and we checked: no page description, title or preview text anywhere on the site still carries the old figures.

Tapping a specialty tab on a phone gave you a wall of text instead of scores. The specialty-index notice rendered

586px tall on a 375px screen, 72% of the viewport, and pushed the leaderboard 742px below the tab bar. A reader tapping Coding to see coding scores got a screenful of methodology and had to scroll to find a single number. Below tablet width these notices now collapse to a summary with a More expander, and the full text is one tap away in place. Collapsed, the notice is 151px and the table sits 307px from the tabs, so it is reachable in one short scroll.

Nothing was deleted and nothing was softened. The visible summary is a compression, not a gentler version: it still says the specialty index is a 0 to 10 scale rather than the AGI Score, that 10 means a perfect score on every benchmark in the set, and that this is

nota human-parity claim. Those are the three things a reader has to know before reading the number, so they stay visible whether or not anybody taps More. The same treatment went to the Value view note about cost being an estimate rather than a measured bill, which kept its estimate caveat in the collapsed line. Switching tabs always re-collapses, so no tab inherits the previous one's expanded state. The expander is a real button with aria-expanded, works without hover, and the Back to AGI Score control stays visible while collapsed. Desktop is unchanged: it has the room, so it shows the full text and no expander at all.

The mobile tap states would not have rendered on an iPhone. The new section pills carry a pressed state, and it verified correctly in desktop Chromium, which is exactly the wrong place to check it. iOS Safari declines to apply the CSS active state to a link unless a touch listener exists on the element or one of its ancestors, so on the only kind of device that has a touch screen the pills would have looked inert when tapped. One empty listener on the page body fixes it for every tap target on the site. Recorded because verifying a touch behaviour in a mouse browser and calling it done is the kind of shortcut that ships a defect, and because the same trap applies to anything interactive we add from here.

The site had no mobile navigation at all. On a phone the header rendered a wordmark and a Contribute button and nothing else. The six section links were desktop-only and the version pill was hidden below tablet width, so a visitor on a phone had no route to Corrections, Methodology or anything else, on a page that is otherwise one continuous scroll, and could not see which release they were reading. Small screens now get a compact sticky bar with the version pill visible and a scrolling row of section pills beneath it, carrying the same six anchors as the desktop nav. A row rather than a hamburger: there is no open state to get stuck, nothing to trap keyboard focus, and every section stays one tap away instead of two. Phones also get a back-to-top button once you are far enough down that scrolling back is a chore.

Section anchors were landing behind the header at every width, because nothing on the page set a scroll margin; jumping to Corrections put its own heading underneath the sticky bar. Fixed for both layouts. Nothing else moved: the desktop bar is unchanged, and no score, weight or benchmark is touched by this release.

We were overstating how much of the scale is anchored to humans. Three places on this site claimed that three or four scored benchmarks carry a measured human ceiling, and named OSWorld, FrontierMath and SimpleBench among them.

The correct number is one. Of the ten benchmarks on the board, only GPQA Diamond (0.81) carries a measured human ceiling; the other nine are scored against the benchmark maximum, where 100 means a perfect score and no human claim is made at all. The other three did pass the same ceiling audit, and none of them is on the board: OSWorld (0.72) was retired in this very release, and FrontierMath (0.35) and SimpleBench (0.837) have no harvested coverage yet. The claim was wrong in the AGI definition modal, in the methodology summary and in the full methodology, and the three did not even agree with one another. This is the same class of error as the correction we published two versions ago: a number that survived because nobody re-derived it after the thing underneath it changed. The AGI Score is a mixed scale, and it is a good deal more mixed than we were saying.

Also in this release: the DeepSeek apology is signed by Barak Laniado, founder and CEO, and its corrections contact is now an email address you can actually write to rather than a link back into the site. The calibration constants box no longer reads “Current as of v1.4” under a v2 banner; the constants are unchanged and were re-confirmed for v2.0.0, which is what it now says. Three hover styles on the corrections card and twelve layout classes on the apology page were also missing from the compiled stylesheet, which is purged to what the leaderboard uses, so they had been failing silently.

The DeepSeek apology gets its own page. The correction we published in v2.0.0, an ARC-AGI-2 score attributed to DeepSeek V4 Pro that carried no source at all, now has a dedicated page at

/corrections/deepseek-arc-agi-2, linked from its entry in the log above. An apology buried as one card among several is easy to walk past, and this one should be readable, citable and linkable on its own. The page sets out what we published, why it was wrong, what it did and did not affect, and what changed in the process so that it cannot recur: under methodology v2 a recorded source is a condition of scoring rather than an expectation, and all 156 scored cells carry one. The wording of the apology itself is unchanged and identical in both places.

One count in that entry was also wrong. It said three further cells were voided in the same pass and then referred to four of them in the next sentence. Four further cells were voided, five in total including the DeepSeek one. Corrected here rather than quietly.

Methodology v2. Every score falls, and no model got worse. Two changes drive it. First, human ceilings. Because a score is normalised as raw divided by a human ceiling, that ceiling is not only an anchor for what counts as human level, it is a volume knob: it multiplies both the level and the spread of a benchmark, and so how hard that benchmark pushes on its component. Most of ours had no measurement behind them. Humanity’s Last Exam sat at 0.50, a number our own internal audit described as a policy floor rather than a finding, and that setting quietly made HLE count double. From now on a ceiling is either a published human result under a protocol comparable to the models’, or it is simply the benchmark maximum and we make no human claim at all. Only four survived the test: GPQA Diamond at 0.81, OSWorld at 0.72 (which we

raised, having carried 0.85 against our own recorded measurement), FrontierMath at 0.35, and SimpleBench at 0.837, where we had rounded a human baseline up. Eleven moved to 1.00. Three component scores that previously sat above 100, on a scale whose 100 is meant to be the human mark, no longer do. Second, Agency was rebuilt on four legs across three independent evaluators: SWE-bench Verified, Terminal-Bench 2.1, τ³-Banking and LiveBench Agentic Coding. Terminal-Bench now comes from vals.ai rather than Artificial Analysis, cutting our reliance on any single evaluator. τ³-Banking enters at benchmark maximum, so its true range shows: the best model completes about a third of these stateful banking workflows. Retired: SWE-bench Pro, OSWorld, BrowseComp, Tau-bench retail and airline, Aider Polyglot, Terminal-Bench 2.0. MCP Atlas demoted to informational. The Tool Use tab is withdrawn until a second clean non-coding benchmark exists. Scores fell by about 6 points from the ceiling change and further from the Agency rebuild; ordering shifted only locally, and the top two are unchanged.

Third, evaluator concentration. Artificial Analysis had been supplying 45% of the whole board, 97% of Knowledge and 100% of Visual Reasoning, so one evaluator changing terms would have taken two components to zero. GPQA Diamond and MMMU-Pro moved to vals.ai alongside Terminal-Bench 2.1, taking Artificial Analysis to 26% of the board and splitting Knowledge 52/48 between two evaluators. Every switch was measured before it was made rather than after, and each offset is published rather than absorbed: GPQA Diamond −0.30pp with 1.38pp scatter across all 20 models, MMMU-Pro +5.71pp with 1.67pp scatter and vals higher on all eleven comparable models, Terminal-Bench 2.1 −7.77pp with 4.54pp scatter and vals lower on 18 of 19.

One row is held rather than used: vals reports Grok 4.5 at 61.8 on MMMU-Pro, below Grok 4.3 at 83.1 and below xAI’s own non-reasoning variants, which is not a capability measurement we can account for, so we do not use it. Grok 4.5 consequently has no Visual Reasoning cell and falls three places. Visual Reasoning is still single-source, having moved from wholly Artificial Analysis to wholly vals.ai: it rests on one benchmark, so this relocates the dependency rather than removing it, and we would rather say so than imply it is solved.

Also in this release: the Reasoning specialty tab is withdrawn. It ran on ARC-AGI-2 and AIME 2025, but AIME 2025 is scored for a single model and sits at 98-99% where it appears, so the tab was ranking on ARC-AGI-2 alone while labelled as using two benchmarks. Withdrawn under the same rule as Tool Use rather than kept with a misleading label; it returns when AIME 2026 is harvested. Cursor Composer 2.5 is removed from the board: it is a product system rather than a model, so under v2 it scores no cells at all, and listing it implied we were waiting on evidence that was never coming. Specialty tabs now report a

0 to 10 index rather than a 0 to 100 score. On the AGI Score, 100 means the genesis of AGI; a specialty score never meant that, and showing both on the same scale implied they were the same kind of claim. Ten now means a perfect score on every benchmark in that set, with no human comparison implied, because most of those benchmarks have no published human study to compare against. The

Multimodal component is renamed

Visual Reasoning. It rests on a single scored benchmark, MMMU-Pro, which is an exam built on diagrams and charts, so calling 11% of the score “multimodal” claimed a breadth of perception we do not measure. Visual Reasoning says what it is: interpreting image data, which underpins reading a chart, extracting a figure from a document, or working from a screenshot. The AGI definition was revised in the same pass and for the same reason. It previously excluded “sensory perceptions whatsoever”, which contradicted our own scoring of a diagram-based benchmark. Rather than bolt vision on, which would have raised the obvious question of why sight and not hearing or smell, we removed the sensory clause and let the physical-body clause carry the weight: what reaches a model as data is in scope, what needs a body to acquire is not. The definition got shorter. The corrections summary counters have also been replaced. They had been hardcoded since May and one of them read “1 model entered scoring” for two months without meaning anything; they are now derived from the scoring ledger on every release.

Corrections to our corrections. We audited this log against our own commit history and found that two entries were wrong. Both times we compared two different configurations of the same benchmark and reported the difference as a laboratory overstating its results. DeepSeek did not overstate GPQA Diamond: the 72.9% we published as an independent contradiction is DeepSeek’s own non-thinking-mode figure, reaching us through a secondary aggregator and set against their thinking-mode result. Measured like for like, their self-report sits 1.3 points above independent measurement rather than 17.2 points above. Moonshot did not overstate Humanity’s Last Exam: their 54.0% was explicitly labelled as a with-tools result, and the same post published 36.4% without tools, within half a point of independent measurement. Our source policy had excluded them on the strength of that misreading. Both entries are retracted and both originals are retained rather than deleted. We have also withdrawn a laboratory promotion test that the policy described but that was never run, and corrected an entry which presented source relabelling as new independent measurement. No score, weight, human ceiling or source tier changed in this release.

Claude Opus 5 and Anthropic roster correction. Added Claude Opus 5 on five independent, variant-distinct frozen-v1 cells: ARC-AGI-2 90.4%, SWE-bench Verified 97.0%, HLE 52.59%, GPQA Diamond 93.23%, and MMMU-Pro 84.74%. The unchanged eligibility engine places it Provisional at 97.31 because one Agency cell and no accepted LiveBench row leave two thin components. The available LiveBench EAP row remains held pending public-model identity resolution; Terminal-Bench 2.1, ARC-AGI-1 and ARC-AGI-3 remain outside frozen-v1 scoring. Anthropic’s active latest-two roster is now Opus 5 and Opus 4.8. Fable 5 leaves the current roster under an identity-integrity exception because fallback-enabled, fallback-disabled and generic rows cannot be combined into one standalone model; Opus 4.7 thinking and base retire as older generations. All historical evidence remains preserved, with no transfer or imputation. No formula, weight or source tier changed.

Grok 4.5 ARC-AGI-2 update. Added ARC Prize’s independently verified Grok 4.5 High result on the 120-task ARC-AGI-2 Public Eval at the official displayed precision of 52.6%. High is the single frozen-v1 scoring configuration; the equal-scoring Medium result and the Low result remain supporting provenance only. ARC-AGI-1 and ARC-AGI-3 remain informational, and no formula, weight, source tier, roster entry or unrelated evidence changed. No imputation.

Capability Heatmap isolated from custom weights. The heatmap now reads a separately stored canonical snapshot of default-weight AGI Score, component values, coverage and tier. Explorer sliders and presets continue to recompute AI Score and leaderboard order without changing any heatmap value, order, colour, tooltip or accessible label. This fixes the state-dependent path left open by v1.11.24, whose direct score accessor still pointed to mutable custom-weight state. No score, rank, component, benchmark, evidence, eligibility or methodology changed. No imputation.

Capability Heatmap AGI Score binding locked. The rightmost heatmap column now declares its full-precision value, colour and accessible label directly from the same canonical final score field used by the leaderboard, model details and Explorer; its default order is explicitly canonical AGI Score descending. A no-cache v1.11.23 preflight already rendered the canonical values, so this release changes no score, rank, component, benchmark, evidence, eligibility or methodology. It adds a machine-verifiable binding contract and rendered-DOM regression coverage against an Agency-column regression.

Inkling and July evidence refresh. Added Thinking Machines’ open-weight Inkling on six independent cells: ARC-AGI-2 36.5%, SWE-bench Verified 77.6%, LiveBench Overall 71.68%, HLE 29.70%, GPQA Diamond 87.17%, and MMMU-Pro 73.47%. The unchanged eligibility engine places Inkling Ranked #17 at 70.97; it remains Coding-ineligible at 1/3 because SWE-bench Pro and Terminal-Bench 2.0 are absent. Refreshed every exact mapped row from the current LiveBench snapshot and current Artificial Analysis HLE, GPQA, and MMMU-Pro tables, including newly available Gemini 3.5 Flash cells; Flash moves from Provisional to Ranked #8. Removed Claude Fable 5’s HLE cell because AA explicitly identifies that row as an Opus 4.8 fallback configuration; routed evidence does not score as pure Fable. Vals SWE-bench Verified and ARC-AGI-2 rows were audited across the active roster. AA Intelligence Index, Terminal-Bench 2.1, SciCode, routed results, and incompatible variants remain unscored. No formula, weight, source tier, or historical artifact changed. No imputation.

Kimi K3 SWE-bench Verified update. Added Kimi K3’s independently evaluated 93.4% SWE-bench Verified result from Vals AI using the Mini-SWE-agent harness. The fifth scoreable cell adds Agency evidence and moves K3 from Provisional to Ranked #2 under the existing frozen-v1 eligibility engine; its AGI Score changes from 89.87 to 88.38 because the new Agency component is below its prior four-component aggregate. K3 has one of three Coding cells and remains Coding-ineligible because SWE-bench Pro and Terminal-Bench 2.0 are missing. Terminal-Bench 2.1 remains held as a distinct version, and ARC-AGI-2 remains missing. No formula, weight, source tier, roster, or existing K3 cell changed. No imputation.

Kimi K3 Frontier Integration. Kimi K3 joins as Kimi’s latest general-purpose flagship, alongside previous general-purpose generation Kimi K2.6. Kimi K2.7 Code is a specialist coding branch; it remains historically documented but leaves the active latest-two roster because it currently qualifies for neither the main Ranked leaderboard nor the frozen-v1 Coding leaderboard. K3 enters Provisional #2 at 89.87 on four independent T1 cells: LiveBench Overall 77.9%, GPQA Diamond 93.5%, Humanity’s Last Exam 44.3% (AA text-only, no tools), and MMMU-Pro 81.0% (AA multimodal, 10-option, no tools). K2.6 remains Ranked #9 at 79.98 and Coding-eligible. Moonshot’s conflicting 56.0% HLE with-tools self-report is retained as rejected provenance, not averaged or scored. Terminal-Bench 2.1 at 85.0% is held because v1 scores Terminal-Bench 2.0; SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.0 and ARC-AGI-2 remain missing. The specialist-slot rule is provider-neutral. No scoring formula or benchmark weight changed. No imputation.

Visual Identity Patch. Added official model-family and provider marks across AGI Ranker’s model surfaces, with local assets, accessible fallbacks and no scoring changes.

Frontier fairness refresh. Muse Spark 1.1 gains four scoreable cells from the approved source harvest: LiveBench 76.23%, Scale SWE-bench Pro 61.5%, Terminal-Bench 2.0 80.0%, and OSWorld-Verified 80.8%. It is now Ranked #4 at 86.76 and places #5 in Coding at 79.64. Gemini 3.5 Flash enters as Provisional at 81.75 using LiveBench, SWE-bench Verified, ARC-AGI-2, and OSWorld-Verified. Gemini 3 Pro leaves the active board under the latest-two Google roster rule, with its historical snapshot retained. GPT-5.6 Sol remains #1 at 93.33 after the frozen roster-sensitive scoring mechanics recompute. Terminal-Bench 2.1 and unsupported Gemini benchmark variants remain held outside production. No methodology or weight change. No imputation.

Muse Spark 1.1 and Sol pricing. Added Muse Spark 1.1 as a separate Meta model with three scoreable independent cells: SWE-bench Verified 82.0% (vals.ai, Mini-SWE-agent), GPQA Diamond 89.8%, and Humanity’s Last Exam 45.1% (Artificial Analysis thinking configuration; HLE no-tools). It enters as Provisional at an AGI Score of about 87.15. AA Intelligence Index v4.1 at 51 and preliminary LM Arena Text Elo 1490 ±10 are retained as informational evidence only. The existing Muse Spark generation remains separate, and its MMMU-Pro 81% result was not transferred to 1.1. Added official GPT-5.6 Sol API pricing of $5 input, $0.50 cached input, and $30 output per million tokens. No benchmark methodology or weights changed. No imputation.

Composer benchmark identity corrected. Removed Cursor Composer 2.5’s 47% result from the SWE-bench Pro column after confirming that the underlying Artificial Analysis result belongs to the distinct SWE-Bench-Pro-Hard-AA system benchmark. Composer remains Coding-eligible through SWE-bench Verified and Terminal-Bench. Because Coding scores renormalise over eligible available cells, its score changes from 70.50 to 80.56 and its rank moves from #10 to #4. No other Coding score values change, though ranks #4–#9 shift down one position. The removed evidence is retained internally for future Coding Systems consideration. The Ranked AGI leaderboard is unchanged. No imputation.

GPT-5.6 Sol added. Added GPT-5.6 Sol from six screenshot-confirmed current-schema cells: ARC-AGI-2, GPQA Diamond, HLE, LiveBench, MMMU-Pro and SWE-bench Verified. Sol enters the Ranked leaderboard immediately at #1 (AGI Score about 92.6 after roster re-equilibration) using Max-effort rows across sources. With Sol now Ranked, the GPT-5.4 family was retired under AGI Ranker’s latest-two-OpenAI-generations roster rule. GPT-5.6 Terra/Luna and Muse Spark 1.1 were not added pending separate scoreable coverage. AA Intelligence Index remains informational and unscored.

LiveBench official-board refresh. Updated seven LiveBench Global Average cells from the official board and added Grok 4.5’s LiveBench result. Grok 4.5 now clears the Ranked floor at #4. GLM 5.2 also moves from Provisional to Ranked via saturation/coverage recomputation; no new GLM benchmark cell was added. GPT-5.6 Sol is not yet included pending Sol-distinct scoreable coverage.

Grok 4.5 added (Provisional). xAI Grok 4.5 enters on four independent T1 cells: AA HLE 0.40 and GPQA Diamond 0.93 (source rows “Grok 4.5 (high)” / effort knob, not a separate SKU), AA MMMU-Pro 0.80, and vals.ai SWE-bench Verified 0.866 (Mini-SWE-agent). Both cores present; stays Provisional (fifth independent cell still required for Ranked; agency and language coverage remain thin). Grok 4.3 retained under the latest-two-versions scope rule. GPT-5.6 family not added (no Sol-distinct row with scoreable coverage). AA Intelligence Index tracked as informational only (not in score). No imputation. (HLE corrected 0.38→0.40 after bar-label re-check.)

MiniMax M2.7 LiveBench harvest. MiniMax M2.7 gains LiveBench Global Average 0.6349 from livebench.ai (source row “Minimax M2.7”). Fourth independent T1 cell; stays Provisional (one more independent cell needed for Ranked). No AA Intelligence Index ingested. No imputation.

Kimi K2.7 Code harvest + Composer SWE-V. Kimi K2.7 Code enters as Provisional on four independent T1 cells: LiveBench Global Average 0.7189, GPQA Diamond 0.896, AA-standardised no-tools HLE 0.328, and vals.ai SWE-bench Verified 0.782 (Mini-SWE-agent harness). Provisional AGI Score about 75.8 on 4/14 benchmarks across all four components; one more independent cell needed for Ranked. Cursor Composer 2.5 gains vals.ai SWE-bench Verified 0.796 (Cursor CLI harness), lifting its Coding specialty from about 65.6 to 70.5. No AA Intelligence Index ingested. No imputation.

GLM 5.2 gains ARC-AGI-2. ARC Prize now lists GLM-5.2 with ARC-AGI-2 at 22.8% (dated 2026-06-13), so the cell is added from the independent leaderboard. GLM 5.2 now appears on the Reasoning tab at about 26.8, while its overall Provisional score settles around 73.5; the lower score is expected because a hard missing reasoning benchmark is now measured rather than absent. It stays Provisional because agency coverage remains just under the core floor and multimodal coverage is still empty. No imputation.

GLM lab label + performance housekeeping. GLM 5.1 and GLM 5.2 now display their lab as Z.AI instead of the generic "other" provider bucket everywhere the public data is served. Also replaced the browser-side Tailwind CDN compiler with a precompiled local stylesheet and lazy-load the optional radar-chart library, reducing startup JavaScript without changing scores or methodology.

GLM 5.1 demoted to Provisional. Two BenchLM cells are voided: SWE-bench Pro was a Provider-exact relay of Z.AI's own figure (not an independent eval), and Tau-bench Retail no longer exists on BenchLM (replaced by tau-Telecom, a different benchmark we do not score). GLM 5.1 now rests on four independent T1 cells only. Provisional AGI Score about 71.9. No imputation.

GLM 5.2 harvest: vals.ai + LiveBench. Five independent T1 cells now ground GLM 5.2: GPQA Diamond 0.86, SWE-bench Verified 0.83, and Terminal-Bench 2.1 0.68 from vals.ai (replacing earlier Artificial Analysis figures on the overlapping benchmarks), HLE 0.40 from our standardized AA source, and LiveBench 0.76. LM Arena Text Elo 1471 is published. Stays Provisional: agency sub-weight coverage sits just below the 40% core floor alongside absent multimodal data. No imputation.

GLM 5.2 graduates to Provisional. Same day it launched, Artificial Analysis published independent results, so GLM 5.2 moves from Awaiting Verification to Provisional on three T1 cells: GPQA Diamond 0.89, Terminal-Bench 2.1 0.75, and HLE 0.40. The independent numbers came in below the lab's launch self-reports (which we had excluded) - exactly why we wait for them. It needs a fifth independent cell to reach Ranked. No imputation.

GLM 5.2 added (Awaiting Verification). Z.AI released GLM 5.2 on June 16. So far only the lab's own launch-day numbers are public (its model card and a benchmark aggregator relaying the same figures), which we exclude as unverified self-reports. GLM 5.2 is listed as Awaiting Verification - no score - until an independent evaluator publishes results. No imputation.

Hero badge. The launch-month pill now reads "Since May 2026" - a founding mark rather than a stale-looking date stamp. No scoring change.

Count consistency. The hero subhead and structured data still read "15 benchmarks" after yesterday's removal of the AA Index; both now read 14, matching the live count. No scoring change.

AA Intelligence Index removed from the score. Artificial Analysis shipped Intelligence Index v4.1, an explicitly agentic composite (GDPval, Terminal-Bench 2.1, τ³-Banking, SciCode, HLE, GPQA and more). Because it re-bundles benchmarks we already score directly, keeping it as our Language signal would double-count those results and mislabel agency as language - so we removed it from the AGI Score (14 live benchmarks now). Language rests on LiveBench; a dedicated language benchmark is on the roadmap. Most scores rise ~0.3-0.5 (the Index had been a slight drag); the ranking is unchanged. We still track the AA Index as an external reference.

Cleaner number type. Scores and stats now render in Inter with tabular figures - more legible in the dense tables than the previous display face, and consistent across the site. Headings and the wordmark are unchanged. No change to any score.

Value view polish. Refined the value-for-money UI and method after testing: the ranking now leads with the picks (Top / Best value / Budget) instead of a raw value index that was hard to read at a glance, the cost column is labelled $/1M tokens for clarity, and the methodology page now spells out exactly how value for money is derived. No change to any score.

Value view. The Value tab ranks models by capability against an estimated API cost. Pick a capability area - Overall (the AGI Score) or Coding, Reasoning, Knowledge, Tool Use - and the table shows the exact same score that area's own tab shows, next to an estimated cost at that area's typical token mix (cache-aware). Sort by capability or by value for money, and read the picks at a glance: top capability, best value, and budget. A value-frontier graph plots the same models. Also renamed Google to Google DeepMind, and renamed the Agentic specialty to Tool Use (coding is also agentic, so the old name was ambiguous; Tool Use covers computer use, web browsing and tool/function calling beyond code). No change to the canonical AGI Score or any benchmark data.

Full data refresh, plus MiniMax M2.7 (Provisional). A complete re-harvest of every model from independent sources corrected several stale or mislabeled cells (a DeepSeek score that was far too low, a Qwen coding number drawn from the wrong benchmark, and more), filled gaps, and upgraded many cells to higher-trust independent measurements. Effect on the board: the top stays a near-tie between ChatGPT 5.5 and Claude Opus 4.8; Qwen 3.7 Max earns a canonical rank as new coverage completes its profile; Gemini 3 Pro and DeepSeek V4 Pro rise on corrected data. MiniMax M2.7 (the prior MiniMax flagship) joins as Provisional. Every cell traces to an independent source. No imputation.

MiniMax M3 added, Ranked at #6. The new MiniMax flagship enters scored entirely from independent sources (Artificial Analysis, vals.ai, benchlm, LiveBench) rather than the lab's own numbers. It is strong on multimodal and general knowledge, weaker on language, and its independently-measured score lands well below its launch claims, which is exactly why we wait for third-party data. Priced low, so it places well in the Value view. No imputation.

Accessibility 100. Underlined the last in-text link that was distinguishable by color alone. Lighthouse: Accessibility 100, SEO 100, Best Practices 100 (desktop Performance 98).

Accessibility and SEO pass. Form controls and the sort menu now carry proper labels, in-text links are underlined rather than color-only, the page has a main landmark for screen readers, and muted text is lightened to readable contrast. Two action links became buttons so search engines crawl every link. Lighthouse accessibility rises from 70 toward the high 90s, with no change to any ranking or score.

Copy fix: 15 live benchmarks. The homepage, methodology, and page metadata now read 15 live benchmarks, matching the count the leaderboard has been computing for a while. The static text had lagged the live data by one. No scoring change.

Sensitivity bands on every score, plus a Claude Fable 5 update. Each score on the main leaderboard now shows a sensitivity band (for example 87.02 ±3.6): we remove each of a model's benchmarks one at a time, re-run the entire scoring pipeline, and report how far the score moves. It is not a confidence interval - it shows how much a score depends on which benchmarks exist, and it widens honestly when coverage is thin. When two models' bands overlap, read their order as a statistical tie. Separately, Claude Fable 5 gains its Humanity's Last Exam result from our standardized independent source, lifting its provisional score to about 95.8. It stays Provisional by explicit editorial hold, with the reason published in the open dataset and shown on its badge: ARC-AGI-2 has not yet been run on Fable 5, and every ranked model's reasoning score includes that benchmark, so ranking it today would compare unlike baskets. It ranks the day that result publishes. No imputation, and no corner-cutting in either direction.

Site visibility upgrade. The full technical methodology now lives at its own address (/methodology) rather than only inside a popup. Every model page now carries a written summary and a static benchmark table with source attribution, readable even without JavaScript - the same numbers as the interactive view. Added robots.txt and per-page structured data so search engines can find and understand every page. The open dataset (models.json) is now formally licensed CC BY 4.0 - cite it freely with attribution.

Claude Fable 5 added, on launch day. Anthropic's new generally-available flagship (its Mythos-class model with production safeguards) enters as Provisional. It posts the strongest agentic-coding and knowledge results on the board, but day-one coverage is uneven across components, so it is not yet canonically ranked. Two launch figures arrived above their benchmarks' human ceilings from a single source; we are holding both pending independent confirmation rather than letting them inflate the score. Fable lifts to ranked once independent evaluations complete its profile. No imputation, even for the year's most anticipated launch.

ARC-AGI-2 cleanup: one effort-tier fix, one version fix. Completing the High-effort standardization, ChatGPT 5.4's ARC-AGI-2 is now read at the same High tier as the rest of the column. Separately, a GLM ARC-AGI-2 result we had attributed to GLM 5.1 in fact belonged to the previous version, GLM 5; with no GLM 5.1 result published on that benchmark, the cell is removed rather than guessed. The correction lifts GLM 5.1 and takes it off the Reasoning leaderboard, where it no longer has a qualifying result. No imputation.

ARC-AGI-2 read at one effort level. Reasoning models can be run at several effort settings, and one model on the board had been carried at a higher tier than the rest. We now standardize ARC-AGI-2 on the High-effort tier so the column compares like-for-like. Claude Opus 4.8 (thinking) gains its ARC-AGI-2 result and joins the Reasoning leaderboard. Effect on the top: ChatGPT 5.5 and Claude Opus 4.8 stay within a fraction of a point, with ChatGPT 5.5 nominally first. No imputation.

SWE-bench Verified standardized on one independent source. Every model's SWE-bench Verified result now comes from the same independent third-party leaderboard (vals.ai), replacing a mix that leaned on a benchmark host whose public table has not refreshed since February. The result is consistent, current agentic-coding measurement across the whole column - and it tightens the top of the board: Claude Opus 4.8 (thinking) and ChatGPT 5.5 now sit within 0.2 points, effectively tied, with Opus 4.8 nominally first. No imputation.

HLE standardized on one independent source; Opus 4.8 joins the ranked board; Opus 4.6 retired. Humanity's Last Exam is now sourced consistently from Artificial Analysis (independent, no-tools) for every model it covers, replacing a patchwork of sources that disagreed by up to 20 points on the same model. Claude Opus 4.8 enters the ranked leaderboard at #2 now that independent reasoning and language scores complete its coverage. Per our scope rule (latest two versions per lab), Claude Opus 4.6 leaves the board. No imputation.

Value view polish. Added a reset for the workload-mix slider (back to the standard 75% input / 25% output), spaced out overlapping model labels on the scatter, and d weight customization while the Value view is open (weights don't change price-performance).

New: the Value view - capability per dollar. A price column (input/output API cost per million tokens) plus a Value tab that plots AGI Score against blended API price and highlights the value frontier - the best score available at each price point. A workload-mix slider weights input vs output cost for your use case. Pricing covers current frontier models; superseded or preview-only models without public pricing are labeled as such. Best value among frontier models - we do not track budget tiers. AGI Score remains the default view.

Source upgrade for Claude Opus 4.8 (thinking). Its SWE-bench Verified result is now drawn from an independent third-party evaluation that corroborates the launch-day figure, replacing the lab self-report we carried at launch. Same value, higher-trust source; the model stays Provisional pending independent reasoning data. No imputation.

Claude Opus 4.8 (thinking) added the day after launch, as Provisional. We hold its GPQA Diamond and AA Intelligence Index (independent T1) plus launch-day agentic results (SWE-bench Verified/Pro, OSWorld); reasoning-component coverage is still too thin for a canonical rank, so no AGI Score position yet. It lifts once independent reasoning and second-source agentic benchmarks publish. Claude Opus 4.6 stays on the board until then. No imputation.

Roster expansion to 18 models. Added Qwen 3.7 Max (Provisional), Cursor Composer 2.5 (System entry - scores reflect the full product harness, not weights alone), and MiMo V2.5 Pro. Lifted Claude Opus 4.7 and Claude Opus 4.6 (thinking) from awaiting-verification to Provisional after new variant-distinct evaluations surfaced.

Coding specialty restructured to reflect agentic-coding reality. Real coding in 2026 happens via tools like Claude Code, Codex, Cursor - not bare LLM completion. Composition: SWE-bench Verified 35 / SWE-bench Pro 30 / Terminal-Bench 35 (replacing Aider Polyglot, which had zero variant-distinct roster coverage after Round 14 cleanup). 9 of 15 roster models now have Terminal-Bench data (Round 25 harvest from vals.ai T1 + benchlm.ai T2 second-source). Added contamination note for SWE-bench Verified and harness disclosure for Terminal-Bench on the methodology page. Top of leaderboard shifts: GPT-5.5 takes #1 in Coding by ~0.7pp over Claude Opus 4.7 (thinking) - within statistical noise of the benchmarks' error bars, which is itself worth surfacing honestly. Each tab now answers its question directly.

/model/{apiName} . The existing detail view opens automatically when the URL is loaded directly, and clicking "View" on a leaderboard row updates the URL via History API. Document title and meta description update dynamically per model. sitemap.xml added with all 15 model URLs.## Independent verification, public corrections

Every cell on the leaderboard cites its source. When we find a mis-attribution, an inflated self-report contradicted by independent measurement, or a stale evaluation that predates the model itself, we correct it - and log the change here.

[Our error · 2026-07-26 We published a DeepSeek score we could not source 22.8%

DeepSeek V4 Pro · ARC-AGI-2 ·
While auditing every cell for methodology v2, we found an ARC-AGI-2 score attributed to DeepSeek V4 Pro with

Two different models holding the same number to three decimals, with one of them sourceless, is the same duplication pattern we found in an earlier audit of a secondary aggregator. We cannot establish that DeepSeek V4 Pro was ever evaluated on ARC-AGI-2, so the cell is voided. The record is retained rather than deleted, as with our other corrections.](/corrections/deepseek-arc-agi-2)

Read the full apology and what we changed

voided

no source at all: no URL, no evaluator, no record of where it came from. It is also identical, to three decimal places, to GLM 5.2’s ARC-AGI-2 score, which is properly sourced to the ARC Prize leaderboard.

This did not change any ranking. The cell had already been excluded from scoring before we found it, so no published score moved, on the current board or the previous one. We are logging it anyway. A correction that costs us nothing is the easiest kind to stay quiet about, and staying quiet is how a value with no source survives long enough to matter.

We owe DeepSeek an apology. We published a number against their model that we could not substantiate. We should have been more careful before publishing it. We will learn from this and will do better in the future. Four further cells were voided in the same pass for weaker provenance defects: a benchmark whose contest year was never pinned, two sourced to a vendor model page rather than the evaluation board that produced them, and one sourced to a commentary article about a different model entirely. None of the five score on the current board.

non-thinkingmode, and the figure reached us through a secondary aggregator that republishes other parties’ numbers. We compared it against their thinking-mode result and read the gap as a 17-point overstatement. Measured like for like, DeepSeek’s self-report sits 1.3 points above independent measurement. The original entry is retained below rather than deleted.

| Model | Benchmark | Change | Note |

|---|---|---|---|
| DeepSeek V4 Pro | GPQA Diamond | 0.901 (T3) → 0.729 (T2) | withdrawn: relayed self-report, not an independent measurement (see retraction) |
| Claude Opus 4.7 (thinking) | SWE-bench Verified | 0.876 (T2) → 0.820 (T1) | private-test-set evaluation (vals.ai) |
Model Benchmark Change Reason
ChatGPT 5.5 (Pro) BrowseComp 0.844 → 0.901 corrected to actual Pro column on release page
ChatGPT 5.5 (Pro) FrontierMath 0.517 → 0.524 corrected to actual Pro column on release page
ChatGPT 5.4 (Pro) SWE-bench Pro 0.577 → null value was from the GPT-5.4 base column
Claude Opus 4.6 (thinking) SWE-bench Pro 0.534 → null source row didn't distinguish thinking from base
Claude Opus 4.6 (thinking) OSWorld 0.727 → null source row didn't distinguish thinking from base
Claude Opus 4.7 (base) GPQA Diamond 0.942 → null source row didn't distinguish base from thinking
ChatGPT 5.5, 5.5 (Pro), 5.4, 5.4 (Pro) AA Intelligence Index 57-60 → null (4 cells) AA family row not variant-distinct
Claude Opus 4.7 (base) AA Intelligence Index 57 → null AA "(max)" attribution ambiguous
Model Benchmark Change Reason
ChatGPT 5.4 SWE-bench Verified 0.728 → null no longer on swebench.com leaderboard
ChatGPT 5.4 (Pro) SWE-bench Verified 0.728 → null no longer on swebench.com leaderboard
Model Benchmark Change Source
ChatGPT 5.5 SWE-bench Verified null → 0.826 vals.ai T1
ChatGPT 5.4 SWE-bench Verified null → 0.782 vals.ai T1 (restores the previously nulled stale-source cell)
Gemini 3.1 Pro (preview) SWE-bench Verified null → 0.788 vals.ai T1

| Claude Opus 4.6 (thinking) | SWE-bench Verified | null → 0.782 | vals.ai T1 (first variant-distinct SWE-V cell for this AWAITING model) | Visual Reasoning is still single-source, having moved from wholly Artificial Analysis to wholly vals.ai. It rests on a single benchmark, so it cannot be diversified until a second visual benchmark exists.

Strengths and weaknesses, at a glance #

Each cell shows a model's score on one of the five cognitive components. The AGI Score (rightmost column) blends these by parent weights (shown beneath each header). 100 marks the AGI threshold; values above are super-human on that dimension.

A component is not the same thing as a specialty tab, even where the names match. The Reasoning component here blends ARC-AGI-2, GPQA Diamond, HLE, AIME and LiveBench; the Coding specialty tab above uses only its three coding benchmarks. The two answer different questions and will not show the same number for the same model.

| Model | Agency 35% | Reasoning 29% | Knowledge 15% | Visual Reasoning 11% | Language 10% | AGI Score | | |---|---|---|---|---|---|---|---| | ... |

By Capability Area #

Top performers in each of the three public capability framings. The overall AGI Score blends these by their parent weights (Thinking 44 / Doing 43 / Communicating 13).

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @agi ranker 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-i-audited-my…] indexed:0 read:47min 2026-07-30 ·