cd /news/artificial-intelligence/hugging-face-agents-fully-reproduced… · home topics artificial-intelligence article
[ARTICLE · art-96667] src=mlq.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Hugging Face agents fully reproduced 266 of 2,226 ICML papers

Hugging Face reported that coding agents fully reproduced every extracted claim in 266 of 2,226 ICML 2026 papers, with 632 more showing partial evidence and no falsified claims. The 19-day challenge, which ran from July 15 to August 2, involved 1,221 participants who published 6,816 logbooks covering 35,908 claims, with missing artifacts being the most common obstacle to verification.

read5 min views1 publishedAug 14, 2026
Hugging Face agents fully reproduced 266 of 2,226 ICML papers
Image: Mlq (auto-discovered)
  • The challenge published 6,816 logbooks covering 2,226 ICML 2026 papers and 35,908 claims. [1] - Only 266 papers had every extracted claim independently verified; 632 more had partial evidence with no claim falsified. [1] - Missing artifacts were the most common reason researchers could not establish a result, while 242 papers produced conflicting reproduction verdicts.
[[1]](https://huggingface.co/blog/icml-2026-open-reproductions) - The strongest results came from human-directed agent workflows, not fully autonomous runs.
[[1]](https://huggingface.co/blog/icml-2026-open-reproductions)

A community experiment using coding agents fully reproduced every extracted claim in 266 of the 2,226 ICML 2026 papers it attempted to test, according to results published Thursday by Hugging Face. [1]

The 19-day challenge produced a less binary picture across the broader sample: 1,103 papers had at least one independently verified claim, while 496 had at least one claim falsified or contested. Another 502 produced only toy-scale evidence and 280 yielded no conclusive evidence. [1]

A claim-level audit, not a conference-wide pass rate #

The project ran from July 15 through August 2 and invited researchers to select papers, identify major claims and use coding agents to rebuild experiments or check mathematical arguments. Participants used tools including Claude Code, Codex, Cursor and OpenResearch. Each attempt generated a public Trackio logbook containing the write-up, code and experiment artifacts, with full agent traces optionally published as datasets. [1][2]

Hugging Face said 1,221 people joined the effort, publishing 6,816 logbooks across 2,226 papers. The attempts covered about one-third of the conference, according to the organizers. [1]

The challenge extracted 35,908 claims and had an automated judge assign one of four labels: verified, falsified, toy or inconclusive. The judge used the open-weights GLM-5.2 model and was instructed to treat each participant’s assessment as untrusted. Hugging Face froze the verdicts in a public dataset at the close of the challenge. [1][3]

The categories should not be read as a single pass-or-fail score. Papers differed in claim count, experiment size, data access and reproduction depth. Some teams substituted synthetic data when original datasets were proprietary or checkpoints were unavailable. A further denominator issue remains: the challenge site described 6,341 accepted papers, while Hugging Face’s retrospective referred to 6,352 accepted papers. That discrepancy does not affect the reported 2,226 attempts, but it should be resolved before using the challenge as a precise conference-wide sample. [1][2]

Missing artifacts were the dominant obstacle #

The organizers said missing artifacts were the most common reason researchers could not establish a result. Unreleased checkpoints, proprietary datasets and incomplete implementation details pushed 280 papers into the inconclusive category, while 502 more could be tested only at toy scale. Code availability therefore did not equal reproducibility: a repository could exist while still lacking the data, weights, configuration or compute needed to recreate the published result. [1]

The challenge also surfaced concrete technical errors. In one theoretical paper on learning-augmented paging, a participant found that an additive term grew with the logarithm of a parameter rather than remaining constant. The organizers extended the test to k = 1,024 and said the result confirmed the failure. In another paper, three teams found counterexamples to a theorem, with violations appearing after 224, about 3,800 and 6,416 steps. The authors acknowledged that finding and were preparing a correction, Hugging Face said. [1]

A separate reproduction found that a transformer evaluation counted padding tokens in roughly 66% of label positions. Correcting that evaluation changed the paper’s reported 3.1% quality cost for a 50% cache reduction to roughly 9.4%, according to the challenge account. In at least one case, however, a claimed falsification was itself wrong: the participant had compared per-trajectory time with per-batch time and, after normalization, reproduced the paper’s reported speedup. [1]

Agents widened the search; people still checked the premise #

Hugging Face’s account says the most reliable attempts involved a human steering the agent, particularly when results depended on scale, units or qualitative judgment. Agents sometimes stopped experiments too early, became trapped in local loops or built an apparent falsification on a units mismatch. [1]

The organizers described a human-in-the-loop reproduction of a quantized image-generation paper. Numerical metrics suggested that image quality had not collapsed, but the final assessment required a person to inspect 128 image pairs. The agent built the review interface and checked the resulting annotations; the human supplied the perceptual judgment. [1]

The challenge’s scoring system awarded two points for either a full reproduction or a full falsification, one point for toy-scale evidence and zero for an inconclusive result. An independent commentary by R. Thompson, who identifies as a Ph.D. holder and AI architect, argued that treating confirmation and falsification symmetrically was one of the project’s most consequential design choices. That commentary was reaction to the challenge, not independent validation of its final statistics. [3][4]

ICML’s public 2026 materials describe its own review and transparency policies, including public reviews for accepted papers, but do not indicate that ICML organized or certified the Hugging Face exercise. The challenge documentation identifies Hugging Face and alphaXiv as its organizers, making this an external reproduction effort rather than an ICML certification. [2][5]

The public record is the result #

The project’s most durable contribution may be the evidence bundle rather than the headline percentages. Each logbook can expose the code, artifacts, execution trace and assumptions behind a verdict, allowing another researcher to challenge a false positive or extend a partial reproduction. Hugging Face said authors had already confirmed findings on multiple papers, with two arXiv corrections in progress and one case of independent convergence after an author had quietly fixed an error in a newer version. [1]

The results depend on which papers participants chose, how claims were extracted, how much compute teams could access and whether an automated judge correctly interpreted each logbook. The final dataset is therefore best read as a public map of where selected ICML claims survived, failed or remained untested—not as a single reproducibility score for the conference.

Further sources #

[[1] Hugging Face, “What We Learned by Reproducing 2,200 papers from ICML,” August 1… ↗](https://huggingface.co/blog/icml-2026-open-reproductions)

[[2] ICML 2026 Open Reproductions challenge site. Describes the July 15–August 2 cha… ↗](https://icml-2026-agent-repro-challenge.static.hf.space/index.html)

[[3] ICML-2026-agent-repro verdict dataset. Public repository for the frozen claim-l… ↗](https://huggingface.co/datasets/ICML-2026-agent-repro/verdicts)

[[4] ICML-2026-agent-repro FAQ. Defines leaderboard scoring: two points for full rep… ↗](https://huggingface.co/spaces/ICML-2026-agent-repro/challenge/blob/main/faq.html)

[5] ICML 2026 Reviewer Instructions and Author Instructions. Official conference ma… ↗

The stories that matter, in one email. Free — unsubscribe anytime.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @hugging face 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/hugging-face-agents-…] indexed:0 read:5min 2026-08-14 ·