WeirdML v3 Håvard Tveit Ihle at the Norwegian Defence Research Establishment (NDRE) released WeirdML v3, an agentic benchmark of 11 complex hand-made tasks designed to test whether models can explore unfamiliar data and build machine learning and data analysis pipelines with limited data, unspecified goals and limited feedback. API costs were supported primarily by EpochAI, secondarily by METR and NDRE. Official scores are 80% normalized area under each model's best-so-far curve on a logarithmic token axis plus 20% of its final value, counting only the 500k to 50M token interval, with shaded bands showing approximate 95% run-uncertainty intervals. WeirdML v3 WeirdML v3 is an agentic benchmark featuring 11 complex hand-made tasks made to challenge the model to explore and understand unfamiliar data, develop machine learning and data analysis pipelines and produce appropriate results despite limited data, unspecified goals and/or very limited feedback. WeirdML v3 was created by me Håvard Tveit Ihle https://htihle.github.io/ at the Norwegian Defence Research Establishment NDRE https://www.ffi.no/en/about-ffi . API costs were supported primarily by EpochAI https://epoch.ai/ , secondarily by METR https://metr.org/ and NDRE https://www.ffi.no/en/about-ffi . Thanks for the support Previous versions: WeirdML v2 weirdml v2.html · WeirdML v1 weirdml v1.html . Open standalone: Interactive plot https://htihle.github.io/weirdml v3 interactive.html · Model summary https://htihle.github.io/weirdml v3 summary.html · Prepared data .json https://htihle.github.io/assets/data/weirdml v3.json · Full data .json https://htihle.github.io/data/weirdml v3 results.json In the Tokens view above, each model’s line shows its average best-so-far effective score across all 11 tasks. That model’s official score is 80% normalized area under its line on a logarithmic token axis, plus 20% of its final value. Only the interval from 500k to 50M tokens contributes to the area; earlier progress is shown for context, and the best score reached before 500k carries into the scoring window. The best-so-far score is carried forward to the full 50M-token limit, even if a run ends early. The curve averages runs within each configuration, then weights each of the 11 tasks equally; hinted and hintless twins each receive half their task’s weight. Ship Detect’s cost-weighted tokens are scaled ×25: its native 20k–2M scoring window maps to 500k–50M on the combined plot. Effective scores include normalization and hint penalties; they are not raw accuracy. Select Per Task to explore any of the 15 configurations, Cost or Date to compare official scores, or Open vs Closed to compare the score frontiers over time. Shaded score bands show approximate 95% run-uncertainty intervals, with variance pooled across configurations and models. Black markers show all 15 configuration means. Cost is mean API cost per run, using the same task weighting as scores. Final Best Score is the equally weighted mean final effective score. Harness shows the agent software and version used for the included runs. Only models with at least one valid run in every configuration are included.