{"slug": "vals-ai-benchmark-international-olympiad-in-informatics", "title": "Vals.ai Benchmark – International Olympiad in Informatics", "summary": "Vals.ai's new International Olympiad in Informatics benchmark found that GPT-6 Astra solved every problem across all three years tested, while GPT-5.6 Sol scored 91.17%, Claude Fable 5.1 scored 90.78% and GPT-5.6 Terra scored 87.61%, with the 26-model field spreading across 90 points down to 9.33%. Cohort mean accuracy fell from 58.38% on 2024 problems to 53.67% on 2025 and 50.82% on 2026, and 15 of the 26 models scored lowest on the 2026 problems, though the two leaders hit 100% on 2026 despite the contest running after their training cutoffs. Each model ran inside the OpenCode coding agent in an isolated sandbox with a C++20 toolchain, seeing only the problem statement, sample grader and one sample, with no internet, official test data or submission feedback.", "body_md": "## Key Takeaways\n\n- Unlike the saturated knowledge benchmarks, IOI still sharply separates models: [GPT-6 Astra](/models/openai_gpt-6-astra) solves every problem in all three years,[GPT-5.6 Sol](/models/openai_gpt-5.6-sol) (91.17%),[Claude Fable 5.1](/models/anthropic_claude-fable-5-1) (90.78%) and[GPT-5.6 Terra](/models/openai_gpt-5.6-terra) (87.61%) follow, and the twenty-six-model field then spreads across ninety points, down to 9.33%, so competitive-programming ability remains a real differentiator.\n- The agent sees only what a contestant sees: the statement, the sample grader and one sample. It has no tests, no internet and no submission feedback, so every point comes from code the model wrote and checked itself.\n- Scores decline on the newest problems: cohort mean accuracy is 58.38% on 2024 and 53.67% on 2025 but 50.82% on 2026, and fifteen of the twenty-six models score lowest on the 2026 problems. [GPT 5.3 Codex](/models/openai_gpt-5.3-codex) drops from 62.83% on 2024 to 41.67% on 2026. The two leaders score 100% on 2026 even though the contest ran after their training cutoffs, so recency alone does not cap performance.\n\n## Why IOI?\n\nRecently, top LLM labs like OpenAI and Google reported that their models achieved gold medals on the International Mathematical Olympiad (IMO). However, advanced models are starting to saturate IMO, meaning it may no longer effectively differentiate between the capabilities of top-performing models.\nReports also suggest the evaluation process faced coordination challenges, with [AI companies seeking expedited validation mid-competition](https://www.google.com/url?q=https://www.reddit.com/r/math/comments/1m6sooc/a_brief_perspective_from_an_imo_coordinator/&sa=D&source=docs&ust=1754700624812439&usg=AOvVaw3bdkgASRiEfR7HSaRDy2r7) that may not reflect standard IMO assessment procedures.\n\nThe International Olympiad in Informatics (IOI) offers several advantages as an LLM benchmark. Unlike the IMO, the IOI is not yet saturated, providing clear differentiation between model capabilities. The competition features standardized and automated grading, ensuring objective evaluation without subjective scoring. Additionally, the IOI has real-world relevance as it tests C++ programming skills that are directly applicable to software development.\n\n## Benchmark Design\n\nWe designed our benchmark to imitate competition conditions as closely as possible.\n\n### Agent Harness\n\nEach model runs inside [OpenCode](https://opencode.ai), a general-purpose coding agent, in an isolated sandbox with a `c++` (v20) toolchain. Its workspace holds exactly what the [contest hands a contestant](https://www.ioi2024.org/gradingenvironment): the problem statement, the task header and solution stub, the sample grader or manager, the compile and run scripts, and the sample input and output from the statement. The official test data, the subtask test lists, and the grading script are withheld until the attempt is over.\nThe agent reads, writes, compiles and tests files freely in its workspace, and writes its answer to `/workspace/solution.cpp`. Like a contestant, it has no access to the public internet: the sandbox reaches only our model gateway, so published editorials and reference solutions are out of reach.\n\nUnlike a contestant, the agent has no submission tool and therefore receives no score feedback during the attempt: it can run the samples and whatever tests it writes for itself, and must decide on its own when its solution is good enough.\n\nSome models spend their entire per-response output budget reasoning about a problem before writing any code. When that happens, the harness continues the cut-off turn on the next step rather than failing the problem: it replays the provider’s own reasoning state where the API returns one, and otherwise restarts the turn with an instruction to work in smaller steps. The same rule applies to every model on this board, and the cost column includes the reasoning spent this way.\n\nResults from our earlier harness, which gave the agent an interactive grading tool, are preserved on [IOI v1](/benchmarks/ioi-v1). The two harnesses are not comparable, so their results are reported on separate boards.\n\n### Scoring\n\nGrading matches the olympiad: after the attempt ends, the solution left in the workspace is copied into a fresh grading directory alongside pristine copies of the official test data and grader, compiled, run against every official test, and scored per *subtask* out of 100 points. Every grading file is verified unchanged after the run.\nA solution that fails to compile, or that the agent never wrote, scores zero.\nA model’s score for a year is its mean score across that year’s six problems, and its overall score is the mean of its three yearly scores.\n\n## Results\n\n[GPT-6 Astra](/models/openai_gpt-6-astra) reaches 100% accuracy, solving all eighteen problems, ahead of [GPT-5.6 Sol](/models/openai_gpt-5.6-sol) at 91.17%, [Claude Fable 5.1](/models/anthropic_claude-fable-5-1) at 90.78%, [GPT-5.6 Terra](/models/openai_gpt-5.6-terra) at 87.61% and [Claude Opus 5](/models/anthropic_claude-opus-5) at 84.33%. Fable 5.1 and Opus 5 both solve every 2024 problem and score 87.17% on 2025; Fable holds 85.17% on 2026 while Opus falls to 65.83%. [Qwen 3.8 Max](/models/alibaba_qwen3.8-max) (68.89%), [GLM 5.3](/models/zai_glm-5.3) (68.44%), [Gemini 3.7 Flash](/models/google_gemini-3.7-flash) (67.83%), [GPT-5.6 Luna](/models/openai_gpt-5.6-luna) (61.78%), [Gemini 3.8 Flash](/models/google_gemini-3.8-flash) (56.94%), [Muse Spark 1.3 Max](/models/meta_muse_spark_1_3_max) (56.56%), [GPT 5.3 Codex](/models/openai_gpt-5.3-codex) (53.83%), [GLM 5.3 Flash](/models/zai_glm-5.3-flash) (52.50%), [DeepSeek V4 Pro 0813](/models/deepseek_deepseek-v4-pro-0813) (51.61%), [Kimi K3](/models/kimi_kimi-k3) (48.94%), [Grok 4.6](/models/grok_grok-4.6) (47.61%), [Claude Sonnet 5](/models/anthropic_claude-sonnet-5) (45.00%) and [Muse Spark 1.3](/models/meta_muse_spark_1_3) (43.94%) form a middle group, while [Grok 4.5](/models/grok_grok-4.5) (40.56%), [DeepSeek V4.1 Flash](/models/deepseek_deepseek-v4.1-flash) (40.28%, the cheapest model on the board), [Qwen 3.8 27B](/models/alibaba_qwen3.8-27b) (39.06%, thirty points behind Qwen 3.8 Max), [Gemini 3.6 Flash](/models/google_gemini-3.6-flash) (35.06%), [DeepSeek V4 Flash 0731](/models/deepseek_deepseek-v4-flash-0731) (32.72%), [Muse Spark 1.2](/models/meta_muse_spark_1_2) (21.78%), [Inkling](/models/thinkingmachines_inkling) (14.94%) and [Inkling Small](/models/thinkingmachines_inkling-small) (9.33%) trail the field. Both Muse Spark 1.3 variants at least double the score of their predecessor Muse Spark 1.2. Claude Sonnet 5 is the most expensive model on the board at $19.33 per problem: on thirteen of its eighteen problems it spent its entire 128k-token output budget reasoning before writing any code, and the harness had to continue the cut-off turn.\n\nWe evaluate three years to check for data contamination.\nThe 2026 problems were released only after the models in this cohort were trained, and they are the hardest set for fifteen of the twenty-six: [GPT 5.3 Codex](/models/openai_gpt-5.3-codex) scores 62.83% on 2024 and 57.00% on 2025 but 41.67% on 2026, and cohort mean accuracy falls from 58.38% (2024) and 53.67% (2025) to 50.82% (2026). [GPT-6 Astra](/models/openai_gpt-6-astra) and [GPT-5.6 Terra](/models/openai_gpt-5.6-terra) both score 100% on 2026, so the newest set is not out of reach.", "url": "https://wpnews.pro/news/vals-ai-benchmark-international-olympiad-in-informatics", "canonical_source": "https://www.vals.ai/benchmarks/ioi", "published_at": "2026-09-12 15:05:04+00:00", "updated_at": "2026-09-12 15:16:52.819755+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents", "developer-tools"], "entities": ["Vals.ai", "International Olympiad in Informatics", "GPT-6 Astra", "GPT-5.6 Sol", "Claude Fable 5.1", "GPT-5.6 Terra", "GPT 5.3 Codex", "OpenCode"], "alternates": {"html": "https://wpnews.pro/news/vals-ai-benchmark-international-olympiad-in-informatics", "markdown": "https://wpnews.pro/news/vals-ai-benchmark-international-olympiad-in-informatics.md", "text": "https://wpnews.pro/news/vals-ai-benchmark-international-olympiad-in-informatics.txt", "jsonld": "https://wpnews.pro/news/vals-ai-benchmark-international-olympiad-in-informatics.jsonld"}}