cd /news/large-language-models/findstatbench-evaluating-large-langu… · home topics large-language-models article
[ARTICLE · art-67988] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis

Researchers introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis, containing 2,329 tasks across 24 collections and 5.52 million hidden instances. The benchmark reveals that the strongest open- and closed-source systems converge within 1 percentage point instance accuracy, examples can hurt performance, and exact symbolic rule induction remains brittle.

read1 min views1 publishedJul 22, 2026

arXiv:2607.18260v1 Announce Type: new Abstract: We introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis. Built from FindStat, it contains 2,329 tasks across 24 collections and 5.52M hidden instances, covering statistic synthesis, which maps objects to integers, and map synthesis, which maps objects to objects. Each task gives a mathematical description and at most five public input-output examples; a model must emit one Python solve function with no retrieval, tool use, execution feedback, voting, or reranking. Submissions are scored by exact sandboxed execution on held-out combinatorial objects. We evaluate eleven systems: four closed-source production models and seven open-weight models served through one inference provider. FindStatBench reveals three main patterns. First, the strongest open- and closed-source systems converge within 1 pp instance accuracy, and both an oracle over all systems and five-way sampling from one mid-tier model yield only limited task-accuracy gains. Second, examples can hurt: several classical bijections are solved perfectly with zero examples but fail under five-example prompts. Third, some failures reflect output-budget mechanics, as reasoning can exhaust the visible response before code is emitted. Overall, statistic synthesis is much easier than map synthesis, some collections remain near-zero, long prompts cause a sharp accuracy cliff, and exact symbolic rule induction remains brittle.

── more in #large-language-models 4 stories · sorted by recency
── more on @findstatbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/findstatbench-evalua…] indexed:0 read:1min 2026-07-22 ·