CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers? CheckerBench was introduced as a benchmark for evaluating whether long-horizon coding agents can synthesize static-analysis checkers, a task requiring agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. The benchmark's authors note that existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability work rather than checker synthesis. Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulne