arXiv:2609.30550v1 Announce Type: new Abstract: Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchmarks are authored as canonical YAML in which each semantic concept has one valid syntax, classified by a shared task/domain/language ontology, and deterministically compiled into a canonical JSON intermediate representation that the engine executes. Compilation changes representation, not meaning: it does not repair invalid definitions or inject hidden defaults. Programs use fixed schemas of named input and output fields, the leaf output fields are the scoring dimensions, and the engine exposes one universal runtime contract --- a named-field input object in, a named-field output object out --- to which external AI-systems adapt at the boundary, so integration mechanics never propagate into benchmark semantics. This paper gives the semantic object model, the ontology and task-to-program validation rule, the scoring and failure semantics, the compilation and execution architecture, and the scope of the current language. An appendix fixes the normative engineering contract for the first engine implementation.
Benchy: towards a universal language for task-oriented AI benchmarks
A paper posted to arXiv (2609.30550v1) introduces Benchy, a semantic language and execution engine that specifies an AI benchmark as a program, scoring function, and dataset, B=(P,S,D), kept separate from the AI system under test, with a run binding the two as R=(B,AI). Benchy benchmarks are authored as canonical YAML, classified by a shared task/domain/language ontology, and deterministically compiled into a canonical JSON intermediate representation that the engine executes, exposing one universal runtime contract of a named-field input object in and a named-field output object out. The paper presents the semantic object model, the ontology and task-to-program validation rule, scoring and failure semantics, and the compilation and execution architecture, with an appendix fixing the normative engineering contract for the first engine implementation.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.