cd /news/ai-agents/rules-to-tools-executable-checks-for… · home › topics › ai-agents › article
[ARTICLE · art-143630] src=arxiv.org ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Rules to Tools: Executable Checks for LLM Agents in Scientific Computing

A study posted to arXiv (2610.00313v1) found that scientific coding agents given prepared executable checks completed 29 of 30 repair tasks, versus 26 of 30 when given written requirements alone. In the eight-task-ID cohort the tool group scored 15/16 against 13/16 for text, with a task-cluster bootstrap 95% interval of [-12.5, 43.75] percentage points, while the larger shared-definition SciCode cohort tied at 13/24 per group. In a matched PDE comparison, detailed text scored 23/24 and checks scored 24/24, with 31.2% lower reported model output for the checks group, though public CPU use rose in both task-ID cohorts.

by read1 min views4 publishedOct 2, 2026

arXiv:2610.00313v1 Announce Type: new Abstract: Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requirements. Matched SciCode repair groups share written checks, starting programs, model, and budgets; the tool group receives a callable implementation. Across two task-ID cohorts, complete repair is 26/30 with text and 29/30 with the prepared checks. Three task IDs favor tools, one favors text, and eleven tie. The eight-ID cohort scores 13/16 versus 15/16, with a task-cluster bootstrap 95% interval of [-12.5, 43.75] percentage points for the difference. The larger shared-definition SciCode cohort ties at 13/24 per group. Five development-exposed tasks with alternate starting programs score 3/10 versus 7/10. The tool group favors tasks 17, 77, and 11; initial checks flag task 17 and report no violation for tasks 77 and 11. Task 37 favors text and has no initial reported violation. A fresh source-through-Python arm also reaches 15/16, matching the dedicated command's aggregate. In a matched PDE comparison, detailed text scores 23/24 and checks score 24/24, with 31.2% lower reported model output for checks. Agent-side output savings vary by cohort, while public CPU use rises in both task-ID cohorts. These results measure task-dependent repair outcomes and agent-side costs with prepared checks.

── more in #ai-agents 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/rules-to-tools-execu…] indexed:0 read:1min 2026-10-02 · —