cd /news/artificial-intelligence/veriphy-agentic-physical-reasoning-f… · home topics artificial-intelligence article
[ARTICLE · art-121120] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

VeriPhy, an auditable physical-verification system that compiles prompts into typed physical obligations and uses frozen low-level experts to produce provenance-carrying evidence, accounts for 228 of 304 human-annotated flaw records on a 149-clip core, outperforming a published question-decomposition evaluator (164) and matching a monolithic backbone (222) while retaining full traceability. The system, introduced in arXiv paper 2609.03153v1, aims to evaluate and refine world models by localizing generation failures in prompt reference, space, and time.

read1 min views1 publishedSep 4, 2026

arXiv:2609.03153v1 Announce Type: new Abstract: Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of indicating the obligation a clip violates or the moment it fails. We present VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed. During execution, observations gate and scope only declared calls to frozen low-level experts (e.g., segmentation and tracking, counting, eleven typed physical measurements over the resulting tracks, depth, OCR, and audio-event detection). Each action returns a provenance-carrying evidence record whose payload, when usable, is either a typed measurement or an explicitly tagged learned state. Typed resolvers and fixed composition map usable records to a three-valued state (supported, contradicted, or unknown, surfaced as plausible, implausible, or abstain) with full provenance, so that every verdict is traceable to the evidence that produced it. We anchor evaluation in a 1,500-clip corpus of human-annotated flaw records that localize real generation failures in prompt reference, space, and time. On a 149-clip core carrying 304 such records, VeriPhy accounts for 228, against 164 for a published question-decomposition evaluator given the same clips and the same claims. Recall alone does not separate it from prompting the same backbone monolithically, which reaches 222; what separates them is that each decision retains its evidence record and provenance, making the traces auditable one verdict at a time and usable as the interface through which a critic verdict could be written back into generation.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @veriphy 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/veriphy-agentic-phys…] indexed:0 read:1min 2026-09-04 ·