04:00
2026-07-22
arxiv.org
large-language-models
When JSON Is Not Enough: Semantic Reliability of Schema-Constrained LLM Ordering Agents
A new study from arXiv introduces OrderBench, a benchmark for restaurant ordering agents, finding that JSON schema-constrained outputs from large language models (LLMs) can achieve 100% syntactic valiβ¦