cd /news/artificial-intelligence/esq-bench-a-multi-tier-enterprise-or… · home topics artificial-intelligence article
[ARTICLE · art-111227] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence

Researchers introduced ESQ-Bench, an Oracle-first NL2SQL benchmark with six populated schemas (465 tables, 164,682 rows) and 550 gold-validated question-query pairs, finding that GPT-4o schema-linked execution accuracy drops from 79.8% to 57.2% across complexity tiers, while Claude Sonnet 4.6 achieves 87.4%, 74.9%, and 68.7% EX, outperforming GPT-4o on every tier. The benchmark reveals silent semantic divergence rates of 73–99% among execution-passing queries and highlights a gap between closed API models and open-weight baselines like Llama 3.2, which reaches only 13.3% bank-wide EX.

read1 min views1 publishedAug 26, 2026

arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD. However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environments. We introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema complexity tiers. We constructed and released six populated schemas (465 tables, 164,682 rows, zero empty tables) with identical seed data on Oracle, PostgreSQL, MySQL, and SQL Server, a four-metric evaluation harness (EM, EX, SR, SD), and 550 gold-validated question-query pairs (Tier-1: 95; Tier-2: 228; Tier-3: 227). Schema-linked prompting with GPT-4o shows monotonic execution-match degradation across tiers: 79.8, 60.3, and 57.2 percent EX on executed queries (June 2026), versus 75.6, 80.4, and 95.8 percent on an earlier 142-question pilot slice. EM stays below 7 percent tier-wide; operational silent-divergence reaches 73 to 99 percent among EX-passing queries. Failure analysis shows wrong-result semantics dominate at higher tiers. Claude Sonnet 4.6 with schema-linked prompts reaches 87.4, 74.9, and 68.7 percent EX (executed queries), exceeding GPT-4o schema-linked on every tier. GPT-4o zero-shot EX on executed queries (78.7, 73.5, and 77.8 percent) inverts schema-linked at Tiers 2 to 3 due to lower execution rates and survivor bias in the zero-shot versus schema-linked analysis. Local Llama 3.2 schema-linked reaches only 13.3 percent bank-wide EX (73 out of 550), underscoring the gap between closed API models and open-weight baselines on enterprise Oracle schemas.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @esq-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/esq-bench-a-multi-ti…] indexed:0 read:1min 2026-08-26 ·