cd /news/large-language-models/format-sensitivity-index-token-contr… · home › topics › large-language-models › article
[ARTICLE · art-58279] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

A new study from arXiv introduces the Format Sensitivity Index (FSI) and Parseability Sensitivity Index (PSI) to measure how prompt wrapper formatting affects LLM benchmark scores. Across 140,000 OpenRouter generations with 7 QA tasks, 5 wrapper families, and 4 models from 7B to 72B parameters, mean FSI varied by over 30x across models, largely due to compliance failures. The authors argue that reporting accuracy without wrapper variance is statistically fragile and offer recommendations for benchmarking and structured-output deployments.

read1 min views24 publishedJul 14, 2026

arXiv:2607.09665v1 Announce Type: new Abstract: Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions. We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced by wrapper choice, and the Parseability Sensitivity Index (PSI), the corresponding range in answer parseability. Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by over 30x across models and is largely explained by compliance failures. A fixed-effects regression shows that parseability remains a strong predictor of accuracy even after controlling for task, model, and wrapper. We argue that reporting accuracy without wrapper variance and compliance is statistically fragile, and we give practical recommendations for both benchmarking and structured-output deployments.

── more in #large-language-models 4 stories · sorted by recency
dev.to · · #large-language-models
ChatGPX
── more on @openrouter 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/format-sensitivity-i…] indexed:0 read:1min 2026-07-14 · —