cd /news/large-language-models/skip-expert-personas-in-prompts-no-a… · home › topics › large-language-models › article
[ARTICLE · art-144440] src=vibeleaderboard.ai ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Skip expert personas in prompts: no accuracy gain, up to 4.5x the cost

A study testing 503 profession-specific system prompts across nine science benchmarks found no clear accuracy gain over minimal prompts, while expert personas produced up to 2.3x more output tokens and raised per-call cost 2.2 to 4.5x. The finding argues against adding expert personas to prompts for accuracy, since the practice mainly increases token usage and cost.

read1 min views1 publishedOct 3, 2026

A study tested 503 profession-specific system prompts on nine science benchmarks. Matched expert profiles showed no clear accuracy gain over minimal prompts, yet produced up to 2.3x more output tokens and raised per-call cost 2.2 to 4.5x. Read: A study tested 503 profession-specific system prompts on nine science benchmarks. Matched expert profiles showed no clear accuracy gain over minimal prompts, yet produced up to 2.3x more output tokens and raised per-call cost 2.2 to 4.5x. Read: OpenRouter launched Model Router Benchmarks, scoring seven routers including NVIDIA Switchyard on six benchmarks with a blended Router Index, and explained why routing often loses to a single model: cache rebuilds, weak complexity signals and latency. Read: Ai2 released AstaBrief 8B, a Qwen3-8B fine-tune that turns a research question and literature excerpts into a cited report. Weights and training data are open, and it cut report latency from 178.5s to 51.1s against the Claude-based mode in Asta. Read: A study of reviewer models auditing 411 coding-agent traces found grounding in execution evidence lifted defect catch and cut over-rejection, while reviewer size predicted little. A cascade using generated tests worked without official tests. Read: Uber detailed its MCP Gateway, a proxy that exposes existing HTTP, gRPC and TChannel services as MCP tools, with a registry, API-crawling discovery and a control plane. It now hosts more than 800 MCP servers and 5,000 tools. Read: Shopify described ShopGym, which converts live storefronts into self-contained sandbox shops called ShopArena and generates grounded shopping tasks, so shopping agents can be benchmarked repeatably despite changing prices and bot detection.

── more in #large-language-models 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/skip-expert-personas…] indexed:0 read:1min 2026-10-03 · —