{"slug": "skip-expert-personas-in-prompts-no-accuracy-gain-up-to-4-5x-the-cost", "title": "Skip expert personas in prompts: no accuracy gain, up to 4.5x the cost", "summary": "A study testing 503 profession-specific system prompts across nine science benchmarks found no clear accuracy gain over minimal prompts, while expert personas produced up to 2.3x more output tokens and raised per-call cost 2.2 to 4.5x. The finding argues against adding expert personas to prompts for accuracy, since the practice mainly increases token usage and cost.", "body_md": "A study tested 503 profession-specific system prompts on nine science benchmarks. Matched expert profiles showed no clear accuracy gain over minimal prompts, yet produced up to 2.3x more output tokens and raised per-call cost 2.2 to 4.5x.\nRead: A study tested 503 profession-specific system prompts on nine science benchmarks. Matched expert profiles showed no clear accuracy gain over minimal prompts, yet produced up to 2.3x more output tokens and raised per-call cost 2.2 to 4.5x.\nRead: OpenRouter launched Model Router Benchmarks, scoring seven routers including NVIDIA Switchyard on six benchmarks with a blended Router Index, and explained why routing often loses to a single model: cache rebuilds, weak complexity signals and latency.\nRead: Ai2 released AstaBrief 8B, a Qwen3-8B fine-tune that turns a research question and literature excerpts into a cited report. Weights and training data are open, and it cut report latency from 178.5s to 51.1s against the Claude-based mode in Asta.\nRead: A study of reviewer models auditing 411 coding-agent traces found grounding in execution evidence lifted defect catch and cut over-rejection, while reviewer size predicted little. A cascade using generated tests worked without official tests.\nRead: Uber detailed its MCP Gateway, a proxy that exposes existing HTTP, gRPC and TChannel services as MCP tools, with a registry, API-crawling discovery and a control plane. It now hosts more than 800 MCP servers and 5,000 tools.\nRead: Shopify described ShopGym, which converts live storefronts into self-contained sandbox shops called ShopArena and generates grounded shopping tasks, so shopping agents can be benchmarked repeatably despite changing prices and bot detection.", "url": "https://wpnews.pro/news/skip-expert-personas-in-prompts-no-accuracy-gain-up-to-4-5x-the-cost", "canonical_source": "https://www.vibeleaderboard.ai/intel/brief/2026-10-03", "published_at": "2026-10-03 11:08:36+00:00", "updated_at": "2026-10-03 12:36:33.635479+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-tools"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/skip-expert-personas-in-prompts-no-accuracy-gain-up-to-4-5x-the-cost", "markdown": "https://wpnews.pro/news/skip-expert-personas-in-prompts-no-accuracy-gain-up-to-4-5x-the-cost.md", "text": "https://wpnews.pro/news/skip-expert-personas-in-prompts-no-accuracy-gain-up-to-4-5x-the-cost.txt", "jsonld": "https://wpnews.pro/news/skip-expert-personas-in-prompts-no-accuracy-gain-up-to-4-5x-the-cost.jsonld"}}