A study tested 503 profession-specific system prompts on nine science benchmarks. Matched expert profiles showed no clear accuracy gain over minimal prompts, yet produced up to 2.3x more output tokens and raised per-call cost 2.2 to 4.5x. Read: A study tested 503 profession-specific system prompts on nine science benchmarks. Matched expert profiles showed no clear accuracy gain over minimal prompts, yet produced up to 2.3x more output tokens and raised per-call cost 2.2 to 4.5x. Read: OpenRouter launched Model Router Benchmarks, scoring seven routers including NVIDIA Switchyard on six benchmarks with a blended Router Index, and explained why routing often loses to a single model: cache rebuilds, weak complexity signals and latency. Read: Ai2 released AstaBrief 8B, a Qwen3-8B fine-tune that turns a research question and literature excerpts into a cited report. Weights and training data are open, and it cut report latency from 178.5s to 51.1s against the Claude-based mode in Asta. Read: A study of reviewer models auditing 411 coding-agent traces found grounding in execution evidence lifted defect catch and cut over-rejection, while reviewer size predicted little. A cascade using generated tests worked without official tests. Read: Uber detailed its MCP Gateway, a proxy that exposes existing HTTP, gRPC and TChannel services as MCP tools, with a registry, API-crawling discovery and a control plane. It now hosts more than 800 MCP servers and 5,000 tools. Read: Shopify described ShopGym, which converts live storefronts into self-contained sandbox shops called ShopArena and generates grounded shopping tasks, so shopping agents can be benchmarked repeatably despite changing prices and bot detection.
Is your AI agent worth its tokens? We measured it with TigerGraph