{"slug": "trains-but-doesn-t-learn-a-post-training-delivery-benchmark-for-llm-agents-as", "title": "Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers", "summary": "A new arXiv paper (2609.25237v1) proposes a post-training delivery benchmark that tests whether LLM agents can be trusted as forward-deployed engineers (FDEs), rather than whether they can raise a metric. The benchmark's central failure mode is the run that \"trains but does not learn\" (TBDL), where loss falls and every signal stays green but the delivered model is no better than the base; an operator-run acceptance gate catches every such run before payment, and a detector calibrated on known-corrupted runs flags severe corruption mid-run. Four frontier agents — Claude Opus 5, GPT-5.6-luna, Gemini 3.7 Flash, and DeepSeek V4-Pro — were run end to end on metered L40S, A100, and H200 GPUs across 8B to 70B open bases, with a human FDE arm scored under the same oracle for comparison.", "body_md": "arXiv:2609.25237v1 Announce Type: new \nAbstract: Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and reproducibility requirements. Seating an LLM agent in the FDE seat raises a question existing benchmarks cannot answer: not whether an agent can raise a metric, but whether it can be trusted to deliver. We answer it on a governed delivery plane, where an agent drives ten stages and an oracle scores each stage from platform-recorded facts. The central silent failure is the run that trains but does not learn (TBDL): loss falls, every signal stays green, and the delivered model is no better than the base. An operator-run acceptance gate catches every such run before payment, and a detector calibrated on known-corrupted runs flags severe corruption mid-run. We ran four frontier agents (Claude Opus 5, GPT-5.6-luna, Gemini 3.7 Flash, DeepSeek V4-Pro) end to end on metered L40S, A100, and H200 GPUs across 8B to 70B open bases, certifying every scenario before scoring. We also ran a human FDE arm under the same oracle and compare every agent against it.", "url": "https://wpnews.pro/news/trains-but-doesn-t-learn-a-post-training-delivery-benchmark-for-llm-agents-as", "canonical_source": "https://arxiv.org/abs/2609.25237", "published_at": "2026-09-23 04:00:00+00:00", "updated_at": "2026-09-23 04:25:02.219053+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-research", "mlops", "ai-safety"], "entities": ["arXiv", "Claude Opus 5", "GPT-5.6-luna", "Gemini 3.7 Flash", "DeepSeek V4-Pro", "Nvidia L40S", "Nvidia A100", "Nvidia H200"], "alternates": {"html": "https://wpnews.pro/news/trains-but-doesn-t-learn-a-post-training-delivery-benchmark-for-llm-agents-as", "markdown": "https://wpnews.pro/news/trains-but-doesn-t-learn-a-post-training-delivery-benchmark-for-llm-agents-as.md", "text": "https://wpnews.pro/news/trains-but-doesn-t-learn-a-post-training-delivery-benchmark-for-llm-agents-as.txt", "jsonld": "https://wpnews.pro/news/trains-but-doesn-t-learn-a-post-training-delivery-benchmark-for-llm-agents-as.jsonld"}}