cd /news/ai-agents/trains-but-doesn-t-learn-a-post-trai… · home topics ai-agents article
[ARTICLE · art-137789] src=arxiv.org ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

A new arXiv paper (2609.25237v1) proposes a post-training delivery benchmark that tests whether LLM agents can be trusted as forward-deployed engineers (FDEs), rather than whether they can raise a metric. The benchmark's central failure mode is the run that "trains but does not learn" (TBDL), where loss falls and every signal stays green but the delivered model is no better than the base; an operator-run acceptance gate catches every such run before payment, and a detector calibrated on known-corrupted runs flags severe corruption mid-run. Four frontier agents — Claude Opus 5, GPT-5.6-luna, Gemini 3.7 Flash, and DeepSeek V4-Pro — were run end to end on metered L40S, A100, and H200 GPUs across 8B to 70B open bases, with a human FDE arm scored under the same oracle for comparison.

by read1 min views2 publishedSep 23, 2026

arXiv:2609.25237v1 Announce Type: new Abstract: Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and reproducibility requirements. Seating an LLM agent in the FDE seat raises a question existing benchmarks cannot answer: not whether an agent can raise a metric, but whether it can be trusted to deliver. We answer it on a governed delivery plane, where an agent drives ten stages and an oracle scores each stage from platform-recorded facts. The central silent failure is the run that trains but does not learn (TBDL): loss falls, every signal stays green, and the delivered model is no better than the base. An operator-run acceptance gate catches every such run before payment, and a detector calibrated on known-corrupted runs flags severe corruption mid-run. We ran four frontier agents (Claude Opus 5, GPT-5.6-luna, Gemini 3.7 Flash, DeepSeek V4-Pro) end to end on metered L40S, A100, and H200 GPUs across 8B to 70B open bases, certifying every scenario before scoring. We also ran a human FDE arm under the same oracle and compare every agent against it.

── more in #ai-agents 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/trains-but-doesn-t-l…] indexed:0 read:1min 2026-09-23 ·