{"slug": "tencent-researchers-say-they-can-create-agent-training-tasks-for-0-05", "title": "Tencent researchers say they can create agent training tasks for $0.05", "summary": "Tencent HY LLM Frontier researchers and collaborators demonstrated that AI agent training tasks can be generated synthetically for $0.05 each, producing 1,000 accepted tasks for $50. Fine-tuning open-weights models on these trajectories improved Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench 2, Terminal-Bench Hard, and Long-Horizon Terminal Bench, according to the paper 'Recursive Synthesis for Long-Horizon Terminal Tasks'.", "body_md": "[AI](/tag/ai/)\n\nAI agents rely on good training exercises to produce consistent results. Humans have needed to build these complex exercises by hand and validate the agent's work, and humans are expensive.\n\nBut what if you could just magic up increasingly difficult tasks synthetically, through cheap recursion, to train agents to string together a whole lot of terminal commands into a workflow that solves a real-world problem?\n\nThat, said Tencent HY LLM Frontier researchers and collaborators in a new paper, is not only possible, but absurdly cheap, with a demonstrated case of 1,000 accepted tasks for $50, or $0.05 each.\n\nThe researchers said that fine-tuning open-weights models using the agent trajectories, the steps the agent took working on a task, saw meaningful improvements in the base model: \"Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench 2, Terminal-Bench Hard, and Long-Horizon Terminal Bench.\"\n\nThe effect is to move the bottleneck for building terminal agents from \"can we afford the data\" to \"how far do we want to push the recursion\", [ said](https://www.linkedin.com/feed/update/urn:li:activity:7491546477144477696/?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAAvEzcBadaY7R05tXOdNbPBs4EFPcOniRY) PwC Austria head of AI engineering Pascal Biese.\n\n**No-ceiling recursion**\n\nYou can think of the approach described in __the paper, Recursive Synthesis for Long-Horizon____ __[ Terminal Tasks](https://arxiv.org/abs/2608.05466?ref=thestack.technology), as a reversal,\n\n[joint lead author Yucheng Shi.](https://www.linkedin.com/posts/yucheng-shi_building-synthetic-long-horizon-terminal-share-7491366305699373057-hwL6/?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAAvEzcBadaY7R05tXOdNbPBs4EFPcOniRY)\n\n__said__## Get the full story: Subscribe for free\n\nJoin peers managing over $100 billion in annual IT spend and subscribe to unlock full access to The Stack’s analysis and events.\n\n[Subscribe now](https://www.thestack.technology/membership/)\n\nAlready a member? [Sign in](https://www.thestack.technology/signin/)", "url": "https://wpnews.pro/news/tencent-researchers-say-they-can-create-agent-training-tasks-for-0-05", "canonical_source": "https://www.thestack.technology/agent-training-tasks-cheap-recursion/", "published_at": "2026-08-12 10:07:47+00:00", "updated_at": "2026-08-12 10:36:28.363491+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research"], "entities": ["Tencent HY LLM Frontier", "Qwen3.5-27B", "Qwen3.5-122B-A10B", "Terminal-Bench 2", "Terminal-Bench Hard", "Long-Horizon Terminal Bench", "Pascal Biese", "Yucheng Shi"], "alternates": {"html": "https://wpnews.pro/news/tencent-researchers-say-they-can-create-agent-training-tasks-for-0-05", "markdown": "https://wpnews.pro/news/tencent-researchers-say-they-can-create-agent-training-tasks-for-0-05.md", "text": "https://wpnews.pro/news/tencent-researchers-say-they-can-create-agent-training-tasks-for-0-05.txt", "jsonld": "https://wpnews.pro/news/tencent-researchers-say-they-can-create-agent-training-tasks-for-0-05.jsonld"}}