{"slug": "ai-system-locus-posttrains-a-model-better-than-qwen3", "title": "AI system Locus posttrains a model better than Qwen3", "summary": "Locus, an automated AI research system developed by Sakana AI, achieves state-of-the-art results on PostTrainBench and surpasses the human post-trained Qwen3 1.7B model in the extended PostTrainBench+ setting, using thousands of H100 hours. In a 16-day test across all live Kaggle competitions with prize money, Locus ranked 4th highest on average among all participants. The system is already in production, post-training models for millions of users, and has been used with Bubble, a no-code app development platform, to fine-tune an open-source model.", "body_md": "The models are improving the models.\nLocus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human post-trained Qwen3 model.\nToday, LLMs post-trained end-to-end by Locus are in production to millions. 🧵👇\nPostTrainBench evaluates agents' ability to post-train models on various domains given 10 H100 hours.\nWe extend PostTrainBench via PostTrainBench+, which has a greatly expanded compute budget that provides clearer signal on automated post-training capabilities. We find that thousands of H100 hours help distinguish methods' performance post-training Qwen3 1.7B-Base models, and that Locus scales best. In this setting, modes trained by Locus collectively surpass the perforamce of the offical human post-trained Qwen3 1.7B model.\nIn a test of generalization, we ran Locus on all live Kaggle competitions with prize money and public leaderboards. After 16 days, Locus achieved the 4th highest average rank among all participants.\n\n- We further analyze methods’ performance on PostTrainBench+ and find that large-scale experimentation is a crucial capability for complex post-training domains, like competition math. Locus is the most effective at deriving significant performance gains from large-scale trainingLocus is already translating these results to real business value. Earlier this year, we began working with\n[@bubble](https://x.com/bubble), the leading no-code app development platform. Locus, our autonomous research system, discovered and executed a post-training recipe that fine-tuned an open-sourceWe would like to thank the lead PostTrainBench authors ([@hrdkbhatnagar](https://x.com/hrdkbhatnagar),[@full__rank](https://x.com/full__rank), and[@maksym_andr](https://x.com/maksym_andr)) for supporting us in performing independent verification of Locus’ work in both the PostTrainBench and PostTrainBench+ settings. All results on the PostTrainBench+ setting # Join the conversation", "url": "https://wpnews.pro/news/ai-system-locus-posttrains-a-model-better-than-qwen3", "canonical_source": "https://twitter.com/intology/status/2084319121332965804", "published_at": "2026-08-03 18:34:08+00:00", "updated_at": "2026-08-03 18:52:29.135073+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-agents", "machine-learning"], "entities": ["Locus", "Sakana AI", "Qwen3", "PostTrainBench", "PostTrainBench+", "Kaggle", "Bubble"], "alternates": {"html": "https://wpnews.pro/news/ai-system-locus-posttrains-a-model-better-than-qwen3", "markdown": "https://wpnews.pro/news/ai-system-locus-posttrains-a-model-better-than-qwen3.md", "text": "https://wpnews.pro/news/ai-system-locus-posttrains-a-model-better-than-qwen3.txt", "jsonld": "https://wpnews.pro/news/ai-system-locus-posttrains-a-model-better-than-qwen3.jsonld"}}