cd /news/artificial-intelligence/ai-system-locus-posttrains-a-model-b… · home topics artificial-intelligence article
[ARTICLE · art-85102] src=twitter.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

AI system Locus posttrains a model better than Qwen3

Locus, an automated AI research system developed by Sakana AI, achieves state-of-the-art results on PostTrainBench and surpasses the human post-trained Qwen3 1.7B model in the extended PostTrainBench+ setting, using thousands of H100 hours. In a 16-day test across all live Kaggle competitions with prize money, Locus ranked 4th highest on average among all participants. The system is already in production, post-training models for millions of users, and has been used with Bubble, a no-code app development platform, to fine-tune an open-source model.

read1 min views1 publishedAug 3, 2026
AI system Locus posttrains a model better than Qwen3
Image: source

The models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human post-trained Qwen3 model. Today, LLMs post-trained end-to-end by Locus are in production to millions. 🧵👇 PostTrainBench evaluates agents' ability to post-train models on various domains given 10 H100 hours. We extend PostTrainBench via PostTrainBench+, which has a greatly expanded compute budget that provides clearer signal on automated post-training capabilities. We find that thousands of H100 hours help distinguish methods' performance post-training Qwen3 1.7B-Base models, and that Locus scales best. In this setting, modes trained by Locus collectively surpass the perforamce of the offical human post-trained Qwen3 1.7B model. In a test of generalization, we ran Locus on all live Kaggle competitions with prize money and public leaderboards. After 16 days, Locus achieved the 4th highest average rank among all participants.

  • We further analyze methods’ performance on PostTrainBench+ and find that large-scale experimentation is a crucial capability for complex post-training domains, like competition math. Locus is the most effective at deriving significant performance gains from large-scale trainingLocus is already translating these results to real business value. Earlier this year, we began working with @bubble, the leading no-code app development platform. Locus, our autonomous research system, discovered and executed a post-training recipe that fine-tuned an open-sourceWe would like to thank the lead PostTrainBench authors (@hrdkbhatnagar,@full__rank, and@maksym_andr) for supporting us in performing independent verification of Locus’ work in both the PostTrainBench and PostTrainBench+ settings. All results on the PostTrainBench+ setting # Join the conversation
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @locus 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-system-locus-post…] indexed:0 read:1min 2026-08-03 ·