The models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human post-trained Qwen3 model. Today, LLMs post-trained end-to-end by Locus are in production to millions. 🧵👇 PostTrainBench evaluates agents' ability to post-train models on various domains given 10 H100 hours. We extend PostTrainBench via PostTrainBench+, which has a greatly expanded compute budget that provides clearer signal on automated post-training capabilities. We find that thousands of H100 hours help distinguish methods' performance post-training Qwen3 1.7B-Base models, and that Locus scales best. In this setting, modes trained by Locus collectively surpass the perforamce of the offical human post-trained Qwen3 1.7B model. In a test of generalization, we ran Locus on all live Kaggle competitions with prize money and public leaderboards. After 16 days, Locus achieved the 4th highest average rank among all participants.
- We further analyze methods’ performance on PostTrainBench+ and find that large-scale experimentation is a crucial capability for complex post-training domains, like competition math. Locus is the most effective at deriving significant performance gains from large-scale trainingLocus is already translating these results to real business value. Earlier this year, we began working with @bubble, the leading no-code app development platform. Locus, our autonomous research system, discovered and executed a post-training recipe that fine-tuned an open-sourceWe would like to thank the lead PostTrainBench authors (@hrdkbhatnagar,@full__rank, and@maksym_andr) for supporting us in performing independent verification of Locus’ work in both the PostTrainBench and PostTrainBench+ settings. All results on the PostTrainBench+ setting # Join the conversation