Post-Training Language Models for Gold-Medal Performance in Coding Competitions NVIDIA researchers report that their Nemotron-3-Ultra-CC (550B-A55B) system scored 535.4 out of 600 on the IOI 2026 problem set, exceeding both the gold threshold of 361.12 and the top human score of 498.27, marking the first AI system to outscore the highest-scoring human contestant on an IOI problem set. The pipeline combines 22,000 curated problems, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL), with GenCorrect, a feedback-driven test-time compute strategy, boosting Nemotron-3-Nano-CC (30B-A3B) from 130 to 468 points on IOI 2025, surpassing the gold threshold of 438.3. Computer Science Machine Learning Submitted on 2 Sep 2026 Title:Post-Training Language Models for Gold-Medal Performance in Coding Competitions View PDF /pdf/2609.02849 HTML experimental https://arxiv.org/html/2609.02849v1 Abstract:Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning SFT , and reinforcement learning RL . Using 22,000 curated problems, we train Nemotron-3-Nano-CC 30B-A3B with SFT and RL and Nemotron-3-Ultra-CC 550B-A55B with SFT alone. We further introduce GenCorrect, a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions. On IOI 2025, Nano-CC improves from 130 points to 291 after post-training and to 468 with GenCorrect, exceeding the gold threshold of 438.3 while Ultra-CC reaches 502. Guided by these results, we develop a competition-specific Ultra-CC system and evaluate it prospectively during IOI 2026. Under the same time, internet-access, and submission constraints as human contestants, it scores 535.4 out of 600, exceeding both the gold threshold of 361.12 and the top human score of 498.27. To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set. Current browse context: cs.LG References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender IArxiv Recommender What is IArxiv? https://iarxiv.org/about arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .