Researchers developing the Nemotron model family reached gold-medal scoring levels at both the 2026 International Olympiad in Informatics and the 2026 International Mathematical Olympiad, Hugging Face reported on October 7, 2026. The two competitive milestones came from adapting Nemotron 3 base checkpoints through supervised fine-tuning, reinforcement learning, and structured inference loops rather than creating separate foundation models for each domain.
At IOI 2026, the Nemotron-3-Ultra-CC system scored 535.4 out of 600 points using supervised fine-tuning alongside an iterative generation and refinement strategy called GenCorrect. That total surpassed the competition's gold medal threshold of 361.12 points and exceeded the highest human contestant score of 498.27 points. The IOI run was prospective and took place under identical time, submission, and internet limits as human participants, though organizers did not include the unsupervised benchmark in the official tournament rankings.
Competitive programming results
Engineers built the programming models using a curated set of 22,000 coding problems paired with synthetic reasoning traces. They trained two variants: Nemotron-3-Nano-CC, which contains 30 billion total parameters and 3 billion active parameters, and Nemotron-3-Ultra-CC, which holds 550 billion total parameters and 55 billion active parameters. Nano received both supervised fine-tuning and reinforcement learning, whereas Ultra received only supervised fine-tuning.
Earlier benchmarks on IOI 2025 demonstrated how each training phase adjusted model capability. Nano climbed from 130 points in its base state to 280 points after supervised fine-tuning and 291 points following reinforcement learning, eventually reaching 468 points when paired with GenCorrect against a gold cutoff of 438.3. Ultra reached 502 points under the same test-time setup. One epoch of supervised fine-tuning enabled Ultra to surpass the fully post-trained Nano model on IOI, ICPC, and LiveCodeBench Pro tasks.
Mathematical proof system
The mathematics project adapted Nemotron 3 Ultra to generate natural-language proofs evaluated directly by official IMO graders. Its supervised fine-tuning corpus comprised 414,890 quality-filtered examples spanning 15,818 unique proof problems. This data trained the network to construct arguments, inspect reasoning gaps, answer critiques, and judge mathematical completeness. A separate reinforcement learning checkpoint trained on 9,597 proof problems selected near the model's capability edge.
During IMO 2026 evaluation, the final system coordinated the general model alongside the fine-tuned and reinforcement learning checkpoints in a generate-verify-refine pipeline. Candidate proofs were drafted, scored, critiqued, and revised without formal theorem provers, external software tools, or network connectivity. The pipeline earned 30 out of 42 points, clearing the official 29-point gold medal threshold and collecting full credit across four of the six exam problems.
Released benchmarks and code
Project contributors released the underlying software, data, and model checkpoints on Hugging Face. The Nemotron Labs IMO 2026 collection includes both fine-tuned checkpoints, the complete training datasets, and Nemotron-IMO-Bench, an evaluation suite of 200 olympiad problems. The team also uploaded Nemotron-3-Ultra-CC to the platform and published inference pipelines, prompt records, and reproducible quickstart configurations in the NeMo-Skills repository.
