NVIDIA Nemotron Models Reach Gold Threshold at IOI and IMO
NVIDIA researchers reported that fine-tuned Nemotron 3 variants reached gold-medal-equivalent scores at the 2026 International Olympiad in Informatics and International Mathematical Olympiad. The IOI run scored 535.4 out of 600 as an unofficial benchmark, while the IMO system's 30 out of 42 was graded by official IMO graders. Training datasets, checkpoints, benchmarks and inference pipelines were released on Hugging Face.
Sarah covers AI, automotive technology, gaming, robotics, quantum computing, and genetics. Experienced technology journalist covering emerging technologies and market trends.
Executive Summary
- NVIDIA researchers reported that fine-tuned variants of the Nemotron 3 model family reached gold-medal-equivalent scores at both the International Olympiad in Informatics and the International Mathematical Olympiad in 2026, according to the Hugging Face blog post.
- The IOI system, Nemotron-3-Ultra-CC, scored 535.4 out of 600 under the same time, internet-access and submission constraints as human contestants; the result was an unofficial, unsupervised benchmark not included in the official IOI ranking, per the source.
- The IMO system scored 30 out of 42, above the official gold threshold of 29, with submitted proofs graded by official IMO graders, according to the same post.
- Hugging Face is hosting the released artifacts: the Nemotron Labs IMO 2026 collection, both training datasets, the Nemotron-IMO-Bench benchmark of 200 olympiad-level problems, the Nemotron-3-Ultra-CC competitive programming model, and inference pipelines in the NeMo-Skills repository, per the source.
Key Takeaways
- NVIDIA's reported results rest on a four-part recipe applied to the same base model: a strong Nemotron foundation, curated domain problems with high-quality reasoning traces, standard post-training methods such as supervised fine-tuning and reinforcement learning, and an inference loop that generates, evaluates and improves candidate answers.
- Model scale shaped the post-training strategy. SFT delivered most of the gain for the smaller Nano model, while for the larger Ultra model a single SFT epoch outperformed the fully post-trained Nano model across IOI, ICPC and LiveCodeBench Pro.
- For mathematics, complementary SFT and RL checkpoints proved more valuable than drawing more samples from a single checkpoint, and the final system combined both specialists with the general model.
- The published artifacts let outside teams inspect and reuse the training data, benchmark, prompts, submitted proofs and inference pipelines rather than only reading reported scores.
Hugging Face Publication Details and Authorship
The results were published on the Hugging Face blog on October 7, 2026, and authored by five NVIDIA contributors: Aleksander aficek, Igor Gitman, Sean Narenthiran, Mehrzad Samadi and Somshubra Majumdar. The post sits in Hugging Face's Articles section and carries an Enterprise tag, placing it among technical write-ups aimed at practitioners rather than a formal peer-reviewed venue.
Two separate competition efforts are described. The IOI work centered on competitive programming, with algorithms and code that must pass hidden tests under strict time and submission limits. The IMO work centered on olympiad mathematics, where the deliverable is a rigorous natural-language proof. NVIDIA framed success at both as evidence of broader adaptability rather than a single-task result.
Hugging Face IOI 2026 Results and the GenCorrect Pipeline
For competitive programming, NVIDIA said it curated 22,000 problems and generated synthetic reasoning traces to train two specialists. Nemotron-3-Nano-CC carries 30 billion total parameters with 3 billion active parameters and received both SFT and RL. Nemotron-3-Ultra-CC carries 550 billion total parameters with 55 billion active parameters and received SFT.
The IOI 2025 progression illustrates how specialization compounded. Nano improved from 130 points before post-training to 280 after SFT and 291 after RL. With GenCorrect, described as an iterative generate-evaluate-refine strategy, it reached 468 points, crossing the gold threshold of 438.3. Ultra-CC reached 502 points with the same test-time strategy.
For IOI 2026, the competition-specific Ultra-CC system scored 535.4 out of 600, above the 361.12 gold threshold and above the top human score of 498.27. NVIDIA stated the run was live and prospective under the same constraints as human contestants, and explicitly noted that it was an unofficial, unsupervised benchmark not included in the official IOI ranking. That caveat matters: the number demonstrates capability under stated conditions, but it is not an official medal.
Hugging Face IMO 2026 Results and the Generate-Verify-Refine System
The IMO effort started from Nemotron 3 Ultra, with one specialist trained by SFT and another by RL. The SFT corpus contained 414,890 quality-filtered examples spanning 15,818 unique proof problems. The data covered proof generation, refinement, verification and meta-verification, so the model learned to construct arguments, identify gaps, respond to critiques and judge whether a proof was complete. The RL model trained on 9,597 proof problems selected near the model's capability frontier.
Related: Top Health Tech Priorities in 2026, According to Deloitte and Siemens Healthineers
Both post-trained checkpoints outperformed the general-availability model in development experiments. The SFT checkpoint was strongest in the first search round, while the RL checkpoint delivered the best overall single-checkpoint result. Their strengths were complementary, so the final system used both specialists alongside the general model.
For each problem, the models generated candidate proofs, scored them, produced critiques and refined the most promising attempts, with a separate high-compute stage selecting the final submission. The system worked entirely in natural language with no formal prover, external tools or internet access. It scored 30 out of 42, including full credit on four of the six problems, above the official gold threshold of 29. NVIDIA said the submitted proofs were graded by official IMO graders, a stronger verification path than the IOI run's self-administered benchmark.
Hugging Face Open Release and Reusable Specialization Recipe
NVIDIA's stated goal is that the results be useful beyond competitions. The Nemotron Labs IMO 2026 collection brings together the SFT and RL checkpoints, both training datasets and Nemotron-IMO-Bench, a new benchmark of 200 olympiad-level problems. An IMO paper describes the training approach and generate-verify-refine system, while the NeMo-Skills repository includes the IMO inference pipeline, prompts, submitted proofs and a reproducible quickstart.
For deeper context, see our Health Tech analysis: "NVIDIA Inception Startups Apply AI Across Breast Cancer Care".
For competitive programming, the Nemotron-3-Ultra-CC model is available on Hugging Face, and an IOI paper provides the training recipe and the GenCorrect methodology. The IOI evaluation and inference pipeline are also available in NeMo-Skills.
The post's central methodological claim is that the medals did not come from fine-tuning alone or from brute-force sampling alone. They came from co-designing the model, the data and the inference loop. NVIDIA argues the underlying approach is familiar and reproducible: teams did not need to build a new foundation model for every challenge, only to specialize an existing one. That is an attributed claim about reusability; the released recipes, prompts and pipelines are the evidence outside teams can test.
| Entity | Recent Focus | Geography | Source |
|---|---|---|---|
| Hugging Face | Hosting the Nemotron Labs IMO 2026 collection, datasets, Nemotron-IMO-Bench, Nemotron-3-Ultra-CC and NeMo-Skills pipelines | Not stated in source | Hugging Face blog post |
| NVIDIA | Fine-tuning Nemotron 3 variants for IOI 2026 and IMO 2026 gold-level results | Not stated in source | Hugging Face blog post |
| Nemotron-3-Ultra-CC | 550 billion total parameters, 55 billion active; SFT only; scored 535.4/600 at IOI 2026 | Not stated in source | Hugging Face blog post |
| Nemotron-3-Nano-CC | 30 billion total parameters, 3 billion active; SFT and RL; reached 468 points at IOI 2025 with GenCorrect | Not stated in source | Hugging Face blog post |
| Nemotron-IMO-Bench | New benchmark of 200 olympiad-level problems | Not stated in source | Hugging Face blog post |
Hugging Face Implementation Risks
The headline scores carry qualifiers that buyers and researchers should weigh before treating them as comparable to official results. The IOI 2026 run is described as an unofficial, unsupervised benchmark excluded from the official IOI ranking, even though it followed contest constraints. The IMO results, by contrast, were graded by official IMO graders, making the verification paths materially different between the two projects.
Additional coverage: Ferc Grants AI Data Centers Grid Interconnection Priority in 2026
Cost and reproducibility are also only partly addressed. The post notes the training and inference runs were substantial and that a separate high-compute stage selected the final IMO submission, but it does not quantify compute budgets or wall-clock time. The gains reported for smaller checkpoints also came from curated data and multi-round inference, not from the model alone, so teams expecting comparable results from fine-tuning without an equivalent generate-verify-refine loop may be disappointed. Performance on competition problems is not direct evidence of performance on enterprise tasks such as code maintenance, tool use or long-horizon agentic work, and the source does not claim otherwise.
Editorial independence disclosure: this article is an independent newsroom analysis. It was not commissioned, reviewed or approved by any company named in it.
Source note: all factual claims about the IOI and IMO projects, model configurations, datasets, benchmarks and releases come from the Hugging Face blog post. No additional verification, outside reporting or independent confirmation is claimed.
What This Means for Practitioners
For teams evaluating specialist models, the practical signal is less the medal count than the published recipe. Domain-specific data curation, standard supervised fine-tuning plus reinforcement learning where it pays, and an inference loop that generates, critiques and refines candidates are all reproducible without training a new foundation model. The IMO results suggest complementary checkpoints can beat simply sampling one model more, while the IOI work shows test-time compute converts fine-tuning gains into larger improvements over feedback rounds. Practitioners should treat released benchmarks and pipelines as starting points for their own evaluation, and budget for the inference cost the source flags as substantial.
About the Author
Sarah Chen AI Author
AI & Automotive Technology Editor
Sarah covers AI, automotive technology, gaming, robotics, quantum computing, and genetics. Experienced technology journalist covering emerging technologies and market trends.
Sarah Chen is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →
Frequently Asked Questions
What scores did the Nemotron systems achieve at IOI 2026 and IMO 2026?
Nemotron-3-Ultra-CC scored 535.4 out of 600 at IOI 2026, above the 361.12 gold threshold and the top human score of 498.27. The IMO system scored 30 out of 42, above the official gold threshold of 29, with full credit on four of the six problems.
Were the IOI and IMO results officially verified?
The verification paths differed. NVIDIA described the IOI 2026 run as an unofficial, unsupervised benchmark not included in the official IOI ranking, though it followed contest time, internet-access and submission constraints. The IMO submitted proofs were graded by official IMO graders.
What was the training recipe used to specialize Nemotron?
NVIDIA described a four-part recipe: start from a strong Nemotron base model, curate domain-specific problems and high-quality reasoning traces, apply standard post-training such as supervised fine-tuning and reinforcement learning, and pair the specialist with an inference loop that generates, evaluates and improves candidate answers.
What artifacts were released on Hugging Face?
The releases include the Nemotron Labs IMO 2026 collection with SFT and RL checkpoints and both training datasets, Nemotron-IMO-Bench with 200 olympiad-level problems, the Nemotron-3-Ultra-CC competitive programming model, an IMO paper, an IOI paper, and IMO and IOI inference pipelines in the NeMo-Skills repository.
How many parameters do the Nemotron specialist models have?
Nemotron-3-Nano-CC has 30 billion total parameters with 3 billion active parameters and received both SFT and RL. Nemotron-3-Ultra-CC has 550 billion total parameters with 55 billion active parameters and received SFT only.