Nemotron Scores 535.4 IOI and 30/42 IMO, Passing Gold in Both
In this article
NVIDIA reports that one Nemotron 3 model family has now passed gold-medal thresholds in both elite reasoning competitions. A specialized Nemotron-3-Ultra-CC system scored 535.4 out of 600 on the International Olympiad in Informatics (IOI) 2026, above the 361.12 gold cut-off and the top human score of 498.27. A separate Nemotron 3 Ultra ensemble scored 30 out of 42 on the International Mathematical Olympiad (IMO) 2026, above the official gold threshold of 29.
For machine-learning teams, the two outcomes demonstrate a repeatable specialization recipe: start with a strong base model, curate domain-specific problems and reasoning traces, apply supervised fine-tuning and reinforcement learning where useful, and pair the result with an inference loop that generates, evaluates, and improves candidate answers.
Checkpoint Specialization and Parameter Scaling
On the coding side, NVIDIA curated 22,000 problems and generated synthetic reasoning traces to train two specialists. Nemotron-3-Nano-CC uses 30 billion total parameters and 3 billion active parameters, the same 30B/3B-active shape as the Nemotron 3.5 Lightning model we covered earlier; Nemotron-3-Ultra-CC uses 550 billion total parameters and 55 billion active parameters.
NVIDIA's IOI 2025 measurements show Nano improving from 130 points before post-training to 280 after supervised fine-tuning and 291 after reinforcement learning. With GenCorrect, an iterative generate-evaluate-refine strategy, the same checkpoint reached 468 and crossed the 438.3 gold threshold. Ultra-CC reached 502 on IOI 2025 with the same test-time loop. One SFT epoch was enough for Ultra-CC to outperform the fully post-trained Nano across IOI, ICPC, and LiveCodeBench Pro, which led NVIDIA to use an SFT-only Ultra-CC checkpoint for the 535.4 IOI 2026 run.
For mathematics, NVIDIA started from Nemotron 3 Ultra and trained two specialist checkpoints. The SFT model processed 414,890 quality-filtered examples across 15,818 unique proof problems, covering proof generation, refinement, verification, and meta-verification. The RL model trained on 9,597 proof problems selected near the model's capability frontier. In NVIDIA's development experiments, the SFT checkpoint performed best in the first search round, while the RL checkpoint achieved the strongest overall single-checkpoint result.
Inference Loops That Cross the Thresholds
The checkpoints did not reach gold on their own. For the IOI 2026 result, NVIDIA paired Ultra-CC with the GenCorrect loop in a live, prospective run under the same time, internet-access, and submission constraints as human contestants. NVIDIA describes the evaluation as unofficial and unsupervised, and it was not included in the official IOI ranking.
The IMO system used a generate-verify-refine ensemble of the general, SFT, and RL checkpoints. It operated entirely in natural language, with no formal prover, external tools, or internet access. The models generated candidate proofs, scored them, produced critiques, and refined the strongest attempts; a separate high-compute stage selected the final submission. Official IMO graders evaluated the submitted proofs, and the system received full credit on four of the six problems.
| Competition Track | Model and Post-Training Configuration | Inference Architecture | Vendor-Reported Result |
|---|---|---|---|
| IOI 2026 (Coding) | Nemotron-3-Ultra-CC (550B total / 55B active, SFT only) | GenCorrect (Iterative generate-evaluate-refine loop) | 535.4/600 (Gold threshold: 361.12) |
| IMO 2026 (Math) | Nemotron 3 Ultra (Ensemble: General, SFT, and RL checkpoints) | Multi-stage generate-verify-refine with final high-compute selection | 30/42 (Gold threshold: 29) |
AI Mastery analysis
NVIDIA's results are further evidence that pipeline architecture, not just model weights, now drives frontier benchmark gains. The IMO setup splits generator and verifier roles across specialized SFT and RL checkpoints, avoiding the failure mode where a single model's verifier misses semantic errors in its own generated logic.
The scale-specific post-training gap is also instructive. A 550B-parameter checkpoint needed only one SFT epoch to beat a fully post-trained 30B-parameter model, while the smaller model's SFT-plus-RL run still produced most of its gain from SFT with a smaller but consistent RL improvement. That suggests large base models may mainly need domain formatting, while smaller active-parameter models benefit from additional reinforcement.
Finally, NVIDIA's release of the 200-problem Nemotron-IMO-Bench, the IMO checkpoints and datasets, the NeMo-Skills inference pipeline, and the Ultra-CC model weights on Hugging Face shifts the practical bottleneck for developers from access to proprietary models toward inference-time search orchestration.
Sources
Frequently asked questions
What score did NVIDIA Nemotron achieve on IOI 2026?
NVIDIA reports Nemotron-3-Ultra-CC scored 535.4 out of 600, above the 361.12 gold threshold and the top human score of 498.27. The system paired SFT with the GenCorrect iterative loop.
How did NVIDIA reach IMO 2026 gold with Nemotron?
A Nemotron 3 Ultra ensemble using general, SFT, and RL checkpoints in a generate-verify-refine system scored 30 out of 42, above the official gold threshold of 29. It received full credit on four of six problems and used no formal prover, external tools, or internet access.
What was the Nemotron-3-Nano-CC progression on IOI 2025?
NVIDIA reports Nano improved from 130 points before post-training to 280 after SFT and 291 after RL. With GenCorrect it reached 468, crossing the 438.3 gold threshold.
What did NVIDIA release alongside these results?
The Nemotron Labs IMO 2026 collection includes SFT and RL checkpoints, both training datasets, and the 200-problem Nemotron-IMO-Bench. NeMo-Skills provides the IMO inference pipeline, and the 550B-parameter Ultra-CC model is available on Hugging Face.
Related Reading
ThinkingCap-Qwen3.8-27B Cuts Thinking Tokens 37.2% for 0.86pp Accuracy
BottleCap AI's fine-tune of Qwen3.8-27B drops thinking tokens 37.2% across 12 benchmarks, losing just 0.86pp of macro accuracy at xhigh effort.
Nvidia's AVO Harness Takes Claude Opus 5 from 30% to 100% on ARC-AGI-3
Nvidia's custom AVO harness lifted Claude Opus 5 from 30% to 100% on ARC-AGI-3 — without changing the model at all.
pplx-embed-v2-late: 0.6B Queries 9B Index at 63.5% ViDoRe
Perplexity's MIT-licensed pplx-embed-v2-late pairs a 0.6B edge encoder with a 9B indexer in one embedding space; the 9B scores 92.4% on MADQA.