Nvidia posted on Hugging Face this week detailing how its Nemotron 3 family performed at this year's International Olympiad in Informatics (IOI 2026) and International Mathematical Olympiad (IMO 2026): both results cleared the gold-medal threshold. Compared with similar results posted by closed models over the past year, what's different this time is that the model, training data, and inference pipeline were all released publicly.
How the two competitions were run
For informatics, the competing system was Nemotron-3-Ultra-CC, a model with 550 billion total parameters and 55 billion active per token, fine-tuned with supervised fine-tuning and layered with a correction pipeline called GenCorrect. It finished with 535.4 points out of 600, clearing the 361.12 gold-medal line. Nvidia stressed this was a "live" test: the model ran during the actual contest window under the same time, internet-access, and submission-count limits as human contestants. The score is unofficial, and no human intervened during the run.
On the data side, the informatics model was trained on roughly 22,000 curated problems plus synthetic reasoning traces. A smaller sibling, Nemotron-3-Nano-CC (30 billion total parameters, 3 billion active), used supervised fine-tuning plus reinforcement learning and scored 468 on last year's IOI 2025 problem set; its score on this year's contest wasn't disclosed in the post.
On the math side, the score was 30 out of 42, above the 29-point gold-medal line, with four of six problems solved perfectly. Answers were graded by official IMO graders, and the whole process used no formal theorem prover, no external tools, and no internet access — just natural-language proofs. The system combined several checkpoints of Nemotron 3 Ultra — a general version, a supervised-fine-tuned version, and a reinforcement-learned version — cycling through a "generate, verify, revise" loop to rewrite proofs. Training drew on 15,818 problems and 414,890 quality-filtered proof samples.
Compared with a year ago
In July 2025, Google DeepMind's Gemini Deep Think and an experimental OpenAI reasoning model each scored 35 on the IMO, crossing that year's gold-medal line. Both were closed models, and outsiders could only see the results.
Nvidia's score this time is 5 points lower than those two from last year, just scraping past the line. But more has been released: Nemotron-3-Ultra-CC's weights are live on Hugging Face, the training dataset sits in Nemotron Labs' IMO 2026 collection, the inference pipeline is in the NeMo-Skills repository with a reproducible quickstart guide, and there's also a 200-problem proof benchmark called Nemotron-IMO-Bench.
The post didn't disclose the sampling count or compute budget used at inference time, nor did it make a direct comparison with models from OpenAI, Google, or DeepSeek. Generate-verify-revise loops are typically compute-hungry, and how many rounds or GPU-hours each problem takes will only become clear once someone reproduces the pipeline from the repository.
For model teams in China, the 400,000-plus filtered competition-grade proof samples are ready-to-download material that's easier to put to direct use than the weights themselves. Deploying the 550-billion-parameter Ultra version locally has a high bar, so most teams are more likely to start from the dataset and the benchmark.
Sources: Nvidia's Hugging Face technical blog post, the NeMo-Skills code repository, CocoLoop, and earlier IMO result announcements from Google DeepMind and OpenAI; the blog post was used to verify the scores, gold-medal lines, parameter counts, and training-data scale for both competitions.