How Did OpenAI’s IMO Team Reach Gold-Level Math Performance?

TL;DR
OpenAI’s IMO team reached gold-level performance by improving reinforcement learning algorithms, scaling test-time reasoning, and making progress on hard-to-verify tasks. The core team, Alex Wei, Sheryl Hsu, and Noam Brown, built on broader OpenAI work and completed the final sprint in only a couple of months. Read on to see how the proofs were verified and why the model’s response to problem six mattered.
Transcript
the pace of progress is is really I think you see it so clearly in math and I think Alex tweeted about this where you know even a few years ago these models were struggling with like grade school math and you know and then we we you know I remember even in 2024 that like GSMAK was used as like the standard eval when everybody would release a model ... Read More
Key Insights
- Mathematical benchmark progress has accelerated from grade-school evaluations to increasingly difficult competition problems. The discussion traces a rapid sequence through GSM8K, MATH, AIME, and USAMO-style difficulty, presenting IMO gold-level performance as the latest result in that progression.
- The core IMO team consisted of only three researchers, Alex Wei, Sheryl Hsu, and Noam Brown. Their focused effort was nevertheless built on work from many OpenAI groups, including teams responsible for inference, scaling, pre-training, and reinforcement learning.
- The final push toward the year’s IMO lasted roughly two months. The underlying reinforcement learning ideas had started coming together about six months earlier, but the decision to prepare a serious attempt for that specific competition came considerably later.
- The research approach targeted hard-to-verify tasks rather than limiting reinforcement learning to problems with straightforward verifiable rewards. Early improvements on these difficult tasks supplied the strongest evidence that the proposed technique might support a credible attempt at gold-level IMO performance.
- The published proofs were raw model outputs chosen for transparency. The team acknowledged that their style was difficult to read and could have used ChatGPT to rewrite them more clearly, but it released the originals so observers could inspect what the model actually produced.
- Proof correctness was assessed by external former IMO medalists. Three medalists independently graded each proof, and the graders reached unanimous agreement on correctness, providing expert human verification for material that some members of the research team could not personally evaluate.
- Problem six remained unsolved despite the team assigning substantial computation to it. The researchers described the problem as exceptionally difficult, with many possible directions and a narrow path to the proof, and the model ultimately returned no answer.
- Refusing to fabricate a solution was an important behavioral improvement. Earlier reasoning models could produce persuasive but incorrect answers when they lacked a solution, forcing experts to inspect every detail for subtle errors such as a reversed inequality.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: How did OpenAI’s IMO team achieve gold-level math performance?
The team improved reinforcement learning algorithms and developed a technique that showed progress on hard-to-verify tasks. It also used general-purpose methods for scaling test-time computation, allowing the model to reason for much longer periods. The result depended on broader OpenAI work in inference, scaling, pre-training, and reinforcement learning.
Q: Who was on OpenAI’s core IMO team?
The core team consisted of Alex Wei, Sheryl Hsu, and Noam Brown. Alex had worked on the central technique for a while, while Sheryl and Noam helped turn it into a serious IMO attempt. They also relied on contributions from other OpenAI teams.
Q: How long did OpenAI’s IMO gold effort take?
The reinforcement learning ideas began coming together about six months before the result. The focused final sprint to prepare for that year’s IMO lasted only a couple of months. The broader goal of achieving IMO gold had been considered for years.
Q: What early evidence convinced the team to pursue the IMO attempt?
The technique began showing improvement on hard-to-verify tasks. This mattered because much of the team’s earlier work had focused on tasks with verifiable rewards. The stronger results encouraged more people to support the initially small and uncertain effort.
Q: How were OpenAI’s IMO proofs verified?
External former IMO medalists evaluated the model’s original proofs. Three medalists independently graded each proof, and they unanimously agreed on correctness. OpenAI published the raw outputs so observers could inspect what the model actually produced.
Q: Why were the published IMO proofs difficult to read?
The team prioritized correct mathematical reasoning rather than optimizing the proof style for readability. Although the researchers considered using ChatGPT to rewrite the proofs more clearly, they released the original model outputs for transparency. This preserved a direct record of what the model produced.
Q: Why did the model not solve IMO problem six?
The researchers described problem six as exceptionally difficult, with many possible directions but a narrow path to the correct proof. They assigned substantial computation to it without finding a solution. The model ultimately returned no answer.
Q: Why was returning no answer on problem six significant?
The model did not fabricate a persuasive but unsupported proof after failing to solve the problem. Earlier reasoning models could produce convincing yet incorrect answers containing subtle errors, such as a reversed inequality. Returning no answer therefore represented an important behavioral improvement, even though the problem remained unsolved.
Summary & Key Takeaways
-
OpenAI’s IMO effort emerged from research ideas developed over roughly six months, followed by a focused final sprint lasting only a couple of months. Although the core team consisted of Alex Wei, Sheryl Hsu, and Noam Brown, their work depended on broader contributions involving inference, scaling, pre-training, and reinforcement learning.
-
The team’s approach emphasized general-purpose reinforcement learning methods that could improve performance on hard-to-verify tasks. Early evidence of progress encouraged continued investment despite skepticism. Instead of relying solely on easily verifiable rewards, the researchers explored techniques intended to handle outputs such as mathematical proofs, where evaluating correctness requires substantial expertise.
-
External former IMO medalists evaluated the model’s original proofs, with three graders assigned to each proof and unanimous agreement on correctness. The team published the raw outputs for transparency despite their poor readability. On problem six, the model returned no answer rather than producing a convincing but unsupported proof after extensive computation.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Sequoia Capital 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator