How Did OpenAI Reach Gold-Level IMO Math?

TL;DR
OpenAI reached gold-medal performance on International Mathematical Olympiad problems by scaling reasoning time and applying general-purpose reinforcement learning techniques to hard-to-verify tasks. A three-person core team completed the final sprint in roughly two months, while external former IMO medalists unanimously graded the proofs and the model notably declined to invent an answer for problem six.
Transcript
the pace of progress is is really I think you see it so clearly in math and I think Alex tweeted about this where you know even a few years ago these models were struggling with like grade school math and you know and then we we you know I remember even in 2024 that like GSMAK was used as like the standard eval when everybody would release a model ... Read More
Key Insights
- Mathematical benchmark progress has accelerated from grade-school evaluations to increasingly difficult competition problems. The discussion traces a rapid sequence through GSM8K, MATH, AIME, and USAMO-style difficulty, presenting IMO gold-level performance as the latest result in that progression.
- The core IMO team consisted of only three researchers, Alex Wei, Sheryl Hsu, and Noam Brown. Their focused effort was nevertheless built on work from many OpenAI groups, including teams responsible for inference, scaling, pre-training, and reinforcement learning.
- The final push toward the year’s IMO lasted roughly two months. The underlying reinforcement learning ideas had started coming together about six months earlier, but the decision to prepare a serious attempt for that specific competition came considerably later.
- The research approach targeted hard-to-verify tasks rather than limiting reinforcement learning to problems with straightforward verifiable rewards. Early improvements on these difficult tasks supplied the strongest evidence that the proposed technique might support a credible attempt at gold-level IMO performance.
- The published proofs were raw model outputs chosen for transparency. The team acknowledged that their style was difficult to read and could have used ChatGPT to rewrite them more clearly, but it released the originals so observers could inspect what the model actually produced.
- Proof correctness was assessed by external former IMO medalists. Three medalists independently graded each proof, and the graders reached unanimous agreement on correctness, providing expert human verification for material that some members of the research team could not personally evaluate.
- Problem six remained unsolved despite the team assigning substantial computation to it. The researchers described the problem as exceptionally difficult, with many possible directions and a narrow path to the proof, and the model ultimately returned no answer.
- Refusing to fabricate a solution was an important behavioral improvement. Earlier reasoning models could produce persuasive but incorrect answers when they lacked a solution, forcing experts to inspect every detail for subtle errors such as a reversed inequality.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: How did OpenAI achieve gold-level performance on IMO problems?
OpenAI combined improving reinforcement learning algorithms with a technique designed to make progress on hard-to-verify tasks. The core researchers developed the ideas over roughly six months and then conducted a focused final sprint lasting only a couple of months. They also scaled the model’s ability to reason for much longer periods and relied on broader OpenAI work in inference, scaling, pre-training, and reinforcement learning.
Q: How large was OpenAI’s core IMO research team?
The core effort involved three researchers: Alex Wei, Sheryl Hsu, and Noam Brown. They described the project as small and scrappy, with Alex having worked on the central technique for some time before Sheryl and Noam helped turn it into an IMO attempt. The result also depended on contributions from other OpenAI teams working on inference, scaling, pre-training, and reinforcement learning.
Q: How long did OpenAI work on its IMO gold attempt?
The reinforcement learning ideas behind the result began coming together about six months before the competition effort concluded. The specific decision to make a serious push for that year’s IMO led to a final sprint of only a couple of months. The broader ambition had existed for years, but the immediate preparation was comparatively short and focused.
Q: How were OpenAI’s IMO proofs checked for correctness?
OpenAI hired external former IMO medalists to evaluate the model’s proofs. Each proof was graded by three medalists, and the graders reached unanimous agreement about correctness. This expert review was necessary because the outputs were technically demanding and sometimes beyond the personal grading ability of team members, even though one researcher had studied mathematics.
Q: Why were the model’s published proofs difficult to read?
The team concentrated on achieving correct mathematical reasoning and did not optimize the outputs as aggressively for human readability. The researchers said the raw proof style was poor but believed it could be improved using the same kinds of methods that make ChatGPT readable. They considered rewriting the proofs through ChatGPT, then published the original versions to preserve transparency.
Q: Why did the model fail to solve IMO problem six?
The researchers described problem six as an exceptionally difficult problem with many possible approaches and a very narrow path to the correct proof. They applied substantial computation but still found no successful solution. Alex Wei said that even months of personal effort, including receiving a major hint about the central idea, would probably not enable him to solve it.
Q: Why was returning no answer on problem six significant?
Returning no answer showed that the model did not automatically fabricate a plausible-looking proof after failing to solve the problem. The team considered the outcome disappointing because the model had used substantial computation, but also encouraging. Earlier models could provide convincing yet incorrect responses, leaving mathematicians to search carefully for subtle errors hidden inside apparently sound reasoning.
Q: Why does the IMO result matter beyond competition mathematics?
The underlying approach used general-purpose techniques for scaling test-time computation and handling hard-to-verify tasks rather than tools designed only for formal mathematical verification. The team’s broader hope is that extending sustained reasoning could eventually help models address difficult unsolved problems in mathematics, science, and other fields, although they distinguished competition success from genuine research breakthroughs.
Summary & Key Takeaways
-
OpenAI’s IMO effort emerged from research ideas developed over roughly six months, followed by a focused final sprint lasting only a couple of months. Although the core team consisted of Alex Wei, Sheryl Hsu, and Noam Brown, their work depended on broader contributions involving inference, scaling, pre-training, and reinforcement learning.
-
The team’s approach emphasized general-purpose reinforcement learning methods that could improve performance on hard-to-verify tasks. Early evidence of progress encouraged continued investment despite skepticism. Instead of relying solely on easily verifiable rewards, the researchers explored techniques intended to handle outputs such as mathematical proofs, where evaluating correctness requires substantial expertise.
-
External former IMO medalists evaluated the model’s original proofs, with three graders assigned to each proof and unanimous agreement on correctness. The team published the raw outputs for transparency despite their poor readability. On problem six, the model returned no answer rather than producing a convincing but unsupported proof after extensive computation.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Sequoia Capital 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator