Will AI Reach Einstein-Level Reasoning in 9 Years?

TL;DR
Extrapolating current trends, OpenAI's Dan Roberts estimates AI could make Einstein-level scientific discoveries in about nine years. The length of tasks AI agents can complete is doubling every seven months, and reaching the roughly eight years of reasoning it took Einstein to derive general relativity would require about 16 more doubling times.
Transcript
Dan Roberts uh is a former Sequoia team member and has been spreading the word about reasoning for many years to us the past two two and a half three years. We sat across from each other for about a year year and a half and I've learned so much from from Dan. I'm I'm really excited for you to share more broadly. I'll share one memory which is last ... Read More
Key Insights
- Test-time compute represents a new scaling dimension beyond training. OpenAI's o1 model, released last September, improved not only with train-time compute but also the more time it spent thinking at test time, letting it reason, iterate, and zoom in on problems.
- Reasoning models can now reproduce advanced textbook physics in about a minute. OpenAI's o3 solved a quantum electrodynamics problem involving a Feynman diagram in roughly one minute, a calculation that took Dan Roberts about three hours to manually verify against four textbooks.
- Reinforcement learning compute is the contrarian bet OpenAI is scaling up. GPT-4o used only pre-training compute, o1 added some RL compute, o3 a bit more, and the goal is a future where training becomes totally dominated by reinforcement learning compute.
- The classic cake metaphor is being deliberately inverted. Yann LeCun's 2019 slide framed pre-training as a big cake with RL as a small cherry on top; OpenAI wants the same-size cake crushed by a giant reinforcement learning cherry.
- Scaling science makes model performance predictable in advance. For GPT-4, small-scale experiments plotted on a log scale let OpenAI predict the final loss of a model bigger than anything seen before, and the prediction was nailed exactly.
- Current models feel like idiot savants that cannot yet discover general relativity. Possible reasons include asking the wrong sorts of questions and training on too many competition math problems, making models jaggedly good at different things.
- OpenAI's stated plan is simply scaling compute at massive scale. The company aims to raise $500 billion, buy land in Abilene, Texas, construct buildings full of computers, generate revenue, and repeat the buildout to train ever-larger models.
- The length of tasks AI agents can perform is doubling every seven months. Agents can currently handle roughly hour-long tasks, and extrapolating the exponential line suggests two to three hour tasks within a year.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: When could AI reach Einstein-level scientific discovery?
Dan Roberts estimates roughly nine years. His reasoning: it took Einstein about eight years of thinking to discover general relativity, and the length of tasks AI agents can complete is doubling every seven months. Reaching that eight-year reasoning capability requires about 16 doubling times, which extrapolates to around nine years from now. He cautions that predictions in AI are dangerous and everyone is always wrong, but presents this as an extrapolation of the current exponential trend line.
Q: What is test-time compute in AI reasoning models?
Test-time compute is a scaling dimension where a model improves the more time it spends thinking at inference, rather than only during training. When OpenAI released o1 last September, one plot showed performance improving with train-time compute, familiar to anyone who trains models, but the exciting plot showed performance also improving with test-time compute. The model was taught to reason, and the longer it thought, the better it performed. Roberts called this a totally new dimension for scaling, important enough to put on a t-shirt.
Q: How did OpenAI's o3 model handle a physics calculation?
OpenAI's o3 was given a quantum electrodynamics problem posed on a sheet of paper that included a Feynman diagram, a way of representing these calculations. The model could see the paper, then thought about it, iterated, and zoomed in before answering correctly in about a minute. Roberts, who was asked to check the calculation before the blog post went up, said verifying it took him about three hours, even though the calculation appears in four textbooks he had to track down to confirm every minus sign was correct.
Q: Why is OpenAI prioritizing reinforcement learning over pre-training?
Roberts frames this as a contrarian bet on scaling reinforcement learning compute. GPT-4o, from a year ago, used only pre-training compute. Then o1 added some RL compute, and o3 a little more. He projects a future with a lot of RL compute, and eventually one totally dominated by it. He illustrates this with Yann LeCun's 2019 cake metaphor, where pre-training is a big cake and RL is a small cherry, and says OpenAI wants to invert it: the same-size cake crushed by a giant reinforcement learning cherry.
Q: What is OpenAI's plan for scaling compute?
Although Roberts joked that the comms team redacted his plan slide, he said the plan is stated clearly: scaling compute. Concretely, OpenAI intends to raise $500 billion, buy land in Abilene, Texas, construct buildings, and fill them with computers. The goal is to train models, generate substantial revenue from them, then build more buildings with more computers and repeat. In concert with this physical buildout, OpenAI is also developing scaling science to understand how to apply that compute effectively.
Q: What is scaling science and why does it matter?
Scaling science is the discipline of predicting how model performance changes as compute increases, and it is one of Roberts's focuses at OpenAI. He cites the GPT-4 blog post, which predates his time: small-scale experiments plotted on a log scale produced a dotted-line prediction of GPT-4's final loss, and the team nailed it for a model bigger than anything previously built. With new directions like test-time compute and RL training, he says they must reinvent what it means to scale up compute predictably.
Q: Why do current AI models still feel limited despite their reasoning ability?
Roberts references a point from podcaster Dwarkesh Patel that today's models feel like idiot savants, capable of impressive calculations but not discovering general relativity. He offers possible explanations: researchers may be asking the wrong sorts of questions, since how you ask a question often matters more than the process and answer. Alternatively, models may be trained on too many competition math problems, making them jaggedly good at different things. In either case, he suggests the fix is scaling up further.
Q: How fast are AI agents' task-completion abilities growing?
According to a plot Roberts referenced, there is exponential growth in the length of tasks that AI agents can perform, doubling every seven months. At present, agents can handle tasks of about an hour. Extrapolating forward, he suggested that within a year they might manage tasks between two and three hours long. He extends this same doubling trend to argue that reaching the roughly eight years of reasoning Einstein needed for general relativity would take about 16 doubling times, or roughly nine years.
Summary & Key Takeaways
-
OpenAI's o1, released last September, introduced test-time compute: the model improved the more time it spent thinking, a new scaling dimension beyond ordinary train-time compute. The newer o3 model can iterate, zoom in, and reason through problems at test time to reach correct answers.
-
To demonstrate reasoning depth, o3 solved a quantum electrodynamics problem with a Feynman diagram in about a minute, which took Roberts three hours to verify. A GPT-4.5-authored general relativity exam question stumped GPT-4.5 but was answered correctly by o3.
-
Roberts argues the future is dominated by reinforcement learning compute, inverting Yann LeCun's cake-and-cherry metaphor. OpenAI plans to raise $500 billion, build compute in Abilene, Texas, and develop scaling science. With task length doubling every seven months, Einstein-level AI discovery could arrive in roughly nine years.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Sequoia Capital 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator