How Can Reinforcement Learning Unlock AI Agents?

TL;DR
Reliable AI agents require greater depth, not just broader language-model capabilities. Misha Laskin argues that combining reinforcement learning and search with LLMs could produce models that plan, act, and solve difficult tasks more reliably, following the path demonstrated by DQN, AlphaGo, AlphaZero, and MuZero while addressing the substantial gap between current agent performance and truly capable autonomous systems.
Transcript
I think someone needs to solve the kind of depth problem the field as a whole I think is or large Labs have been have been really working on on on the breadth that's amazing and there there's a big market for that and um a lot of very useful things that get unlocked but um someone needs to solve the depth problem too hi everyone welcome to training... Read More
Key Insights
- AI research should target foundational problems that can be unlocked in the present era. Laskin left professional physics after realizing that physics had been a defining foundational problem roughly 100 years earlier, while intelligent agents appeared to offer a more timely opportunity for transformative research.
- Mastery can develop through sustained curiosity and growing competence rather than existing passion. Laskin says his parents initially pursued chemistry PhDs partly because Israel offered stipends to Russian immigrants, then came to love their craft as they studied deeply and became increasingly accomplished in it.
- AlphaGo was a decisive demonstration of creative superhuman agency for Laskin. Its famous move 37 appeared poor to Lee Sedol and observers, yet about 10 moves later it proved optimal for establishing a winning position, showing behavior that did not look like simple brute force.
- DQN was the first successful agent of the deep learning era according to Laskin. It learned to act reliably in Atari video-game environments from raw sensory inputs, providing an early proof that deep neural systems could autonomously acquire useful behavior through interaction with an environment.
- The AlphaGo research line showed that reinforcement-learning agents could scale to remarkable depth. Laskin points to AlphaGo, AlphaZero, and MuZero as evidence that relatively small models can become exceptionally capable within a specific domain when learning, search, and action are tightly integrated.
- Current LLM development emphasizes breadth across many useful tasks. Laskin recognizes that this breadth unlocks valuable applications and supports a large market, but argues that another major research challenge is depth, meaning models that can reason, plan, and execute reliably on demanding tasks.
- Current coding agents remain far from dependable general agency despite measurable progress. The description reports high-teens percentage scores on SWE-bench, compared with a previous unassisted baseline of 2 percent and an assisted baseline of 5 percent, indicating both improvement and substantial unresolved difficulty.
- Reflection AI’s strategy is to combine reinforcement-learning search with large language models. Laskin and Ioannis Antonoglou draw on work spanning AlphaGo, AlphaZero, MuZero, and Gemini to pursue reliable models for developers who are building workflows in which agents must plan tasks and execute actions.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: Why are current LLMs still insufficient for reliable AI agents?
Current LLMs provide impressive breadth and unlock many useful applications, but Laskin argues that dependable agents also require depth. Agents must plan tasks, choose actions, interact with environments, and execute reliably rather than merely demonstrate broad digital intelligence. The description’s coding example illustrates the gap: high-teens SWE-bench performance exceeds earlier baselines but still leaves most benchmark issues unresolved.
Q: How could reinforcement learning improve LLM-based agents?
Reinforcement learning can teach a system to select actions through interaction with an environment, while search can help it explore possible decisions before acting. Reflection AI’s stated approach is to bring those capabilities together with LLMs. The goal is to retain broad language-model intelligence while developing the deeper planning, action selection, and reliability required for developers to build effective agentic workflows.
Q: What did AlphaGo teach Misha Laskin about AI agents?
AlphaGo showed Laskin that a computer could outperform humans while producing solutions that appeared creative rather than merely mechanical. He highlights move 37 against Lee Sedol, which initially looked like an error but about 10 moves later created a winning position. That experience convinced him that solving agency was a profound foundational problem and motivated his move into artificial intelligence.
Q: Why was AlphaGo move 37 considered evidence of creative behavior?
Move 37 looked like a bad decision when AlphaGo played it, and both Lee Sedol and observers were perplexed. Its value became apparent about 10 moves later, when it proved to be the optimal move for putting AlphaGo in a winning position. For Laskin, this delayed and surprising effectiveness suggested that the system had found a solution people had not previously considered.
Q: What role did DQN play in the development of deep reinforcement learning?
Laskin describes the deep Q-network, or DQN, as the first successful agent of the deep learning era. The system learned to play Atari video games and acted reliably from raw sensory input. It served as a proof point that neural systems could autonomously learn behavior inside an environment, helping catalyze research in deep reinforcement learning across video games and robotics.
Q: What is the difference between breadth and depth in AI models?
Breadth refers to the wide range of useful capabilities that large language models can offer across many tasks. Depth refers to becoming exceptionally capable, reliable, and strategic within demanding tasks that require planning and action. Laskin says large labs have made remarkable progress on breadth, but believes someone must also solve the depth problem to unlock stronger agents.
Q: How capable are current coding agents on SWE-bench?
The description states that coding agents score in the high-teens percentage range on SWE-bench, a benchmark involving the resolution of GitHub issues. That result is well above the earlier unassisted baseline of 2 percent and assisted baseline of 5 percent. However, the high-teens result also shows that current systems remain far from resolving most benchmark tasks reliably.
Q: Why did Misha Laskin move from physics into artificial intelligence?
Laskin was drawn to physics because he wanted to understand how the world works at a foundational level, and he eventually earned a PhD in the field. He later concluded that researchers should pursue foundational problems suited to their own time. Deep learning’s rise, especially AlphaGo’s demonstration of creative superhuman agency, convinced him that building intelligent agents was such a problem.
Summary & Key Takeaways
-
Misha Laskin traces his interest in foundational problems from an immigrant childhood shaped by his parents’ persistence, an isolated upbringing in Washington state, and Richard Feynman’s physics lectures. After earning a physics PhD, he concluded that researchers should pursue foundational problems that are especially ready to be unlocked in their own time.
-
AlphaGo persuaded Laskin to enter AI because it demonstrated a superhuman agent capable of apparently creative behavior. The famous move 37 initially looked mistaken but later created a winning position. That episode suggested the system was doing more than brute-force computation, motivating his ambition to build agents from the beginning of his AI career.
-
Reflection AI aims to improve agent models by combining the search capabilities of reinforcement learning with LLMs. Laskin distinguishes the breadth delivered by large language models from the depth needed for reliable task execution. Current coding-agent benchmark gains are meaningful, but their high-teens SWE-bench scores still leave considerable room for progress toward dependable autonomous workflows.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Sequoia Capital 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator