Why Can’t LLMs Make Novel Connections? Dwarkesh Patel’s AMA with Sholto Douglas & Trenton Bricken

57.4K views
March 25, 2025
by
Dwarkesh Patel
YouTube video player
Why Can’t LLMs Make Novel Connections? Dwarkesh Patel’s AMA with Sholto Douglas & Trenton Bricken

TL;DR

LLMs may struggle to make novel cross-field connections because pre-training provides broad knowledge without necessarily teaching the skill of research and discovery. In this Ask Me Anything episode, Sholto Douglas argues that significant reinforcement learning on similar tasks may be required, while Trenton Bricken highlights primitive memory scaffolding and the tension between exact recall and generalization. Read on for their reasoning and Dwarkesh Patel’s introduction to The Scaling Era.

Transcript

Today, this is going to be an Ask  Me Anything episode. I'm joined   by my friends Trenton Bricken and Sholto  Douglas. You guys do some AI stuff, right? Yeah. We dabble. They're researchers at Anthropic. Other news; I have a book launching today, it's called The  Scaling Era. I hope one of the questions ends up being why you should buy this book. ... Read More

Key Insights

  • The Scaling Era is a curated book from Stripe Press compiling the most insightful snippets across Dwarkesh Patel's interviews with AI lab CEOs, researchers, economists, and philosophers, sliced by topic across conversations for a digestible read.
  • LLMs memorize all of human knowledge yet rarely make novel connections across fields, unlike humans who occasionally link disparate facts, such as noticing that magnesium deficiency's effect on the brain mirrors migraine structure.
  • The pre-training objective imbues flexible general knowledge about the world but does not necessarily imbue the skill of making novel connections or research, which people acquire through PhD programs and interacting with the world.
  • Significant reinforcement learning on similar tasks is likely the minimum needed for models to approach making novel scientific discoveries, and the field has not yet done this in a meaningful or scaled way.
  • Memory scaffolding for models is very primitive right now; models cannot construct summaries to retain new lessons the way a human learner deliberately would, and most training is just predicting the next word.
  • LLMs may be idiot savants, comparable to Kim Peek who had encyclopedic memory but social debilitations, being amazingly good at niche topics while totally failing at others.
  • Perfect, unforgettable memory can be debilitating, like a transformer context window of trillions of tokens where attending to every past detail prevents extracting generalizable insights.
  • Humans learn best as children yet have total amnesia of childhood, while LLMs sit at the opposite end, capturing exact Wiki text phrasing but failing to generalize in obvious ways.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: Why can’t LLMs make novel connections across different fields?

Sholto Douglas argues that pre-training gives models flexible general knowledge but does not necessarily teach research or the skill of making novel connections. He suggests that significant reinforcement learning on similar tasks may be the minimum required, and says this has not yet been done meaningfully at scale.

Q: Can scaling or thinking models solve the cross-field connection problem?

The discussion does not claim that scale or thinking models alone will solve it. Sholto instead emphasizes reinforcement learning on tasks resembling research and discovery, because possessing knowledge is different from learning how to generate new connections from it.

Q: How do humans and LLMs differ in making discoveries?

Humans have occasionally connected facts from separate domains to produce discoveries, while the speakers know of no comparable example from an LLM. They discuss the example of noticing a similarity between a brain after magnesium deficiency and the structure seen during a migraine.

Q: What role does memory scaffolding play in current AI limitations?

Trenton Bricken describes current memory scaffolding as very primitive. Humans can deliberately summarize a new lesson so it remains useful, whereas most model training centers on predicting the next word and remembering particular facts.

Q: Why does Trenton Bricken compare LLMs to idiot savants?

He compares their uneven abilities to Kim Peek’s combination of encyclopedic memory and serious limitations in other areas. Likewise, an LLM can perform remarkably well on a niche topic while failing completely on another task.

Q: Why might perfect memory make generalization harder?

The speakers argue that remembering everything can trap a learner in details instead of helping it extract useful patterns. They compare this to a transformer with a context window of trillions of tokens that spends its effort attending to the past rather than forming generalizable insights.

Q: What is Dwarkesh Patel’s book The Scaling Era about?

The Scaling Era is a Stripe Press book compiling curated excerpts from Dwarkesh Patel’s interviews with AI lab CEOs, researchers, economists, philosophers, and other scholars. Organized across topics and conversations, it examines intelligence, scaling, AI development, and the economics of billions of additional workers.

Q: Why is The Scaling Era intended to be accessible to ordinary readers?

The book includes an introduction, diagrams, side captions, definitions, context, and commentary explaining concepts such as parameters and models. Its topic-based structure also turns material from many technical interviews into a more digestible, page-by-page collection, including two interviews that had not previously been released publicly.

Summary & Key Takeaways

  • Dwarkesh Patel announces his new book, The Scaling Era: An Oral History of AI 2019-2025, made with Stripe Press. It compiles curated snippets from his interviews with lab CEOs, researchers, economists, and philosophers, addressing questions like the nature of intelligence and the economics of billions of extra workers.

  • A listener question raises why LLMs, despite memorizing all human knowledge, fail to make cross-field connections. Scott Alexander notes humans also lack logical omniscience, but the hosts emphasize humans have demonstrably made such discoveries while no LLM example is known.

  • Sholto argues novel discovery needs significant RL beyond pre-training, since pre-training gives general knowledge but not research skill. Trenton points to primitive memory scaffolding and the idiot-savant analogy of Kim Peek, plus the trade-off between perfect memory and useful generalization.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Dwarkesh Patel 📚