How Does OpenAI's Sora 2 Generate Realistic AI Video?

314.2K views
•
November 6, 2025
by
Sequoia Capital
YouTube video player
How Does OpenAI's Sora 2 Generate Realistic AI Video?

TL;DR

Sora 2 uses diffusion transformers that generate an entire video at once from space-time tokens, letting properties like object permanence and physics emerge from scale rather than being programmed. When a simulated action fails, the model defers to physics instead of forcing the requested outcome, a distinction the team calls model failure versus agent failure.

Transcript

for OpenAI across the board, it's really important that we kind of like iteratively deploy technology in a way where we're not just like dropping bombshells on the world when there's like some big research breakthrough. We want to co-evolve society with the technology. And so that's why we really thought it was important to like do this now and lik... Read More

Key Insights

  • Diffusion transformers generate video by adding noise to a signal and training a neural network to predict that noise, then removing it one step at a time, unlike autoregressive models that produce one token at a time conditioned on previous tokens.
  • Space-time tokens (also called space-time patches) are the minimal building block of visual generative models, small cuboids spanning both spatial X and Y dimensions and a temporal dimension, treated almost voxel by voxel.
  • Object permanence emerges because every space-time patch shares information with all others through the attention mechanism, giving the network full global context of everything in the video at every position in space-time.
  • Generating the whole video simultaneously fixes prior systems' problem of quality degrading or changing over time, which is a key reason diffusion transformers proliferated across video models in both the US and China.
  • Sora 2's intelligence is largely a function of scale, similar to how object permanence emerged in Sora 1 only after crossing a critical compute (flops) threshold during pre-training.
  • The team distinguishes model failure from agent failure: if a simulated basketball star misses a free throw, Sora lets the ball rebound off the backboard rather than magically guiding it in to satisfy the prompt, deferring to the laws of physics.
  • OpenAI intentionally released Sora to iteratively deploy technology and co-evolve society with it, framing this as a GPT-3.5 moment for video rather than dropping a research breakthrough on the world all at once.
  • World models emerged internally in language models scaling from GPT-1 to GPT-3 despite very simple tokenizers, suggesting that enough compute and data on next-token prediction produces rich internal representations.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What is a diffusion transformer and how does it differ from autoregressive models?

A diffusion transformer generates data using diffusion rather than autoregressive modeling. Instead of producing tokens one at a time while conditioning on all previous ones, diffusion takes a signal such as video, adds a large amount of noise, and trains neural networks to predict the noise that was applied. Generation happens by gradually removing that noise one step at a time, producing the whole video simultaneously rather than token by token.

Q: What are space-time tokens in Sora?

Space-time tokens, also called space-time patches, are the minimal building block used to construct visual generative models. Each is a small cuboid spanning both spatial dimensions (X and Y) and a temporal dimension, considered almost voxel by voxel. Where characters are a fundamental building block for language, this space-time patch is the fundamental unit for vision, and video is patchified into many of these tokens for the model to process.

Q: How does Sora achieve object permanence?

Object permanence emerges because all the space-time patches communicate with one another through the attention mechanism. This gives the neural network full global context of everything happening in the video at every position in space-time. By representing data as space-time tokens and properly using attention to transfer information throughout the entire video at once, properties like object permanence fall out naturally rather than being explicitly programmed.

Q: Why do diffusion transformers generate the whole video at once?

Generating the entire video simultaneously is a powerful inductive bias for video because it solves problems where quality could degrade or change over time, a major issue for prior video generation systems. Because the model has global context across all space-time positions, it maintains consistency throughout. This advantage is why diffusion transformers have proliferated across video generation stacks, powering most competitor models in both the United States and China.

Q: Is Sora 2 purely a result of scaling up the model?

According to the team, Sora 2's improvements come largely from scale combined with core generative modeling research done since Sora 1. Intelligent agent-like behavior is mostly implicit from scale, similar to how object permanence began to emerge in Sora 1 only after crossing a critical compute threshold measured in flops. As they push the next frontier, agents act more intelligently and the laws of physics are respected in ways they aren't at lower compute scales.

Q: What is the difference between model failure and agent failure in Sora 2?

Model failure means the generation system itself produces something wrong, while agent failure refers to the agent Sora implicitly simulates making a realistic mistake. For example, if the prompt describes a basketball star shooting a free throw and he misses, Sora will not magically guide the ball into the hoop to satisfy the request. Instead it defers to the laws of physics and lets the ball rebound off the backboard, a new semantic failure case not seen in prior video models.

Q: Why did OpenAI decide to release Sora now?

The team says it is important for OpenAI to iteratively deploy technology rather than dropping bombshells on the world when a big research breakthrough happens. They want to co-evolve society alongside the technology. Framing this as a GPT-3.5 moment for video, they wanted to make the world aware of what is now possible and help society get comfortable figuring out the rules of the road for a longer-term vision where copies of yourself run tasks in Sora.

Q: How do internal world models emerge in these AI systems?

World models emerged internally as language models scaled from GPT-1 to GPT-2 to GPT-3. What is notable is that the tokenizers creating the training data are incredibly simple, using basic representations like byte-pair characters. Despite this simplicity, putting enough compute and data into the systems to solve next-token prediction causes rich internal world models to emerge, a pattern the Sora team draws on for video generation with space-time tokens.

Summary & Key Takeaways

  • The Sora team (Bill Peebles, Thomas Dimson, Rohan Sahai) discuss OpenAI's approach of iteratively deploying technology to co-evolve society with it, treating Sora 2 as a GPT-3.5 moment for video and preparing the world for a future of Sora copies running tasks and reporting back.

  • Bill, inventor of the diffusion transformer, explains that video is patchified into space-time tokens and generated all at once by predicting and removing noise, using attention to share information across space-time so object permanence and coherent quality emerge naturally.

  • Sora 2's leap comes largely from scale and core generative modeling research: agents implicit in the video behave more intelligently and respect physics at higher compute, failing in a new semantic way where the model defers to physics rather than forcing the requested outcome.


Read in Other Languages (beta)

Share This Summary 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator

Explore More Summaries from Sequoia Capital 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator