What Is AV 2.0? Wayve's End-to-End Self-Driving Approach

231.3K views
•
November 18, 2025
by
Sequoia Capital
YouTube video player
What Is AV 2.0? Wayve's End-to-End Self-Driving Approach

TL;DR

Autonomous Driving 2.0 replaces the hand-engineered perception, planning, mapping, and control stack with a single end-to-end neural network that learns to drive directly from sensor inputs to motion outputs. Wayve, founded in 2017, uses this approach to generalize across different vehicles, sensor architectures, and cities without HD maps or heavy infrastructure.

Transcript

You know, if you're building a vertically integrated robotic solution, maybe you can go deep. But our ambition is to be the embodied AI foundation model for all of the best fleets and manufacturers around the world. And uh and you know, to do that, you know, unless we want to overload the company by building a separate neural network for each appli... Read More

Key Insights

  • AV 1.0 breaks autonomy into separately hand-engineered components (perception, planning, mapping, control) and relies on infrastructure like high-definition maps, while AV 2.0 replaces that whole stack with one end-to-end neural network that learns to drive from data.
  • Wayve's core architecture is simple at a high level: sensor inputs go in, motion outputs come out, with a large neural network in the middle handling the driving decisions onboard the vehicle rather than depending on external maps.
  • The company's ambition is to be the embodied AI foundation model for the world's best fleets and manufacturers, amortizing cost over one large intelligence that adapts quickly to each application instead of building a separate neural network per use case.
  • Autonomous driving adds constraints absent from large language models: the system must be safe by design, since you cannot simply add more data and hope hallucinations disappear, and it must run in real time within onboard compute and sensor limits.
  • Wayve trains its embodied AI model on diverse sensor permutations so it can understand camera-only, camera-radar, and LiDAR configurations, choosing whichever sensor approach fits, rather than committing to a single fixed hardware setup.
  • The interpretability objection to end-to-end deep learning is now outdated, as strong tools exist to understand how these systems reason; expecting to trace an intelligent machine's outcome to a single line of code is naive given their complexity.
  • Mass-produced cars from leading manufacturers increasingly ship with onboard GPUs, surround cameras, surround radar, and sometimes a front LiDAR, creating software-defined infrastructure that opens these platforms to AI and reached a tipping point in the last couple of years.
  • Hybrid approaches that bolt hard constraints or a rules-based stack onto an end-to-end learned stack often get the worst of both worlds, adding cost and complexity rather than improving safety, according to Kendall.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What is the difference between AV 1.0 and AV 2.0?

AV 1.0 is the classical robotics approach that breaks autonomy into separate components such as perception, planning, mapping, and control, largely hand-engineering each and relying on infrastructure like high-definition maps. AV 2.0, promoted by Wayve since 2017, replaces that entire stack with one end-to-end neural network that gives the robot onboard intelligence to make its own decisions. Wayve describes it as sensor inputs going in and motion outputs coming out through a large neural network in the middle.

Q: How does Wayve's autonomous driving system work?

At a simple level, Wayve's system takes sensor inputs and produces motion outputs, with a gigantic end-to-end neural network in the middle. Rather than depending on hand-engineered interfaces or high-definition maps, the network learns to drive directly from data. It must also be safe by design, meaning the architecture is engineered to be functionally safe and support a robust behavioral safety case, and it must run in real time within the vehicle's onboard compute and sensor limitations.

Q: Why was the end-to-end deep learning approach considered contrarian in 2017?

When Wayve started in 2017, most self-driving stacks were massive hand-coded C++ codebases covering every possible edge case, like navigating around double-parked cars. Industry critics said end-to-end deep learning would never work, arguing it was not safe, not interpretable so you could not understand what it was doing, or simply that they had not heard of the AI approach. Kendall says he could count hundreds of such objections, though he considers the interpretability argument outdated today.

Q: What sensor approach does Wayve use for autonomous driving?

Wayve deliberately avoids committing to a single sensor setup. It wants to build an AI that understands all kinds of sensor architectures, sometimes camera-only and sometimes camera-radar-LiDAR, training its embodied AI model on all those permutations from diverse data sources. The car Kendall drove in during the interview was a camera-only stack, while other partner cars use radar and LiDAR. He notes mass-produced cars increasingly carry onboard GPUs, surround cameras, surround radar, and sometimes a front LiDAR.

Q: Why can Wayve launch in hundreds of cities without building HD maps?

Because Wayve's end-to-end neural network gives the vehicle onboard intelligence to make its own driving decisions, it does not depend on high-definition maps or heavy infrastructure the way AV 1.0 companies do. Those incumbents often need to go out and build an HD map city by city. Wayve's generalization-first approach let it expand within a year from central London to highways in Europe, Japan, and North America, including New York City.

Q: How does Wayve plan to scale its business across manufacturers?

Wayve's ambition is to be the embodied AI foundation model for all the best fleets and manufacturers around the world. Rather than overloading the company by building a separate neural network for each application, it aims to amortize cost over one large intelligence that generalizes and adapts quickly to each customer's application. It sells its stack to auto OEMs, similar to Tesla FSD but for non-Tesla vehicles, with major manufacturers like Nissan choosing Wayve.

Q: Why does Wayve argue against hybrid rules-based and learned approaches?

Kendall notes that even today some people accept end-to-end AI but insist on adding hard constraints or safety guarantees, believing a hybrid approach combining a rules-based stack with an end-to-end learned stack is safest. He argues these approaches often get the worst of both worlds or just add cost and complexity. He sees the market as split between those leaning in and moving fast and those with catching up to do.

Q: What role did large language models play in autonomous driving's shift?

Kendall credits large language model breakthroughs for making deep learning world-changing and mainstream, inspiring the world and opening the market's mind to be curious about the technology. He observes the same narrative playing out in robotics that appeared in language and gameplay agents: an end-to-end data-learned solution out-competes anything that can be hand-coded. This shift helped move AV 2.0 from contrarian to closer to consensus over the last two or three years.

Summary & Key Takeaways

  • Alex Kendall founded Wayve in 2017 on a contrarian bet: replace the hand-coded C++ autonomous driving stacks, which covered every edge case like double-parked cars, with one end-to-end neural network. He wagered on synthetic data and world models as the path to generalization and scaling across physical AI.

  • Wayve sells an autonomous driving stack to auto OEMs, similar to Tesla FSD but for non-Tesla vehicles, with manufacturers like Nissan choosing it. Its aim is to be an embodied AI foundation model that generalizes across many vehicles, sensor architectures, and use cases rather than building a separate network per application.

  • Once contrarian, AV 2.0 has moved toward consensus, inspired partly by large language model breakthroughs. Wayve expanded from driving in complex central London to highways across Europe, Japan, and North America, including New York City, launching in hundreds of cities without needing to build HD maps first.


Read in Other Languages (beta)

Share This Summary 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator

Explore More Summaries from Sequoia Capital 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator