Why ElevenLabs Beat Foundation Models in AI Voice

TL;DR
ElevenLabs stayed competitive against big foundation model labs by staying narrowly focused on audio, both in research and product. Its early breakthrough was building text-to-speech models that understand the context of text and deliver audio with far better tonality and emotion, bringing transformer and diffusion techniques into a domain most researchers had ignored.
Transcript
late 2021 the inspiration came from Peter was about to watch a movie with his with his girlfriend she didn't speak English so they turned it up in Polish and that kind of brought us back to something we grew up with where every po every every movie you watch in Polish um every foreign movie watched in Polish has all the voices so whether it's male ... Read More
Key Insights
- Staying focused on audio is how ElevenLabs out-competed larger labs. As a company, in both research and product, they concentrated narrowly on audio rather than spreading across modalities, which co-founder Mati Staniszewski credits as the usual but genuinely true advice for surviving foundation models.
- Audio was under-researched when ElevenLabs started. Most researchers focused on language models or images because results were easier to see and more exciting, so innovations like diffusion and transformer models had not been applied efficiently to the audio domain.
- The company's core differentiation was true research innovation. For the first time, their text-to-speech models could understand the context of the text and deliver audio with much better tonality and emotion, which set their work apart from earlier approaches.
- Product delivery matters as much as the model itself. Staniszewski stresses it is not only the model that matters but how you deliver that experience to the user, spanning audiobooks, voiceovers, movie translation, and conversational agents.
- The inspiration came from dubbing in Poland. Foreign movies watched in Polish are narrated by a single monotonous voice for every character, male or female, which Staniszewski calls a horrible experience that technology could finally change.
- The founders met roughly 15 years ago in high school in Warsaw. Mati and co-founder Peter took the same IB classes, bonded over a shared love of mathematics, and remain best friends after living, studying, traveling, and working together.
- Early hack-weekend projects preceded ElevenLabs. The pair built a recommendation algorithm that optimized to user selections, a crypto risk analyzer that did not fully work, and in early 2021 an audio project analyzing speech and giving tips on how people speak.
- The open-source Tortoise TTS repo signaled voice cloning was possible. Created around 2022, it replicated a voice and generated speech with incredible though unstable results, giving the founders early glimpses that human-quality voice generation was achievable.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: How did ElevenLabs stay competitive against foundation models?
ElevenLabs stayed competitive by staying focused on audio as a company, in both research and product. Staniszewski says the usual and genuinely true advice is staying focused, and in their case that meant concentrating narrowly on audio rather than spreading across modalities. They built some of the best research models and out-competed big labs, while also building the product layer that delivers the research to users across audiobooks, voiceovers, translation, and agents.
Q: What was the inspiration behind starting ElevenLabs?
The inspiration came in late 2021 when co-founder Peter was about to watch a movie with his girlfriend, who did not speak English, so they turned it up in Polish. In Poland, every foreign movie is narrated by one single monotonous voice for all characters, male or female, which Staniszewski describes as a horrible experience that still happens today. That aha moment convinced them technology could let people enjoy content in its original voice delivery.
Q: Why was audio under-researched when ElevenLabs started?
When ElevenLabs started, there was very little research done in audio because most people focused on language models and some on images, where results were easier to see and frequently more exciting for researchers. As a result, the set of innovations that happened in prior years, including diffusion models and transformer models, had not been applied to the audio domain in an efficient way, leaving an opening that ElevenLabs was able to exploit with its research.
Q: What made ElevenLabs' text-to-speech models different?
For the first time, ElevenLabs' text-to-speech models were able to understand the context of the text and deliver that audio experience with much better tonality and emotion. This true research innovation differentiated their work from other approaches. Staniszewski describes reaching another level of human quality where you could actually feel like it was a human voice, achieved by bringing transformer and diffusion techniques into the audio space and starting from scratch on new innovations.
Q: How did the founders of ElevenLabs meet?
Mati Staniszewski and co-founder Peter met about 15 years ago in high school in Warsaw, Poland. They started an IB class together and took all the same classes. They hit it off quickly over mathematics, which they both love, and began sitting together and spending a lot of time together, which morphed into time outside school. Over the years they did it all, living, studying, working, and traveling together, and remain best friends today.
Q: What early projects did the founders build before ElevenLabs?
While Peter was at Google and Mati was at Palantir, they did hack-weekend projects to explore new technology for fun. They built a recommendation algorithm where selecting one of a few presented options optimized the next set closer to previous selections. They also tried building a risk analyzer for crypto, which did not fully work. Then in early 2021 they created an audio project that analyzed how people speak and gave tips on speaking.
Q: What role did Tortoise TTS play in ElevenLabs' development?
Tortoise TTS was an incredible open-source text-to-speech model created around the time ElevenLabs was exploring what was possible, roughly a year into the company in 2022. It provided incredible results replicating a voice and generating speech, though it was not very stable. It gave the founders glimpses that human-quality voice generation was possible and served as another element confirming their direction before they innovated further from scratch with transformer and diffusion approaches.
Q: Why does building the product matter as much as the model at ElevenLabs?
Staniszewski stresses that it is not only the model that matters but also how you deliver that experience to the user, something they have seen many times. After the first research breakthrough, ElevenLabs fast-followed by building all the product around it so people could actually use the research. That product layer, spanning narrating audiobooks, voiceovers, turning movies into other languages, and building conversational experiences and agents, keeps helping them win across foundation models and hyperscalers.
Summary & Key Takeaways
-
ElevenLabs was expected to become roadkill for foundation models expanding into voice, but survived by staying focused on audio in both research and product. Staniszewski credits co-founder Peter, a genius who made early innovations and assembled a rockstar team, for continually pushing what is possible in audio.
-
When the company started, little research was done in audio because language models and images offered easier, more exciting results. ElevenLabs brought diffusion and transformer models into audio efficiently, producing text-to-speech that understood context and delivered better tonality and emotion, then built product layers on top to reach users.
-
The origin traces to a 15-year high school friendship in Warsaw and hack-weekend projects in recommendations, crypto, and speech analysis. The aha moment came in late 2021 when Peter watched a movie dubbed into Polish with monotonous single-voice narration, convincing the founders that technology could transform how content is voiced.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Sequoia Capital 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator