Rohit Prasad: Solving Far-Field Speech Recognition and Intent Understanding | AI Podcast Clips

December 15, 2019
by
Lex Fridman
YouTube video player
Rohit Prasad: Solving Far-Field Speech Recognition and Intent Understanding | AI Podcast Clips

TL;DR

Amazon made Alexa’s far-field speech recognition useful by combining large-scale data, deep learning, distributed GPU training, algorithmic advances, and engineering improvements. Beginning after Rohit Prasad joined in April 2013, the team tackled wake-word detection and accurate recognition of requests spoken from 20 or even 40 feet away in noisy homes. Read on to see why false wake-ups remain difficult and how the team approached them.

Transcript

the inspiration was the Star Trek computer so when you think of it that way you know everything is possible but when you launch a product you have to start with someplace and when I joined we the product was already in conception and we started working on the far field speech recognition because that was the first thing to solve by that we mean tha... Read More

Key Insights

  • 😯 Far field speech recognition was initially considered an unsolvable problem, but through the combination of deep learning, large-scale data, and engineering advancements, the team at Amazon was able to overcome the challenges.
  • 😯 Accurately detecting the wake word "Alexa" and recognizing speech accurately in a noisy household setting were major hurdles that needed to be addressed.
  • 👻 Deep learning played a crucial role in improving accuracy by allowing the system to learn from vast amounts of data and continuously improve over time.
  • 😯 Setting high standards for accuracy and usability was important in creating a delightful customer experience with speech recognition.
  • 😯 The team had to overcome skepticism and doubters within the company, but their conviction and belief in the potential of far field speech recognition led to the successful launch of Alexa.
  • 🤔 The process of thinking about a product in terms of a press release and FAQs helped the team stay focused and prioritize the right problems to solve.
  • 👥 Feedback from users and continuous learning ensured that the team could further refine and improve the speech recognition capabilities of Alexa.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How did Amazon solve far-field speech recognition for Alexa?

The team combined large-scale data, progress in deep learning, distributed GPU training, algorithmic improvements, and substantial engineering work. This made far-field recognition useful in home settings, although Rohit Prasad emphasizes that speech recognition is still not perfectly solved.

Q: What does far-field speech recognition mean for Alexa?

It means users can speak to the device from a distance rather than through a phone or close-talking microphone. The transcript describes Alexa answering people who are 20 feet away and, in some homes, as far as 40 feet away.

Q: Why was detecting the Alexa wake word difficult?

Alexa must detect its wake word with very high accuracy amid household noise and ordinary conversation. Similar-sounding words and names such as “I like you,” “Alec,” and “Alex” make it difficult to distinguish a command from unrelated speech.

Q: How does Alexa avoid waking when its name is mentioned but no command is intended?

The goal is to detect when “Alexa” is directed at the device rather than merely mentioned in statements such as “I love my Alexa.” The team also developed technology to distinguish human speech from other sources and filter marketing content, but Prasad says false wake-ups remain an unsolved problem.

Q: Could an Alexa device wake up while this podcast repeatedly says “Alexa”?

Yes. Prasad says that if this kind of podcast is playing aloud, the device may wake up a few times despite the filtering technologies Amazon has developed.

Q: What problem came after wake-word detection?

The next challenge was large-vocabulary speech recognition for the many requests users might make. Alexa had to recognize words accurately from about 20 feet away even when music, children, and other household conversations created background noise.

Q: What role did deep learning play in Alexa’s speech recognition?

The team doubled down on deep learning immediately because it needed to improve accuracy quickly and expected a successful device to generate large volumes of learning data. It developed distributed GPU training alongside algorithmic and engineering improvements so models could train on thousands and thousands of hours of speech.

Q: When did Amazon develop and launch this far-field speech-recognition work?

Rohit Prasad joined in April 2013, when promising neural-network research in speech recognition was still at an early stage. He describes the combination of data, deep learning, GPUs available through AWS, and engineering work coming together across 2013 and 2014, when Echo launched.

Summary & Key Takeaways

  • The team started by focusing on far field speech recognition, which allows users to interact with Alexa from a distance, but it was considered an unsolvable problem at the time.

  • They first had to solve the challenge of accurately detecting the wake word "Alexa" in a noisy environment, where other similar words could be mistaken.

  • Another major challenge was recognizing various requests accurately in a large vocabulary speech recognition problem, especially in a busy household setting with background noise.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Lex Fridman 📚