C5W3L10: How Do You Build a Trigger Word Detection System?

9.0K views
•
February 5, 2018
by
DeepLearningAI
YouTube video player
C5W3L10: How Do You Build a Trigger Word Detection System?

TL;DR

A trigger word detection system can be built by extracting audio features such as spectrogram features, passing them through an RNN, and training it to output 1 immediately after the trigger word ends. Because this labeling produces many more zeros than ones, several consecutive positive labels can make training easier. Read on for the labeling process, practical examples, and the method’s main tradeoff.

Transcript

you've now learned so much about deep learning and sequence models that we can actually describe a trigger word system quite simply just on one slide as you see in this video but when the rise of speech recognition have been more and more devices you can wake up with your voice and those are sometimes called trigger word detection systems so let's ... Read More

Key Insights

  • 🛝 Deep learning and sequence models have simplified trigger word detection systems to just a single slide explanation.
  • 🔑 Trigger word systems have become increasingly popular with devices like Amazon Echo, Apple Siri, and Google Home.
  • 👋 The development of trigger word detection algorithms is still ongoing, with no widely accepted best algorithm.
  • ⌛ An imbalanced training set can be addressed by outputting multiple ones over a fixed period of time after the trigger word is spoken.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How do you build a trigger word detection system?

Compute audio features, such as spectrogram features, from an audio clip and pass the resulting sequence through an RNN. Define target labels so the RNN outputs 1 immediately after the trigger word finishes and 0 at other times.

Q: What is a trigger word detection system?

A trigger word detection system listens for a particular word or phrase that wakes a device or initiates an action. For example, saying “Alexa” can wake an Amazon Echo so it can respond to a voice query.

Q: What are examples of trigger words and devices that use them?

The examples given are Amazon Echo with “Alexa,” Apple Siri with “Hey Siri,” and Google Home with “Okay Google.” Each system wakes when it detects its designated phrase.

Q: What audio features can be used for trigger word detection?

The approach can compute spectrogram features from an audio clip. These sequential audio features, represented as x1, x2, x3, and so on, are then passed through an RNN.

Q: How should target labels be assigned for trigger word detection?

Set the target labels to 0 before and during the trigger word. Immediately after the speaker finishes saying it, set the target label to 1; repeat this pattern whenever the trigger word occurs again.

Q: Why does trigger word labeling create an imbalanced training set?

Using a single positive label after each trigger word creates far more 0 labels than 1 labels. This imbalance is a disadvantage of the basic labeling scheme, even though the approach can still work reasonably well.

Q: How can the imbalance between positive and negative labels be reduced?

Instead of assigning 1 at only one time step, assign several consecutive 1 labels for a fixed period immediately after the trigger word. The labels then return to 0, slightly evening out the ratio of ones to zeros and potentially making the model easier to train.

Q: Is there a widely accepted best algorithm for trigger word detection?

No. The trigger word detection literature is still evolving, so there is not yet wide consensus on the best algorithm; the RNN method described here is one example.

Summary & Key Takeaways

  • Deep learning and sequence models have made it possible to create trigger word detection systems that can wake up devices using specific words.

  • Examples of trigger word systems include Amazon Echo (Alexa), Apple Siri (Hey Siri), and Google Home (Okay Google).

  • The literature on trigger word detection algorithms is still evolving, but one example involves using audio features and RNNs to identify the trigger word and set target labels.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from DeepLearningAI 📚