Stanford XCS224U: NLU I Fantastic Language Models and How to Build Them, Part 1 I Spring 2023

August 17, 2023
by
Stanford Online
YouTube video player
Stanford XCS224U: NLU I Fantastic Language Models and How to Build Them, Part 1 I Spring 2023

TL;DR

ELECTRA improves Transformer pretraining by teaching a discriminator to identify original and replaced tokens in corrupted input sequences. Unlike BERT, which masks about 15% of tokens and learns only from those selected positions, ELECTRA’s discriminator never sees MASK tokens and becomes the pretrained model used for downstream fine-tuning. Read on to understand its generator-discriminator design, contrastive objective, efficiency gains, and key differences from BERT.

Transcript

all right welcome everyone welcome back let's get started we have another action-packed day for you time's a wasting to start here I'm going to finish up our big uh slide deck on contextual word representations there are just a few more small things to cover and then Sid is going to get us help us get Hands-On with training really big models so the... Read More

Key Insights

  • 💡 The Transformer architecture combines ideas from RNNs and CNNs to create a powerful language model.
  • 🤕 Self-attention and multi-head attention mechanisms enable the learning of contextual representations for each token.
  • 💁 The addition of an MLP layer introduces non-linearity and helps the model forget irrelevant information.
  • 👻 The Transformer architecture addresses the limitations of RNNs and CNNs, allowing for parallel processing and capturing token relationships.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How does ELECTRA work?

ELECTRA first masks some input tokens and uses a typically small masked-language-model generator to sample replacements, creating a corrupted sequence. A discriminator then learns which tokens are original and which were replaced. Training combines the generator’s BERT-style masked language modeling loss with a weighted ELECTRA loss.

Q: What problem with BERT does ELECTRA address?

ELECTRA addresses the mismatch caused by BERT using MASK tokens during pretraining even though those tokens are absent during fine-tuning. It also targets BERT’s data inefficiency: BERT masks or replaces about 15% of tokens, and only those tokens contribute to its learning objective.

Q: What is the role of the generator in ELECTRA?

The generator is typically a small model with a masked language modeling objective and can be a BERT model. It produces candidate tokens for masked positions, which are sampled to replace some original tokens and form a corrupted input sequence.

Q: What does the ELECTRA discriminator learn?

The discriminator determines which tokens in a corrupted sequence are originals and which are replacements. Its task is described as a contrastive learning objective that asks which words belong in the sequence and which do not.

Q: Which part of ELECTRA is used for downstream fine-tuning?

After pretraining, the generator can be discarded. The discriminator becomes the pretrained ELECTRA artifact used for downstream fine-tuning, much as a downloaded BERT model serves as a pretrained artifact.

Q: How is ELECTRA’s objective different from BERT’s objective?

BERT learns to recover what is missing from surrounding context through masked language modeling. ELECTRA’s discriminator instead learns to detect which words in a sequence were replaced and which are original.

Q: Why does the ELECTRA discriminator avoid the MASK-token mismatch?

The discriminator never receives MASK tokens as its input. It sees corrupted sequences containing sampled replacement tokens, so the model retained after pretraining was not trained to depend on the special MASK token.

Q: How do generator and discriminator size and parameter sharing affect ELECTRA?

When the generator and discriminator are the same size, they can share all their Transformer parameters, and the lecture reports that more sharing is better. However, the best results described in the lecture come from using a generator that is small relative to the discriminator.

Summary & Key Takeaways

  • The Transformer architecture emerged as a combination of ideas from recurrent neural networks (RNNs) and convolutional neural networks (CNNs).

  • Self-attention and multi-head attention mechanisms were introduced to enable each token to serve as its own query, key, and value, allowing for the learning of contextual representations.

  • To incorporate non-linearity, an MLP layer was added to the end of the Transformer block, providing a way to forget irrelevant information and crystallize the structure of features.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Stanford Online 📚