How Does Post-Training Turn Pretrained Models into ChatGPT? Stanford CS224N Spring 2024 Lecture 10 by Archit Sharma

21.6K views
•
March 4, 2025
by
Stanford Online
YouTube video player
How Does Post-Training Turn Pretrained Models into ChatGPT? Stanford CS224N Spring 2024 Lecture 10 by Archit Sharma

TL;DR

Post-training turns pretrained language models into more useful assistants through prompting, instruction fine-tuning, RLHF, and the simpler preference-based DPO approach. The lecture traces this progression after explaining that recent Llama 3 models were trained on roughly 15 trillion tokens and that pretraining runs can cost hundreds of millions of dollars. Read on to understand how each optimization method works and why reward hacking remains a concern.

Transcript

Good evening, people. How are you guys doing? All right. My name is Richard Sherman. I'm a PhD student at Stanford, and I'm very, very excited to talk about posttraining generally speaking for large language models, and I hope you guys are ready to learn some stuff because this has been one of the last few years in machine learning have been very, ... Read More

Key Insights

  • Pretrained language models learn by predicting text tokens and can be enhanced through fine-tuning.
  • Instruction fine-tuning involves training models on a wide range of tasks to better align with user intent.
  • Reinforcement Learning from Human Feedback (RLHF) optimizes models based on human preferences.
  • Direct Preference Optimization (DPO) provides a simpler alternative to RLHF by using preference data directly.
  • DPO models can perform comparably to RLHF models without complex reinforcement learning procedures.
  • Larger models tend to perform better and respond more effectively to instruction fine-tuning.
  • Reward hacking can occur if models exploit learned reward models, underscoring the need for careful optimization.
  • ChatGPT and similar models are the result of advanced techniques like RLHF and DPO, improving user interaction.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How does post-training turn pretrained language models into models like ChatGPT?

The lecture describes a progression through prompting, instruction fine-tuning, and RLHF. These methods move beyond next-token prediction by training or guiding models to follow user intent and better reflect human preferences.

Q: How does instruction fine-tuning improve language models?

Instruction fine-tuning trains models on instruction-output pairs spanning a wide range of tasks. This helps align responses with user intent and improves the model’s ability to perform tasks that were not directly represented during pretraining.

Q: What is Reinforcement Learning from Human Feedback (RLHF)?

RLHF optimizes a language model according to human preferences. It uses preference data to train a reward model, then applies reinforcement learning to adjust the language model toward outputs that receive higher predicted rewards.

Q: What is Direct Preference Optimization (DPO)?

DPO trains a language model directly from preference data by expressing the reward model in terms of the language model itself. It uses a classification loss and avoids the separate, complex reinforcement-learning procedure required by RLHF.

Q: How does DPO compare with RLHF?

DPO is simpler because it directly optimizes from preference data without running the full reinforcement-learning process used in RLHF. The existing lecture summary reports that DPO can perform comparably to RLHF, making it suitable for open-source and production models.

Q: Why is reward hacking a problem in RLHF?

Reward hacking occurs when a model exploits weaknesses in a learned reward model to obtain high rewards without producing the intended behavior. It shows why preference optimization requires careful reward design and constraints.

Q: What does a language model learn during pretraining besides next-token prediction?

Pretraining teaches models syntax, semantics, factual knowledge, languages, mathematics, and code generation. The lecture also presents evidence that next-token prediction may produce representations of agents, beliefs, actions, and human intent.

Q: How much data and compute are used to pretrain recent language models?

The lecture says recent Llama 3 models were trained on roughly 15 trillion tokens, compared with about 1.4 trillion tokens discussed for 2022. It also says current pretraining exceeds 10^26 floating-point operations and that major runs cost hundreds of millions of dollars.

Summary & Key Takeaways

  • Pretrained models predict text tokens and can be enhanced through instruction fine-tuning, which involves training on diverse tasks, aligning models with user intent. This process can be further optimized using RLHF, which directly optimizes for human preferences.

  • Direct Preference Optimization (DPO) offers a simpler, more accessible alternative to RLHF, allowing models to be trained effectively without complex reinforcement learning processes. DPO uses preference data directly, making it suitable for open-source and production models.

  • The transition from pretrained models to advanced models like ChatGPT involves optimizing for human preferences, enhancing model performance and user satisfaction. Techniques like RLHF and DPO are crucial in achieving this, with DPO providing a simpler implementation path.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Stanford Online 📚