Aligning Language Models to Follow Instructions and Learning is a Lifelong Process

Kazuki Nakayashiki

Hatched by Kazuki Nakayashiki

Aug 03, 2023

4 min read

0

Aligning Language Models to Follow Instructions and Learning is a Lifelong Process

Learning is a lifelong process, and it is through collective learning that humans have become smarter across generations. No one can learn anything by themselves; even when studying alone, we rely on the knowledge and insights curated by others for teaching purposes. The ability to explain a concept simply is a sign that someone has truly learned it. This is why it is important to constantly expose ourselves to new information and ideas.

One way to enhance our learning process is by reading articles and actively engaging with the content. As we read, we can highlight important passages and take notes. These notes serve as a reference point and can be revisited whenever we feel the need. By doing this consistently, we are constantly feeding our brains with valuable knowledge that will benefit us throughout our lives.

But learning is not just about consuming information; it also involves making connections with our own experiences and ideas. We "borrow" ideas from different sources and use them as building blocks to create our own understanding. This is where tools like Glasp come into play. Glasp is a web highlighter that allows users to leave a digital legacy by making their collected insights available to others. By sharing our notes and highlights, we can contribute to the learning process of others and foster a community of collective growth and knowledge.

However, aligning language models to follow instructions is a challenge that needs to be addressed. Models like InstructGPT and GPT-3 have shown significant differences in their ability to follow instructions. InstructGPT models are much better at following instructions and generate fewer made-up facts compared to GPT-3. This misalignment with users is due to the training process, where GPT-3 is trained to predict the next word based on a large dataset of internet text, rather than focusing on performing specific language tasks as desired by the user.

To make language models safer and more aligned with users, reinforcement learning from human feedback (RLHF) is used. By fine-tuning models on a small curated dataset of human demonstrations, harmful outputs can be reduced. Additionally, human evaluations are conducted to ensure that generated outputs are appropriate and factual. Despite these efforts, there is still progress to be made, as these models can generate toxic, biased, sexual, and violent content without explicit prompting.

One important aspect of aligning language models is considering cultural values and biases. Currently, InstructGPT is biased towards the cultural values of English-speaking people since it is trained to follow instructions in English. To address this, research is being conducted to understand the differences and disagreements between labelers' preferences, allowing models to be conditioned on the values of more specific populations. This ensures that the outputs generated by these models are not biased towards a particular culture or group.

In conclusion, learning is a lifelong process that requires continuous engagement with new ideas and information. By actively reading, highlighting, and taking notes, we can feed our brains with valuable knowledge that will stay with us throughout our lives. Additionally, tools like Glasp allow us to share our insights and contribute to the learning process of others. However, aligning language models to follow instructions poses its own challenges, but through techniques like reinforcement learning from human feedback, progress is being made. Three actionable advice to enhance the learning process and align language models are:

  1. Actively engage with the content you consume by highlighting important passages and taking notes. Revisit these notes regularly to reinforce your learning.
  2. Share your insights and notes with others using tools like Glasp. By contributing to the collective learning process, you can help others grow and gain new knowledge.
  3. Support research and initiatives that aim to align language models with specific populations and cultural values. This ensures that the outputs generated by these models are not biased towards a particular group.

By implementing these actions, we can foster a culture of continuous learning and contribute to the development of safer and more aligned language models.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣
Aligning Language Models to Follow Instructions and Learning is a Lifelong Process | Glasp