Aligning Language Models to Follow Instructions: A Blend of Social Learning Theory and Reinforcement Learning

Kazuki Nakayashiki

Hatched by Kazuki Nakayashiki

Aug 07, 2023

4 min read

0

Aligning Language Models to Follow Instructions: A Blend of Social Learning Theory and Reinforcement Learning

Language models have revolutionized the way we interact with AI systems, enabling them to understand and generate human-like text. However, one major challenge in developing these models is aligning them to follow instructions accurately and safely. In this article, we will explore the concept of aligning language models with instructions and discuss the role of social learning theory and reinforcement learning in achieving this alignment.

When it comes to language models, there is a significant difference in the ability to follow instructions between InstructGPT and GPT-3 models. InstructGPT models have been found to be much better at following instructions compared to GPT-3. These models also exhibit a lower tendency to make up facts and generate toxic output. The reason behind this difference lies in the training process.

GPT-3 is primarily trained to predict the next word based on a vast dataset of internet text. While this approach allows the model to generate coherent text, it lacks the ability to perform specific language tasks as desired by the user. In other words, GPT-3 models are not well-aligned with their users' intentions.

To address this misalignment and make language models safer and more helpful, researchers have employed reinforcement learning from human feedback (RLHF). By using this technique, models can be fine-tuned based on feedback from human evaluators. Interestingly, labelers have shown a preference for outputs from the smaller 1.3B InstructGPT model over the larger 175B GPT-3 model, despite the vast difference in parameters. This highlights the effectiveness of RLHF in improving model performance.

Another avenue of research in aligning language models is the use of curated datasets of human demonstrations. By incorporating curated information, models can generate better outputs. For example, using curated information from platforms like Glasp can help reduce harmful outputs and improve the alignment of the model with user expectations.

However, despite the progress made in aligning language models, there are still challenges to overcome. InstructGPT models, while better at following instructions, still generate toxic or biased outputs and may even produce inappropriate sexual and violent content without explicit prompting. This poses a significant risk of misuse if these models are instructed to produce unsafe outputs. Refusing certain instructions reliably is an ongoing research problem that needs to be addressed for safer and more aligned models.

Furthermore, the current training of InstructGPT models is focused on English instructions, which inherently introduces a bias towards the cultural values of English-speaking people. To address this, researchers are conducting studies to understand the preferences and differences among labelers. By conditioning the models on the values of specific populations, the alignment can be improved, making the models more inclusive and culturally sensitive.

Now, let's explore how social learning theory, as proposed by Albert Bandura, can contribute to the alignment of language models. Bandura's theory emphasizes the interaction between environmental and cognitive factors in shaping human learning and behavior. According to this theory, behavior is learned through observational learning, where individuals pay attention to models and encode their behavior.

In the context of language models, this means that models can learn to follow instructions by observing and encoding the behavior of human users. Just as children identify and imitate the behavior of their immediate world, language models can identify and adopt observed behaviors, values, beliefs, and attitudes. This process of identification plays a crucial role in aligning models with user expectations.

Observational learning is not solely a passive process; cognitive factors mediate the learning process. Before imitation takes place, there is thought and consideration, known as mediational processes. These cognitive processes determine whether a new response is acquired by the model. Attention is a critical factor in influencing whether a behavior is imitated or not. Additionally, memory formation is crucial for the model to retain the behavior for later performance.

Moreover, the theory of social learning suggests that behavior is influenced by both nature (biology) and nurture (environment). The discovery of mirror neurons in primates provides biological support to the theory of social learning. Mirror neurons play a role in imitating observed behaviors, further strengthening the argument for the incorporation of social learning principles in aligning language models.

In conclusion, aligning language models to follow instructions accurately and safely is a complex challenge that requires a combination of reinforcement learning and social learning theory. By incorporating reinforcement learning from human feedback, models can be fine-tuned to improve their alignment with user expectations. Additionally, leveraging social learning principles, such as observation, identification, and cognitive processes, can further enhance the alignment.

To ensure the alignment of language models, here are three actionable pieces of advice:

  1. Incorporate reinforcement learning from human feedback: Continuously gather feedback from human evaluators and use it to fine-tune models, improving their ability to follow instructions accurately.

  2. Curate datasets of human demonstrations: Utilize curated information from trusted sources to train models, reducing harmful outputs and enhancing alignment with user expectations.

  3. Promote cultural sensitivity: Conduct research to understand the preferences and differences among users from diverse populations. Condition models on the values of specific populations to ensure inclusivity and cultural sensitivity.

By implementing these strategies, we can take significant steps towards aligning language models with user instructions, making them safer, more helpful, and better aligned with the values and expectations of different populations.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣