Aligning Language Models to Follow Instructions: The Brief History of Artificial Intelligence and What Might Be Next

Kazuki Nakayashiki

Hatched by Kazuki Nakayashiki

Sep 08, 2023

4 min read

0

Aligning Language Models to Follow Instructions: The Brief History of Artificial Intelligence and What Might Be Next

Artificial intelligence (AI) has come a long way in a relatively short period of time. Just a decade ago, machines were unable to match human capabilities in language processing or image recognition. However, the rapid advancements in AI systems have led to significant breakthroughs, with these systems now surpassing human performance in various domains.

One of the key factors driving the capabilities of AI systems is training computation. Training computation refers to the number of floating point operations (FLOP) required to train the system. Each FLOP involves a basic arithmetic operation, such as addition, subtraction, multiplication, or division. As AI systems heavily rely on machine learning, training computation, along with algorithms and input data, plays a crucial role in enhancing their capabilities.

For the past six decades, training computation has consistently increased in line with Moore's Law, which states that computing power doubles approximately every 20 months. However, since 2010, this exponential growth has accelerated even further, with training computation now doubling every six months. This rapid growth has paved the way for groundbreaking advancements in AI.

Looking ahead, there is a strong possibility that "transformative AI" will be developed by 2040, according to Cotra's estimation. This term refers to artificial general intelligence (AGI) that can perform any intellectual task a human can do. In fact, many AI experts believe that human-level AGI could be achieved within the next few decades, while some even argue that it may be realized much sooner.

Despite the remarkable progress in AI, there are still significant challenges to overcome. One of the key issues is aligning language models with user instructions. The InstructGPT models, which are specifically designed to follow instructions, have demonstrated superior performance compared to the more general GPT-3 models. InstructGPT models exhibit better adherence to instructions, generate fewer fabricated facts, and show decreased toxic output generation.

To enhance the safety, usefulness, and alignment of AI models, reinforcement learning from human feedback (RLHF) is employed. By fine-tuning models on a small curated dataset of human demonstrations, harmful outputs can be reduced. Additionally, human evaluations on API prompt distribution have shown that InstructGPT models produce more appropriate outputs and hallucinate facts less frequently. However, despite these advancements, there is still work to be done as InstructGPT models occasionally generate toxic, biased, sexual, or violent content without explicit prompting.

Solving these challenges requires AI models to be able to refuse certain instructions, thereby preventing the generation of unsafe outputs. This, however, poses an important open research problem. Furthermore, the current InstructGPT models are biased towards the cultural values of English-speaking individuals, highlighting the need to understand and address the differences and disagreements among labelers' preferences. Research is underway to condition models on the values of more specific populations, promoting inclusivity and reducing bias.

As we reflect on the brief history of AI and its rapid evolution, it is clear that the world has changed dramatically. AI systems now possess capabilities that were unimaginable just a few years ago. However, we stand at the cusp of an even more transformative era. The development of human-level AGI holds immense potential but also raises significant ethical and societal concerns.

In conclusion, here are three actionable pieces of advice to navigate the future of AI:

  1. Foster alignment and safety: Continuously strive to align AI models with user instructions and promote safe outputs. Invest in reinforcement learning from human feedback to reduce harmful outputs and improve adherence to instructions.

  2. Address biases and inclusivity: Recognize and address biases in AI systems, particularly in language models. Conduct research to understand and incorporate the values and preferences of diverse populations, ensuring inclusivity and reducing cultural biases.

  3. Ethical considerations: As AI continues to progress, it is crucial to prioritize ethical considerations. Engage in open discussions and collaborations with experts, policymakers, and stakeholders to establish guidelines and regulations that govern the development and deployment of AI systems.

By embracing these actions, we can shape a future where AI aligns with human values, fosters inclusivity, and acts as a powerful tool for positive change.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣