Unlocking Real World Value through Human Feedback and Public Learning

Kazuki Nakayashiki

Hatched by Kazuki Nakayashiki

Jul 30, 2023

4 min read

0

Unlocking Real World Value through Human Feedback and Public Learning

In the world of artificial intelligence, language models have become increasingly powerful and prevalent. However, there are significant challenges associated with training these models, particularly when it comes to their accuracy and potential for harmful applications. This has led to a collaborative effort between Humanloop and Stability AI to build the first open-source InstructGPT, a language model that can be trained using reinforcement learning from human feedback (RHLF).

LLMs trained solely on next word prediction often produce factually inaccurate or offensive output. By incorporating RHLF techniques, models can become considerably more aligned and easier to use. This approach has already been successfully employed by industry leaders like OpenAI, DeepMind, and Anthropic to create LLMs that follow instructions or act as helpful assistants.

The potential of RLHF-tuned models extends beyond academia, hobbyists, and industry. We envision a future where these models are applied and adapted to every domain and task, unlocking significant real-world value. To achieve this, partnerships are crucial. Carper AI is collaborating with Humanloop and Scale to collect and apply human feedback data, which will be used to improve the underlying language model that Carper trains. Humanloop specializes in adapting LLMs based on human feedback, while Scale is a leader in data annotation. The final trained model will be hosted by Hugging Face, ensuring its accessibility to a wide range of users.

While the focus of this collaboration is on improving language models, there is a broader concept at play here - learning in public. Traditionally, learning has been seen as an individual pursuit, with knowledge hoarded and shared only through formal processes like workshops or case-based events. However, this approach may not be sufficient to meet the demands of a rapidly evolving world.

Imagine a workplace where most learning is in public, where knowledge flows freely and is accessible to all. This "socialized knowledge management system" has the potential to revolutionize organizations and make them more effective. Instead of relying solely on sporadic workshops or events, organizations can tap into the collective intelligence of their employees by creating a culture of public learning.

One framework that can enhance knowledge flows in organizations is Personal Knowledge Management (PKM). PKM focuses on the needs and desires of individuals, making each person's knowledge flow public through the process of Seek-Sense-Share. By openly seeking new information, making sense of it through reflection and sense-making, and sharing their insights with others, individuals contribute to the collective knowledge of the organization.

Learning in public is not always easy. It requires transparency and vulnerability, as well as a willingness to receive feedback and support. However, the benefits far outweigh the challenges. When we learn in public, our work becomes transparent, allowing others to build upon our ideas and contribute their own perspectives. This collective effort can lead to significant improvements and innovations in our increasingly complex workplaces.

Transparency is a key element in creating new management frameworks for a networked world. By learning in public, we not only share and develop knowledge, but also foster a culture of trust and collaboration. It allows us to collectively develop critical next practices that can drive success in our organizations.

To harness the power of public learning, here are three actionable pieces of advice:

  1. Foster a culture of transparency: Encourage employees to share their knowledge and insights openly. Create platforms or spaces where individuals can easily share their work, ideas, and reflections with others.

  2. Embrace feedback and collaboration: Encourage a feedback-rich environment where employees feel comfortable giving and receiving feedback. Encourage collaboration and the sharing of diverse perspectives to foster innovation.

  3. Provide resources and support for learning: Invest in training programs, resources, and tools that enable individuals to enhance their learning and share their knowledge effectively. Encourage continuous learning and provide opportunities for growth and development.

In conclusion, the collaboration between Humanloop and Stability AI to build an open-source InstructGPT highlights the importance of human feedback in training language models. By incorporating reinforcement learning from human feedback, these models can become more accurate, aligned, and useful. Additionally, the concept of learning in public has the potential to transform organizations by fostering transparency, collaboration, and collective knowledge development. By embracing public learning and implementing actionable advice, organizations can unlock real-world value and drive success in an increasingly complex and interconnected world.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣