"What 5 years at Reddit taught us about building for a highly opinionated user base" & "Aligning Language Models to Follow Instructions"

Glasp

Hatched by Glasp

Aug 04, 2023

3 min read

0

"What 5 years at Reddit taught us about building for a highly opinionated user base" & "Aligning Language Models to Follow Instructions"

Building and managing a product for a highly opinionated user base can be both challenging and rewarding. This is something that Reddit, a popular online community, has learned over the course of five years. One of the key lessons they have discovered is that you can't ask passionate users not to love your product. Instead of trying to destroy their passion, it is important to fill the Trust Vault and harness their passion productively.

However, it's crucial to remember that just because someone is loud doesn't mean you should act on their complaints. Identifying whom you should pay attention to is essential. The loudest people may not represent your general user base or your ideal customer profile. It's important to analyze the feedback and determine whether it represents a significant portion (10% or more) of your user base and whether the users giving feedback can influence the opinions of others.

Another interesting insight is that customers typically won't go out of their way to provide feedback on useful features. They treat these features as table stakes, expecting them to be there, and therefore, not commenting on them. This can make it difficult to gauge the attitudes of most users based solely on negative feedback.

On the other hand, when it comes to language models, aligning them to follow instructions is crucial. Reddit's experience with their user base sheds light on the importance of understanding the values and preferences of your users. In the case of language models, models like InstructGPT have proven to be significantly better at following instructions compared to models like GPT-3.

GPT-3, although trained on a large dataset of internet text, is not aligned with the specific language tasks that users want to perform. This misalignment can lead to outputs that are less helpful and potentially generate harmful or toxic content. To address this, techniques like reinforcement learning from human feedback (RLHF) have been implemented. By fine-tuning models like InstructGPT and incorporating curated datasets of human demonstrations, harmful outputs can be reduced.

However, it's important to note that even with these improvements, InstructGPT models are still far from fully aligned or fully safe. They can still generate biased, toxic, and inappropriate content without explicit prompting. Solving this issue requires models to refuse certain instructions, which is a challenging problem that requires further research.

In terms of cultural biases, InstructGPT is currently biased towards the cultural values of English-speaking people. To address this, research is being conducted to understand the differences and disagreements between labelers' preferences. By conditioning models on the values of more specific populations, biases can be minimized.

In conclusion, building and managing products for highly opinionated user bases requires a deep understanding of user feedback and preferences. It's important to identify the key individuals to pay attention to and consider feedback that represents a significant portion of your user base. Additionally, aligning language models to follow instructions is crucial for ensuring their safety and effectiveness. By incorporating techniques like reinforcement learning from human feedback and addressing cultural biases, models can become more helpful, aligned, and safe.

Actionable advice:

  1. Prioritize feedback from users who represent your ideal customer profile and make up a significant portion of your user base.
  2. When seeking feedback, only ask for it if you genuinely intend to consider and act upon it. Failing to do so can lead to a loss of trust.
  3. Continuously work towards aligning language models with user instructions and reducing biases. This can be achieved through techniques like reinforcement learning and understanding specific population values.

By implementing these actions, you can foster a stronger relationship with your passionate user base and ensure that your products and language models are better aligned with their needs and expectations.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣