The Fundamental Attribution Error and the Scaling of Language Models: Understanding Cognitive Bias and Breakthrough Performance

Glasp

Hatched by Glasp

Sep 03, 2023

4 min read

0

The Fundamental Attribution Error and the Scaling of Language Models: Understanding Cognitive Bias and Breakthrough Performance

In the field of psychology, the fundamental attribution error, also known as correspondence bias or over-attribution effect, has been a subject of interest. This phenomenon refers to the tendency for people to overemphasize dispositional explanations for others' behavior while underemphasizing situational factors. In other words, we have a cognitive bias to attribute someone's actions to their personality rather than the social and environmental influences that may be at play.

One possible reason behind this bias is perceptual salience. When we observe others, we tend to focus on their visible characteristics, which are often more apparent than the situational factors that may have influenced their behavior. Additionally, our lack of detailed information about what causes their actions further strengthens our inclination to attribute their behavior to their inherent qualities.

Now, let's shift our focus to the fascinating advancements in language models, particularly the Pathways Language Model (PaLM). Recent developments in language models, such as GLaM, LaMDA, Gopher, and Megatron-Turing NLG, have pushed the boundaries of performance by scaling model size, incorporating sparsely activated modules, and training on larger and more diverse datasets.

PaLM represents a significant leap in scale compared to its predecessors. It is a single model that can generalize across domains and tasks while maintaining high efficiency. The training process of PaLM involved a combination of English and multilingual datasets from various sources, including web documents, books, Wikipedia, conversations, and even GitHub code.

One notable achievement of PaLM is its impressive training efficiency, achieving a hardware FLOPs utilization of 57.8%. This is the highest efficiency recorded for language models of this scale. PaLM's remarkable performance extends to a range of tasks, including arithmetic problem-solving and commonsense reasoning.

For instance, when prompted with grade-school math problems, PaLM combined with chain-of-thought prompting demonstrated strong performance. By breaking down the problem into intermediate steps, similar to how a person would approach it, PaLM outperformed previous models in solving challenging grade-school level math questions. In fact, with just 8-shot prompting, PaLM was able to solve 58% of the problems in the benchmark, surpassing the prior top score of 55%.

What makes PaLM even more impressive is its scalability. The Pathways system can effectively utilize thousands of accelerator chips across two TPU v4 Pods to train a 540-billion parameter model. This breakthrough in model scale enables PaLM to achieve exceptional few-shot performance across various natural language processing, reasoning, and code-related tasks.

So, what can we learn from these two seemingly unrelated topics? While they may appear distinct, there are common threads that tie them together. Both the fundamental attribution error and the scaling of language models highlight the importance of understanding context and considering multiple factors when making judgments or evaluations.

In psychology, recognizing the influence of situational factors is crucial for overcoming the fundamental attribution error. By acknowledging that someone's behavior is not solely determined by their personality, we can gain a more accurate understanding of their actions.

Similarly, in the field of language models, the scaling of models like PaLM emphasizes the significance of incorporating diverse datasets and considering various sources of information. By training on a wide range of data, these models can achieve breakthrough performance and generalize across different domains and tasks.

In light of these insights, here are three actionable pieces of advice:

  1. Challenge your assumptions: When observing someone's behavior, question your initial inclination to attribute it solely to their personality. Consider the situational factors that may have influenced their actions and strive for a more comprehensive understanding.

  2. Embrace diversity in data: Whether you're working with language models or conducting research, make an effort to incorporate diverse datasets that encompass different perspectives and sources of information. This will enhance the model's performance and ensure a more comprehensive understanding of the subject matter.

  3. Context matters: Remember that context plays a crucial role in both psychology and language models. Be mindful of the broader circumstances surrounding a situation or problem, as they can significantly impact the outcomes and interpretations.

In conclusion, the fundamental attribution error and the scaling of language models provide intriguing insights into cognitive biases and breakthrough performance. By recognizing the influence of situational factors and embracing diverse datasets, we can develop a more nuanced understanding of human behavior and enhance the capabilities of language models. So let's strive for a more comprehensive and context-aware approach in our endeavors.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣