A Test Score Is Never Just a Test Score

Mark Erdmann

Hatched by Mark Erdmann

Aug 31, 2026

10 min read

92%

0

What if the thing you are measuring is partly created by the conditions under which you measure it?

A language model can appear brilliant because it has encountered the questions before. A human can appear less capable because the room is too cold. In both cases, the score seems to describe an underlying ability, but it also records an interaction between ability and circumstance.

This creates a problem for anyone who relies on tests, benchmarks, interviews, dashboards, or performance reviews: performance is not a fixed property that simply reveals itself. It is an event produced by a mind operating inside an environment.

That distinction sounds philosophical, but it has practical consequences. It changes how we should evaluate artificial intelligence, design workplaces, interpret research, and make decisions about people.

The hidden variable inside every score

Imagine two runners competing in a race. One runs on a dry track, the other on wet pavement. If the first runner wins, we might conclude that she is faster. That may be true. But the race result also reflects traction, footwear, wind, and the timing of the race.

Most evaluation systems quietly pretend these variables do not matter. They treat the result as a clean signal of competence. A benchmark score becomes intelligence. An interview becomes judgment. A productivity number becomes effort. A test grade becomes learning.

Yet every measurement has at least two components:

  1. The capability of the thing being measured
  2. The conditions that allow or obstruct that capability from appearing

A useful conceptual model is:

Observed performance equals underlying capability multiplied by environmental fit, plus noise.

The formula is not meant to be precise mathematics. It is a reminder about structure. If the environment is poorly matched to the task, a strong performer may look mediocre. If the environment is unusually favorable, a mediocre performer may look exceptional.

This is why the design of a test matters as much as the score it produces. A benchmark that uses old questions may reward memorization rather than reasoning. A workplace that maintains a temperature uncomfortable for much of its staff may measure physical tolerance along with verbal performance. A test taken under severe time pressure may measure speed, anxiety management, or familiarity with the format in addition to the knowledge supposedly being assessed.

The central question is not merely, “Who scored higher?” It is:

What combination of ability, context, and exposure produced this score?

Fresh questions and comfortable rooms solve the same problem

At first, artificial intelligence benchmarks and office thermostats seem to belong to different intellectual worlds. One concerns machine reasoning. The other concerns human comfort and demographic differences. But they expose the same weakness: a measurement can become contaminated by factors that are not the target of measurement.

Consider a benchmark built from publicly available questions. Over time, those questions may enter training data, evaluation discussions, or online solution repositories. A model can then produce a correct answer through recognition or memorization rather than genuine problem solving. The score rises, but the meaning of the score changes.

This is not necessarily deliberate cheating. It is a natural consequence of repeated measurement. Once a test becomes visible, the test itself becomes part of the environment. Systems adapt to it. People teach to it. Models absorb it. The instrument no longer observes capability in a neutral way because it has begun to shape the behavior it measures.

Now consider a room that is kept unusually cold. The room is not intended to test verbal reasoning. Yet the physical environment can affect concentration, comfort, and cognitive performance. If those effects differ across groups or tasks, the final score contains information about the thermostat as well as the test taker.

In both situations, the evaluator may believe it is measuring one thing while actually measuring a bundle of things:

  • reasoning ability plus exposure to prior questions
  • verbal skill plus thermal comfort
  • knowledge plus test familiarity
  • creativity plus psychological safety
  • productivity plus interruption load

The benchmark contamination problem is obvious once stated: an old question may no longer be a new test of reasoning. The environmental problem is subtler: a physical condition may be invisible precisely because everyone in the room has become accustomed to treating it as normal.

Normal conditions are not necessarily neutral conditions.

Capability is a surface, not a point

A better mental model treats ability as a performance surface. Instead of assigning someone or something a single score, imagine a landscape showing performance across different conditions.

For a language model, the surface might vary according to:

  • whether the task is novel or familiar
  • whether the prompt is precise or ambiguous
  • whether tools are available
  • whether the task requires recall, reasoning, planning, or communication
  • whether the answer is judged by exact match or by a human evaluator

For a human worker, it might vary according to:

  • temperature and lighting
  • time of day
  • interruptions
  • sleep and stress
  • social pressure
  • task type
  • the degree of autonomy available

A single score samples only one point on this surface. The danger comes when we mistake that point for the entire landscape.

This model explains why two evaluators can disagree while both appear to have good evidence. One may test performance under familiar conditions. The other may test performance under novel conditions. One may reward speed. The other may reward accuracy. One may create comfort. The other may introduce friction. Their results are not necessarily contradictory. They may be observing different slices of the same capability surface.

The implication is especially important for comparisons. Suppose Model A outperforms Model B on a public benchmark. If the questions have circulated widely, the result may favor the model with more relevant exposure. Suppose a male and female group perform differently on a verbal test in a particular room. The result may reflect a difference in how the environment affects performance, rather than a broad difference in verbal ability.

A responsible comparison therefore requires more than ranking outcomes. It requires understanding the transfer conditions: will the advantage survive when the questions change, the temperature changes, the task changes, or the evaluation format changes?

A useful test of an evaluation is this:

If the conditions change slightly, does the ranking remain stable?

If not, the evaluation may still be useful, but it is measuring a narrower and more conditional property than its headline score suggests.

The danger of optimizing for the instrument

Once a metric becomes important, subjects begin adapting to it. This is familiar in schools and businesses. If employees are judged by the number of tickets closed, they may close tickets quickly rather than solve difficult problems. If students are judged by standardized questions, teaching may shift toward test patterns. If a model is trained on a public benchmark, its apparent improvement may reflect increased familiarity with the benchmark rather than a general increase in reasoning.

The metric becomes a target, and the target becomes part of the system.

This creates a feedback loop:

  1. An evaluator chooses a metric.
  2. People or models adapt to the metric.
  3. The metric becomes easier to optimize directly.
  4. The evaluator mistakes optimization for broad capability.
  5. The system is redesigned around a distorted picture of performance.

A fresh question set helps interrupt this loop because it restores novelty. It does not solve every problem. New questions can contain their own biases, may be poorly calibrated, and can still reward narrow forms of competence. But freshness protects one crucial property: the test remains at least partially independent of the training process.

Environmental variation plays a similar role. If an organization wants to know whether a policy improves productivity, it should not measure performance under only one narrow set of conditions. It should ask whether the improvement persists across times, tasks, teams, and settings. Otherwise, the organization may simply be optimizing for one favorable measurement environment.

This suggests a distinction between score maximization and capability discovery.

Score maximization asks: “How can we achieve a higher result on this test?”

Capability discovery asks: “What can this system do reliably when the test itself changes?”

The first is useful when the test is the task. The second is essential when the test is supposed to stand in for something larger.

An evaluation protocol for the real world

If performance depends on context, should we abandon benchmarks and tests? No. We should make them more honest.

The first improvement is to separate measurement conditions from performance claims. Instead of saying, “This system is better,” say, “This system performs better on novel verbal tasks under these timing and tool conditions.” Precision about the claim is not bureaucratic caution. It is protection against overgeneralization.

The second improvement is to test for robustness. A strong evaluation should vary the conditions that are not supposed to matter and observe whether the result survives. For an AI system, this might mean using newly written questions, alternate phrasings, unfamiliar domains, and tasks that require explanation rather than answer selection. For a workplace study, it might mean measuring performance at several temperatures, times, and task types rather than treating one office setting as universal.

The third improvement is to measure the environment itself. Temperature, noise, interruptions, prompt quality, available tools, and prior exposure should not be treated as background scenery. They are explanatory variables.

The fourth is to report distributions, not just averages. An average can hide a large interaction. If one group improves substantially under warmer conditions while another changes little, the overall average may conceal the practical importance of the environment. Similarly, a model average score may hide excellent performance on novel problems and weak performance on familiar ones, or the reverse.

Finally, evaluators should ask whether the evaluation resembles the world in which the result will be used. A system designed to assist with unfamiliar problems should be tested on unfamiliar problems. A workplace designed for sustained collaboration should measure more than isolated individual output. Ecological validity matters: the closer the test conditions are to real use, the more meaningful the inference.

Key Takeaways

  • Treat every score as conditional. Ask which abilities, environmental factors, and forms of familiarity contributed to the result.
  • Protect novelty when measuring reasoning. Rotate questions, use fresh tasks, and test alternate formulations to reduce contamination by prior exposure.
  • Measure the environment, not just the subject. Temperature, noise, interruptions, timing, and social conditions can alter observed performance.
  • Test robustness across conditions. If a ranking disappears when the task or setting changes, describe the capability narrowly rather than making a general claim.
  • Optimize for transfer, not test performance alone. The most valuable result is one that survives contact with unfamiliar questions and realistic working conditions.

The deeper lesson is not that tests are unreliable. It is that tests are relationships. They connect a subject with an instrument, a task, and a setting. Change any part of that relationship and the result may change too.

A benchmark is not a window through which ability passively shines. It is more like a lens. Its shape determines what becomes visible, what is distorted, and what remains outside the frame. A thermostat is not merely a comfort control. It can become an invisible component of an experiment. A public question set is not merely a neutral challenge. It can become training data.

Once we understand this, better evaluation becomes less about finding the one perfect test and more about mapping the conditions under which performance is stable. We should stop asking whether a person or model is simply smart, capable, or productive. We should ask a harder and more useful question:

Under which conditions does this capability reliably appear, and who or what is responsible for creating those conditions?

That question shifts evaluation from judgment to diagnosis. It also shifts responsibility. When performance falls, the problem may be the performer. But it may also be the room, the prompt, the timing, the benchmark, or the assumptions built into the instrument.

The fairest measurement is not the one that pretends conditions do not matter. It is the one that makes them visible.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣