When the Test Becomes the Product

Mark Erdmann

Hatched by Mark Erdmann

Jul 24, 2026

10 min read

88%

0

The Strange New Incentive Problem

What happens when the thing you are trying to measure becomes the thing people learn to optimize, instead of the underlying reality?

That question sits at the center of two seemingly different stories. One is about modern benchmarks, where models can look brilliant by being trained on the answer sheet, then re sampled until they land on the right response. The other is about management practices, where a simple intervention in manufacturing plants created real gains that still lingered a decade later. At first glance, one story is about cheating and the other about genuine improvement. But together they reveal something deeper and more unsettling: measurement systems do not just observe performance, they shape the behavior that produces it.

The difference between a contaminated benchmark and a persistent management reform is not just technical. It is philosophical. In one case, the metric was gamed so thoroughly that it stopped being a measure. In the other, the intervention became a durable part of how work was done, which means the metric and the practice stayed aligned long enough for true capability to compound.

The real question is not whether something works once. The real question is whether it changes the system that produces future performance.

That is the hidden thread between AI evaluation and organizational improvement. Both are about feedback loops. Both can be fooled. Both can create lasting effects. The challenge is learning which kind of effect you are seeing.


Why Short Term Wins Can Be Lies, and Long Term Effects Can Be the Truth

In many domains, we reward the easiest evidence to collect. A model passes a benchmark. A team hits quarterly targets. A plant raises output after consultants arrive. These are all useful signals, but they are not equally trustworthy.

The danger is that a short term bump can come from at least three very different mechanisms:

  1. Real capability change: the underlying system got better.
  2. Temporary performance boost: people behaved differently because they knew they were being watched.
  3. Metric exploitation: the system learned how to satisfy the test without improving the thing we care about.

Only the first mechanism is genuinely productive. The second may be valuable if it nudges behavior into a better routine. The third is pure illusion, and it often thrives in environments obsessed with leaderboard rank, quarterly optics, or management theater.

This is why benchmark contamination matters so much. A benchmark is supposed to estimate general capability, but once the answer leaks into training, the benchmark stops being an independent check. It becomes a mirror. The score rises, but the meaning of the score collapses. The model has not necessarily become smarter in the way we hoped. It has become more familiar with the exam.

The same principle applies to organizations, just more slowly. A plant may improve after a management intervention not because the plant memorized a test, but because the intervention changed daily habits: how supervisors track bottlenecks, how workers escalate problems, how meetings are run, how accountability is made visible. Ten years later, if half the effect still remains, that is powerful evidence that the change was not cosmetic. It altered the operating system.

That is the crucial distinction: tests can be gamed, systems can be upgraded.


The Benchmark Trap: When Selection Becomes Substitution

One way to understand contaminated benchmarks is to think about studying for an exam. If you genuinely learn the material, the test is a proxy for understanding. But if you memorize the answer key, the proxy breaks. You can get a perfect score while remaining unable to solve the real problem in the wild.

This is what happens when selection turns into substitution. The benchmark is no longer selecting for capability. It is selecting for proximity to the test distribution.

This logic shows up everywhere:

  • A sales team optimizes for leads that are easy to close, not customers who remain loyal.
  • A hospital optimizes for discharge speed, not long term recovery.
  • A startup optimizes for vanity metrics, not product market fit.
  • A model optimizes for benchmark accuracy, not robust reasoning.

In each case, the measured outcome starts replacing the underlying goal. Once that happens, the organization may look better while becoming less capable.

The deeper problem is that humans love visible scores. Scores simplify complexity. They create a feeling of objectivity. But a score is only as good as the process that keeps it connected to reality. If the connection weakens, the score becomes a performance of competence rather than competence itself.

This is why benchmark contamination is not merely an evaluation problem. It is a warning about any world where incentives get too close to the metric. The closer the metric is to the reward, the more likely people are to optimize the metric directly. The more they do that, the less the metric tells you.

A metric that can be perfectly optimized too easily is often not a metric of the thing you wanted, but a loophole in disguise.

That is the paradox. The more precise the test seems, the more vulnerable it may be to substitution. Precision can create the illusion of truth while hiding a broken relationship between signal and reality.


Management Practices as Anti Contamination

Now consider the management example. Consulting teams introduced basic managerial practices into manufacturing plants, with others left as a control group. The initial gains were strong. Ten years later, about half of those effects still remained.

Why does that matter so much?

Because it suggests the intervention was not just a cosmetic nudge or a one off burst of attention. It changed behavior in ways that endured. That persistence is the telltale sign of a real causal mechanism, not a statistical mirage.

A useful way to think about this is to distinguish between surface performance and deep capability.

Surface performance is what a system does when the spotlight is on it. Deep capability is what the system can reliably do because its routines, incentives, and habits have changed.

A temporary management initiative can raise surface performance. Better meetings, clearer task assignment, improved inventory tracking, and basic accountability can generate a quick jump in output. But if those changes are merely superficial, they fade as soon as the external pressure disappears.

Persistent effects imply something else happened. The intervention likely created new habits, new expectations, and perhaps even new norms among workers and supervisors. It became embedded. The plant did not just perform better for a quarter. It learned a different way of operating.

This is the mirror image of benchmark contamination. In contaminated benchmarks, the test gets absorbed into training, but the absorption is illegitimate because it bypasses the real objective. In durable management change, the intervention gets absorbed into the system, and the absorption is legitimate because it truly improves the underlying process.

The difference is whether the system is becoming more capable, or just more test savvy.


A Better Mental Model: Coupling, Not Just Measurement

The most useful framework here is to stop thinking about evaluation as a simple check and start thinking about coupling.

Every measurement system has a degree of coupling to the reality it is supposed to reflect. If the coupling is strong, changes in the score correspond to changes in the underlying capability. If the coupling is weak, the score can drift away from reality. If the coupling becomes adversarial, the score can be actively manipulated.

You can picture this as three states:

  • Coupled: the metric tracks the goal closely.
  • Loose: the metric is an imperfect but useful proxy.
  • Broken: the metric can be maximized without improving the goal.

Benchmarks often begin in the coupled state, then drift into the loose state, and eventually become broken once they are widely studied, scraped, and optimized against. Management practices can do the opposite. They begin as external inputs, then become internal routines, strengthening coupling between daily work and long term performance.

This is why the best interventions are not the ones that merely increase output. They are the ones that rewire the process that produces output.

Think of a factory. If a consultant shows up and orders workers to move faster for one week, output may rise. But if the consultant introduces visual dashboards, daily problem solving, clearer handoffs, and a habit of surfacing bottlenecks early, then the plant gains a new nervous system. The improvement persists because the organization has changed how it perceives and responds.

Think of AI evaluation. If a model gets a better score because it has seen the benchmark, that score is informationally weak. But if training or fine tuning genuinely improves general reasoning, then the benchmark score is a byproduct of a deeper shift.

The question, then, is not whether a system can score well. The question is whether the scoring process is still coupled to reality, and whether the intervention improved the machinery behind the score.


The Hardest Part of Progress Is Preserving Meaning

There is a subtle irony in progress. The more valuable a benchmark or metric becomes, the more attention it attracts. The more attention it attracts, the more likely it is to be optimized against. And the more it is optimized against, the less meaningful it may become.

That creates a permanent race between measurement and manipulation.

In AI, the response cannot simply be “make the benchmark harder,” because every benchmark can eventually be studied, duplicated, or indirectly learned. Instead, systems need rotating tests, hidden sets, real world tasks, and evaluation methods that reward transfer rather than memorization.

In organizations, the response cannot simply be “measure more.” Too many metrics create bureaucracy and encourage box ticking. The real task is to design feedback loops that remain connected to meaningful outcomes. That means measuring not just output, but process quality, adaptation, and persistence.

A company that wants genuine improvement should ask:

  • Did the change survive after the consultants left?
  • Do frontline workers still use the new routine when no one is watching?
  • Can the system handle new problems, or only repeat old ones?

An AI lab should ask:

  • Does the model perform well on fresh, unseen tasks?
  • Does performance hold under slight distribution shifts?
  • Is the model learning principles, or pattern matching familiar test items?

In both cases, the answer you want is not just a higher number. It is a stronger relationship between the number and the world.


Key Takeaways

  1. Do not confuse score improvement with real improvement. Ask whether the underlying capability changed or whether the system learned to satisfy the test.

  2. Treat persistence as evidence. If an intervention still matters months or years later, it is more likely to have changed deep routines rather than just surface behavior.

  3. Design metrics that resist substitution. The best measures are hard to game because they stay tied to outcomes that cannot be easily rehearsed or memorized.

  4. Look for changes in operating systems, not just outputs. In factories, teams, or AI models, the most durable gains come from better processes, not just better results on a single scoreboard.

  5. Assume every metric attracts optimization. The moment a metric becomes important, it risks being gamed. Build evaluation systems with that in mind.


The Real Lesson: Improvement Must Outlive the Test

The deepest connection between benchmark contamination and persistent management reform is this: both force us to ask whether progress is real enough to survive contact with time.

A contaminated benchmark gives you the wrong kind of confidence. It tells you the system has improved when in fact the system may have only learned the exam. A durable management change gives you the right kind of confidence. It suggests that performance has been embedded in routines, habits, and structure, so that the gains remain even after attention fades.

That is the standard worth aiming for in any serious domain. Not just can you win the test, but can you still perform when the test is gone?

The highest form of improvement is not making the metric go up. It is making the world itself better in a way the metric can only imperfectly describe.

Once you see this, evaluation looks different. A benchmark is no longer the goal, only a stress test. A management practice is no longer just a policy, but a way of changing the system’s memory. And real progress becomes something rarer and more valuable than a good score: a change that keeps paying dividends after the spotlight moves on.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣