The Best System Metric May Be a Sign of Failure

Thomas Hirschmann

Hatched by Thomas Hirschmann

Aug 30, 2026

10 min read

92%

0

What if a falling employment rate could make a community healthier, while a rising productivity score could make its users miserable?

These are not merely statistical curiosities. They expose a central problem in how we evaluate modern systems: the thing we can measure most easily is often not the thing we care about most. Economic output can rise while air becomes harder to breathe. A product can achieve excellent engagement numbers while solving the wrong problem. A service can satisfy its designers and frustrate the people who depend on it.

The deeper connection is this: both public policy and technology fail when they treat indirect signals as substitutes for human understanding. The remedy is not to abandon metrics. It is to place metrics inside a disciplined process of listening, interpretation, and revision.

A metric tells you what changed. It does not necessarily tell you what happened to people.

The strange case of harmful activity producing a good outcome

During the Great Recession, some American counties experienced sharp job losses. Ordinarily, this would be read as an unmistakable deterioration in social welfare. Unemployment means lost income, disrupted routines, anxiety, and reduced opportunity. Yet mortality declined in those same places, and the improvement was substantial.

The explanation was not that unemployment is secretly beneficial. It was that economic contraction reduced pollution. With fewer factories operating, fewer trucks moving goods, and less fuel being consumed, concentrations of fine particulate matter fell. Cleaner air accounted for more than a third of the decline in mortality.

This is a powerful example of causal entanglement. One event can contain both damage and benefit. Losing a job can harm a household while reducing emissions across a region. A recession can weaken a local economy while temporarily relieving the lungs of people who live near industrial activity.

The mistake would be to look at the mortality data and conclude that unemployment improves health. That conclusion confuses a side effect with a solution. It is equivalent to praising a power outage because electricity consumption fell, or praising a broken delivery network because traffic congestion disappeared.

The meaningful lesson is more uncomfortable: our prosperity may be generating health costs that ordinary economic indicators conceal. When economic activity slows, the hidden cost becomes visible through cleaner air. The recession acts like an accidental experiment, revealing a relationship that normal growth makes difficult to see.

This should change the questions we ask. Instead of asking only whether the economy is expanding, we should ask what kinds of activity produce that expansion, who benefits, who bears the costs, and which harms are being exported into the environment or into private lives.

The same discipline is necessary when designing artificial intelligence systems, digital products, and public services. A numerical improvement can be real and still be misleading. More clicks may mean greater usefulness, or greater confusion. Faster completion may mean efficiency, or that users are being pushed through a process without understanding it. Fewer support requests may mean a clearer product, or that people have given up trying to get help.

The measurement trap in technology

Before anyone designs a useful system, they must understand the people who will use it. This sounds obvious, but it is routinely violated. Teams often begin with a solution, a technical capability, or a business objective, then search for users whose behavior can be made to fit the plan.

That reversal creates a predictable risk: misunderstanding the problem while optimizing the answer.

Imagine a hospital introducing an automated appointment system. The team defines success as reducing average booking time. After deployment, the average falls from ten minutes to four. On paper, the project is a success. But field visits reveal that older patients are abandoning the process because the system assumes they know the difference between a referral appointment and a diagnostic appointment. The people who complete the process are faster. The people who need assistance have disappeared from the metric.

Or consider an educational application that celebrates daily engagement. Its designers add reminders, streaks, and rewards. Usage rises. Yet diaries and experience sampling reveal that students open the application mainly to preserve their streak, not to learn. The system has increased activity while weakening attention.

These are not unusual failures. They arise because designers confuse observable behavior with underlying need. The user clicks, but the click may represent interest, confusion, habit, fear, or an attempt to escape an interface. The worker stays late, but that may indicate commitment, or an unusable process. A customer submits fewer complaints, but that may reflect satisfaction, or resignation.

The environmental example and the design example share the same structure. In both cases, the headline metric is incomplete because it does not represent the full system.

Economic output leaves pollution partly invisible. Product engagement leaves frustration partly invisible. Employment statistics leave household stress partly invisible. Mortality statistics leave the distribution of health gains partly invisible. The central challenge is not measurement itself. It is measurement without interpretation.

Listen before you optimize

Human centered research offers a practical answer to this problem. Interviews, focus groups, field visits, diaries, surveys, and experience sampling are not decorative steps added to a design process. They are methods for discovering the variables that the initial model omitted.

Each method reveals a different layer of reality.

A nondirected interview can expose the language people use to describe a problem before designers impose their own categories. A field visit can reveal workarounds that users no longer mention because they seem ordinary. A diary can show how an experience changes over time rather than capturing a single polished account. Experience sampling can record what people feel at the moment a problem occurs, instead of relying on a retrospective explanation shaped by memory.

Together, these methods perform a kind of reality check on abstraction. They ask whether the system's formal description corresponds to the lived experience of the people inside it.

This matters far beyond software. A city planner may define transport success as average travel speed. Residents may define it as arriving safely, knowing when a bus will come, and not having to choose between reliability and affordability. A manager may define workplace flexibility as the ability to work remotely. Employees may experience it as an expectation to remain permanently available. A health service may define access as the number of appointments offered. Patients may experience access as the ability to understand what kind of appointment they need.

The formal metric is not necessarily false. It is simply partial. The danger begins when partial truth is treated as complete truth.

The purpose of research is not to collect opinions after a system has been designed. It is to discover what the system must be designed to notice.

This principle also helps explain why unintended benefits and unintended harms are so difficult to manage. A system's designers typically observe the direct output they intended to create. They may not observe the secondary effects that occur elsewhere, later, or among people who are absent from the design conversation.

Pollution is a delayed and distributed consequence of production. Stress is often a delayed and private consequence of unemployment. Exclusion is frequently a silent consequence of automation. None appears clearly in the first dashboard.

A three layer model for better decisions

A useful way to connect these ideas is to evaluate every major decision across three layers: output, experience, and externality.

1. Output: What did the system produce?

This is the layer most organizations already measure. It includes revenue, speed, completion rates, employment, energy use, accuracy, and mortality. Output metrics matter because they tell us whether something changed.

But outputs are only the beginning. They are the visible surface of a much larger process.

2. Experience: What did the change feel like for the people involved?

Experience includes comprehension, effort, dignity, confidence, anxiety, autonomy, and trust. It is often difficult to quantify, which makes organizations tempted to ignore it. Yet difficulty of measurement is not evidence of unimportance.

A system that completes a task quickly but leaves users confused has not necessarily succeeded. A labor market that creates jobs but makes workers physically ill has not achieved unqualified progress. Experience asks whether the change improved people's ability to live and act, not merely whether it altered a number.

3. Externality: What costs or benefits appeared outside the stated goal?

Externalities are effects borne by people, places, or future periods that were not included in the original objective. Air pollution is an externality of production. Cognitive overload can be an externality of poorly designed software. Increased administrative burden can be an externality of a supposedly efficient policy. The time spent correcting an automated mistake is an externality of automation that counts as zero in the system's official completion rate.

This third layer is where the most surprising discoveries often occur. A recession can lower mortality by reducing pollution. A new interface can improve average speed by excluding people who struggle with it. A productivity tool can raise output by transferring invisible work from an organization to its employees.

The three layer model prevents a common analytical error: treating one layer as a proxy for all three. It also creates a practical sequence for investigation.

First, measure the output. Second, ask people how the change affected their experience. Third, look for costs and benefits that have moved outside the boundary of the original metric.

This sequence is especially important when a result appears counterintuitive. If a harmful event is associated with a positive outcome, do not celebrate or dismiss the result immediately. Investigate the mechanism. What changed in the environment? Which groups experienced the benefit? Which groups suffered? Can the beneficial pathway be preserved without preserving the harmful event?

That last question is the mark of mature reasoning. The goal is not to recreate the recession in order to obtain cleaner air. It is to identify the pollution producing activities and redesign them so that health does not depend on economic collapse.

From accidental experiments to deliberate learning

Unexpected outcomes should be treated as invitations to learn, not anomalies to be suppressed. A community that becomes healthier during an economic downturn has learned something about the cost of its normal activity. A product whose usage rises while satisfaction falls has learned something about the difference between engagement and value.

Organizations need mechanisms for capturing these lessons. One practical approach is to create a countermetric review whenever a primary metric changes significantly. Alongside the question, “Did we achieve the target?” ask:

  • Who was missing from the measurement?
  • What behavior did the metric reward?
  • What might people have done to adapt to the system?
  • Which costs were shifted to households, workers, communities, or the environment?
  • What evidence would show that the apparent success was actually failure in disguise?

These questions can be embedded in ordinary research practices. Before launch, conduct field visits and nondirected interviews to identify hidden needs. During use, gather short experience reports at meaningful moments, not merely at the end of a process. After launch, compare behavioral data with qualitative accounts and environmental or health indicators.

The aim is not to replace quantitative analysis with anecdotes. A single interview cannot establish that pollution caused a change in mortality, just as a dashboard cannot explain why users abandon a service. Strong decisions come from triangulation, the deliberate comparison of different kinds of evidence.

Numbers reveal patterns. Conversations reveal meanings. Observation reveals workarounds. Diaries reveal time. External indicators reveal consequences beyond the product or policy boundary. When these forms of evidence converge, confidence rises. When they conflict, the conflict itself is valuable because it identifies an assumption that requires examination.

Key Takeaways

  • Separate the metric from the mechanism. A positive outcome does not prove that the event associated with it was beneficial. Investigate what caused the change before recommending it.
  • Measure output, experience, and externality. Ask what happened, how it felt, and who paid costs that the official metric ignored.
  • Research the people who disappear from the data. Users who abandon a process, workers who absorb hidden labor, and communities exposed to pollution are often absent from success dashboards.
  • Use mixed evidence deliberately. Combine interviews, field observation, diaries, surveys, behavioral data, and environmental or health measures. Each reveals a different part of the system.
  • Turn surprising results into design questions. If a harmful event produces a benefit, preserve the beneficial mechanism while removing the harm.

The most dangerous systems are not those with no metrics. They are systems with persuasive metrics that describe only the part of reality they were built to see.

A society that measures growth without pollution may mistake illness for prosperity. A technology team that measures engagement without understanding may mistake compulsion for value. In both cases, the remedy begins with humility: the recognition that the model is smaller than the world.

The next time a number improves, pause before calling it progress. Ask who experienced the improvement, who was excluded from it, what hidden condition made it possible, and what changed beyond the boundary of measurement. Sometimes the most important evidence is not the success signal itself, but the human and environmental story that the signal leaves out.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣