Why Data Stops Being a Normal Input Once It Starts Changing Behavior

SEAN SYLVIA

Hatched by SEAN SYLVIA

May 19, 2026

11 min read

86%

0

The strange thing about modern markets: the more you know, the less standard the game becomes

What happens when a market learns from itself? At first, the answer sounds straightforward: collect more data, improve the model, make better decisions. But a deeper question hides inside that optimism. When information becomes an input that changes what people do, is it still just information, or does it become a force that reshapes the market itself?

That question connects two worlds that are usually kept apart. In one, researchers study how to allocate scarce resources, target the poor, measure treatment effects, and detect quality in messy human systems like health care and tenancy contracts. In the other, economists of digitization ask whether data creates market power, whether data feedback loops are real, and whether digital platforms alter welfare through their ability to observe, predict, and steer behavior.

The common thread is not simply “data.” It is the economics of response. Once people react to what they are measured on, targeted by, or recommended into, data stops being a passive record of reality. It becomes part of the mechanism that creates reality.


When measurement changes the thing being measured

The oldest lesson in applied economics is that incentives matter. A tenant may work harder when the contract shifts. A provider may improve when there is accountability. A patient may report higher satisfaction when attention rises, even if clinical quality barely changes. A household may self-select differently when a program is designed around visible thresholds.

These are not separate case studies. They are variations on one foundational idea: people adapt to the system that observes them. The moment observation becomes consequential, behavior becomes strategic. Then measurement no longer reveals a stable underlying state. It helps produce the state.

That is why the classic aspiration to “just use more data” is so seductive and so incomplete. In many domains, data behaves like sand in your hands. The tighter you grip, the more it slips into a new shape. A doctor who knows she is being audited may change practice. A farmer offered a contract may adjust effort. A welfare recipient may respond to a targeting rule. A consumer on a platform may be nudged by rankings and recommendations.

This is the first deep connection: the social world is reflexive. The act of measurement can alter the object measured, which means economics cannot treat data as a neutral substance the way physics treats weight or temperature. Data in human systems is often closer to a conversation than a snapshot.

In human markets, data is not just a mirror. It is also a lever.


Why more data does not always mean stronger power

At first glance, digital platforms seem to confirm the fear that data creates domination. If a firm learns from billions of interactions, surely its advantage should compound. More users generate more data, more data improves predictions, better predictions attract more users, and the loop keeps spinning. That is the familiar data feedback loop story.

But there is an important wrinkle: data may not behave like a conventional input. In ordinary production, doubling an input can often meaningfully raise output, at least up to a point. Yet learning from data may have diminishing returns. The first thousand observations can be transformative. The next million may refine the model only slightly.

That matters because market power built on data depends on more than just quantity. If returns to data are nonincreasing, then scale alone cannot explain dominance. A platform may still have power, but not because raw data acts like oil pumped from a well. Instead, power may come from how data interacts with product design, switching costs, network effects, and the ability to shape behavior.

This distinction changes the policy conversation. If data were a conventional input with strong increasing returns, then “whoever has the most wins” would be a plausible shorthand. But if the returns flatten, then the real question becomes more subtle: what kinds of data matter, for which decisions, under what institutional conditions? A platform with lots of noisy data may still learn less than a smaller platform with cleaner feedback, sharper incentives, or better classification schemes.

Think of the difference between a doctor and a shopkeeper. A shopkeeper may know what sells. A doctor needs to know what works. The latter requires not just more observations, but causal clarity, trustworthy implementation, and protection against distortion. Similarly, a digital platform can observe endless clicks and still fail to understand welfare.


The hidden common problem: classification under strategic behavior

The deepest bridge between development economics and digitization is not scale. It is classification.

A government trying to target poor households must classify people into groups. A provider must classify symptoms into diagnoses. A landlord or tenant must infer effort from incomplete signals. A platform must classify preferences from behavior. In each case, the institution is not merely observing a world that already comes sorted. It is trying to create a usable map from ambiguous traces.

Classification is hard because people know they are being classified. Once the rule is visible, they may adapt to it. That creates three intertwined problems:

  1. Noise: the signal is incomplete or error prone.
  2. Strategic response: people change behavior to fit or evade the category.
  3. Feedback: the category itself alters future data.

This is why targeting the poor is so difficult. If a transfer rule is too blunt, it misses many who need help. If it is too transparent, people may manipulate indicators. Self targeting uses the insight that people respond to design, but then the design must be built around incentives, not just eligibility. The same tension appears in health care audits. Observed quality may improve when the audit is active, but what is being measured then, quality or compliance with inspection?

Now translate that to digitization. A platform learns what users click. But users click within a menu of what the platform shows them. A recommendation system is not merely discovering preferences. It is co-producing preferences and the data that claim to reveal them. This is why the learning problem in digital markets is structurally similar to the targeting problem in development economics. Both are classification systems operating in a world where the classified agents react.

The more visible the rule, the less innocent the data.

That does not make the data useless. It makes interpretation harder. The central error is to treat observed behavior as if it were a pure readout of latent demand, effort, or need. Often it is demand under constraint, effort under incentives, or need under an institution that itself has already changed the field.


A better mental model: data as a thermostat, not a thermometer

A thermometer records temperature. A thermostat changes it.

That is the mental model that best unifies these ideas. In many markets, data functions less like passive measurement and more like a thermostat. It detects deviations, triggers action, and thereby changes the environment. Once we see that, several puzzles become clearer.

First, why treatment effects differ so much across people and settings. A policy is not a single object. It interacts with local incentives, expectations, and constraints. A welfare program may work in one district and fail in another not because the program changed, but because the environment reacted differently. This is exactly why methods that estimate heterogeneous treatment effects matter. They search for which groups respond, rather than assuming one average effect tells the whole story.

Second, why accountability can improve quality but also produce gaming. If a clinic knows it is audited, behavior may improve along measured dimensions. Yet the system can also drift toward what is visible, documented, or easy to verify. The measured world becomes better in one sense and more distorted in another.

Third, why digital platforms can be so powerful without unlimited increasing returns to data. A platform does not need data to scale forever if it can use data to alter the rules of interaction. Better ranking, tighter matching, and targeted nudges can change the market topology itself. That is a different source of power from simply having more observations.

The thermostat model also clarifies the role of experimentation. Randomized experiments are valuable not only because they identify causal effects, but because they reveal how systems respond when we intervene. The point is not merely to estimate average treatment effects. It is to learn where the feedback loops are strongest, where incentives bend behavior, and where the institution starts to behave like a live organism instead of a fixed machine.


The real economic question: who controls the feedback loop?

Once you accept that data can reshape behavior, the crucial issue is not “Who has data?” It is who controls the feedback loop.

If a health system uses audits, who defines quality? If a government targets transfers, who sets the eligibility rule? If a platform recommends products, who decides what counts as relevance? If a tenancy contract changes incentives, who bears the risk of effort being hard to observe? These are questions of power disguised as questions of measurement.

This is where the development and digitization literatures meet most sharply. In development settings, institutions are often trying to solve a problem of missing or manipulated information with limited administrative capacity. In digital markets, platforms often have abundant information but still face strategic behavior and welfare tradeoffs. In both settings, the institution that controls measurement can shape the set of feasible actions for everyone else.

That creates a moral and political dimension to data that purely technical discussions miss. A targeting rule can improve efficiency while hardening exclusion. A recommendation system can increase convenience while narrowing choice. A monitoring system can reduce shirking while crowding out intrinsic motivation. Data governance is never just about accuracy. It is about which adaptations are encouraged, which are punished, and which groups get trapped by the rule itself.

This is why “better data” is not an endpoint. Better data can simply mean better control. The more useful question is: better for whom, and toward what behavior?


A practical framework: the three tests for any data system

To think clearly about data in markets, policy, or platforms, use three tests.

1. The observability test

What is actually being observed, and what is missing?

A health audit may observe documentation more easily than care quality. A platform may observe clicks more easily than satisfaction. A poverty targeting algorithm may observe assets more easily than vulnerability. If the observable proxy is thin, the system will optimize the proxy, not the goal.

2. The adaptability test

How can people change behavior once they know the rule?

If agents can easily game the metric, then the data is already part of the incentive structure. Any design must anticipate strategic response. This is why self targeting can work in some cases: it deliberately turns adaptation into a sorting mechanism.

3. The feedback test

Does learning from the data alter future data generation?

If yes, then the system is recursive. More use of the data changes the population, the platform, or the institution. That means the model will drift unless it is continuously re-evaluated. It also means historical data may lose value as the environment adapts.

These three tests separate naïve data optimism from serious institutional design. A system passes when it not only predicts well, but stays honest under adaptation.


Key Takeaways

  • Do not treat data as passive evidence in human systems. If people react to being measured, data is part of the intervention.
  • Ask what the metric changes. A good metric can improve behavior, but it can also distort effort toward what is easiest to observe.
  • Beware of assuming scale equals advantage. More data may help, but returns can flatten quickly, so power often comes from control of the feedback loop, not sheer volume.
  • Use the three tests: observability, adaptability, feedback. Before trusting any data system, ask what it sees, how people can game it, and how it changes future behavior.
  • Focus on institutional design, not just prediction. The goal is not only to infer what people want or need, but to build systems that remain fair and useful once people start responding to them.

The deepest lesson: information is never only informational

The seductive promise of the data age is that once we know enough, we can reduce uncertainty and make cleaner decisions. But the better lesson is more unsettling and more useful: in social systems, information changes the system it studies.

That is why audits improve some things and distort others. That is why targeting can help and exclude. That is why platforms can learn a great deal without data behaving like a normal input. That is why empirical work in messy human environments keeps returning to treatment heterogeneity, incentives, and strategic response. The world is not just observed. It is negotiated.

So the next time someone says that more data will solve the problem, a better response is not skepticism for its own sake. It is a sharper question: What behavior will this data create once it becomes part of the environment?

That question turns data from a forecasting tool into a design problem. And once you see that, you stop asking only how to measure the world more accurately. You start asking how measurement itself is remaking the world you thought you were merely describing.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣