When Signals Fail: Rethinking Merit, Moral Posturing, and Risk in the Age of Real Disruption

Profuse Habits

Hatched by Profuse Habits

Apr 15, 2026

9 min read

72%

0

What if both moral outrage and apocalyptic warnings are symptoms of the same breakdown?

In boardrooms, on social platforms, and in op-eds, debates about hiring, fairness, and technological risk often look like battles over values. One side accuses policies meant to diversify institutions of being another form of discrimination. Another side celebrates the discovery of untapped talent by looking where others do not. Meanwhile, on a different front, generations of technologists have warned that the next breakthrough will upend work and power structures; some of those warnings were exaggerated, and some were not. The same rhetorical patterns repeat: certainty where there is none, simple categories substituted for complex tradeoffs, and moral language used to close the argument.

What ties these arguments together is not the ideology behind them. It is the condition they are trying to manage: signal failure. When the signals we rely on to find skill, value, and risk stop working, people adopt blunt substitutes: slogans, moral absolutes, or catastrophic metaphors. Those substitutes can be comforting and rhetorically powerful. They are rarely accurate.

This essay argues that the central question of our moment is not whether diversity policies are moral or whether AI is the end of work. The real question is how to construct better systems for discovering talent and assessing risk when old signals are eroding. If we treat both debates as problems of signal design, we can move past binary fights and toward practices that are robust, experimental, and humane.


Signal Compression: why institutions replace nuance with slogans

Organizations cannot consider everything about every person, product, or technology. They compress information into practical signals: diplomas, resumes, brand affiliations, race and gender categories, patents, press coverage, and platform follower counts. These signals are heuristic shortcuts. They are useful until they are not.

When a signal grows noisy because of systemic change, people stop trusting it. There are two ways institutions respond. The first is to layer on new signals intended to correct perceived failures. For example, policies that expressly prioritize designated groups are an attempt to change what counts as an acceptable hiring signal. The second response is to reject those new signals as illegitimate and to assert a purer signal of merit: the single metric that explains everything.

Both responses aim to reduce uncertainty. The first reweights which attributes matter. The second insists on a different compression entirely. Both are vulnerable to the same problems: they can be gamed, they can obscure other important attributes, and they can become identity markers detached from the processes they were meant to improve.

Think of hiring as gold panning versus sieve selection. Traditional resumes and credentials are a sieve: they filter out large amounts of dirt quickly but can miss nuggets that are small, rare, or nonconformist. Newer policies that push recruiters to look in different neighborhoods are like sending prospectors to unexplored streams. They can uncover overlooked talent, but they also invite accusations that the prospector is skipping the sieve because standards were lowered. The rhetorical clash is about method versus moral claim: is the new stream a legitimate place to look, or a shortcut to good PR?

This conflict is intensified when technological change alters the underlying geology of the stream. When tools change what counts as useful work, the shape of talent and the signals of skill change too. That is where the AI question fits into this pattern.


The wolf, the sheep, and the new information terrain

Warnings about new technology face two perennial problems. First, technologies trend toward exaggeration because attention rewards bold claims. Second, most technological change is incremental and domain constrained. These facts together make it hard to know when a claim is a false alarm and when it is a genuine inflection.

We can think of the difference as a pragmatic test: a real inflection will change the signal space, not just the distribution of outcomes within the old space. When an innovation makes prior signals obsolete, institutions must invent new selection mechanisms. When it merely improves productivity while preserving existing evaluation methods, institutions can adapt more gradually.

Recent systems of generative artificial intelligence appear to be shifting the signal space. Tasks that once provided reliable, observable signals of cognitive work can now be mimicked by models trained on large data sets. Writing samples, formulaic problem solutions, and even some creative outputs are now less reliable as indicators of an individual agent's capabilities. This is what separates a wolf from a false alarm: the wolf alters how you test for competence.

A useful thought experiment is the blind audition. In classical music, orchestras discovered that auditioning behind a curtain increased the selection of women because subjective biases linked to visual cues were removed. If AI can produce the audition itself, then the curtain is no longer enough. You now have to test for capabilities in a live, interactive, adversarial context where improvisation, iteration, and adaptive thinking are observable. The criterion for competence must be something AI cannot easily fake in the given setting.

That is also why debates about diversity policies and debates about AI are not opposites. Both are reactions to the same core problem: how to find the right people when the usual tests are compromised. One reaction is to insist on new categorical priorities. The other is to redesign the tests themselves.


A practical framework for rebuilding selection and evaluation

If the goal is to make institutions better at finding talent and assessing risk, we need a few complementary principles and concrete experiments. I propose a four-part framework: Reduce Compression, Red Team Signals, Short Bets, and Field-Proof Tests.

  1. Reduce Compression. Stop treating any single signal as decisive. Replace binary cutoffs with layered, diverse evidence. Use a portfolio approach to evaluation where different kinds of evidence carry different weights. For hiring, combine objective micro-tasks, structured interviews, work trials, and references, each with explicit scoring rubrics. The portfolio approach reduces the chance that any single corrupted signal will determine the outcome.

  2. Red Team Signals. Adopt adversarial testing for your evaluation metrics. If a small set of practices or policies can be gamed to meet your desired signal, assume it will be gamed. Create teams or exercises that attempt to game hiring or risk metrics. For instance, if a writing sample is valued, run a controlled experiment where some candidates use AI-assisted writing while others do not, then test for differences in downstream performance. The red team will expose where your signals are fragile.

  3. Short Bets. Replace high-stakes, irreversible decisions with time-limited, scalable commitments. Instead of hiring for life based on a single evaluation, hire for a 3 to 6 month paid milestone. Use those periods as explicit discovery phases with measurable outputs. In the context of organizational strategy, prefer modular investments that can be reversed without catastrophic cost. This protects you from false positives when the signal environment is uncertain.

  4. Field-Proof Tests. Design evaluations that require sustained, adaptive performance in realistic conditions. If a task can be done by a snapshot or by a model trained on static data, require interactive and iterative tasks that reveal process and judgment. For creative work, require a live brainstorm, a sequence of revisions, or a collaboration that yields evidence of interpersonal and iterative skill. Field-proof tests are closer to real-world performance and harder for simple algorithms to fake.

These principles apply to both human capital and technological risk assessment. When evaluating the claims of a new technology, use red team exercises, short bets in contained environments, portfolios of evidence that include real-world outcomes, and field-proof tests that stress the system in realistic scenarios.

When the old signals break, do not clutch slogans. Build experiments instead. Experiments reveal whether a change is substitutional, superficial, or genuinely structural.


Concrete examples and analogies to make this useful now

Example 1: Talent discovery in a changing labor market. A company faces two competing pressures: a desire to diversify channels of recruitment and a fear of lowering standards. Apply the framework. Replace absolute resume thresholds with a set of micro-tasks that emulate on-the-job work. Run a red team exercise where some applicants are coached on optimizing resumes while others are evaluated solely on task performance. Hire a small cohort for a paid three month project with clear metrics. Compare retention and productivity across cohorts. This approach reveals whether alternate pipelines actually deliver durable performance.

Example 2: Evaluating a new AI tool. A group of analysts claim a generative model will automate 50 percent of a job category. Instead of accepting the headline, design field-proof tests: pair the model with human reviewers in realistic workflows, measure time saved, error rates, and downstream client outcomes. Run short bets in different departments with clear rollback plans. Have a red team attempt to exploit failure modes. This shows whether the model creates structural change or only superficial shifts in task completion.

Analogy: Gold panning versus mapmaking. Old signals were like maps. They told you where gold had been historically. New methods that expand search are like panning in unexplored streams. Artificial intelligence is like a new kind of metal detector that sometimes beeps at nuggets and sometimes at fool's gold. The solution is not to shout that detectors are immoral or to declare they are omniscient. The solution is to combine detectors with tests that verify an actual nugget when you dig.


Key Takeaways

  • Use a portfolio of signals: never let one metric, credential, or policy decide high-stakes outcomes. Mix small task-based evaluations with structured interviews and time-limited trials.

  • Red team your signals: intentionally try to game your own hiring or evaluation process to reveal weaknesses before others exploit them.

  • Prefer short bets to irreversible commitments: hire or invest for discovery periods with clear deliverables and exit rules.

  • Design field-proof tests: require interactive, iterative, and context-rich demonstrations of capability that are difficult for static or automated proxies to fake.

  • Treat technology claims as hypotheses: validate them in the environment where they will matter, not in isolated benchmarks or media narratives.


Conclusion: the politics of method, not the politics of moral purity

The debates that dominate headlines about institutional fairness and technological risk often feel like moral trials. People take sides, marshal slogans, and deliver definitive pronouncements. Those fights can be historically important. They can also obscure the more mundane but urgent task: rebuilding methods for seeing truth in noisy environments.

When signals fail we choose between two poor options: we either fall back on high-certainty moral claims or we embrace apocalyptic metaphors that promise clarity through fear. Both are attractions, because they make complexity feel manageable. But neither substitutes for the hard work of redesigning evaluation systems.

If we treat these debates as problems of signal design, we can ask better questions: What tests do we use? How are those tests gamed? How can we run cheap experiments that reveal whether a policy or a technology truly changes outcomes? Those are practical, sometimes boring questions. They are also the questions that determine whether institutions will adapt well to real disruption, or simply replace one illusion of certainty with another.

The future will not be decided by slogans or by panic. It will be decided by systems that can detect the difference between noise and novelty, that reward discovery without abandoning rigor, and that prefer experiments over dogma. If you want to change who gets in and what gets built, start by changing how you look.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣