When a Company Can Move Faster Than Its Purpose
Hatched by Peter Buck
Jun 22, 2026
9 min read
2 views
84%
The strange problem behind speed
What happens when an organization becomes better at judging itself than at knowing what to build?
That is the deeper tension hiding inside both the collapse of a chaotic company culture and the rise of scalable machine judges. In one case, a business is stripped down, made leaner, harder, and more execution focused, yet still faces a brutal question: execution toward what, exactly? In the other, a new kind of evaluator can score thousands of open ended outputs at astonishing speed, but the harder challenge is not judging more quickly. It is deciding what judgment is actually for.
This is one of the defining problems of modern systems, whether human or artificial: speed of evaluation is rising faster than clarity of purpose. We are becoming very good at producing ratings, rankings, signals, and verdicts. We are less good at knowing which ones matter, or what future they are intended to create.
The deepest bottleneck is often not execution. It is the absence of a coherent destination.
That idea applies to companies, models, teams, and institutions. Once you see it, a lot of apparent progress starts to look more ambiguous.
Execution is easy to optimize, direction is not
In a broken organization, the first instinct is usually to fix execution. Cut bloated layers. Rebuild accountability. Make the engineers hard core again. Reduce noise. This can absolutely be necessary. A company can become so overstaffed, so internally fragmented, and so politically exhausted that it loses the ability to ship anything at all.
But there is a trap here. A sharper machine still needs a destination. A team that can move quickly without knowing where it is going does not become visionary by default. It becomes dangerous, or at best, erratic. The modern obsession with operational excellence often assumes that if you restore discipline, the right strategy will somehow emerge from the motion itself. Sometimes that happens. Often it does not.
Think of a ship with a rebuilt engine but no map. It can now sail faster into fog.
This is why cultural resets so often feel thrilling in the short term and unsettling in the long term. They answer a real problem, namely, that broken institutions cannot execute. But they leave open the larger question: what is the institution for? If that question remains unresolved, speed merely amplifies uncertainty.
A company can be ruthlessly efficient at producing outputs that are strategically meaningless. A model can be highly optimized at generating answers that are confidently irrelevant. The form changes, but the structure of the problem stays the same.
The rise of machine judges and the illusion of certainty
The new world of LLM evaluation makes this tension unusually visible. Open ended tasks are hard to measure because they are not like checking arithmetic or classifying spam. There is no single obvious correct answer. A response can be helpful but incomplete, fluent but shallow, original but unsafe, concise but misleading. In that setting, the appeal of a scalable judge is obvious: if humans cannot reliably and cheaply evaluate thousands of outputs, perhaps another model can.
And indeed, a model judge can be extraordinarily efficient. It can score huge batches in minutes and even agree with expert judgments at surprisingly high rates. That sounds like progress, and in many ways it is. When the volume of generated text explodes, evaluation becomes a bottleneck. Without some scalable surrogate for judgment, iteration slows to a crawl.
But the deeper issue is not whether a machine can imitate a human evaluator with impressive consistency. The deeper issue is what kind of evaluation regime we are building.
A judge does not merely measure progress. It shapes what counts as progress in the first place. If your evaluator rewards verbosity, models become verbose. If it rewards safety over usefulness, models become timid. If it rewards surface coherence over substantive reasoning, models learn to sound right rather than be right. In other words, every judge is also an architect of behavior.
Evaluation is never neutral. It is a hidden form of product design.
This is why the promise of scalable judgment is both empowering and unsettling. It lets us move faster, but it also risks hardening our blind spots. Once a judge becomes a throughput engine, the organization may stop asking whether the judge is measuring the right thing.
The same thing happens in companies. When internal metrics become too good, they can create the illusion that strategy is solved. Dashboards fill up. Meeting cadence improves. Performance reviews get tighter. Yet the organization may still be asking the wrong question. We end up with a culture of legibility, not necessarily of wisdom.
The real challenge is not measurement, but alignment of purpose
There is a deeper connection between the company and the model: both are systems where judgment and execution have been separated.
In older, slower systems, the people who built things often also knew why they were building them. Feedback was messy and local. A product manager sat near the engineers, a customer complaint carried moral force, and strategic confusion was harder to hide because progress itself was slow. As systems scale, that closeness breaks. Builders become specialized. Evaluators become specialized. Metrics become abstractions.
That specialization is powerful, but it introduces a new failure mode: the organization can optimize one layer while losing coherence at another. Engineers can ship. Judges can score. Yet the meaning of success drifts.
This is especially dangerous in AI, because a model can appear to improve along a benchmark while getting worse in ways the benchmark does not notice. A polished answer may hide shallow reasoning. A good score may conceal brittle behavior. A highly rated system may still fail in the wild, where the real world is messy, adversarial, and full of edge cases.
But human organizations do this too. They confuse measurable improvement with meaningful improvement. They reward speed, output, and internal consistency, then later discover that these were proxies all along. The proxies were useful until they became the goal.
A useful mental model here is the distinction between throughput, judgment, and direction:
- Throughput answers: how much can we produce?
- Judgment answers: how well can we assess quality?
- Direction answers: what is worth producing at all?
Most institutions overinvest in the first two and underinvest in the third. That is why they can become both efficient and lost.
What scalable judging teaches us about organizations
The temptation is to think that better evaluation solves everything. It does not. It only makes the system more responsive to its chosen criteria. If the criteria are weak, scalable evaluation accelerates mediocrity.
This insight matters because organizations often try to fix confusion with audits, reviews, scorecards, and tighter performance loops. Those tools are valuable. Yet they are only as good as the underlying theory of value. A company can track response times, code velocity, user growth, and headcount efficiency while missing the more strategic question of whether it is creating something that people truly need.
The same goes for LLMs. A judge can tell you which answer is more aligned, more helpful, or more complete, but only if the rubric reflects the outcome you actually care about. Otherwise, the judge is simply industrializing a mistake.
Consider a restaurant. A critic can rate dishes on presentation, freshness, and consistency. Helpful. But if the kitchen starts optimizing entirely for those scores, it may produce beautiful plates that nobody wants to eat twice. The rating system has not merely measured the food. It has begun to shape the cuisine.
That is the hidden danger of scalable judges in any domain: they make optimization cheap. And when optimization becomes cheap, people tend to optimize prematurely.
The best systems keep judgment close to purpose
So what should smart organizations do with this insight?
They should not reject metrics, judges, or operational discipline. That would be romantic but ineffective. Instead, they should build systems where judgment remains tethered to first principles. The key is not to eliminate evaluation. It is to prevent evaluation from becoming self-contained.
A strong institution asks three questions repeatedly:
- What are we optimizing for?
- How do we know our measurement reflects that goal?
- What important outcomes are we still failing to see?
This is true whether the evaluator is a manager, a benchmark, or a model judge. The best systems treat judgment as a tool for learning, not a substitute for thinking. They use metrics to expose reality, not to replace it.
That means creating friction in the right places. Human review for high stakes cases. Periodic benchmark redesign. Red teaming. Qualitative audits. Direct contact with end users. These are not inefficiencies to be eliminated. They are safeguards against overconfidence.
The most mature organizations understand that speed without calibration is just a faster way to be wrong.
A practical framework: the three questions of honest intelligence
If there is one framework that emerges from these ideas, it is this:
1. Can we do it?
This is the execution question. Do we have the capability, the talent, the infrastructure, the discipline?
2. Can we judge it?
This is the evaluation question. Do we have a reliable way to distinguish good from bad, useful from useless, safe from unsafe?
3. Should we do it?
This is the purpose question. Does this move us toward a coherent destination that matters?
Most failures happen when the first two are answered loudly and the third is left vague.
A company may say yes to execution and yes to judgment, then discover that it is building a product nobody needs. A model may score well and benchmark well, then fail in deployment because the benchmark was too narrow. In both cases, the system mistakes local optimization for global success.
The most valuable habit, then, is not just measuring better. It is periodically stepping outside the measurement system and asking whether the game itself still makes sense.
Key Takeaways
- Do not confuse operational clarity with strategic clarity. A team can become faster without becoming more purposeful.
- Every judge shapes the behavior it measures. If you optimize the rubric, you are also redesigning the system.
- Use metrics as probes, not as prophets. They should reveal reality, not replace judgment.
- Keep a live connection to the end user or end goal. The farther evaluation drifts from actual use, the more likely it is to reward the wrong thing.
- Ask the purpose question regularly. Not just what is working, but what is worth working on.
The final lesson: faster systems need deeper questions
The most seductive idea in modern management and AI is that if we can just make systems more efficient at deciding, scoring, and executing, then the hard problems will fade. But the real lesson is almost the opposite. As systems get faster, the cost of being wrong falls on us more quickly and more repeatedly. Speed magnifies intent.
That is why the deepest challenge is not building a company that can execute, or a model that can judge. It is building an organization, or a technology, that knows what kind of world it is trying to create.
A ship with a stronger engine is impressive. A judge with higher agreement is impressive. But neither tells you whether you are sailing toward land or deeper into fog.
In the end, the real test of intelligence is not how quickly it can answer. It is whether it can still ask the better question: what are we actually for?
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣