The Best Questions Are Built Like Rockets
Hatched by Kunal Grover
Aug 22, 2026
11 min read
2 views
92%
What if the most important difference between analytics and data science is not the sophistication of the mathematics, but the willingness to let reality interrupt the plan?
A detective begins with an incident and asks: What happened? A laboratory scientist begins with a possibility and asks: What might happen if we change the conditions? Both work with evidence, but they use it differently. One reconstructs the past. The other creates disciplined encounters with the unknown.
A rocket test reveals why this distinction matters. A launch can achieve its primary ascent, execute a complex stage separation, continue toward space, and still end with a booster failing to complete its return. If the only question is whether the vehicle landed intact, the result looks like failure. If the question is which parts of the system survived contact with reality, the same flight becomes a dense deposit of knowledge.
This is the deeper connection: good inquiry is not merely the collection of facts. It is the design of situations in which facts can change what we believe. Analytics tells us how to investigate what already happened. Science, engineering, and ambitious organizations must also learn how to make the next test informative.
The Difference Between Explaining the Past and Expanding the Possible
Data analytics is often detective work. A business sees a decline in customer retention, a hospital sees an increase in readmissions, or a website sees an unusual spike in traffic. The analyst gathers clues, separates signal from noise, and builds the most credible account of the event.
This is valuable because organizations are full of unexplained outcomes. A dashboard may show that sales fell, but not whether the cause was price, seasonality, product quality, a competitor, or a change in customer behavior. Analytics narrows the field. It turns a vague concern into a defensible explanation.
Data science, at its most ambitious, asks a different kind of question. It may investigate whether a new recommendation system changes purchasing behavior, whether a different medical protocol improves recovery, or whether a forecasting model can identify risk earlier. The work is not simply to inspect reality. It is to construct a test that distinguishes among competing possibilities.
The distinction can be expressed as a movement between two forms of uncertainty:
- Retrospective uncertainty: We know the outcome, but not the causes.
- Prospective uncertainty: We know several possible futures, but not which one will occur.
A detective works backward from an observed outcome. A scientist or engineer works forward through a designed intervention. The first seeks explanation. The second seeks learning through controlled exposure to consequences.
But this distinction is incomplete unless we add a third question: What kind of failure would teach us the most?
That question changes everything. It shifts attention away from whether an experiment succeeded in a simple binary sense and toward the information gained at each stage. A test can fail operationally while succeeding epistemically. It can also succeed operationally while teaching almost nothing if the conditions were too safe, too ambiguous, or too poorly measured.
The value of a test is not determined only by whether the system works. It is determined by how clearly the test reveals what the system can and cannot yet do.
A Rocket Launch as a Lesson in Layered Learning
Consider a demanding flight test involving a heavy launch vehicle. The vehicle ignites all 33 engines on its booster, rises from the launch site, loses one engine during ascent, performs a hot staging maneuver, and sends its upper stage onward with six engines of its own. Later, the booster attempts to turn, perform a boostback burn, reignite its engines, and land. It cannot light all the planned engines, the boostback burn ends early, and the booster ultimately makes a hard splashdown.
A headline might compress this into one word: failure. Yet that word collapses several distinct questions into one verdict.
Did the booster produce enough thrust to leave the ground? Yes. Did the vehicle tolerate an engine shutdown during ascent? It continued. Did stage separation occur? Yes. Did the upper stage proceed toward space? Yes. Did the return sequence work as intended? No. Did the flight identify a failure mode in the restart or boostback chain? Apparently so.
The important analytical move is to decompose the event into claims, rather than treating it as a single performance score. A complex system makes many promises at once. It promises to ignite, lift, separate, navigate, restart, control its trajectory, survive heating, communicate, and land. One flight is therefore not one experiment. It is a bundle of experiments coupled together.
This gives us a more useful model of progress:
Progress equals the number of important uncertainties converted into evidence, not merely the number of missions ending in success.
That does not make outcomes irrelevant. A hard splashdown is materially different from a controlled landing. But the operational result and the learning result should be tracked separately. If they are fused, teams become tempted to hide failures, exaggerate partial victories, or abandon valuable lines of investigation because the final image is disappointing.
Imagine a medical trial in which a treatment improves symptoms but produces an unexpected side effect in a specific subgroup. The treatment is not simply a success or failure. It has a performance profile. The next question is not, “Did it work?” but, “Under which conditions did it work, and where did the system become fragile?”
The same logic applies to a machine learning model. A model may achieve excellent overall accuracy while failing badly on rare but consequential cases. The aggregate metric is analogous to asking whether a rocket flew. The subgroup analysis is analogous to examining which engine shut down, which maneuver failed, and what happened when the system transitioned from ascent to return.
Averages describe the vehicle. Edge cases describe the engineering.
The Hidden Skill Is Designing Informative Failure
Most organizations say they want experimentation. In practice, many want reassurance. They approve tests that are unlikely to threaten existing assumptions, use metrics that conceal tradeoffs, and celebrate favorable outcomes before asking whether the result was actually informative.
This is why a culture of experimentation can still produce little learning. An experiment is not defined by the presence of a control group, a dashboard, or a statistical model. It is defined by whether the outcome can discriminate between plausible explanations.
Suppose an online retailer changes its checkout page and sales rise. That may seem conclusive, but perhaps a major advertising campaign ran at the same time. Perhaps the increase came from returning customers, while new customers experienced no improvement. Perhaps the result is real but temporary. Without a design that isolates the intervention and measures the relevant groups, the organization has an encouraging event, not reliable knowledge.
The strongest experiments therefore have three properties.
They expose a meaningful uncertainty
A test should address a question that could change a decision. Measuring a trivial detail with great precision is not the same as learning something important. Before collecting data, state what belief is at risk.
They create observable stages
Complex systems should be evaluated as sequences, not only as final outcomes. A launch has ascent, separation, propulsion, navigation, reentry, and landing. A product has acquisition, activation, use, retention, and renewal. A clinical intervention has recruitment, adherence, treatment response, side effects, and long term outcomes.
When the stages are visible, a final failure does not erase intermediate knowledge. More importantly, the team can locate where its mental model broke.
They preserve the surprise
An experiment loses much of its value when inconvenient evidence is averaged away or explained after the fact. Unexpected engine behavior, an unanticipated customer response, or a sudden shift in model performance should not be treated as noise merely because it complicates the story.
The surprise is often the payload.
This suggests a practical concept: the surprise budget. Every plan contains assumptions about how the world will behave. The more assumptions a project makes, the more room it must reserve for outcomes that violate them. A team that schedules every hour, allocates every dollar, and defines only one acceptable result has no capacity to learn when reality refuses to cooperate.
A well designed test does not eliminate surprise. It makes surprise legible and survivable.
From Detective Mode to Laboratory Mode
The detective and the laboratory scientist are not rival personalities. They are sequential modes of thought, and strong teams know when to switch between them.
After an unexpected outcome, begin in detective mode. Establish the facts before constructing a narrative. What exactly happened? Which measurements are trustworthy? What changed from the previous attempt? Which events occurred in what order? What explanations are consistent with the evidence?
Then switch to laboratory mode. Turn the leading explanations into competing tests. If an engine failed to reignite because of a temperature condition, a valve issue, a control sequence, or a fuel state, each hypothesis should imply a different observation or intervention. The aim is not to produce the most eloquent explanation. It is to create the next encounter that can prove one explanation more credible than another.
Finally, return to detective mode after the new test. Did the intervention produce the predicted result? If not, was the hypothesis wrong, or was the test incapable of distinguishing it? This cycle can be represented as:
Observe, isolate, intervene, measure, update.
Many teams stop after the first step. They observe an outcome and write a report. Better teams complete the cycle. They use the report to redesign reality, then use reality to revise the report.
This framework also clarifies why metrics alone cannot produce understanding. A metric is an instrument, not a question. A landing rate, conversion rate, error rate, or retention rate becomes meaningful only when connected to a causal model and a decision.
Ask of every important metric:
- What uncertainty does this measure reduce?
- What alternative explanations remain possible?
- What result would cause us to change course?
- Which stage of the system does this metric represent?
- What can this metric hide?
The last question is especially important. A final score can hide timing, distribution, and failure modes. A system may improve overall while becoming dangerously brittle in one transition. It may hit its average target while producing more severe outliers. It may look stable because the measurement is too coarse to reveal deterioration.
How to Build a Learning Architecture
The most useful practical shift is to stop treating learning as something that happens after execution. Learning should be designed into the architecture of the work.
Before a project begins, write a learning contract. It should contain four elements:
- The central uncertainty: What do we not know that matters?
- The observable claims: What must be true at each major stage?
- The failure signatures: What evidence would reveal a specific problem?
- The update rule: What will we do if the evidence supports or weakens each hypothesis?
For example, a team launching a new software feature might define not only a target increase in engagement, but also the expected effects on page speed, support requests, retention, and different customer segments. If engagement rises while support requests surge, the result is not a clean win. It is a tradeoff that requires a decision.
Afterward, conduct a failure decomposition, not a failure autopsy. An autopsy implies that the event is over and the goal is to assign a cause. Decomposition asks which promises held, which broke, and which remain uncertain.
A simple table can help:
| System stage | Intended claim | Observed result | Remaining question |
|---|---|---|---|
| Initiation | The system starts reliably | Confirmed or contradicted | What conditions affect startup? |
| Transition | Control passes between phases | Confirmed or contradicted | Where does instability appear? |
| Recovery | The system returns to a safe state | Confirmed or contradicted | Which component limits recovery? |
This format prevents the final outcome from dominating the entire interpretation. It also creates a bridge between analytics and experimentation. Analytics supplies the careful reconstruction. Experimental design supplies the next intervention.
The deeper organizational lesson is that a test should be judged by its decision value. If the result does not change what the team will do, it may have been measurement theater. The strongest tests narrow the future. They make some paths less plausible and others more attractive.
Key Takeaways
- Separate operational success from learning success. Ask both whether the system achieved its mission and which uncertainties the test resolved.
- Decompose complex outcomes into stages. Track initiation, transitions, performance, recovery, and failure modes instead of relying on one aggregate score.
- Move deliberately between two modes. Use detective mode to establish what happened, then laboratory mode to design the next test that can distinguish competing explanations.
- Write the update rule before collecting evidence. Decide in advance which findings will change the plan, so favorable results do not receive special treatment.
- Treat surprises as assets. An unexpected result is valuable when it is observable, interpretable, and connected to a subsequent intervention.
The temptation in any ambitious project is to ask whether the mission succeeded. That question is necessary, but it is too small to guide serious learning. A rocket that reaches space but cannot return has not demonstrated everything its designers hoped. It has, however, mapped a boundary between what the system can do and what it cannot yet do.
That boundary is where progress begins.
The best analysts do not merely explain yesterday. The best scientists do not merely imagine tomorrow. Together, their habits form a more powerful discipline: designing contact with reality so that even disappointment produces direction.
A failed landing, a broken model, or an unexpected customer response is not automatically valuable. It becomes valuable when the organization has the courage to inspect it without defensiveness and the ingenuity to turn it into a sharper next question. The mature definition of success is therefore not “nothing went wrong.” It is “we learned exactly what went wrong, why it mattered, and what to test next.”
The future belongs to teams that can do more than launch ideas. It belongs to teams that can make every launch, especially the imperfect ones, teach them how to fly farther.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣