Why Good Judgments Need Better Evidence, and Better Evidence Needs Better Notes
Hatched by kaiyan zhang
Jun 07, 2026
10 min read
2 views
34%
The hidden problem with almost every decision system
What do a YouTube note-taking workflow and an oncology approval endpoint have in common? At first glance, almost nothing. One is about capturing ideas from videos into a personal knowledge system. The other is about deciding whether an experimental cancer therapy deserves accelerated approval. Yet both are wrestling with the same deep problem: how do you know a signal is real before time, context, and noise distort it?
That question matters more than it seems. In one case, the signal is an idea worth remembering. In the other, it is a treatment effect worth trusting. In both cases, the temptation is to wait for the most complete evidence possible. But complete evidence is slow, expensive, and sometimes unavailable until the moment it no longer matters. The alternative is to identify the earliest trustworthy marker, the smallest unit of proof that still deserves belief.
This is where the two domains unexpectedly converge. A good note-taking system is not just a storage system. It is an evidence system for your thinking. And a good clinical endpoint is not just a metric. It is an evidence system for action. Both are attempts to answer the same practical question: what is the earliest meaningful proof that something works?
The real challenge is not collecting more information. It is learning which information is already enough to act on.
The illusion of completeness
Most people think better judgment comes from more data. In practice, better judgment often comes from better proxies. The problem is that proxies are easy to misuse. A transcript can look like understanding, just as a tumor shrinkage statistic can look like cure. The danger is mistaking a convenient indicator for the thing itself.
That is why note-taking during video consumption is deceptively difficult. A video can feel rich and memorable in the moment, but memory fades quickly. If you merely watch, you get an experience. If you capture timestamps, summaries, and embedded references, you begin to separate what was interesting from what was actually useful. The note is no longer a passive archive. It becomes a testable record of your interpretation.
Clinical research faces the same trap. Overall survival is intuitively compelling because it is concrete and undeniable. But it is also slow, confounded by subsequent therapies, and difficult to attribute cleanly to one intervention. Progression free survival is earlier, but still imperfect. Objective response rate, by contrast, is narrower but more direct: did the tumor shrink in a way that can reasonably be attributed to the drug? In a single arm setting, that simplicity has power.
This reveals a striking pattern: when the environment is noisy and time is scarce, the best signal is often the one that is most directly attributable. Not the most complete, not the most emotionally satisfying, but the one with the cleanest causal link.
That principle applies well beyond medicine and notes. A startup may obsess over revenue while ignoring activation. A student may collect highlights while never producing recall. A manager may track quarterly outcomes while missing the leading indicator that the team has stopped speaking honestly. In each case, the wrong endpoint feels mature because it is final. The right endpoint feels smaller because it is earlier.
But earlier does not mean weaker. It means more actionable.
What a timestamp and a response rate are really doing
At a technical level, a timestamp in a video note and an objective response rate in a trial both serve the same function: they preserve traceability. They let you return to the original event and inspect whether your interpretation still holds.
A timestamp says, “This idea came from this exact moment, in this exact context.” That matters because meaning changes when detached from its source. A quote without context becomes a slogan. A note without provenance becomes folklore. The timestamp protects against self deception by forcing you to show your work.
Similarly, ORR protects against over interpretation by tying a claim to a visible clinical change. It does not answer every question, but it answers one important question with unusual clarity: did the treatment produce a direct and measurable effect? In a single arm trial, that can be the difference between a plausible intervention and wishful thinking.
The deeper lesson is that good evidence is not just about magnitude, it is about attribution. A big effect that cannot be traced to the cause is less useful than a smaller effect that can. This is true for learning systems too. If you remember a concept but cannot say where it came from, when it mattered, or why it seemed important, you do not really own the idea. You are carrying an unverified impression.
This is why the best note-taking workflows increasingly resemble evidence pipelines rather than notebooks. Embedding the video, pausing to annotate, extracting the transcript, and syncing notes into a second system are not just productivity hacks. They are attempts to increase auditability. They make the path from stimulus to interpretation visible.
In medicine, the same logic is why surrogate endpoints are both useful and dangerous. They accelerate decisions, but only if the surrogate truly predicts what matters. In knowledge work, the equivalent danger is believing that better capture automatically produces better understanding. It does not. It only produces better raw material. Understanding still requires judgment.
A useful mental model: the three layers of proof
To synthesize these ideas, it helps to think in three layers of proof.
1. Capture proof
This is the first layer, where you establish that something happened and can be revisited. In note taking, this is the embedded video, transcript, or timestamped highlight. In trials, this is the observed tumor change or response assessment.
Capture proof answers: Did anything visibly happen?
2. Attribution proof
This is the second layer, where you ask whether the observed change can reasonably be linked to the intervention or input. In note taking, this means knowing which claim came from which minute, which example supported which concept, and how the idea fit into the broader argument. In trials, this means whether the response is likely due to the therapy rather than noise, natural history, or another treatment.
Attribution proof answers: Can I credibly connect the effect to the cause?
3. Meaning proof
This is the hardest layer. It asks whether the signal actually matters for the larger goal. A note can be accurately timestamped and still be trivial. A tumor response can be dramatic and still not translate into longer life or better quality of life.
Meaning proof answers: Does this matter in the system I ultimately care about?
This framework explains why people so often argue past one another about metrics. One person is demanding capture proof, another wants attribution proof, and a third is already thinking about meaning proof. They are not disagreeing about facts. They are operating at different layers.
The same is true when building a personal knowledge system. Many people get stuck because they optimize for capture proof only. They accumulate clips, highlights, and summaries, but never promote them into attribution or meaning. The result is a beautifully indexed archive of things they once thought were interesting.
A better system asks: what should graduate from a momentary signal into durable belief? That question forces a discipline that is rare in both learning and research. It requires you to separate the seductive from the significant.
Not every measurable change deserves attention. The real craft is deciding which changes are evidence and which are noise.
Why the earliest signal is often the most honest one
There is a counterintuitive truth here: the earliest useful measure is sometimes more honest than the final outcome.
Why? Because final outcomes often absorb too many unrelated influences. A patient may live longer for reasons that have little to do with the drug. A student may eventually master a topic after multiple exposures, but the original note may not reveal what actually helped. A business may grow for reasons unrelated to the original product decision. Final outcomes are important, but they are often contaminated by the world after the intervention.
Early signals can be cleaner. A direct response to a therapy, if observed carefully, tells you something about biological activity sooner than waiting years for survival data. A timestamped note that immediately connects an idea to its source tells you something about comprehension sooner than relying on memory months later. The early signal is not the whole story, but it can be the most diagnostically useful part of the story.
That does not mean early signals are sufficient. It means they are decision quality accelerators. They let you iterate before the full truth arrives. In clinical development, that can speed access to promising therapies. In learning, that can help you decide whether a concept deserves elaboration, comparison, or deletion from your system.
Think of it like a weather forecast. You do not wait for the storm to hit the porch before deciding whether to bring in the chairs. You use the best available indicator. But you also know the forecast is not the storm. Likewise, an ORR is not a cure, and a note is not understanding. Both are forecast devices for something more important.
The discipline is to use the forecast well without worshipping it.
From note taking to decision making: the same architecture
If you zoom out, both of these domains are really about building an interface between uncertainty and action.
When you watch a video and take notes, you are converting flowing language into structured memory. When you assess a trial endpoint, you are converting biological complexity into a decision variable. In both cases, you need a filter that is strict enough to reject noise and loose enough to detect real change.
This is why the best systems are not maximalist. They are layered. A good workflow might look like this:
- Raw capture: embed the source, record the moment, store the transcript.
- Selective extraction: pull only the claims, examples, or numbers that matter.
- Causal tagging: annotate why you think this mattered, not just what was said.
- Review loop: revisit the note or endpoint later to test whether the signal still holds up.
That same architecture applies to clinical reasoning. First you observe the response. Then you ask whether it is attributable. Then you decide whether it changes practice. Then you revisit those assumptions when later data arrive.
This is a useful antidote to one of the most common intellectual mistakes: confusing visibility with validity. Just because a signal is easy to see does not mean it is trustworthy. Just because it is hard to see does not mean it is unimportant. Good systems are designed to make the right signals visible at the right time.
For individuals, that means resisting the urge to make your note system a junk drawer. For researchers, it means resisting the urge to over promise on endpoints that are convenient but weak. For everyone, it means asking a more uncomfortable question: what are we treating as evidence simply because it is available?
Key Takeaways
- Choose signals that are directly attributable. When possible, favor evidence that can be traced back to a specific cause rather than a vague downstream effect.
- Separate capture from understanding. Saving a transcript or highlight is not the same as making sense of it. Add interpretation, not just storage.
- Use early indicators, but do not confuse them with final truth. Early signals are useful for decision making, not for declaring victory.
- Build systems with auditability. Timestamps, provenance, and clear endpoints reduce self deception and make later review possible.
- Ask which layer of proof you are actually seeking. Capture proof, attribution proof, and meaning proof are different problems and require different standards.
The real lesson: evidence is a design choice
We often talk about evidence as though it simply exists, waiting to be found. But evidence is also something we design. We choose what to measure, when to measure it, and how much confidence to grant it. A note taking system and a clinical endpoint are both engineered answers to the same human limitation: we cannot wait forever, and we cannot perceive everything at once.
That is why these two seemingly unrelated practices belong in the same conversation. Both ask us to replace vague faith with calibrated confidence. Both ask us to identify the smallest trustworthy unit of truth. Both remind us that action usually happens before certainty.
The deepest shift is this: better judgment is not about having perfect information. It is about having the right evidence at the right time, in a form you can trust.
So the next time you watch a video and mark a timestamp, or read a trial result and ask whether ORR is enough, you are doing the same intellectual work. You are deciding whether a signal is merely interesting or actually decision worthy.
And that may be the most important skill in modern life: not collecting more input, but learning how to recognize the first honest proof that something matters.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣