The Hidden Pattern in Biology Is Not the Gene, It Is the Local Context

genken

Hatched by genken

May 29, 2026

10 min read

88%

0

What if the same cell means something different every time you look at it?

Most biology still begins with a familiar assumption: if you can measure the right molecules, you can understand the system. But there is a harder question hiding underneath that assumption. What if the meaning of a molecular profile depends less on the profile itself than on the neighborhood around it? A cell can be stressed, activated, dormant, malignant, or healing, and yet the same internal signals may point in different directions depending on which other cells, spatial structures, and tissue conditions surround it.

This is why two ideas that seem, at first glance, to live in different parts of the computational toolkit actually belong together. One is the rise of large scale models that learn from single cell and spatial omics, where the key unit is not just a molecule but a cell in context. The other is the logic of single sample scoring of molecular phenotypes, where the key challenge is to infer biological meaning from one sample at a time, without relying on a large cohort as a crutch.

Together they point to a deeper shift in biology: from asking, “What is this cell made of?” to asking, “What kind of situation is this cell in?”


The old game: measuring parts, hoping the whole will appear

Biology has long been driven by reduction. Break the system into genes, proteins, pathways, and markers, then rebuild the story from those pieces. That strategy has been enormously successful, but it carries an invisible cost: it often treats biological signals as if they were portable, as if a molecular pattern meant the same thing everywhere.

That assumption works only up to a point. A marker of immune activation in one tissue may mean healing in another and pathology in a third. A proliferative signature in a developing organ is not the same as a proliferative signature in a tumor. Even within a single tumor, a cell near blood vessels can live in a very different reality from one trapped in a hypoxic pocket. The same internal readout can therefore be less like a diagnosis and more like a sentence fragment, meaningful only when you know the surrounding paragraph.

This is the first tension connecting the two methods: biology is local, but many analytical methods are global. Traditional cohort based analysis averages away the very heterogeneity that gives biological signals their meaning. It asks for stable patterns across many samples, then uses those patterns to define phenotypes. But what happens when the phenotype is not stable, or when the main unit of meaning is an individual sample rather than the population average?

That is where single sample scoring becomes crucial. It asks whether a molecular phenotype can be estimated within one sample, on its own terms, without waiting for a control group to rescue interpretation. This is not merely a technical convenience. It is a philosophical correction. It acknowledges that biology often presents itself as a unique configuration, not as a representative instance of a class.

A biological signal is not always a trait. Sometimes it is a relation.


Single sample scoring solves one problem, but reveals a bigger one

At first, single sample scoring seems like a statistical refinement. Instead of comparing across groups, you score a signature within a single profile and ask whether that profile expresses a pathway, phenotype, or program. That is valuable in clinical settings, rare diseases, and heterogeneous tumors, where cohort assumptions can be weak or misleading.

But the deeper implication is more radical. Once you can score a phenotype within one sample, you begin to treat the sample itself as a miniature ecosystem. A sample is no longer just a vector of expression values. It becomes a structured situation containing competing programs, local signals, and context dependent interactions.

This is where the spatial turn changes everything. Single cell measurements tell you what each cell is doing. Spatial omics tells you where it is doing it. Together, they reveal that phenotype is often not an intrinsic property of a cell alone, but an outcome of placement, proximity, and neighborhood effects. A single sample score can tell you that a tissue contains an inflammatory program. Spatial and single cell analysis can tell you whether that program is centralized, peripheral, immune confined, or woven through the entire tissue architecture.

Imagine reading a city by counting only the number of restaurants. You might know something useful, but not enough to understand the city’s function. Are those restaurants clustered in a wealthy district? Near a transit hub? In a tourist zone? Are they serving residents or visitors? Location changes meaning. Biology is the same. A pathway score without spatial and cellular context is like a heat map without a map.

The real tension, then, is not simply cohort versus single sample. It is summary versus situated meaning. A score is a summary. A foundation model trained on single cell and spatial omics is a way of learning the situations that make those summaries interpretable.


Why foundation models matter: they learn the grammar of context

What makes a foundation model relevant here is not just scale. It is the possibility of learning a grammar of biological context.

A traditional analytical pipeline often starts with a predefined feature set, then asks whether a sample matches a known phenotype. A foundation model does something more flexible. It learns patterns across many cells, tissues, and spatial arrangements so that it can represent biological states as relationships, not just as isolated measurements. Instead of treating each gene as a standalone clue, it learns how genes, cell types, neighborhoods, and spatial arrangements combine to produce meaning.

This is especially powerful for omics because biological meaning is hierarchical. A gene participates in a module, a module in a cell state, a cell state in a tissue pattern, and a tissue pattern in an organismal outcome. If you only look at one layer, you may get the right numbers and the wrong interpretation.

Here is the most useful mental model: single sample scoring tells you whether a pattern is present, while foundation models help you understand the kind of context in which that pattern is operating. One is a detector. The other is a contextual interpreter. The combination is what turns noisy molecular data into actionable biology.

Think of a music app that can tell you a song’s tempo and key. Useful, but limited. Now imagine the app also understands genre, instrumentation, harmonic tension, and how the same chord progression can signal joy in one style and melancholy in another. That second layer is what context does. It does not replace the first measurement. It makes it intelligible.

In single cell and spatial omics, this means a model can begin to learn that the same phenotype signature may correspond to different biological stories depending on the neighborhood. An immune activation score in a region crowded with suppressive myeloid cells is not the same as the same score in a region rich in antigen presenting cells and infiltrating T cells. The score is the note. Context is the chord.


The real breakthrough: from phenotype to ecology

The strongest synthesis of these ideas is this: molecular phenotypes are not just states, they are ecological events.

An ecology is defined not merely by what exists, but by interactions, densities, gradients, boundaries, and dependencies. That is exactly how single cell and spatial data should be interpreted. A cell state becomes meaningful when seen as part of a local ecology of signaling, competition, cooperation, and physical arrangement.

Single sample scoring contributes a crucial capability here: it lets us estimate whether a phenotype exists in a specific specimen, especially when that specimen is unusual. Foundation models contribute a second capability: they let us compare that specimen not only to other samples, but to a learned space of cellular and spatial relationships. Together, they create a bridge from measurement to meaning.

This matters because many of the most important biological questions are ecological rather than purely molecular:

  1. Why does the same tumor behave differently across patients? Because the local cellular ecology differs, not just the mutational profile.

  2. Why does a biomarker predict outcome in one cohort but fail in another? Because the biomarker is not standalone. It may depend on tissue architecture, immune composition, or disease stage.

  3. Why do some signatures appear strong in bulk data but weak at the cell level? Because bulk data collapses ecological structure into averages, then mistakes compression for clarity.

Once biology is viewed ecologically, scoring and modeling become complementary rather than competing tasks. Scoring tells you what kind of ecological program is present. Context modeling tells you where it lives, what sustains it, and what it depends on.

The future of molecular interpretation is not better averages. It is better situational awareness.


A practical framework: the three layers of biological meaning

To make this synthesis useful, it helps to use a simple framework for interpretation. Every molecular phenotype can be read at three layers:

1. Presence

Is the phenotype detectable at all in this sample? This is the domain of single sample scoring. It answers the basic question of whether a pathway, signature, or program is active.

2. Placement

Where in the tissue or cellular network is the phenotype located? Is it diffuse, localized, peripheral, clustered, or adjacent to a specific cell type? This is where spatial omics matters. Placement often reveals whether a phenotype is pathological, compensatory, or merely incidental.

3. Dependency

What context makes the phenotype what it is? Which cells, signals, and structural features appear to sustain it? This is where a foundation model can help by learning contextual dependencies across many examples.

This framework prevents a common mistake: confusing a score with an explanation. A high score may tell you that a biological program is active, but not whether it is driving disease, responding to disease, or being mimicked by another process. Placement and dependency help distinguish those possibilities.

Consider an inflammatory signature in tissue. At the presence level, the signature is high. At the placement level, it may be concentrated around necrotic regions. At the dependency level, it may correlate with a particular immune neighborhood or vascular pattern. Only when the three layers align can you begin to talk confidently about mechanism.

This layered view also suggests a better way to build predictive models. Instead of asking one model to do everything, let scoring provide the local alarm, let spatial analysis provide the map, and let a contextual foundation model infer the grammar connecting them. That division of labor is not just elegant. It is robust.


Key Takeaways

  • Treat molecular scores as situational indicators, not universal truths. A phenotype score is most useful when interpreted in context.
  • Add placement to presence. If you only know that a program exists, you do not yet know whether it matters biologically.
  • Use context models to learn dependencies, not just correlations. The best models should reveal which local environments make a phenotype interpretable.
  • Think ecologically. Many disease states are better understood as interacting neighborhoods than as isolated abnormal cells.
  • Design analyses in layers. First detect, then localize, then explain.

The deeper lesson: biology is becoming a theory of context

The convergence of single sample scoring with single cell and spatial foundation models points toward a larger intellectual shift. Biology is moving away from the dream of a single decisive measurement and toward a more mature view: meaning emerges from context, and context is itself learnable.

That changes how we think about precision medicine. Precision is not just about finer measurement. It is about better inference under local conditions. A patient does not need a generic profile of what usually happens in a disease. They need an interpretation of what is happening in this tissue, at this time, in this neighborhood.

It also changes how we think about biomarkers. The best biomarkers may not be the strongest signals, but the signals whose meaning is most stable once context is accounted for. In other words, the ideal biomarker is not merely predictive. It is context aware.

This is the hidden connection between scoring a phenotype in one sample and training models on single cell and spatial data. Both reject the assumption that biological truth is easiest to find in averages. Both insist that the unit of meaning is not the isolated molecule, but the molecule embedded in a situation.

The most important question in modern biology may therefore not be, “What is the signal?” but, “What kind of world does this signal inhabit?” Once you start asking that, the data stops looking like a list of variables and starts looking like a living system.

And that is the real breakthrough: not a better catalog of parts, but a better understanding of how parts become meaningful only when they are in the right place, among the right neighbors, under the right conditions.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣