When the Signal Is Also the Noise: What Single Cell Statistics Reveal About Immune Checkpoints
Hatched by Miyabi
Aug 21, 2026
11 min read
3 views
88%
What if the cells most clearly marked as “dysfunctional” are actually the cells most actively responding to danger?
That question sounds biological, but it is also statistical. In single cell RNA sequencing, the measurements we treat as evidence of a cellular state are shaped by sampling, dropout, sequencing depth, normalization, and the error model used to interpret sparse counts. In autoimmune disease, the same problem appears in a different form: immune checkpoint molecules such as PD 1 and TIGIT may look like signs of restraint, yet their presence can identify CD4 positive T cells that are abnormally activated.
These are not separate puzzles. They are instances of the same deeper problem: a measurement can be both a consequence of a state and part of the mechanism producing that state. The marker is not merely a label attached to biology. It is an event inside biology, observed through an imperfect measurement system.
The practical consequence is profound. Whether we are analyzing a cell atlas or designing an intervention, we should stop asking only, “What does this marker mean?” We should ask three harder questions:
- What process produced the marker?
- What distortions affect our observation of it?
- How does changing the marker feed back into the system?
Once these questions are combined, statistical modeling and immunological intervention begin to look like versions of the same discipline: learning how to distinguish a meaningful signal from a context dependent response, without destroying the system that generated the signal.
The dangerous comfort of a clean label
Biology tempts us to think in labels. A cell is “activated,” “exhausted,” “regulatory,” or “dysfunctional.” A gene is “on” or “off.” A cluster is assigned a name, and the name quietly becomes an explanation.
But biological states are rarely categorical in the way our figures are categorical. A T cell can be activated and restrained at the same time. It can express checkpoint receptors because it is exhausted, because it has recently encountered antigen, because it is attempting to prevent tissue damage, or because several of these processes are occurring together.
The distinction matters in systemic lupus erythematosus. PD 1 and TIGIT are commonly associated with inhibitory regulation, so it is tempting to interpret their expression as evidence that a population is already suppressed. Yet PD 1 positive and PD 1 positive TIGIT positive CD4 positive T cells can represent abnormally activated cells in the lupus context. Their inhibitory receptors do not erase their activation history. They may instead record the immune system’s attempt to place limits on a response that has already become pathological.
This is a general pattern in complex systems. A fever is not simply “heat”; it is evidence of an active regulatory response to a threat. A financial market’s volatility is not merely random noise; it can be the visible consequence of participants reacting to uncertainty. A warning light on a machine is not proof that the machine is inactive. It may be proof that the machine is operating under dangerous stress.
A restraint marker is not necessarily evidence of restraint. It may be evidence that restraint has become necessary.
The same caution applies to single cell data. Sparse gene counts do not present us with direct access to a cell’s identity. They provide a partial, noisy sample of a rapidly changing molecular state. A gene absent from the observed transcriptome may still be active. A gene present in a few reads may be biologically important, technically favored, or both.
The label is therefore downstream of two processes: the cell’s biology and the measurement system’s limitations. If we ignore either one, we risk turning an observation into a false mechanism.
Error models are theories of what we cannot see
Statistical error models are often treated as technical plumbing. Researchers choose a model, run an analysis, and focus on the biological result. But an error model is more than a mathematical convenience. It is an explicit or implicit theory about why the observed data differ from the underlying biological reality.
In single cell RNA sequencing, counts are sparse and uneven. Some cells receive more sequencing reads than others. Some transcripts are captured more efficiently than others. Many genes appear as zero not because they are biologically silent, but because the relevant molecules were not sampled. Replicate structure, cell quality, and library preparation introduce further variation.
Different statistical models respond to these features in different ways. A model may treat the data as arising from a Poisson process, a negative binomial process, or a more elaborate mixture that accounts for excess zeros and variation across cells. These choices influence which genes appear variable, which cells appear similar, and which clusters seem biologically distinct.
Imagine looking at a city through a set of security cameras. One camera records every passing car but has poor color resolution. Another captures vivid colors but misses vehicles at night. A third records only intersections with heavy traffic. If we cluster neighborhoods by the footage, our map will reflect both urban life and the cameras’ biases.
Single cell analysis has the same structure. Observed variation is a mixture of biological variation and measurement variation. The central statistical task is not to eliminate all variation. It is to estimate which part is likely to be biology and which part is likely to be the observation process.
This becomes especially important when studying immune activation. Activation is not a single gene. It is a coordinated, time dependent program that may be distributed across many genes, and its expression can vary with cell cycle, cytokine exposure, tissue location, and recent antigen encounter. If the model overreacts to technical zeros, it may fragment one continuum into artificial clusters. If it smooths too aggressively, it may erase a meaningful rare population.
The problem is not solved simply by collecting more cells. More observations can increase confidence in a biased conclusion. A large dataset analyzed under a poor model can make an artifact look inevitable.
The correct mental model is a measurement funnel:
- A biological process occurs in the cell.
- That process produces molecules in a particular abundance and distribution.
- The experimental platform samples some of those molecules.
- The computational pipeline transforms the samples into features, clusters, and labels.
- The researcher interprets those labels as biological states.
At every stage, information can be lost or distorted. The final cluster is not the cell. It is the endpoint of a chain of transformations.
The checkpoint paradox: when inhibition signals activation
The PD 1 and TIGIT example adds a crucial layer to this statistical perspective. It shows that even a perfectly measured marker can be misinterpreted if we assume that its meaning is fixed across contexts.
A receptor is not a dictionary entry with one definition. Its meaning depends on what cell expresses it, what other receptors are coexpressed, what signals preceded its expression, and what functional consequences follow when the receptor is engaged or manipulated.
Consider two simplified cells. Cell A expresses PD 1 after prolonged antigen exposure and has lost much of its effector capacity. Cell B expresses PD 1 and TIGIT after intense, abnormal activation in lupus, while remaining part of a pathogenic immune response. The same surface marker appears in both cells, but its biological meaning differs because the surrounding state differs.
The error is a form of semantic confounding. We infer function from a name, rather than from a network. In machine learning terms, the marker is predictive only within a particular distribution. Change the disease context, tissue, activation history, or cell lineage, and the relationship between marker and function may change.
This is why interventions that target checkpoint pathways cannot be evaluated solely by asking whether receptor expression rises or falls. The more important questions concern system level outcomes: Does the balance between effector T cells and regulatory T cells shift? Does inflammatory activity decrease? Does the intervention reduce pathogenic function without collapsing protective immunity?
A dual activating membrane nanoparticle directed toward PD 1 and TIGIT, combined with dexamethasone, can be understood through this lens. The goal is not simply to remove a marker or silence every cell carrying it. The goal is to reshape a network in which activated effector cells and regulatory cells influence one another, while corticosteroid treatment changes the broader inflammatory environment.
The word “synergy” should therefore be used carefully. Two interventions are not synergistic merely because both reduce a measured marker. They are synergistic when their combined action produces a system level effect that is greater, more selective, or qualitatively different from what either could achieve alone.
Dexamethasone may broadly suppress inflammatory activity. A targeted membrane based intervention may alter signaling more selectively in receptor bearing cells. Their combination can be viewed as a division of labor: one changes the background conditions, while the other changes the behavior of a strategically positioned subpopulation.
This is analogous to statistical correction. A normalization procedure adjusts for broad technical differences, while a carefully chosen model preserves meaningful structure. The broad correction and the local inference work together, but neither is sufficient by itself.
A new framework: read the state, not the label
The connection between single cell error modeling and immune checkpoint biology suggests a practical framework for interpreting complex biological data. Call it state aware inference.
1. Separate observation from interpretation
First record what is directly observed. For a single cell experiment, that may be a sparse count matrix and a surface protein measurement. For an immune intervention, it may be receptor abundance, cytokine production, proliferation, or the ratio of effector to regulatory cells.
Only afterward should we assign a biological interpretation. “PD 1 positive” is an observation. “Suppressed” is an interpretation. “Cluster enriched for interferon response genes” is an observation. “Pathogenic cell type” is a stronger interpretation that requires additional evidence.
This separation sounds obvious, but it prevents many analytical errors. It forces us to distinguish what the data contain from what we want the data to mean.
2. Model the process that generates the signal
Ask why the marker or expression pattern appeared. Is it caused by stable lineage, recent stimulation, chronic exposure, cell migration, stress, technical sampling, or a combination?
For example, a checkpoint receptor may be induced as a protective feedback response. A zero count may arise from true absence or from failed capture. These possibilities should not be treated as interchangeable.
3. Look for conditional meaning
A marker’s value often lies in its interaction with other features. PD 1 alone may be less informative than PD 1 together with TIGIT, activation genes, regulatory markers, cytokine signatures, and functional assays. Similarly, a gene’s differential expression is more credible when it persists across donors, batches, tissues, and reasonable model specifications.
The key question is not “Is this marker associated with the state?” It is “Under what conditions does this association hold?”
4. Test intervention, not just classification
A classification becomes biologically persuasive when it predicts what happens after perturbation. If a population is truly pathogenic, altering its relevant pathway should change pathogenic function. If a checkpoint receptor is merely a passive label, targeting it may produce little selective effect.
Perturbation is the bridge between description and mechanism. It converts a static map into a causal test.
5. Preserve uncertainty as information
Uncertainty is not an embarrassing gap to hide in a polished figure. It identifies where the system is ambiguous and where additional measurement will have the highest value.
A rare cluster supported by one weak marker should be treated differently from a population replicated across donors and confirmed by independent assays. Likewise, a therapeutic effect should be evaluated with uncertainty intervals, dose response, and comparisons against each component alone.
The purpose of an error model is not to make uncertainty disappear. It is to make uncertainty legible.
What this changes in practice
For researchers analyzing single cell immune data, the framework has immediate consequences. Begin with sensitivity analysis: compare reasonable count models, normalization strategies, and clustering parameters. If a biological conclusion exists only under one set of assumptions, report it as model dependent rather than definitive.
Next, treat dropout and technical variability as hypotheses about the data generating process, not merely nuisances. Examine sequencing depth, cell quality, donor effects, and batch structure before interpreting a rare population. Validate important markers with orthogonal measurements, such as flow cytometry, protein assays, spatial information, or functional experiments.
For immunologists, avoid treating checkpoint expression as a one dimensional scale from “active” to “inhibited.” Examine coexpression, activation history, lineage, tissue, and function. A receptor can be both a brake and a footprint of acceleration. Its therapeutic significance depends on which role dominates in the disease context.
For clinicians and translational teams, evaluate combination therapies at the level of cellular balance and tissue outcome, not only molecular abundance. A useful treatment may leave some markers elevated while changing the behavior of the cells that express them. Conversely, a dramatic molecular change may have little clinical meaning if the pathogenic network remains intact.
Key Takeaways
- Treat every biological label as an inference, not a fact. “Activated,” “exhausted,” and “regulatory” describe interpretations that require contextual evidence.
- Distinguish biological variation from measurement variation. In single cell data, zeros, clusters, and differential expression patterns are shaped by both cell state and the data collection process.
- Interpret markers conditionally. PD 1 and TIGIT expression can indicate attempted restraint, recent activation, chronic stimulation, or pathological immune activity, depending on context.
- Use perturbation to test meaning. A marker becomes mechanistically convincing when changing its pathway changes the predicted cellular function.
- Define synergy at the system level. A combination is valuable when it reshapes the balance of immune populations or functions more effectively and selectively than either intervention alone.
The deepest lesson is that biological signals are not passive labels pasted onto cells. They are traces left by dynamic processes. Some traces record what a cell has experienced. Others reveal what the cell is trying to do next. Many do both.
That is why the most sophisticated analysis does not ask whether a signal is real in isolation. It asks how the signal was generated, how it was distorted during observation, and how the system will respond if the signal is changed.
A cell atlas is therefore not just a map of expression. It is a map of evidence, uncertainty, and hidden causes. An immune checkpoint is not just a brake. It may be the dashboard light warning that the engine has been running too hard. And a successful therapy is not necessarily the one that makes the warning light disappear. It is the one that changes the underlying motion of the machine.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣