The Molecule Is Not the Target: What Protein Design Teaches Us About Disease Biomarkers

Emil Funk Vangsgaard

Hatched by Emil Funk Vangsgaard

Aug 30, 2026

11 min read

94%

0

What if the most important mistake in molecular biology is also the most intuitive one: assuming that a disease target is a single molecule?

A protein is often treated as though it were a lock waiting for the right key. In practice, it is more like a city. It has districts, temporary structures, hidden rooms, shifting boundaries, and different populations depending on the time of day. A drug must reach the right neighborhood. A biomarker must reveal which version of the city exists, where it is located, and what has happened to it.

This distinction connects two seemingly separate activities: designing proteins from scratch with machine learning and developing biomarkers for diseases involving TDP 43, a protein central to amyotrophic lateral sclerosis and frontotemporal lobar degeneration. Both are exercises in interface engineering. Both ask which molecular features can be recognized, stabilized, blocked, or measured. And both expose the same deeper truth:

The real object of molecular design is not a sequence. It is a biologically meaningful state that a sequence can occupy.

Once this is understood, computational protein design and biomarker development stop looking like opposite directions. One constructs new molecular states. The other learns to distinguish pathological states from healthy ones. The common challenge is deciding which features matter, which are accidental, and which are visible only under particular conditions.

From sequences to states

Modern protein design workflows create an extraordinary illusion of directness. Begin with a sequence model, generate hundreds or thousands of candidates, predict their structures, and rank them according to properties such as thermal stability, aggregation tendency, electrostatic potential, hydrophobicity, or binding geometry. The process can move from an abstract sequence to a plausible molecular object in a few hours.

But each step is really a translation between different descriptions of the same candidate. A sequence is translated into a predicted fold. A fold is translated into a surface. A surface is translated into a possible interaction. The final question is not whether the sequence looks protein like. It is whether the molecule performs a particular function in a particular environment.

That last qualification is decisive. A designed enzyme that folds correctly but aggregates is not a useful enzyme. A binder that touches the intended target but binds the wrong surface is not a useful therapeutic. A candidate with excellent predicted stability may still fail because a flexible loop closes the active site, because a cellular partner competes for the same surface, or because the protein never reaches the compartment where it is needed.

The computational pipeline therefore works best when viewed as a series of uncertainty reductions, not as a machine that discovers truth. Sequence generation expands the search space. Structure prediction removes implausible folds. Stability and aggregation filters remove additional failures. Interface analysis identifies candidates that deserve experimental attention. Every stage narrows possibility, but none establishes biological meaning by itself.

The same logic applies to TDP 43. Its identity cannot be reduced to one uninterrupted string of amino acids. It contains a nuclear localization signal, a nuclear export signal, two RNA recognition motifs, a glycine rich domain, and multiple sites of modification. Disease associated material can also appear as C terminal fragments of different lengths, including fragments associated with insoluble fractions from affected brains.

If one asks, “What is the TDP 43 biomarker?” there may be no single correct answer. The relevant signal could be full length protein, a particular fragment, a phosphorylation pattern, a change in cellular location, a change in solubility, or a combination of these features. The biomarker is not simply TDP 43. It is TDP 43 in a disease relevant state.

The hidden common problem: choosing the right interface

A drug designer often begins with a functional hypothesis: block a protein interaction, stabilize a vulnerable structure, or create a new catalytic activity. In the PCSK9 example, the proposed intervention focuses on a region involved in binding the LDL receptor. Specific residues in that region become hotspots. A computational scaffold is then designed around the desired contact zone, sequences are generated to fit the scaffold, and predicted complexes are inspected.

This is not merely a search for a protein that binds PCSK9. It is a search for a protein that binds the right place on PCSK9, with sufficient stability and appropriate physical properties to remain useful. Binding is necessary, but it is not sufficient. A molecule that binds an irrelevant patch may be biologically inert. A molecule that binds the correct patch weakly may be clinically useless. A molecule that binds strongly but aggregates may create a new problem.

Biomarker development has an analogous interface problem. An antibody or assay does not measure “the protein” in the abstract. It recognizes an epitope, a structural arrangement of residues, a fragment boundary, or a chemical modification. The measured signal depends on whether that feature is exposed, preserved during sample preparation, present in the relevant compartment, and sufficiently different from healthy background.

This suggests a useful framework with three layers:

  1. Identity: Which molecular entity is present? Full length TDP 43, a C terminal fragment, or a modified form?
  2. Geometry: What three dimensional or local structural feature is exposed? An antibody may recognize a sequence in one conformation but not another.
  3. Context: Where is the entity, what is it bound to, how soluble is it, and what cellular process produced it?

Many failed molecular tools solve only the first layer. They identify a protein name but not the state that carries the biological information. The more valuable tools solve all three.

A successful binder does not merely recognize a molecule. It recognizes a molecule under the conditions that make the molecule matter.

This is why the computational evaluation of aggregation, charge, hydrophobicity, and stability is so important. These properties are not decorative annotations. They determine whether a designed object can exist long enough, in the right environment, to reach its intended interface. In diagnostic work, solubility, fragment composition, localization, and chemical modification play a similar role. They determine whether a disease state can be detected rather than merely imagined.

Why pathological proteins defeat simple recognition

TDP 43 makes the problem especially vivid because disease can alter not only how much protein exists, but also what kind of object it becomes. A protein that normally moves between cellular compartments may be redistributed. A protein that normally participates in RNA regulation may become concentrated in abnormal assemblies. A protein that exists as a complete chain may be processed into fragments. Post translational modifications may change its behavior and its detectability.

A conventional assay often assumes that more signal means more protein. But disease biology may instead involve a redistribution of forms. The total amount of TDP 43 could be less informative than the ratio between soluble and insoluble material, nuclear and cytoplasmic material, full length and fragmented material, or unmodified and modified material.

This is a general principle of biomarker design: abundance is only one coordinate in molecular state space. A better mental model is a vector:

a = quantity b = molecular form c = chemical modification d = location e = physical behavior

A disease signature may emerge not from any one coordinate, but from their combination. For example, a modest increase in a specific C terminal fragment accompanied by altered phosphorylation and increased insolubility may be more informative than a large but nonspecific increase in total protein.

Machine learning can help here, but only if the task is framed correctly. A sequence model trained to generate plausible enzymes learns patterns in biological sequences. A structure model estimates whether a sequence can adopt a particular shape. A classifier trained on biomarker data may identify combinations of features associated with disease. None of these systems automatically knows which features are causal, which are correlated, and which are artifacts of sample handling.

The danger is particularly acute when the target is dynamic. A model may learn that a certain fragment is associated with disease because all disease samples were processed in one laboratory and all control samples in another. It may learn a technical signature rather than a biological one. Similarly, a structure prediction may look convincing because a candidate resembles known proteins, while the actual functional interface remains wrong.

The remedy is not to distrust computation. It is to design the workflow around orthogonal evidence. A candidate biomarker should survive changes in assay format, sample preparation, cohort composition, and measurement method. A candidate therapeutic protein should survive tests of folding, aggregation, binding, activity, and performance in the relevant cellular environment.

In both cases, confidence should come from independent agreement among different measurements, not from a single impressive prediction.

Design the experiment around failure, not novelty

The excitement of generative biology naturally pulls attention toward what can be created: a new enzyme, a new binder, a new diagnostic reagent. Yet the highest value often comes from anticipating how the creation could fail.

A protein engineer can ask:

  • Does the sequence fold into the intended structure?
  • Does the structure remain stable at the working temperature?
  • Does the surface promote unwanted self association?
  • Does the candidate bind the intended hotspot rather than a nearby decoy site?
  • Does it function in the chemical environment where it will be used?

A biomarker developer can ask parallel questions:

  • Is the measured form genuinely associated with disease rather than tissue damage in general?
  • Is the epitope accessible in the native sample?
  • Does the assay distinguish full length protein from disease associated fragments?
  • Does sample handling preserve the relevant soluble or insoluble state?
  • Does the signal improve diagnosis, prognosis, patient stratification, or treatment monitoring?

These questions reveal a broader design principle: the best workflow is organized by failure modes. Instead of ranking candidates only by predicted success, assign each candidate a failure profile.

For a designed binder, the profile might include poor folding confidence, high aggregation risk, weak interface complementarity, and uncertain expression. For a TDP 43 assay, it might include epitope masking, cross reactivity, degradation during collection, and inability to distinguish disease associated forms.

This approach changes the role of computation. Models are no longer treated as oracles that select a winner. They become instruments for allocating experimental effort. A candidate with moderate predicted binding but low aggregation risk may be more valuable than one with spectacular interface scores and severe developability concerns. A biomarker with a weaker raw signal but excellent specificity across sample conditions may be more useful than a high signal marker that collapses outside one carefully controlled experiment.

The practical objective is not maximum prediction. It is maximum decision quality per experiment.

A field guide to molecular state engineering

The intersection of generative protein design and neurodegenerative biomarker research yields a compact operating system for molecular discovery.

1. Define the state before designing the tool

Do not begin with “I need a binder to this protein” or “I need an assay for this biomarker.” Specify the state that matters: location, fragment, modification, conformation, oligomeric status, or interaction partner. A target defined only by name is usually underspecified.

2. Separate recognition from usefulness

Recognition asks whether a molecule can attach to something. Usefulness asks whether that attachment changes a relevant outcome or measures a clinically meaningful state. Evaluate both explicitly. A binder can be specific but therapeutically irrelevant. An assay can be sensitive but diagnostically unhelpful.

3. Treat physical properties as part of function

Stability, aggregation, charge, hydrophobicity, solubility, and expression are not secondary engineering details. They govern whether the intended interaction happens in reality. A molecular tool that cannot survive its environment has no practical function.

4. Build orthogonal validation into the first design cycle

Use different methods to interrogate different claims. Structure prediction should be paired with biophysical measurements. Binding should be paired with functional inhibition or activation. A TDP 43 signal should be compared across fragment sensitive assays, modification sensitive assays, and measures of solubility or localization.

5. Preserve uncertainty instead of hiding it

A ranked list is more honest and more useful when it displays why each candidate is uncertain. Record predicted structure confidence, interface ambiguity, aggregation risk, assay accessibility, and experimental gaps. This creates a better feedback loop for the next round of design.

The next revolution will be about context

Generative models have made it easier to propose molecular objects. The harder and more consequential task is learning which objects remain meaningful when exposed to biology.

That requires a shift from designing molecules to designing molecular states and their measurements. The future therapeutic may not be selected because it has the most elegant predicted fold. It may be selected because its fold, surface chemistry, stability, localization, and interaction profile remain useful in the messy environment of a cell. The future biomarker may not be the most abundant disease associated molecule. It may be the feature that most reliably distinguishes a harmful state from a harmless one across time, tissue, and treatment.

The apparent gap between creating a protein that does not exist in nature and detecting a fragment produced by a diseased brain is therefore narrower than it seems. Both demand the same discipline: identify the biologically decisive interface, model the surrounding context, anticipate physical failure, and test the claim in the world rather than only in silico.

Biology does not reward molecules for being plausible. It rewards them for being consequential in context.

The deepest promise of machine learning in molecular science is not that it will eliminate experiments. It is that it will make experiments more intelligently targeted. And the deepest lesson from biomarker development is that the target is rarely the molecule alone. It is the molecule, its form, its location, its history, and the consequences of its interactions.

Once we begin to design for that complete state, molecular discovery becomes less like finding a key and more like learning the conditions under which a lock exists at all.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣