Why Measuring Empathy Is an Experimental Problem: Treating the Body as a Variable, Not a Costume
Hatched by Thomas Hirschmann
Apr 16, 2026
9 min read
8 views
82%
Opening question: can you run a lab experiment that truly measures empathy?
Imagine two people reading the same heartbreaking story. One says they feel moved. The other says they do not. Which answer is true? Now imagine testing both people on a new interface that claims to increase empathy. You randomize them into two groups, run the task, compute a p value, and publish the result. The finding is either that the interface works or it does not. End of story.
That neat story hides a collision between two realities. On one side is the modern experimental toolkit: protocols, random assignment, within subjects designs, counterbalancing, and statistical thresholds. On the other side is the messy fact that empathy is rooted in bodily signals: heartbeat awareness, breathing, gastrointestinal cues, the felt sense we call interoception. These internal signals shape perspective taking in ways that make classic experimental controls feel inadequate.
This article argues that to study and design systems that influence empathy we must stop treating the body as noise. We should treat it as an experimental variable. Doing so requires a new experimental grammar: one that pairs rigorous controls with direct measurement and manipulation of interoceptive states, and that anticipates learning, transfer, and asymmetry in human experience. The result is not only more valid science; it is a pathway to building technology that genuinely supports perspective taking.
Why the body matters for perspective taking: the tension at the heart of empathy research
We tend to speak about empathy and perspective taking as cognitive operations: one imagines another person, maps their beliefs, infers emotions. That view assumes the mind is a clean information processor and that empathy can be measured like reaction time. But human social understanding is not purely top down. Interoception, the perception of internal bodily states, is a core component of empathy. Heart rate, breathing rhythm, and visceral sensations help disambiguate emotions and orient attention toward another person.
This creates an uncomfortable tension for experiments. Experimental design asks for control: hold everything constant except the independent variable, randomize, and replicate. Yet interoception is variable across people and across moments, and it can be altered by the very interventions we test. A decision support system, interface, or training program might not simply change cognitive strategy. It might alter arousal, breathing, or bodily attention, which in turn changes perspective taking. If we ignore those signals, we risk mistaking bodily learning for cognitive learning, and we risk building systems that look effective in a lab but fail in real life.
Concrete example: imagine testing a new empathy training interface that provides subtle haptic pulses meant to simulate another person’s heartbeat. If participants in the first condition receive the heartbeat cue and then perform better on a perspective taking task, what caused the improvement? The interface may have taught strategy. It may have taught the story structure, which leads to transfer. Or it may have shifted interoception, making participants more attuned to bodily cues and thereby better at reading emotional states. Standard experimental choices about assignment and counterbalancing will shape how we interpret that improvement, and a naive design will conflate these possible mechanisms.
The experimental problem made concrete: three ways classic designs can mislead
-
Learning confounds that mimic treatment effects. When people repeat tasks, they learn. Within subjects designs are powerful because they control for between person variability. But they are vulnerable when learning is asymmetrical. If exposure to condition A teaches participants domain knowledge that transfers to condition B more than the reverse, counterbalancing fails. The measured effect may be a skill transfer effect rather than an effect of the manipulation.
-
Unobserved bodily state changes as hidden mediators. Experiments that measure only task performance and self report miss internal shifts. An interface that increases physiological arousal may create short term improvements that do not reflect durable perspective taking. Without measuring physiology we cannot distinguish learning from transient bodily entrainment.
-
Binary notions of success driven by p value thresholds. The p value summarizes how surprising the data is given the null hypothesis at a predetermined confidence level. It does not tell us whether the observed change came from a shift in internal state, a change in strategy, or a measurement artifact. Relying on p value thresholds without richer measurement encourages false closure and undermines replication.
These three problems are not theoretical curiosities. They are the practical reasons why many interventions in social and human computer interaction research fail to replicate when taken out of the lab.
A framework for experiments that take the body seriously: Signal, Context, Interpreter
To reconcile experimental rigor with embodied complexity I propose a minimal experimental grammar called Signal, Context, Interpreter. It is a way to design, run, and interpret experiments that aim to change or measure empathy.
Signal: the experimental manipulation you apply. This can be an interface feature, a training module, a haptic cue, a breathing exercise, or a feedback channel. Be explicit about whether the signal targets cognition, the body, or both. If the signal is bodily, measure the body.
Context: the background conditions that determine how signals are received. These include prior task learning, order effects, social framing, and environment. Counterbalancing, washout periods, and randomization are tools to manage context, but they must be applied with the expectation that learning can be asymmetrical.
Interpreter: the internal processes through which people transform signal plus context into behavior. Interoception is a core interpreter. Cognitive strategies, prior knowledge, and affective state are others. The interpreter is what we aim to observe and, if possible, manipulate.
Use this grammar as a checklist when designing an experiment:
- Explicitly state which part of the system you consider the dependent variable: performance on a task, self reported empathy, physiological synchrony, or some combination.
- Decide if the signal targets the body or the cognition. If the body is targeted or plausibly involved, add physiological measurement and manipulation to the protocol.
- Choose a design that anticipates learning and transfer. When skills or bodily states can carry across conditions, consider between subjects designs or include washout periods and asymmetric counterbalancing strategies tuned to the likely direction of transfer.
Protocol prescription: a practical experimental recipe for embodied empathy interventions
Below is a practical protocol you can adapt. It balances the need for replication with the reality that bodies change.
-
Preregister the protocol and the analysis plan. State the primary outcome, whether you will treat interoception as a mediator, and what constitutes a successful replication.
-
Choose your basic design using the Signal, Context, Interpreter lens. If the manipulation plausibly teaches the task, favor between subjects designs or include long enough washout intervals. If the manipulation targets bodily attention and bodily effects are short lived, a within subjects design with randomized order and physiological baseline control may be preferable.
-
Measure the body. At minimum record heart rate and breathing rate. Better still, include measures of vagal tone, skin conductance, or simple heartbeat detection tasks. These measures help distinguish cognitive learning from bodily entrainment.
-
Include manipulation checks that probe the interpreter. If you hypothesize increased interoceptive awareness, include both objective heartbeat detection and subjective interoception scales. If you hypothesize strategy change, include tasks that reveal reasoning steps rather than only final accuracy.
-
Counterbalance with intent. Standard counterbalancing assumes symmetric transfer. Assess likely asymmetry before choosing an order. When asymmetry is expected, structure orders to isolate the asymmetrical transfer, or use a fully between subjects design.
-
Use richer statistical reporting. Report effect sizes and confidence intervals, not just p values. Consider Bayesian models that explicitly estimate the probability of the hypothesis given the data. When bodily measures are included, use mediation analysis to test whether interoception explains the treatment effect.
-
Plan for replication that varies context. If your protocol depends on subtle bodily cues, test it across lab environments and time delays. A robust effect should survive changes in context or reveal where context matters.
Concrete example of a study using this protocol
Goal: test whether heartbeat feedback improves perspective taking.
Signal: a wearable that provides participants with a tactile pulse matched to a confederate's heartbeat while they read short social vignettes.
Context: two sessions separated by one week. Randomize participants to either receive heartbeat feedback in session one and no feedback in session two, or the reverse. Include a between subjects control group with no feedback in either session.
Interpreter measures: objective heartbeat detection task, self report interoception scale, task performance on perspective taking items, reaction time, and physiological synchrony with the confederate.
Design choices: include washout tasks between conditions, counterbalance order only when preliminary pilots show symmetric transfer, otherwise use a mixed design with a between subjects control arm.
Analysis: preregister the primary outcome as change in perspective taking from baseline to session, use mediation analysis to test whether change in heartbeat detection explains the effect, and report effect sizes and Bayesian posterior probabilities rather than only p values.
Why this matters for designers, researchers, and practitioners
Treating the body as a variable changes both the claims you can make and the products you can build. If you design empathy tools without measuring or controlling bodily states you risk two failures:
- False positives: the tool appears to work because it temporarily entrains bodily signals, but it does not lead to durable perspective taking or transfer to other contexts.
- False negatives: the tool actually changes bodily attunement, but your outcome measure did not capture the right interpreter, leading you to conclude the tool failed.
Designers can use embodied experimental design to produce clearer evaluations and better products. Researchers can produce results that are more replicable and interpretable. Practitioners who care about long term social change can know whether an intervention actually shifts how people tune into each other, or simply trains them to game a test.
Empathy is not a black box. It is an ecology of signals, contexts, and interpreters. To change it responsibly we must attend to the body as both source and measure of knowledge.
Key Takeaways
- Treat the body as an experimental variable: If your intervention plausibly alters arousal or bodily attention, measure physiology and include it in your analysis.
- Choose design based on transfer risk: Use within subjects when participant variability is the main threat; use between subjects or include washout when asymmetrical learning is likely.
- Make mediation explicit: Preregister whether interoception is a hypothesized mediator and include objective and subjective measures to test that claim.
- Move beyond binary statistics: Report effect sizes, confidence intervals, and consider Bayesian analyses to express the strength of evidence in context.
- Preregister and replicate across contexts: Replication that systematically varies environment and time reveals whether an effect is body dependent or generalizable.
Closing thought: rebuilding trust between rigor and lived experience
Experimental science and lived social life are often spoken about as if they occupy different domains. That separation explains why many promising empathy interventions sputter when scaled or transferred. The solution is not less rigor. It is a richer rigor: experimental protocols that accept embodied signals as legitimate data, that anticipate learning and asymmetric transfer, and that make hypotheses about inner states explicit and testable.
If you want to build systems that help people take another person’s perspective, design your experiments as if the body mattered. Your results will be truer, your claims will be more robust, and the systems you build will be more likely to work when someone uses them outside the lab.
Imagine the alternative: a future where empathy technologies are optimized only to perform well on narrow tests. That future looks efficient but brittle. The alternative is a science that respects the stubborn complexity of being human. That science begins by measuring the body, not ignoring it.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣