Glasp Research

What do millions of highlights reveal about how people actually read and learn? Glasp publishes research based on its own platform data, a public study series on social highlighting at a scale lab studies cannot reach. Every paper ships with a reproducibility bundle — the code and the derived data behind each number — and most of them with a public repository too.

Disentangling Answer Engine Optimization from Platform Growth: A Difference-in-Differences Natural Experiment

A natural experiment measuring whether AEO changes (llms.txt, structured data) move AI-assistant referrals independent of a platform's underlying growth.

Personal Salience: Highlighting Is Social, but Individuality Lives in Selection

What people highlight within a document is mostly shared across readers; an individual's extra within-document signal is tiny (≈ +0.017 AUC).

Selection, Not Salience: The Shape and Limits of Personalization in Social Highlighting

Individuality lives in which documents a reader chooses, not which sentences: document-selection personalization gains ≈ +0.13 versus ≈ +0.017 for in-document salience.

Factions Within, Uncertain Across: Within-Document Reader Sub-Groups in Social Highlighting

Readers split into strong within-document factions (z ≈ +6.3), but faction membership barely carries across documents, so it cannot be recovered from a user's history.

The Long Tail, Not the Front Page: Cold-Start Prediction of Crowd Highlight Salience

Crowd highlight salience on brand-new documents is predictable (+0.04 AP over lead-position baselines), and the advantage concentrates on long-tail passages.

Trait, Not State: The Durability of Reading Identity in Social Highlighting

A reader's interest profile frozen after six months keeps its full predictive power years later: reading identity behaves like a durable trait, not a passing state.

Measuring Alignment With Reader Highlights Net of Position and Length

Crowd highlights sit early in a page and run long, so anything that favours early or long sentences looks aligned with readers. Netting both out, a language model's importance ranking reaches +0.196 excess agreement with the crowd — and a single human reader reaches +0.182.

Language Models Agree With Each Other, Not With Readers

Two language models pick the same sentences far more often than two readers do: across 18 model arms from 11 vendors, the median model pair reaches +0.093 excess agreement against a human yardstick of +0.040, and no model agrees with readers more than a reader does. The human side is 2,523 mark sets from people highlighting for their own reasons, not for a study.

Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?

A benchmark score means nothing without a floor and a ceiling, so we measured both. Half the crowd predicts the other half at +0.203 AP over lead position; surface features recover 5% of that headroom, the best zero-shot frontier model 35–53%, and an unweighted cross-vendor fusion 60% — confirmed on 217 independent documents. An 8B model distilled from that fusion reaches statistical parity with the strongest single frontier model.

Why we publish

Most research on highlighting comes from small lab studies. Glasp's platform captures how people highlight in the wild, at scale, across years. Publishing what we find, including the negative results, keeps our product decisions honest and gives the learning-science community data it cannot get anywhere else. Questions or collaboration ideas: [email protected].