Learning

Desirable Difficulties: Why Hard Learning Sticks

When studying feels smooth and confident, you're often learning less than you think. The moves that feel awkward, slow, and slightly frustrating are usually the ones building memories that last.

25 min read
Key Takeaways
    • Desirable difficulties are conditions that slow learning down and improve it: Robert Bjork coined the term in 1994 for manipulations that depress performance during practice while raising long-term retention and transfer.
  • Five have the strongest evidence, and a sixth is close behind: spacing, interleaving, retrieval practice, generation, and varied practice, plus productive failure (attempting a problem before you're taught the method).
  • Fluency is a liar: The feeling that you're learning (smooth re-reading, confident recognition, "I've got this") doesn't predict whether you'll remember the material next week.
  • Storage strength and retrieval strength aren't the same thing: A memory can be deeply stored yet temporarily hard to access. The struggle to access it is what strengthens future access.
  • Difficulty is only desirable inside a window: For novices on complex material, generating reverses and studying worked examples wins. Interleaving even runs backwards when the things you're mixing aren't confusable.
  • AI removes the difficulty by default: In a 2025 PNAS trial, students given an unrestricted GPT-4 tutor gained 48% during practice, then scored 17% below students who never had AI once it was taken away.

What Are Desirable Difficulties?

Desirable difficulties are learning conditions that make practice feel harder and slower while you're doing it, yet produce better retention and transfer weeks later. The paradox is that they look like they're failing while you're inside them, which is exactly why almost nobody chooses them.

The term comes from Robert A. Bjork at UCLA, who named the idea in 1994 and developed it over the following decades with Elizabeth Ligon Bjork. Their lab kept hitting the same result: the manipulations that depressed performance during study were often the ones that improved it after. The best-known summary is a 2011 chapter by Elizabeth and Robert Bjork whose title states the principle outright, "Making things hard on yourself, but in a good way."

Five conditions have decades of replication behind them:

  • Spacing: Distributing study across separate sessions instead of massing it into one.
  • Interleaving: Mixing different problem types or topics within a session instead of blocking them.
  • Retrieval practice: Pulling information out of memory instead of re-reading it in.
  • Generation: Producing an answer before you're shown it.
  • Varied practice: Practicing under changing conditions instead of identical ones.

A sixth, productive failure, grew up in a separate research tradition and belongs in the same family. It gets its own section below.

What unites them is that each one forces your brain to reconstruct something rather than recognize it. Recognition is cheap and feels like knowledge. Reconstruction is expensive and builds it.

Adding friction is not the goal, though, and this is where the idea gets misapplied. A difficulty only counts if it triggers extra processing you're able to complete. Difficulty that simply blocks you isn't desirable. It's an obstacle. Here's what doesn't qualify:

Looks like a difficultyActually desirable?Why not
Studying in a noisy cafe you can't concentrate inNoThe difficulty is unrelated to the material. It consumes attention rather than encoding anything
Reading a badly written textbookNoThe struggle is with the prose, not the concept
A problem several levels beyond your trainingNoYou can't produce anything, so there's no retrieval to strengthen
Interleaving topics you don't yet understand individuallyNoNothing to discriminate between yet
Re-reading a chapter for the third timeNoNot a difficulty at all. It raises fluency and leaves storage untouched

The Fluency Illusion

Open a textbook chapter you've already read twice. Run your eyes down the page. The sentences feel familiar. Each paragraph clicks into place. You close the book convinced you know the material.

A week later, you can barely retrieve the main argument.

Cognitive psychologists call this the fluency illusion or illusion of competence, and it's the single biggest reason most study time is wasted. When information processes smoothly, your brain interprets the smoothness as evidence of mastery. It isn't. Smoothness is just smoothness.

The data is brutal. Dunlosky and colleagues' 2013 review in Psychological Science in the Public Interest sorted ten common study techniques into three tiers by how useful the evidence showed them to be. Highlighting and re-reading, two of the most popular methods on Earth, landed in the lowest tier. Practice testing and distributed practice landed in the highest. The methods learners love produce weak learning. The methods learners avoid produce strong learning.

Re-reading is the clearest case. Callender and McDaniel (2009) ran a series of experiments on immediately re-reading a text and found little or no benefit on later tests, with only a few exceptions. Familiarity isn't memory. Recognition isn't recall.

The illusion is structural. Your brain uses ease of processing as a heuristic for "I know this." The heuristic just happens to be wrong for predicting future retrieval. Learning to distrust your sense of mastery is the first step toward studying things that stay studied.


What Bjork Actually Discovered

Robert A. Bjork and Elizabeth Ligon Bjork have spent forty years untangling why this happens. Their 1992 chapter "A New Theory of Disuse and an Old Theory of Stimulus Fluctuation" introduced the framework that explains all of it.

The theory splits memory into two separate dimensions.

Storage strength measures how deeply a memory is wired in. It's a function of how much you've engaged with the material across different contexts and over time. Storage strength can only go up.

Retrieval strength measures how easily you can pull that memory up right now. It fluctuates wildly. It rises when you've just studied something, falls without use, and depends heavily on the cues available in the current moment.

The fluency illusion lives in the gap between these two. When you re-read a chapter, retrieval strength shoots up because the material is right in front of you. Storage strength barely moves. The moment retrieval strength fades, there's nothing underneath to support recall.

Bjork's second move gave the field its name. In 1994 he argued that increases in storage strength come specifically from retrieving information when retrieval is difficult, not from re-presenting information when it's easy. Difficulty during practice, paradoxically, is what creates lasting learning. Hence: desirable difficulties.

Soderstrom and Bjork's 2015 paper, "Learning versus Performance: An Integrative Review," tightened the distinction further. Performance is what you can do during practice, today. Learning is the relatively permanent change in your ability to do it later, in a new context. Most study habits maximize performance and undermine learning.

Two decades of classroom work later, the Bjorks returned to the topic in a 2020 article in the Journal of Applied Research in Memory and Cognition, and their tone had shifted. The mechanisms were well established. Getting anyone to act on them was not. They admitted they had been "prone to assuming, unrealistically, that simply telling learners and teachers about relevant research findings is enough," when learners arrive with naive theories about studying that have to be dislodged first. Quoting Rohrer and Hartwig in the same issue: "Too often, the classroom is where promising interventions go to die."

The takeaway sits in the center of everything that follows. If your session feels effortless, retrieval strength is doing the work and storage strength is going nowhere. If your session feels effortful in a productive way (you're searching, failing, recovering, reaching), storage strength is climbing. The discomfort is the deposit.

This is what ties together the methods Glasp readers already know. Active recall and spaced repetition, along with the Feynman technique, the protégé effect and the blurting method, aren't unrelated tricks. They're different ways of forcing your brain into the effortful work that builds storage strength. Desirable difficulties is the meta-principle. Everything else is implementation.

Feels Easy but Doesn't WorkFeels Hard but Works
Re-reading the same chapterClosing the book and writing what you remember
Massed practice (cramming)Distributed practice over days or weeks
Studying one topic to mastery before moving onInterleaving multiple topics in a single session
Reading worked examples on material you already graspGenerating answers before checking
Practicing the exact same problem typeMixing problem types and contexts
Highlighting whole paragraphsHighlighting sparsely and writing your own marginalia

Spacing: Why Distributed Study Beats Cramming

The spacing effect is the oldest and most replicated finding in this whole literature. Hermann Ebbinghaus described it in 1885. It still works.

The claim is simple. If you have ten total minutes to study a piece of material, you'll remember more of it on a delayed test by splitting the ten minutes across several sessions than by spending all ten in one block. The total time's identical. The distribution is what matters.

Cepeda and colleagues' 2006 synthesis in Psychological Bulletin pulled together 839 assessments from 317 experiments across 184 articles, all of them verbal recall tasks. Spacing beat massing consistently, and the gap that produced the best retention grew as the retention interval grew. A follow-up experiment by the same group in 2008 pinned down the ratio: the optimal gap is a shrinking proportion of how long you want to remember something, running near 20 to 40% of a one-week horizon but only around 5 to 10% of a one-year horizon. In plain terms, to hold something for a month, review every few days. To hold it for a year, let the gaps stretch into weeks.

The mechanism connects directly to the storage/retrieval split. When you study something twice in a row, the second exposure happens while retrieval strength is still high, so the practice is essentially free and storage gains are minimal. When you study it again three days later, retrieval strength has decayed. Pulling the memory up takes work, and that work is what buys the storage.

In practice, spacing fights two enemies: cramming and unscheduled review (which never happens). The fix is calendar-level. Pick the things you actually want to retain, give each a recurring review slot, and trust the schedule over your sense of what needs more attention. That sense is, again, mostly fluency talking.


Interleaving: Why Mixing Beats Blocking

Interleaving means mixing different topics or problem types within a single study session, rather than studying one to mastery before starting the next. If you're learning algebra, you don't do twenty quadratic problems in a row. You do a quadratic, then a system of equations, then a function transformation, then back to a quadratic.

It feels worse. Performance during the session drops. Learners regularly rate interleaved practice as less effective even after they've measurably learned more from it. This metacognitive mismatch is one of the strongest illusions in the field.

Doug Rohrer and Kelli Taylor demonstrated it in 2007 with a study whose title says the finding: "The Shuffling of Mathematics Problems Improves Learning." College students learned to solve several kinds of problem, then practiced either blocked by type or randomly mixed. Tested a week later, the mixed group's performance was vastly superior. Rohrer, Dedrick and Stershic replicated it in a real classroom in 2015, with 126 seventh graders working through interleaved or blocked assignments over three months, and got the same result on a delayed test.

Why does it work? Two mechanisms. First, interleaving forces discrimination: you can't just apply the same procedure on autopilot, you have to figure out which procedure each problem calls for. That discrimination is the skill you actually need on a test or in real work. Second, interleaving spaces each topic by definition: the gap between problem-type-A items is filled with problem-type-B items, so each return to A involves real retrieval.

The discrimination mechanism also tells you where interleaving stops working, and this is the one place the popular write-ups get sloppy. Brunmair and Richter's 2019 meta-analysis in Psychological Bulletin, "Similarity matters," pooled 59 studies and 238 effect sizes. Interleaving won overall, moderately (Hedges' g = 0.42), but the benefit split hard by material: biggest for visual categories like paintings, smaller for math problems, and for plain word lists it reversed, with blocking winning outright. The pattern is what the theory predicts. If the categories you're mixing are confusable, juxtaposing them teaches you to tell them apart. If they have nothing in common, there's nothing to discriminate and you pay the switching cost for free.

For self-directed learners, interleaving is easier than it sounds and harder than it looks. You don't need a fancy scheduler. You just need to refuse the urge to "finish" a topic before moving on. Read a chapter on stoicism, then a chapter on probability, then a chapter on UI design, then back to stoicism. Your brain will protest. Your brain is wrong. Our full guide to interleaving practice covers the scheduling mechanics.


Retrieval Practice: The Testing Effect

If you only adopt one desirable difficulty, make it this one. Retrieval practice (the testing effect) is the act of pulling information out of memory rather than pushing it back in. It's the engine behind active recall, flashcards, the blurting method, and most of what works in education.

Henry Roediger III and Jeffrey Karpicke's 2006 paper in Psychological Science, "Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention," is the canonical demonstration. Students studied prose passages, then either restudied them or took free-recall tests, with no feedback either way. Final tests came at five minutes, two days, or one week. At five minutes, the restudiers won. By one week, the testers had roughly 50% higher retention. Same total time, different work: every retrieval, even a failed one, modified the underlying memory in a way re-presentation couldn't.

The mechanism is retrieval-induced reconsolidation. When you successfully pull a memory up, the neural pattern that supports it gets re-encoded with whatever cues are currently active. That re-encoding strengthens the pattern and broadens the cue set that can trigger it later. When you fail and then study the answer, the failed search itself primes you to encode the correction more deeply. Richland, Kornell and Kao named and demonstrated that in 2009, calling it the pretesting effect.

Retrieval works in dozens of forms: closed-book recall, flashcards, free-form summary writing, teaching aloud, the Feynman technique. What they share is the underlying move: pull it out of memory before you look.

After you've highlighted a chapter with Glasp's web highlighter, close the source and try to reconstruct the argument from your highlights alone. Then open the page and check. The gap between what you produced and what's there is your storage-strength deficit. Closing that gap is the work.


Generation: Produce the Answer First

The generation effect is a close cousin of retrieval practice, distinct enough to deserve its own line. Slamecka and Graf (1978) showed across five experiments that learners who generated material themselves remembered it better than learners who simply read the same material.

Their paradigm was rule-based and almost embarrassingly simple. One group read word pairs outright, like "lamp - light." The other was handed a rule (associate), the stimulus word "lamp," and the single letter "l," and had to produce "light" themselves. The generators won by a wide margin, even though the reading group had seen the answer.

The principle scales up. Reading a worked proof builds less competence than attempting the proof and then comparing. Reading someone's book summary teaches less than writing your own. Watching someone code teaches less than typing the code yourself, hitting an error, and figuring out why.

Generation is desirable difficulty in pure form. It deliberately withholds information you'd be happy to receive, forcing you to produce it. The frustration of half-remembering, half-guessing is the mechanism. By the time you check your answer, you've already done the work that converts exposure into encoding.

Highlighting can be a generation tool or a passive one. Yellow-bombing every paragraph is passive: you're outsourcing judgment to a future re-read. Sparse, deliberate highlighting is generative: you're forcing present-you to commit to "this matters more than that," which is itself synthesis. The science of highlighting goes deeper.


Varied Practice: Change the Conditions

Varied practice means practicing the same skill in slightly different forms, contexts, and conditions, rather than under identical conditions every time.

Kerr and Booth's 1978 study with children throwing beanbags became the textbook example. They ran two age groups, 8-year-olds and 12-year-olds. One group practiced from a single distance. The other practiced from a mix of distances that bracketed the target but never included it. The varied group outperformed the fixed group on the test throw they'd never practiced. They'd built a more general motor representation.

Note the design detail, because it's usually dropped: the varied distances surrounded the target. Variability isn't a license for arbitrary noise. It works when the variations span the space the skill has to generalize across.

Cognitive learning shows the same pattern. Vary the wording of definitions you study. Vary the contexts in which you encounter a new term. Read about a concept in two different fields rather than two different chapters of the same textbook.

Variability supports transfer: the ability to use what you've learned outside the conditions in which you learned it. Transfer is what most learners want and what most study habits actively prevent. If you practice a skill in one form, you encode the form along with the skill, and you'll struggle when the form changes. Variability decouples the skill from any single form.


Productive Failure: Struggle Before the Explanation

Manu Kapur at ETH Zurich calls it productive failure: give learners a hard problem before teaching them the method, let them generate their own flawed attempts, and only then deliver the instruction.

The attempts usually fail. That's the design. What the failed attempts do is build a map of the problem space, so that when the canonical solution arrives, the learner has somewhere to put it and can see what their own approach was missing.

Sinha and Kapur's 2021 meta-analysis in Review of Educational Research pooled 166 experimental comparisons. Problem-solving before instruction beat instruction before problem-solving on conceptual understanding and transfer (Hedges' g = 0.36), without costing anything on procedural knowledge, and studies that followed Kapur's design principles most closely reached g = 0.58. The evidence base is heavily STEM, and the effect tracks age: it was strongest from about sixth grade upward and slightly negative for the youngest children, who don't yet have the problem-solving footing to fail productively.

The reading version is easy to run. Before you open an article on a topic you half-know, spend two minutes writing what you think the answer is. Then read. Your version is almost always wrong in an interesting way, and the gap between it and the author's is what you'll remember.

The six at a glance:

DifficultyMechanismConcrete ExampleKey Study
SpacingForces real retrieval as memory decaysReview notes on day 1, 3, 7, 21 instead of four times todayCepeda et al. (2006)
InterleavingForces discrimination between problem typesMix algebra, geometry, and stats in one sessionRohrer & Taylor (2007); Brunmair & Richter (2019)
Retrieval PracticeRe-encodes memory with new cuesClose the book, write what you remember, then checkRoediger & Karpicke (2006)
GenerationProducing forces deeper encoding than readingPredict the answer before reading the explanationSlamecka & Graf (1978)
Varied PracticeBuilds context-independent representationsSolve the same concept in 3 different domainsKerr & Booth (1978)
Productive FailureFailed attempts build a map the explanation lands inAttempt the problem before you're taught the methodSinha & Kapur (2021)

When Difficulties Become Undesirable

Not all difficulty helps, and the research on when it stops helping is more specific than most summaries admit.

A difficulty becomes undesirable when the learner can't engage with it productively. If you can't decode the words at all, slowing your reading further won't help. If you don't know what a derivative is, interleaving derivative and integration problems is noise. If your retrieval practice is so far past your storage strength that you produce nothing, you're not retrieving, you're flailing.

Cognitive load theory gives this a mechanism. Working memory can juggle only a handful of interacting elements at once. When material already has high element interactivity (lots of pieces that only make sense in relation to each other), adding generation or interleaving on top pushes total load past capacity, and the extra difficulty buys nothing because there's no spare capacity to process it with.

Chen, Kalyuga and Sweller demonstrated the crossover directly in a 2015 study in the Journal of Educational Psychology, using geometry instruction. On low-element-interactivity material, the generation effect showed up as expected: generating answers beat being shown them. On high-element-interactivity material it flipped, and studying worked examples beat generating. Then in a second experiment with more knowledgeable learners, the effective complexity dropped and generation won on everything. That's the expertise reversal effect. The correct study strategy is not fixed. It moves as you get better.

Bjork's framing holds all of this together: there has to be enough scaffolding underneath for the struggle to land.

The right zone is where you can produce something, even if it's incomplete. You're reaching, not falling. You should be failing on some attempts and succeeding on most. Getting everything right means the material is too easy or the gap too short. Producing almost nothing means you've crossed the boundary and need scaffolding, not grit.

So adjust in both directions. When something's genuinely beyond you, don't romanticize the struggle: read the explanation, get the foundation in place, then come back. When something feels too easy, don't let the comfort fool you: increase the spacing, mix in harder variants, or move to a more generative format.

Picture storage strength on one axis and retrieval strength on the other. High storage / high retrieval is recently practiced and well-learned. High storage / low retrieval is where retrieval practice does the most good, because the search is hard but the deposit is real. Low storage / high retrieval is the dangerous one: stuff you just re-read and feel you know but haven't built. Cramming lives here. Low storage / low retrieval is genuinely new material, where you need scaffolding first.


Do AI Tools Remove Desirable Difficulties?

By default, yes. An AI assistant is an extraordinarily efficient machine for deleting exactly the effort that builds storage strength, and the first good evidence on what that costs arrived in 2025.

Hamsa Bastani, Osbert Bastani and colleagues ran a randomized controlled trial, published in PNAS, with nearly a thousand high school math students at a large school in Turkey. Students got one of two GPT-4 tutors during practice sessions. "GPT Base" behaved like the ordinary ChatGPT interface. The researchers prompted "GPT Tutor" with safeguards, most importantly to offer teacher-designed hints rather than hand over answers.

During assisted practice, both worked spectacularly. GPT Base students improved 48% over controls. GPT Tutor students improved 127%.

Then the researchers took the tools away and ran an exam.

Students who'd practiced with GPT Base scored 17% worse than students who never had AI at all. Not worse than they'd looked with the tool in hand: worse than the control group that never got one. The GPT Tutor group came out statistically indistinguishable from controls, with a point estimate near zero.

That's Soderstrom and Bjork's performance/learning dissociation, reproduced with a chatbot, at school scale. The tool that helped most during practice hurt most afterwards. And notice the fix wasn't banning it. The fix was making the AI withhold the answer, which is the generation effect wearing a different hat.

Adult knowledge work shows the same shape. At CHI 2025, Lee and colleagues at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real examples of AI-assisted tasks. Most reported reduced cognitive effort, and the correlation that matters ran this way: higher confidence in the AI predicted less critical thinking, while higher confidence in themselves predicted more. Trust in the tool substitutes for thought.

A widely shared MIT Media Lab study by Kosmyna and colleagues points the same direction, with EEG recordings showing the weakest neural connectivity in participants who wrote essays using an LLM, and those participants struggling to quote work they'd just produced. Treat it as suggestive rather than settled: it's still a preprint, only 54 people took part (18 of them in the final session), and other researchers have already posted a detailed methodological critique of its sample size and analysis.

None of this is an argument against using AI. It's an argument about where in the loop you put it.

Use that deletes the difficultyUse that preserves it
"Summarize this article for me" before reading itRead it, write your own summary, then ask the AI to find what you missed
"Solve this problem""I got this answer this way. Where does my reasoning break?"
Asking for the explanation firstPredicting the answer first, then asking for the explanation
Generating flashcards and reading themGenerating flashcards and answering them closed-book
Accepting the first responseAsking it to argue the opposite case, then judging between them

The rule of thumb: let the AI grade you, not answer for you. Do the retrieval yourself, produce something, and only then bring in the model to compare your output against the source. You keep the difficulty and outsource the marking. Our piece on AI and learning works through the broader tradeoff, and AI study modes compared looks at which assistants now withhold the answer by default.


Designing Your Routine for Desirable Difficulties

Knowing the principles doesn't make them automatic. You have to engineer them in, because the path of least resistance will always be the comfortable, low-storage option. Here's how to bake the difficulties into a self-directed learning system.

Highlight to generate, not to mark. Mark sentences sparingly, then write a brief note in your own words explaining why this passage matters and how it connects to what you already know. The sparse selection is discrimination. The note is generation. Highlighting without notes is the passive trap.

Use AI chat for retrieval, not for explanations. The wrong way to use an AI assistant is to ask it to summarize a chapter you haven't read. The right way is to read the chapter, close it, write your own summary, then paste that into Glasp's AI chat and ask it to grade your reconstruction against the source. You did the retrieval. The AI does the comparison.

Space with Kindle highlights, not memory. Your sense of when to review is broken: it's run by recency and emotion, not by storage curves. Schedule reviews on a calendar, pull up highlights from a book you read three weeks ago, and try to reconstruct the argument before scrolling. The mechanical schedule is what protects spacing from your fluency-driven instincts.

Interleave with the community feed. Don't read four articles in a row from the same person on the same topic. Cycle: cognitive science, then entrepreneurship, then writing, then back to cognitive science a day later. Discrimination across domains is a stronger workout than depth in one.

Watch a YouTube Summary, then test yourself. Close the page and write three things you'd want to remember in a year. Compare to the transcript. The gap is where your work is.

Predict before you read. Run the productive-failure move on anything you open cold. It changes what you notice.

Vary the format. Read the same idea in a book, a paper, a thread, and a video. Each medium encodes the idea differently, and your brain has to abstract across them. That's the point of varied practice.

The system view lives in our companion piece on building a learning OS. Desirable difficulties is the principle. The OS is the daily mechanics.


Frequently Asked Questions

What are the five desirable difficulties?

Spacing (distributing study over time), interleaving (mixing topics or problem types), retrieval practice (recalling instead of re-reading), generation (producing an answer before seeing it), and varied practice (changing the conditions of practice). Productive failure, attempting a problem before you're taught the method, is a closely related sixth. All six share one mechanism: they force reconstruction rather than recognition.

What is an example of a desirable difficulty?

Closing a book and writing a summary from memory is the cleanest example. So is reviewing on a spaced schedule instead of cramming, mixing problem types in one session, or predicting an answer before you read the explanation. The test is whether the difficulty comes from the material and how you engage with it, and whether you can still produce something. Studying in a distracting cafe fails that test, because the difficulty is unrelated to the content. So does a problem far beyond your current level, because you can't engage with it productively.

Where did the term "desirable difficulties" come from?

Robert A. Bjork coined it in 1994. The accessible summary most people cite is a 2011 chapter by Elizabeth and Robert Bjork, "Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning," in Psychology and the Real World. Its core argument: conditions that slow acquisition and make performance look worse during practice often produce better long-term retention and transfer, and learners systematically misread that short-term dip as failure. The Bjorks revisited the topic in the Journal of Applied Research in Memory and Cognition in 2020, mostly to note how hard the ideas are to get adopted.

Is highlighting a desirable difficulty?

It depends entirely on how you do it, and most highlighting isn't. Marking whole paragraphs just defers the decision to whoever re-reads the page later. Highlighting sparsely, one or two passages per page, and writing a note in your own words for each one is both a discrimination task and a generation task. Treat your highlights as future retrieval cues rather than a substitute for reading.

Should I struggle on every problem?

No. The goal is productive struggle, not flailing. If you're producing nothing, you've gone past the desirable zone into overload. Back up, get the foundation in place, then return to the difficulty. You want to be failing sometimes and succeeding most of the time. And for genuinely complex material you're new to, the research says studying worked examples first beats generating: Chen, Kalyuga and Sweller (2015) found that on high-element-interactivity content, worked examples beat problem-solving until learners had enough expertise for the effect to flip back.

Does using ChatGPT hurt learning?

It depends entirely on where you put it in the loop. In the 2025 PNAS trial, students who practiced with an unrestricted GPT-4 tutor later scored 17% below students who never had AI, while students whose tutor gave hints instead of answers showed no such penalty. Ask a model to do the retrieval for you and you lose the learning. Do the retrieval yourself and use the model to check your work, and you keep it.

How long should I space my reviews?

Longer than feels comfortable, and longer each time. The gap that maximizes retention grows with your target horizon while shrinking as a proportion of it, so a month-long horizon wants a few days between reviews and a year-long one wants weeks. When a review feels easy, lengthen the next gap. When it feels hard, shorten it. A schedule beats your gut either way.

Do desirable difficulties apply to skills as well as facts?

Yes, possibly more strongly. Variability research came from motor learning. Interleaving was studied in sports and music before textbooks. Any skill with discrimination, transfer, and retrieval components benefits: coding, writing, design, language, instruments, athletic technique. The form of practice changes; the principle doesn't.

Why does my school still teach the easy way?

Because the easy way looks better in the short run. Massed practice and re-reading produce higher scores on quizzes given right after instruction, and worse scores a month later. Most schools don't measure delayed retention. Soderstrom and Bjork (2015) made exactly this point: confusing performance with learning is structural, not personal. As a self-directed learner, you don't have to wait for institutions to catch up.


Conclusion

The principle to walk away with: if your studying feels easy, you're probably not studying. Effort is the price of durable learning, and the brain pays out only when you've earned it through spacing, interleaving, retrieval, generation, variability, or productive failure.

This doesn't mean grinding harder, it means grinding smarter. Most learners are already putting in the time. They're just spending it on activities that maximize fluency and minimize storage. Swap one re-read for a closed-book reconstruction. Swap one cram session for four spaced reviews. Swap one block of identical problems for a mix. Each swap trades short-term comfort for long-term retention.

The stakes went up in 2026. Every tool on your desk now offers to do the effortful part for you, and the evidence says taking that offer costs you the learning even while it flatters your performance. The good news from the same research: a tool that withholds the answer produces no such penalty. You get to choose which one you're using.

Glasp was built around this trade: sparse highlighting, marginal notes, AI-graded reconstructions, spaced reviews of past reads, and a community feed that interleaves topics by default. Each is a small piece of friction designed to convert exposure into storage.

Tomorrow, pick the easiest thing in your routine and replace it with the harder version. That swap is the whole game.

Start building your knowledge library

Highlight what matters as you read across the web. Save insights from articles, books, and YouTube videos in one place.

Get Started Free

Or highlight this page as you read it