Learning

Cognitive Load Theory: A Reader's Guide

Cognitive load theory is the best-tested account of why some material teaches you and some just exhausts you. Almost every guide to it addresses teachers. This one addresses the person reading something nobody designed.

18 min read
Key Takeaways
    • Working memory is the bottleneck, but only for new material: Cowan's 2001 review put the limit near four chunks, not seven, and Sweller's contribution was to notice it binds only on novel information. Anything already in long-term memory costs almost nothing, which is why the same page is easy for an expert and impossible for a beginner.
  • There are two loads that matter, not three: Sweller, van Merriënboer and Paas revised the theory in 2019. Germane load stopped being a third category and became the redistribution of resources from wasted work toward the actual learning.
  • Extraneous load is the only load you can cut: Intrinsic difficulty is fixed by the material and your prior knowledge. Everything you can actually fix is presentation: split attention, redundancy, and information that vanishes before you've processed it.
  • You can run the theory on yourself: In one accounting experiment, learners taught to reorganize badly designed material outscored learners handed a professionally integrated version (Sithole et al., 2017). Nobody is optimizing the web for you, so this is the version that counts.
  • Highlighting is a load-management move, not a memory trick: Across 103 studies and 12,201 learners, signaling improved retention and transfer, and cut measured load in the studies that tested for it (Schneider et al., 2018), though in most of those studies the designer applied the signal, not the reader. It fails when you mark so much that nothing stands out.

What Cognitive Load Theory Actually Says

Cognitive load theory says that learning fails when the working memory demands of a task exceed what working memory can handle. Your working memory is severely limited when it processes new information, and essentially unlimited when it draws on what you already know. So the job of good instruction, and of good self-directed reading, is to spend that narrow capacity on the material itself instead of on the work of navigating how it was presented.

John Sweller laid this out in a 1988 paper in Cognitive Science called "Cognitive Load During Problem Solving: Effects on Learning." His argument was counterintuitive at the time. Conventional problem solving, the kind where you work backward from the goal and keep chipping away at the gap between where you are and where you need to be, eats so much processing capacity that almost nothing is left for building the mental structures that constitute learning. Students could solve the problems and still learn nothing from them.

The capacity limits themselves weren't new. Miller published his magical-number-seven paper in 1956, and Peterson and Peterson showed how fast unrehearsed material is lost in 1959. Nelson Cowan's 2001 reconsideration put the working figure closer to four chunks than seven. Sweller's team supplied the part everyone had left implicit: those limits apply only to novel information. Material you've already organized and stored in long-term memory moves into working memory as a single unit and costs almost nothing.

That asymmetry is the engine of the whole theory. It explains why an expert skims a paper you have to reread four times, and it explains why the advice that works for the expert is often actively bad for you.


The Three Types of Cognitive Load, and Why One Got Demoted

If you've read anything about cognitive load, you've met the three categories: intrinsic, extraneous, and germane. Most explanations online still describe them the way the 1998 version of the theory did, as three things that add up to a total. That version is out of date, and the correction matters.

Intrinsic load is the difficulty built into what you're learning. Not word count, not page count, but how many pieces of the idea have to be held in mind simultaneously because they only make sense in relation to each other. You can't reduce intrinsic load without either changing what you're trying to learn or changing how much you already know.

Extraneous load is everything the presentation makes you do that has nothing to do with understanding. Hunting for the figure a paragraph refers to. Rewinding a video because a number went by too fast. Reading a caption that repeats the sentence above it. This is imposed by design choices, and it's the one you can attack.

Germane load was originally defined as the effort that goes into actually building understanding, treated as a third quantity contributing to the total. In "Cognitive Architecture and Instructional Design: 20 Years Later" (Educational Psychology Review, 2019), Sweller, van Merriënboer and Paas changed that. In their words, germane load "refers to the working memory resources that are devoted to dealing with intrinsic cognitive load rather than extraneous cognitive load."

The reason for the change is a nice piece of scientific housekeeping. Under the old model, if germane load simply replaced extraneous load whenever extraneous load fell, total load should stay flat. But study after study measured total load dropping when extraneous load was cut. So the authors reframed germane load as having "a redistributive function from extraneous to intrinsic aspects of the task rather than imposing a load in its own right."

Here's what that means for you. Germane load takes care of itself the moment you stop spending capacity on presentation. You don't manufacture deep processing on top of a badly built page. You clear the junk out of the way and the capacity flows to the material.

Load typeWhat causes itCan you change it?Your move
IntrinsicComplexity of the material relative to what you already knowOnly by learning prerequisites or narrowing scopeBuild background first, or split the topic into smaller pieces
ExtraneousHow the material is presented and what it makes you doYes, and this is where all the leverage isReorganize, integrate, annotate, convert transient to permanent
GermaneNot a separate load since 2019Not directlyCut extraneous load and this takes care of itself

Element Interactivity: Why Difficulty Is Personal

The technical name for what makes something hard is element interactivity: the number of pieces that have to be processed together because you can't understand any of them in isolation.

Sweller's team uses a lovely example. Take the word "characteristics." For you, reading this sentence, it's one element. You pull it from long-term memory whole, without noticing. For someone learning to read English, that same word can be, in the authors' phrasing, "multiple squiggles" that must be held and combined in working memory at once, "a very high element interactivity task that overwhelms working memory."

Same ink on the same page, wildly different load, and the only thing that changed is the reader.

People consistently get the consequence wrong. Difficulty is not a property of a document. Any measure of complexity that ignores who is reading is, as the 2019 paper puts it bluntly, "largely useless." When a paper feels impossible, that is real information about the gap between its element interactivity and your current schemas, and not a verdict on your intelligence or on the writing.

Element interactivity also produces the most practically useful finding in the whole theory: the expertise reversal effect (Kalyuga, Ayres, Chandler and Sweller, 2003). Support that helps a novice starts to shrink in value as expertise grows, then disappears, then goes negative. Worked examples, detailed scaffolding, and step-by-step guidance all follow this curve. The summary that made a topic click for you in week one is redundant clutter by week ten, and processing it costs you.

So the honest answer to "what's the best way to study this?" is that it depends on where you already are. That is the finding, not a dodge. It's also why the desirable difficulties research and cognitive load theory sound contradictory but aren't. Desirable difficulties add load that does learning work. Extraneous load adds work that does nothing. The skill is telling them apart.


The Load Effects That Change How You Read

Over about forty years, cognitive load researchers named a long list of effects. Most came out of classrooms and instructional design labs. A handful transfer directly to somebody reading a PDF at 11pm, and those are worth knowing by name.

EffectWhat it saysWhere you meet itWhat to do
Split-attentionLearning drops when related information sits apart and you must integrate it yourselfFigure on page 4, discussion on page 7. Footnotes. Definitions in an appendixBring them together: annotate in place, put the note on the passage
RedundancySaying the same thing twice in two formats hurts more than it helpsSlides read aloud verbatim. A summary that restates the paragraph above itSkip one version instead of dutifully processing both
Worked exampleStudying a full solution beats attempting the problem, for novicesTutorials, walkthroughs, annotated codeStudy examples early, switch to problems once they feel obvious
ModalityNarration plus visuals beats on-screen text plus visualsVideo explainers, lecture recordingsListen while watching the diagram, don't read a caption over it
Transient informationMaterial that disappears imposes extra load because you must hold itVideo, audio, animation, anything autoplayingConvert it to something that stays put: a transcript, notes, stills
Expertise reversalThe effects above weaken, vanish, then reverse as you gain expertiseRe-reading beginner guides in a field you now knowDrop the scaffolding deliberately as you improve

Two of these deserve a number. Schroeder and Cenkci pulled the split-attention and spatial contiguity literature together in a 2018 meta-analysis in Educational Psychology Review: 58 independent comparisons, 2,426 learners, and an overall effect size of g = 0.63 favoring integrated designs. That's a solid medium effect for something as mundane as putting the label next to the thing it labels.

The redundancy effect surprises people, because it violates the intuition that more explanation can't hurt. It can. The 2019 review notes the effect "has been discovered, forgotten and rediscovered over many decades," pointing back to a 1937 report that young children learning to read nouns did better with the words alone than with the words plus similar pictures. Extra material isn't free. Processing it costs the same capacity as processing the thing you needed.


The Self-Management Effect

This is the part the teaching literature skips, and it's the reason this article exists.

Sweller built cognitive load theory to tell instructional designers how to build materials. That assumes somebody is designing your materials. For most of what you read now, nobody is. The 2019 review says so directly. Ideally students would only meet material built with cognitive load in mind. In reality, the authors write, the Internet "enables information to be created and shared by anyone," which makes it likelier that students hit "low-quality learning materials that have not been designed with any consideration of cognitive load."

Their answer is the self-management effect: the finding that learners can be taught to apply cognitive load principles to material themselves, and that doing so works.

The experimental design is clean. Students get material in a split-attention format, the bad version where text and diagram sit apart. One group learns how to reorganize it themselves, a second group receives a properly integrated version somebody else built, and a third gets the bad version with no instruction. Then everyone is tested on new split-attention material in a different subject.

Roodenrys, Agostinho, Roodenrys and Chandler ran an early version in Applied Cognitive Psychology in 2012 and established the basic result: you can teach students to manage split attention themselves, and the skill carries to material they haven't seen. Sithole, Chandler, Abeysekera and Paas pushed it further with 123 undergraduates learning introductory accounting, published in the Journal of Educational Psychology in 2017. Students taught to self-manage their attention, by reorganizing text and diagrams so they didn't have to keep searching between them, outperformed both the conventional split-attention group and the group given the pre-integrated version, on recall and on transfer.

Sit with that second comparison. The students who fixed the bad material themselves scored higher than the students who were handed the good material. If that holds up, doing the reorganizing is worth something in itself.

Two limits are worth stating plainly, because the rest of this article leans on them. The effect has only been studied with split-attention materials, so extending it to every kind of badly presented information is an inference, not a result. And the specific finding that self-managers beat an expert-integrated version rests on that single accounting study.

Even with those caveats, the effect reframes what active reading is for. Annotating is the intervention that carried the effect in these experiments, and it's the one you can run on anything: a whitepaper, a Wikipedia article, a thread, a lecture recording, a manual written by someone who has never once thought about your working memory.


Why Highlighting Is a Load Move, and When It Backfires

Highlighting has a bad reputation in study-advice circles, mostly downstream of the Dunlosky et al. (2013) review that rated it "low utility." We've written about what that review actually said and why the verdict is about how people typically highlight rather than about highlighting itself. Cognitive load theory explains the mechanism underneath both the success and the failure.

In the multimedia learning literature, drawing a reader's attention to what matters is called signaling, and it's one of the better-supported design principles there is. Schneider, Beege, Nebel and Rey's 2018 meta-analysis in Educational Research Review covered 103 studies and 12,201 participants across 46 years of research. Signaling improved retention and transfer, and, importantly for our purposes, measured cognitive load went down.

One honest caveat: in most of those studies the designer applied the signal, not the learner. That's the gap the self-management research starts to fill. When you highlight, you're producing your own signal, and the deciding is part of the benefit.

So what is highlighting doing in load terms? Three things:

  • It reduces search. A page you've marked has a structure you can re-enter without rereading. Search costs are pure extraneous load.
  • It forces classification. Deciding what gets marked, and in which of your multiple highlight colors, is a judgment you can only make by processing the meaning, which is exactly where you want your capacity going.
  • It anchors the note to the source. A note written next to the passage it comments on is the split-attention fix. A note in a separate app is a split-attention problem you built yourself.

The failure mode falls out of the same theory. Marking most of a page signals nothing, since contrast is the whole mechanism. Worse, you've now created a redundant second copy of the text that you'll dutifully reread later, paying full processing cost for material you already covered. Over-highlighting manufactures extraneous load where there was none.


Video Is the Hard Case: Transient Information

Video is the format cognitive load theory is hardest on, and given how much people now learn from it, this is the finding that gets quoted least.

Leahy and Sweller named the transient information effect in Applied Cognitive Psychology in 2011. Transient information is anything that appears and then vanishes: speech, animation, video. With permanent information like a page of text and diagrams, the 2019 review notes, everything "is available to the learner at the same time and may be revisited when needed." With transient information, you have to actively hold what just went by while processing what's arriving now, and that holding is extraneous load.

This is why a twenty-minute explainer can leave you with less than ten minutes of reading the same content, and why the feeling of understanding while watching is so unreliable. Your attention was fine. The format taxed capacity that the text version wouldn't have.

The research also names the fixes:

  • Self-pacing. Mayer and Chandler (2001) found that learners who controlled the pace of an instructional animation scored higher on transfer tests, though not on retention. Pausing is a documented countermeasure, not a sign you're slow.
  • Segmentation. Spanjers, Wouters, van Gog and van Merriënboer (2011) found that breaking an animation into parts with pauses between them was more efficient, meaning the same test scores for less mental effort, for learners new to the material. For learners who already knew it, the benefit disappeared, which is the expertise reversal effect again.
  • Make it non-transient. The most complete fix is to convert the transient stream into something that sits still. A transcript turns a video into a document you can scan, search, mark, and return to.

That last one is why getting a transcript counts as a genuine learning technique and not just a convenience. YouTube Summary pulls the transcript with timestamps and lets you highlight it directly, which converts transient information into permanent information and applies signaling to it in one pass.

Note the interaction, though: effects like modality and segmentation don't reliably show up for information that isn't transient in the first place, and segmentation fades as you gain expertise. If you're watching something easy, or something you half know already, none of this is worth the trouble. Save the machinery for the hard stuff.


Does AI Lower Your Load or Just Move It?

An AI summary looks like the perfect extraneous-load reduction. It takes a sprawling document and hands you a compact one. Sometimes that's exactly what it is, and sometimes it's the intrinsic processing being deleted rather than the extraneous work.

Researchers have started separating the two. Zhu, Li, Dong, Chang and Fan surveyed 589 Chinese university students and early-career knowledge workers across three waves in Frontiers in Psychology in July 2026, distinguishing two ways of leaning on generative AI. Dependent offloading means treating the model as a substitute for your thinking: accepting outputs with minimal scrutiny and shipping them as the product. Autonomous offloading means treating outputs as a starting point, comparing them against your own reasoning, and keeping ownership of the result.

Both groups reported the same short-term relief, around r = 0.22 each. What happened next is where they split. Dependent offloading was associated with handing over cognitive agency (β = 0.35) and lower intrinsic motivation (β = -0.23), while autonomous offloading showed no agency transfer and higher intrinsic motivation (β = 0.28). The authors are careful that this is correlational evidence about self-reported appraisals rather than measured ability, so read it as a hypothesis with numbers attached. Even at that strength, it's a useful one: same tool, same immediate relief, different reported relationship to your own thinking.

Cognitive load theory offers a mechanism. Germane load, in the 2019 sense, is working memory resources aimed at intrinsic load. If AI removes extraneous work, searching, formatting, hunting for the relevant section, your capacity gets redistributed to the material and you learn more. If AI does the intrinsic processing, there's no load left to redistribute, because the part that would have built understanding never ran in your head.

The practical test is a single question: after the AI does its thing, is there still something left for me to think about?

  • Load reduction: using AI to locate the three paragraphs that matter in a forty-page report, then reading those three yourself.
  • Load deletion: reading the summary and skipping the report.
  • Load reduction: asking for the definitions of unfamiliar terms so you can follow the argument.
  • Load deletion: asking what the argument means and taking the answer.

Asking questions against your own highlights is the version that keeps the thinking with you, because you've already done the selection. That's the design of Glasp's AI chat: the corpus is what you chose to mark, so the model helps you interrogate your reading instead of replacing it. For the broader picture on where this goes wrong, see the AI thinking trap.


A Load-Managed Reading Workflow With Glasp

Here's what load management looks like on a normal article that nobody designed for you.

1. Set intrinsic load before you start. Check honestly whether you have the prerequisites. If a paper assumes a concept you don't hold, no amount of clever annotation fixes that; the element interactivity is simply too high. Go get the prerequisite first. This is the one load you handle by changing the material or changing yourself.

2. Cut the presentation costs first. Reader mode, close the other tabs, kill autoplaying video. Every one of those is measurable extraneous load, not productivity theater.

3. Signal as you go. Use Glasp's web highlighter to mark the small number of sentences that carry the argument. Stay selective enough that the marks still mean something, and let color carry a category so the choice requires actual comprehension. If you learn a lot from video, how to learn from YouTube covers the same moves for lectures and explainers.

4. Kill split attention at the moment you notice it. When a passage refers to a figure, a definition, or something ten pages back, write the connection into a note attached to that highlight. That's the self-management move from the accounting experiment, applied to the web.

5. Make transient material permanent. For a video, pull the transcript through YouTube Summary and highlight there. For a book, bring your Kindle highlights in so the passages sit alongside everything else instead of being stranded on a device.

6. Come back to the signals, not the source. A week later, review the highlights instead of rereading the article. Rereading is largely a redundancy tax; retrieving from your own marks is active recall. Adding a visual to a highlight with structure gives you a second memory trace, which is the dual coding move.

7. Borrow other people's schemas. Kirschner, Paas and Kirschner described the collective working memory effect in 2009: a group can be treated as a single information processing system with a larger shared working space, where a gap in one person's knowledge gets filled by another's. That's the underlying logic of reading in public. Glasp's community shows what other readers marked on the pages you're reading, which is a cheap way to see the structure an expert saw. The 2019 review is honest that coordination has "transaction costs," so this pays off on genuinely complex material and not on everything.

None of it requires new habits so much as noticing which load you're currently paying, and refusing to pay the one you don't have to.


Frequently Asked Questions

What is cognitive load theory in simple terms?

Cognitive load theory says learning fails when a task demands more working memory than you have. Working memory only holds a few new items at once, so anything past that produces no learning however hard you try. Good instruction, and good self-directed reading, spend that scarce capacity on the ideas and not on the packaging.

Who developed cognitive load theory?

John Sweller, an Australian educational psychologist at the University of New South Wales. The roots go back to work with Levine in 1982, and the first full statement was his 1988 paper "Cognitive Load During Problem Solving: Effects on Learning" in Cognitive Science. Jeroen van Merriënboer and Fred Paas are his main long-term collaborators, and the three of them co-authored the major 1998 and 2019 reviews.

What is an example of cognitive load theory?

Take the word "characteristics." For a fluent reader it's a single element pulled whole from long-term memory, at almost no cost. For someone learning to read English it's a pile of interacting squiggles that has to be assembled in working memory, which can swamp it. Identical text, opposite difficulty, and the only variable is the reader's prior knowledge.

What are the three types of cognitive load?

Intrinsic load is the inherent complexity of the material given what you already know. Extraneous load comes from how the information is laid out, and you can cut it. Germane load was originally a third category, but Sweller, van Merriënboer and Paas revised it in 2019: it now describes working memory resources being redistributed toward intrinsic load rather than a load in its own right. In practice there are two loads to manage.

What is the difference between intrinsic and extraneous cognitive load?

Intrinsic load comes from the material itself: how many pieces of an idea you have to hold together at once, given what you already know. Extraneous load comes from the packaging: hunting for a figure referenced three pages back, rereading a caption that repeats the paragraph above it, rewinding a video because a number went past too quickly. The difference matters because only one of them is yours to cut. You lower intrinsic load by learning the prerequisites or narrowing the topic. You lower extraneous load by reorganizing the material, and that's where the leverage is.

How many items can working memory hold at once?

Four, not seven, and only for material you don't already know. Miller's famous 1956 figure was seven plus or minus two; Cowan's 2001 review revised it down to roughly four chunks under conditions that stop you grouping items. The qualifier is what people miss, and it's Sweller's contribution rather than Cowan's: an expert reading in their own field is barely using that budget at all.

How do you reduce cognitive load when studying?

Attack extraneous load, because it's the only kind you can cut directly. Put related information together instead of jumping between places (split-attention). Skip duplicated explanations instead of processing both (redundancy). Convert video and audio into transcripts and notes so nothing disappears before you've used it. Study worked examples before attempting problems while you're still a novice.

Is cognitive load theory still valid?

Yes, it's one of the more durable frameworks in educational psychology, and it's still being revised, which is a sign of health, not trouble. The 2019 update in Educational Psychology Review reworked the germane load category, consolidated the theory's grounding in evolutionary psychology, and added effects discovered after 1998. Individual effects like split-attention have held up in meta-analysis (Schroeder and Cenkci, 2018).

Does cognitive load theory contradict desirable difficulties?

No, though they're easy to confuse because both feel identical while you're in them. Use this test: ask where the effort went. Struggling to retrieve a fact you half remember builds the memory, so the effort landed on the material. Struggling to find which page the mislabeled figure lives on builds nothing, so the effort landed on the packaging. Keep the first kind, delete the second.

Does AI reduce cognitive load?

They can reduce extraneous load, and that genuinely helps: finding the relevant section, defining unfamiliar terms, converting a video to a transcript. The risk is deleting the intrinsic processing instead, which is where understanding gets built. A 2026 survey study separated dependent from autonomous offloading and reported comparable short-term perceived gains alongside very different downstream associations for agency and motivation.


The Bottom Line

Cognitive load theory is usually taught as advice for people who build courses. That framing hides the finding most relevant to how anyone actually learns now. Sweller's own team pointed out that most of what you read was never designed with your working memory in mind, and rather than wait for better material they proposed teaching learners to fix it themselves. In the one experiment that tested that comparison head to head, the self-managers came out ahead of the group handed the professionally designed version.

That's the whole argument. Difficulty describes a relationship between a document and what you already know. It is never a property the document has on its own. Intrinsic load you handle by building background. Extraneous load you handle by reorganizing: integrating what's been split apart, ignoring what's been said twice, and pinning down whatever was about to scroll away. Germane processing is simply what your attention does once it isn't busy with the other two.

The tools are unglamorous. Highlight less than you want to. Write the connection down where the passage is, not somewhere else. Turn video into text before you try to learn from it. Look at what other readers marked when the material is genuinely over your head.

Glasp exists to make those moves cheap: highlight in place, keep the note attached to the passage, pull video and Kindle into the same collection, and ask questions against what you selected instead of what a model summarized. Start with the next hard thing you have to read, and this time fix the presentation before you blame yourself for the difficulty.

Start building your knowledge library

Highlight what matters as you read across the web. Save insights from articles, books, and YouTube videos in one place.

Get Started Free

Or highlight this page as you read it