Skip to content
AtomicReps
Atomic Reps · The Evidence / Retrieval practiceLast reviewed: August 2026Every claim linked
A research publication by Atomic Reps

Retrieval Practice. Why it works, and where it fails.

20 sources, every one gradedReviewed semi-annually

You shipped it Tuesday. Could you whiteboard it Friday?

Atomic Reps sells a retrieval-practice tool, so this page argues our own case. Every claim links to its source, the caveat rides next to its number, and the weakest part of the evidence for us - working professionals, coding - gets its own section below.

Key findingsAugust 2026
  • 01In the foundational lab experiment, students who practiced recall remembered 61% of a passage a week later. Students who reread it four times remembered 40% - and had predicted the best performance of any group.
  • 02The effect replicates at scale. A 2021 meta-analysis of 222 classroom studies (48,478 students) puts the average benefit around half a standard deviation, and its own bias checks barely move the number.
  • 03The comparison decides the size. Against rereading or nothing, the effect is medium to large. Against other active strategies like explaining the material to yourself, one meta-analysis measured it near zero.
  • 04Retrieval helps on new problems too, not just repeats - but the lift shrinks as the problems get further from what was practiced, and vanishes for material never practiced at all.
  • 05The gap that matters for us: roughly six studies on working professionals, with a confidence interval crossing zero, and no randomized trial on programming. We build a coding tool on this literature, so that absence is ours to disclose.
[02] The experiment

Read it four times, or read it once and recall it.

The setup

The modern reference point is a pair of experiments at Washington University. Students read short prose passages - TOEFL test-prep material - and then either reread them or took free-recall tests, writing down everything they could remember. No one got feedback or corrections at any point. The authors expect testing with feedback to do better still, though this study did not test that. Every condition got the same four working periods, so the tested students never spent more time on the material. They spent it differently.

The inversion

Five minutes after learning, rereading looked better: 81% versus 75% in the first experiment. The picture inverted as time passed. Two days later the tested group led 68% to 54%; a week later, 56% to 42%. Rereading wins the quiz you take immediately and loses every one after that.

The first experiment also gives that gap in days rather than percentage points. The tested group's score after a full week was 56%. The rereaders' score after only two days was 54%. One recall test, with no feedback, held the material at that level for roughly five extra days.

The second experiment

The second experiment made it starker. One group read the passage in four consecutive periods - about 14 readings in total. Another read it once and then took three recall tests, rereading nothing, about 3.4 readings in total. A week later the repeated readers recalled 40%; the tested group recalled 61%. Against the five-minute score for the same routine, the rereaders had lost 52% of what they could recall, the tested group 14%. Different students sat the five-minute and the one-week tests, so that is a comparison between groups, not the same people measured twice.

The practice score lies

The tested group also looked worse while it was practicing. Across its three back-to-back recall tests it produced 20.9, then 21.2, then 21.1 idea units out of 30. Flat, at roughly 70% each time, with no feedback in between. The group that studied three times and tested once scored 77% on its single practice test. A week later that ordering flipped, 61% to 56%. The better practice score belonged to the group that remembered less.

Where it is weaker

Two honest complications. Across the whole second experiment the testing advantage shows up in the interaction with delay, not as a blanket main effect - at five minutes, repeated reading still won. And the 61% versus 56% gap between three tests and one test was itself only marginal. The clean, significant contrast is either testing routine against four rereads. This is a finding about durable memory, not about tomorrow morning.

Fig. 01 - recall one week later, by study routine% of idea units recalled · times passage was read
Read once, tested 3x3.4 readings total
61%
Read 3x, tested once10.3 readings total
56%
Read 4 study periods14.2 readings total
40%

Effect sizes on this page: Cohen's d and Hedges' g both measure a gap in standard deviations (0.2 small, 0.5 medium, 0.8 large); g is the version corrected for small samples. Redrawn from Figure 2 and the printed results text, Roediger & Karpicke 2006, Experiment 2 (N=180, 30 per cell, one-week delay). Five minutes after learning the order reverses (83% / 78% / 71%): repeated reading wins the immediate test and loses the delayed ones. Free recall of prose passages, scored by idea units; no feedback on any test.

The confidence inversion. The rereaders were also the most confident. Asked to rate how well they would remember the material a week out, on a 1-to-7 scale, the repeated-study group averaged 4.8; the repeatedly tested group, 4.0. The group that predicted the best memory produced the worst. The predictions were ratings, not percentages - retellings that quote predicted percentages are quoting numbers the paper never printed. The same questionnaire also asked how interesting the passage was. The rereaders rated it lowest of the three groups, 3.8 against 4.6 for the tested group.

How big the gaps were. Effect size in Cohen's d - roughly, the size of a gap next to normal person-to-person variation, where 0.5 is the conventional medium effect. First experiment: d = 0.52 in rereading's favor at five minutes, then 0.95 and 0.83 in testing's favor at two days and one week. Second experiment, one week out: three tests beat four rereads by d = 1.26, and even one test beat four rereads by d = 0.82.

Roediger, H. L. & Karpicke, J. D., Psychological Science 17(3) (2006)

[03] The record

Replicated at scale - against the right control.

Why one result is not enough

One striking lab result would not be worth building on. What makes this literature load-bearing is the replication record. Read it with the comparison group in view.

The replication record

The broadest classroom synthesis pooled 222 independent studies - 48,478 students from elementary school to university - and found quizzing lifted achievement by about half a standard deviation (g = 0.499). Hedges' g is the same effect-size scale as the d values above, adjusted for small samples. The authors ran five publication-bias checks and the corrected estimates stayed between 0.43 and 0.48, which is unusually clean for education research.

An earlier meta-analysis of 272 lab and classroom comparisons (118 articles, 15,427 participants) found the same shape: g = 0.93 against doing nothing, g = 0.51 against restudying. Its authors call the restudy comparison the most accurate indicator, because doing nothing is not a study strategy anyone defends.

The number that keeps this honest

The number that keeps this page honest comes from the friendliest source. In the 222-study synthesis, testing beat no activity at 0.61 and restudying at 0.33 - but against other active strategies, such as students explaining material to themselves, it measured 0.095, close to nothing. Retrieval practice reliably beats passive review. It has not been shown to beat every active alternative.

Where it ranks

The technique also holds its rank in Dunlosky and colleagues' 2013 utility review: of the ten study techniques they graded, practice testing and spaced practice were the only two rated high-utility. Rereading and highlighting - the two most popular - were graded low.

Fig. 02 - the comparison group decides the sizeeffect size (Hedges' g)
vs no practiceAdesope 2017
g = 0.93
vs no activity, classroomsYang 2021
g = 0.61
vs restudyingAdesope 2017
g = 0.51
vs restudying, classroomsYang 2021
g = 0.33
vs other active strategiesYang 2021
g = 0.095

Drawn from two meta-analyses: Adesope et al. 2017 (Table 1: g = 0.93 vs filler/no activity, g = 0.51 vs restudy; fixed-effects model) and Yang et al. 2021 (control-condition moderator: g = 0.610, 0.330, 0.095). Same scale (Hedges' g), different study pools - read each bar against its own source. The bottom bar is the bound: against other elaborative strategies the measured advantage is near zero.

Yang, C., Luo, L., Vadillo, M. A., Yu, R. & Shanks, D. R., Psychological Bulletin 147(4) (2021)Adesope, O. O., Trevisan, D. A. & Sundararajan, N., Review of Educational Research 87(3) (2017)Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J. & Willingham, D. T., Psychological Science in the Public Interest 14(1) (2013)

[04] Transfer

Does it teach anything beyond the test?

The obvious objection: maybe testing just teaches the test. A meta-analysis of 122 experiments (10,382 participants) measured what happens when the final test differs from the practice - new formats, new questions, new problems.

The overall answer is d = 0.40 in favor of testing, a moderate effect. But the average hides a gradient, and the gradient is the finding. Switching test formats: d = 0.58. Application and inference questions: d = 0.32. Rearranged versions of practiced problems: d = 0.22, with a confidence interval touching zero. Material that was in the lesson but never practiced: d = 0.16, which the authors read as no compelling evidence of transfer.

Their conclusion is conditional, not triumphant: transfer is most likely when the practice required real retrieval effort, came with feedback, and succeeded at the time. Practice what you need to keep, in a form close to how you will need it. Retrieval strengthens what you retrieve. It does not strengthen everything around it.

Fig. 03 - how far the benefit travelseffect size (Cohen's d), by transfer distance
Related retrieval cuesd = 0.61
d = 0.61
Different test formatd = 0.58
d = 0.58
Application & inferenced = 0.32
d = 0.32
Problem-solving skillsd = 0.29 · CI crosses 0
d = 0.29
Rearranged problemsd = 0.22 · CI crosses 0
d = 0.22
Untested materiald = 0.16 · CI crosses 0
d = 0.16

Redrawn from the category-level meta-analyses in Pan & Rickard 2018 (d = 0.61 mediator/related cues, 0.58 test format, 0.32 application/inference, 0.29 problem-solving, 0.22 stimulus-response rearrangement, 0.16 untested materials; the bottom three confidence intervals cross zero). Overall transfer d = 0.40, 95% CI [0.31, 0.50].

Pan, S. C. & Rickard, T. C., Psychological Bulletin 144(7) (2018)

[05] The classroom

It survives contact with real students.

The lab result survives contact with real classrooms. In a sixth-grade social-studies class, some facts from each chapter went into brief, low-stakes quizzes; matched facts from the same chapters did not. On the multiple-choice chapter exams, quizzed material scored 94%; non-quizzed material, 81%. The teacher called 81% her usual level - roughly a B minus - so the quizzes recovered 13 of the 19 available points.

The advantage persisted. On the end-of-semester exam, one to two months later, quizzed material still led 79% to 67%. And the study ran the control this page keeps insisting on: in its second experiment, material that students reread three times scored no better than material never revisited at all - 83% versus 81%, a difference indistinguishable from noise.

Scope this honestly. That paper is three experiments in one school - about 140 students enrolled per experiment, over a year and a half - and its second experiment's semester-delay advantage largely disappeared. The often-quoted 1,400-students-over-5-years figure describes the wider middle-school program those experiments anchored, reported separately in 2012.

And a 2021 systematic review of 50 classroom experiments (5,374 students) found benefits across education levels and subjects, with 57% of measured effects medium or large. The same review notes that only 6% of those experiments came from outside Western, educated, industrialized, rich, democratic countries.

Roediger, H. L., Agarwal, P. K., McDaniel, M. A. & McDermott, K. B., J. Experimental Psychology: Applied 17(4) (2011)Agarwal, P. K., Bain, P. M. & Chamberlain, R. W., Educational Psychology Review 24 (2012)Agarwal, P. K., Nunes, L. D. & Blunt, J. R., Educational Psychology Review 33 (2021)

[06] Professionals

The honest gap: adults at work, and code.

Our weakest ground

Nearly everything above measured students. We sell to engineering teams, so the professional evidence deserves the harshest light we can put on it.

The best professional result

The best professional result is medical. Residents in pediatrics and emergency medicine learned two topics; for one they took spaced quizzes, for the other they restudied a review sheet, randomized and counterbalanced. On the final exam about 6.7 months later, quizzed topics scored 39%, restudied topics 26% - a 13-point gap, d = 0.91, from 40 analyzed residents. Small trial, strong signal.

The larger trials

The larger professional trials are humbler. A 537-resident randomized trial of spaced email quizzing found the quizzed group ahead on the study's own online tests - and no advantage at all on the specialty's standardized In-Service Examination. A 634-resident cluster trial across 12 residency programs measured a certification-exam difference of 0.8 points, confidence interval spanning minus 1.0 to plus 2.6: a null. Retrieval practice moves the measure near the practice; whether it moves distant, high-stakes outcomes for professionals is unproven.

Our own domain

The 222-study meta-analysis says the quiet part with numbers: for continuing education - working adults - it found six usable studies, pooling to g = 0.314 with a confidence interval crossing zero. Its authors write that no firm conclusion can be made. And for programming specifically, we searched for a randomized trial of retrieval practice on coding skill and found none. The closest items are a teacher-run quasi-experiment and a 60-person unreviewed preprint. The technique with a century of student evidence has essentially not been tested on our own domain. Nobody should claim otherwise, including us.

Larsen, D. P., Butler, A. C. & Roediger, H. L., Medical Education 43(12) (2009)Kerfoot, B. P. et al., Journal of Urology 177(4) (2007)Grad, R. et al., Advances in Health Sciences Education 26(3) (2021)

[07] The beliefs

The people it works on bet against it.

A consistent side-finding runs through this literature: the people it works on do not believe in it.

In the original 2006 experiments, the repeated readers predicted the best one-week memory and produced the worst. In a 2011 Science paper, Karpicke and Blunt pitted retrieval practice against elaborative concept mapping, an active and respected study method. A week later, retrieval scored 67% to concept mapping's 45%, and retrieval won for 84% of individual students. Yet 75% of the students had predicted concept mapping would do as well or better. The method that won the experiment lost the poll.

Rivers' 2021 review of learners' beliefs finds the same pattern outside the lab: students use self-testing mainly to check what they already know, and they stop once retrieval starts to succeed, which is exactly when the memory benefit gets earned. That mismatch is why a tool has to carry the schedule. Left to feel, people reread. It feels smoother, and the feeling is the bug.

Karpicke, J. D. & Blunt, J. R., Science 331(6018) (2011)Rivers, M. L., Educational Psychology Review 33(3) (2021)

[08] Boundaries

Where the effect bends or breaks.

  • Working memory may gate the benefit.

    A 30-person lab study found the testing effect only for learners with working-memory capacity to spare; under high load, low-capacity learners trended worse than restudy. One small study, but a plausible mechanism: a retrieval attempt costs capacity, and the re-encoding that follows needs what's left. It is part of why our reps stay at one short question.

    Zheng, Y., Sun, P. & Liu, X. L., npj Science of Learning 8 (2023)

  • The complexity boundary is contested.

    Van Gog and Sweller argue the effect shrinks or disappears as material gets more complex. Karpicke and Aue reply that the studies behind that claim never manipulated complexity, and that the effect does show up with complex materials. Both sides are arguing over the same evidence, so treat complexity as an open question, not a settled limit.

    van Gog, T. & Sweller, J. / Karpicke, J. D. & Aue, W. R., Educational Psychology Review 27, both sides (2015)

  • Unsupervised online learners produced two nulls.

    A 2026 study ran the standard paradigm on a crowdworking platform and found nothing, twice, with heavy dropout. The authors blame engagement: unmonitored learners may not put in the retrieval effort the effect runs on. That is the closest published analogue to a self-serve learning app, and it cuts against us - a reason our reps are short, social, and in the channel people already watch.

    Sigayret, K., Parmentier, J.-F. & Silvestre, F., Frontiers in Psychology 17 (2026)

  • Nine STEM courses, mostly flat.

    A single-paper meta-analysis embedded spaced retrieval quizzes in nine introductory STEM courses (578 students analyzed). Two courses showed significant gains; pooled across eight, the effect was about 1.5 percentage points and not significant. It turned significant only when calculus rejoined the pool. The authors' own title asks whether the glass is half full or half empty.

    Bego, C. R. et al., International Journal of STEM Education 11 (2024)

  • One study suggests it stress-proofs recall.

    In a Science study, learners who practiced retrieval showed no memory impairment under an acute lab stressor, while restudiers showed the usual drop. A published commentary notes no cortisol was measured and the retrieval group had simply learned better. A same-lab follow-up found protection for item memory but not source memory. Suggestive, not settled.

    Smith, A. M., Floerke, V. A. & Thomas, A. K., Science 354(6315) (2016)

[09] Methods

Every paper, graded.

The grade column is ours. It says how much weight the design can carry, not whether we like the result.

Every paper reviewed, with its sample, its design and our grade.
PaperNDesignTaskGrade
Roediger & Karpicke 2006120 + 180Lab experimentsProse passages, free recall, up to 1 weekPeer-reviewed, foundational
Roediger et al. 2011142/143/132 enrolled3 classroom experiments6th-grade social studies, real examsPeer-reviewed, one school
Agarwal, Bain & Chamberlain 20121,400+ students5-year applied programMiddle-school classroomsProgram report, peer-reviewed
Karpicke & Blunt 201180 + 120Lab experimentsScience texts vs concept mapping, 1 weekPeer-reviewed (Science)
Adesope et al. 201715,427 / 272 effectsMeta-analysis, 118 articlesLab + classroom practice testsPeer-reviewed meta-analysis
Pan & Rickard 201810,382 / 122 experimentsMeta-analysisTransfer to new formats & problemsPeer-reviewed meta-analysis
Yang et al. 202148,478 / 222 studiesMeta-analysisClassroom quizzing, all levelsPeer-reviewed, bias-checked
Agarwal, Nunes & Blunt 20215,374 / 50 experimentsSystematic reviewApplied classroom researchPeer-reviewed review
Dunlosky et al. 201310 techniquesMonograph reviewUtility grades for study techniquesPeer-reviewed review
Larsen et al. 200940 analyzedRandomized crossoverMedical residents, 6-month examRCT, small N
Kerfoot et al. 2007537 residentsRCTSpaced email questions, 27 weeksRCT; distal outcome null
Grad et al. 2021634 residentsCluster RCTFamily-medicine certification examRCT; null result
Zheng, Sun & Liu 202330Lab experimentAssociative learning under WM loadSingle lab study, small N
Smith et al. 2016~120Lab experimentRecall under induced stressContested; partial replication
Bego et al. 2024578 analyzed9-course meta-analysisSpaced retrieval in intro STEMPeer-reviewed; mostly null
Sigayret et al. 202667 + 129 completersOnline experimentsCrowdworkers, delayed testsPeer-reviewed; double null
[10] Limits

Where this page is weakest.

  • The evidence base is students, not engineers.

    The strongest results come from undergraduates and sixth-graders on prose passages and course facts. The professional pool is about six studies with a confidence interval crossing zero, and there is no randomized trial on programming. We are extrapolating, and we say so.

  • Against good active strategies, the edge may be small.

    The headline effect sizes compare testing to rereading or to nothing. The one meta-analytic comparison against other elaborative strategies measured g = 0.095 - near zero. What is settled is that retrieval beats passive review, not that it beats everything.

  • Immediate performance can favor restudy.

    In the foundational experiments, rereading won the test given five minutes later. The benefit is specific to delayed retention. A tool that promised faster learning by tomorrow would be misusing this literature.

  • Transfer thins with distance.

    d = 0.58 to new test formats, d = 0.16 to material never practiced, with the harder transfer categories' confidence intervals crossing zero. Retrieval strengthens what is retrieved; it is not a general thinking-skills upgrade.

  • The big classroom numbers need their scope tags.

    94% vs 81% is three experiments in one school's sixth grade; the durable delayed advantage is from its first experiment, and its second experiment's semester advantage largely disappeared. The 1,400-student, 5-year figure is the surrounding program, reported separately.

  • Recent field results include honest nulls.

    Two crowdworking experiments found no effect, and a nine-course STEM deployment pooled to about 1.5 points and non-significance without calculus. Low engagement is the going explanation, and self-serve tools like ours own that risk rather than footnote it.

  • Boundary mechanisms are early-stage.

    Working-memory gating rests on 30 people; the complexity dispute is argument over old studies; stress protection has a methods critique and a partial replication. We cite these as bounds, not established laws.

What we built on this finding

The literature's clearest lessons: retrieval beats passive review at a delay, something other than feel has to carry the schedule, and long unsupervised sessions are where the effect goes to die. Atomic Reps runs one short retrieval question a day in Slack - the case for that exact dose, and its limits, is on the evidence hub.

Read the full case on the evidence hub
[11] Sources

The receipts.

The letter

One study at a time, from issue one.

The Retrieval is this page in instalments: one study worth knowing about, one idea worth a name, and one question you answer from memory in about thirty seconds. Everyone starts at issue one, so nothing in it assumes you read the last one.

Read issue one

An email address. Unsubscribe from any issue.