Skip to content
AtomicReps
Atomic Reps · The Evidence / The doseLast reviewed: August 2026Every claim linked
A research publication by Atomic Reps

The Dose. What the spacing research licenses.

20 sources, every one gradedReviewed semi-annually

You crammed the night before. What was left a month later?

Atomic Reps sells a 30-second daily retrieval rep, so this page argues our own dosing. No study tested exactly that dose. What follows shows which bound each study actually supplies, and marks every place where the assembly is our call rather than a finding.

Key findingsAugust 2026
  • 01The best gap between reviews is not fixed - it scales with how long you need to remember. Measured optima ran from about 20-40% of a one-week goal down to 5-10% of a one-year goal. Practicing at the optimum improved recall 64% over no gap.
  • 02Spacing out retrieval beats massing it: g = 0.74 after publication-bias correction, across 11 studies. The popular refinement - expanding the gaps over time - measured g = 0.034 against plain even spacing: nothing. The spacing does the work, not the schedule.
  • 03For working professionals, spaced digital education now has a real meta-analysis: 23 studies, 19 of them randomized, knowledge gains around a third of a standard deviation, graded moderate certainty - with the authors' own quality warnings attached.
  • 04Low stakes cost nothing measurable. The 222-study classroom meta-analysis found no reliable difference between high- and low-stakes quizzing, and in a 1,408-student survey, 72% said regular low-stakes quizzes made them less nervous. 6% said more.
  • 05Nobody has tested exactly one 30-second question a day. The nearest attempt - a daily text-message question for 293 medical residents - found nothing, because most residents stopped answering. Adherence, not cognition, is where this dose can fail.
[02] The spacing law

The gap is a dial, and it was mapped.

The oldest fact

Start with the oldest and best-measured fact in this literature: when you review matters as much as whether you review. Cepeda and colleagues pooled 839 separate assessments across 317 experiments and found the same shape over and over. Repetitions spread over time beat the same repetitions bunched together, and the best gap grew as the retention goal grew.

The gap-by-delay grid

One study then mapped the whole gap-by-delay grid. Over 1,300 people learned trivia facts, reviewed them after gaps from minutes to 3.5 months, then sat a test up to a year later. For every retention goal there was a sweet spot: review too soon and the second pass adds little, wait too long and the memory is gone before you return. Practicing at the measured optimum improved final recall by 64% over no gap at all (d = 1.1).

A ratio, not a constant

The sweet spot is a ratio, not a constant, and the ratio itself shrinks: roughly 20-40% of a one-week goal, down to 5-10% of a one-year goal. Want to remember something for five weeks? The interpolated optimum was about an 8-day gap. For a year, about 27 days. Kang's 2016 practice guide compressed this into the working rule most tools use - space reviews at about 10-20% of the retention interval.

What this licenses

It licenses spreading practice out and letting the gaps grow with the goal. It does not name 30 seconds, it does not name daily, and it was measured on trivia facts in a lab. Those are the first two places our design goes-beyond-the-data, and the assembly section below marks them.

Fig. 01 - the optimal review gap, by retention goalgap as % of the retention interval · interpolated optimum
Remember for 1 weekoptimal gap ~3 days
43%
Remember for 5 weeksoptimal gap ~8 days
23%
Remember for 10 weeksoptimal gap ~12 days
17%
Remember for ~1 yearoptimal gap ~27 days
8%

Effect sizes on this page: Hedges' g, Cohen's d and SMD all measure a gap in standard deviations (0.2 small, 0.5 medium, 0.8 large); I2 is how much of the spread between studies is real disagreement rather than chance; trim-and-fill re-estimates an effect after allowing for studies that were probably never published; GRADE is the review's own rating of how far its evidence can be trusted. Redrawn from Cepeda, Vul, Rohrer, Wixted & Pashler 2008 (N=1,354, trivia facts, retention intervals of 7/35/70/350 days). Gaps shown are the paper's cubic-spline interpolated optima (3/8/12/27 days, i.e. 43%/23%/17%/8% of the interval); the raw within-study optima were 1/11/21/21 days. At the optimum vs a zero-day gap, final recall improved 64% (d = 1.1). One study, one lab, one material type.

Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T. & Rohrer, D., Psychological Bulletin 132(3) (2006)Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T. & Pashler, H., Psychological Science 19(11) (2008)

[03] Spaced retrieval

Space the recall itself. Skip the fancy curve.

Does timing matter

The retrieval practice page showed that retrieving beats rereading. This literature asks the follow-up: does it matter when you retrieve? Latimier's meta-analysis of 29 studies says yes, a lot. Spacing retrieval attempts out beats massing the same attempts together at g = 0.74. That is the estimate after publication-bias correction pulled it down from 1.02, and the correction is worth naming: the raw literature flatters the effect.

The refinement that failed

The same meta-analysis demolishes a popular refinement. Expanding schedules - short gaps first, then progressively longer ones, the signature feature of several flashcard algorithms - measured g = 0.034 against plain even spacing. Nothing. The authors' own words: the results do not support the wide belief that intervals should be progressively increased. Spacing is the active ingredient; the fancy curve is branding.

Successive relearning

The most practical version of this is called successive relearning: retrieve until correct, then come back days later and do it again. After 533 participants and more than 100,000 scored responses, Rawson and Dunlosky printed a prescription. Recall a concept correctly three times up front, then relearn it three times at widely spaced intervals. In a real upper-level course, students who did this every few days beat their own baseline material by a letter grade or more in both experiments (d = 0.54 and d = 1.10, about 10-16 points). Small classes, one course, but this is the closest published relative of a spaced daily rep.

The efficiency numbers

The efficiency numbers matter for dosing. In the relearning experiments, once students had learned a concept, each spaced relearning pass took about a minute or two per concept. Durable memory does not come from hours. It comes from returns.

Fig. 02 - the spacing does the work, not the scheduleeffect size (Hedges' g)
Spaced vs massed retrieval11 studies, 39 effects · bias-corrected
g = 0.74
Expanding vs even gaps16 studies, 54 effects · null
g = 0.034

Drawn from Latimier, Peyre & Ramus 2021. The spaced-vs-massed estimate is the trim-and-fill corrected g = 0.74 (uncorrected g = 1.02, I2 = 51%); the expanding-vs-uniform comparison is g = 0.034, 95% CI [-0.10, 0.17], p = .62, I2 = 0%. With more than four exposures per item, expanding schedules trended better (g = 0.2, n.s.) - the authors call that tentative.

Latimier, A., Peyre, H. & Ramus, F., Educational Psychology Review 33 (2021)Rawson, K. A. & Dunlosky, J., J. Experimental Psychology: General 140(3) (2011)Janes, J. L., Dunlosky, J., Rawson, K. A. & Jasnow, A., Applied Cognitive Psychology 34(5) (2020)Carpenter, S. K., Pan, S. C. & Butler, A. C., Nature Reviews Psychology 1 (2022)

[04] Professionals

Two decades of pushed questions, nulls included.

Thin ground, again

The retrieval practice page was blunt that professional evidence for the technique is thin overall. For the specific delivery pattern we sell - short spaced questions pushed to working adults - the evidence is thicker, because medicine has been testing it for two decades under the name spaced education.

The synthesis

Martinengo's 2024 synthesis covers 23 studies of spaced digital education for health professionals, 19 of them randomized, 3,371 participants. Spaced beat massed online education on post-intervention knowledge by about a third of a standard deviation (SMD 0.32, CI 0.13 to 0.51), graded moderate certainty, and held a similar edge on retention. The authors attach their own warning - most included studies carry unclear or high risk of bias, and heterogeneity is substantial. An earlier Phillips review of 17 clinician studies found the same direction: knowledge moves, behavior sometimes, patient outcomes almost never measured.

The nulls

The nulls from the retrieval practice page still stand, and one new one joins them. The 537-resident email trial moved its own quizzes but not the specialty exam. The 634-resident cluster trial was null on certification scores. And in 2025, the nearest published relative of our exact product - one multiple-choice question texted to 293 pediatric residents every weekday morning - found no exam effect for the plainest possible reason: residents stopped answering. Median 34 questions answered in eleven months; a fifth never answered one.

Why adherence, not cognition

Read that last null carefully, because we did. It is not evidence that a daily question cannot work. Adherence collapsed before anyone could test the cognition. It is evidence that an opt-in side channel with no social context dies. That is the strongest argument in this literature for delivering the question where the team already is - and it is an argument about plumbing, not proof that our plumbing works.

Fig. 03 - spaced digital education for professionals, meta-analyzedstandardized mean difference (SMD)
Knowledge, post-coursemoderate certainty
SMD 0.32
Knowledge retentionI2 = 0%
SMD 0.38
Surgical skills (simulation)low certainty
SMD 1.15

Redrawn from Martinengo et al. 2024 (23 studies, N=3,371, 19 RCTs): knowledge SMD 0.32, 95% CI [0.13, 0.51], I2 = 66%, GRADE moderate; retention SMD 0.38 [0.10, 0.65], I2 = 0%; simulation-based surgical skills SMD 1.15 [0.34, 1.96], I2 = 74%, GRADE low. On the skills row the paper prints two pooled figures for the same two studies; we take the abstract's 1.15, the one carrying the GRADE rating, over the results text's 1.24 [0.84, 1.64]. The authors flag the low quality of included studies; the skills bar is the least certain and shown muted for that reason.

Martinengo, L. et al., Journal of Medical Internet Research 26 (2024)Phillips, J. L., Heneka, N., Bhattarai, P., Fraser, C. & Shaw, T., Medical Education 53(9) (2019)Nelson, A. et al., Cureus 17(10) (2025)

[05] Brief & low-stakes

Low stakes are supported. Thirty seconds is a bet.

Low-stakes, and one question

Two design choices remain: the rep is low-stakes, and it is one question. The stakes choice is the well-supported one. Across Yang's 222-study classroom meta-analysis, low-stakes quizzing measured g = 0.477 against high-stakes at 0.441 - no reliable difference, and the authors conclude that even low-stakes quizzes reliably promote learning. The anxiety side agrees. Agarwal's team surveyed 1,408 students after years of regular low-stakes clicker quizzes, and 72% said the quizzing made them less nervous about tests. Honesty requires the other rows of that table: 6% said more nervous, and it is a perception survey, not an anxiety instrument.

The brevity choice is weaker

The brevity choice has weaker direct evidence, and we say so. No meta-analysis carries a quiz-length or session-duration moderator - we looked. What exists is circumstantial. In Kim's 10,514-employee workplace dataset, real training sessions already averaged about three minutes and two to three questions, and the spacing effect showed up anyway. Zheng's lab study of retrieval under working-memory load found the benefit only when the task left spare capacity, which argues for small bites (30 people, one study).

Microlearning

The word for tiny lessons is microlearning. The honest read of that literature: it is mostly marketing. The two serious reviews are scoping reviews of about 17 studies each, with one randomized trial between them and zero studies measuring the outcomes that matter most. We do not cite microlearning as evidence. The evidence is the spacing and retrieval literature above; micro is just the size our dose happens to be.

Yang, C., Luo, L., Vadillo, M. A., Yu, R. & Shanks, D. R., Psychological Bulletin 147(4) (2021)Agarwal, P. K., D'Antonio, L., Roediger, H. L., McDermott, K. B. & McDaniel, M. A., J. Applied Research in Memory and Cognition 3(3) (2014)Kim, A. S. N., Wong-Kee-You, A. M. B., Wiseheart, M. & Rosenbaum, R. S., Behavior Research Methods 51(4) (2019)Zheng, Y., Sun, P. & Liu, X. L., npj Science of Learning 8 (2023)

[06] The assembly

The design, bound by bound.

Here is the whole design, bound by bound. The left column is what a study actually licenses; the right column is the part we chose. If a row's right column is empty of evidence, that is the point of the table - you should know which parts are engineering judgment.

  • Practice is spaced, gaps grow with the goal

    What the evidence licenses

    Optimal gap scales with the retention interval; spaced beats massed with total time equal

    Cepeda 2006, 2008 · Kang 2016 (10-20% working rule)

    Where it is our call

    Per-question spacing runs on a deterministic queue, not daily repeats of one item. The daily rhythm is delivery; the science is in the gaps between reps of the same question.

  • The rep is retrieval, spaced

    What the evidence licenses

    Spacing retrieval attempts beats massing them (g = 0.74, bias-corrected)

    Latimier et al. 2021

    Where it is our call

    We use even, floor-bounded gaps, not an expanding curve - supported by the same meta-analysis's null on expanding schedules.

  • Come back until it sticks

    What the evidence licenses

    Retrieve to criterion, then relearn in later spaced sessions; letter-grade course gains

    Rawson & Dunlosky 2011 · Janes et al. 2020

    Where it is our call

    Missed questions requeue sooner and resurface. We do not enforce the printed 3x3 criterion; our cadence is looser than the studied one.

  • Low stakes, never graded publicly

    What the evidence licenses

    No detected penalty vs high stakes; most students report less test anxiety

    Yang et al. 2021 · Agarwal et al. 2014

    Where it is our call

    Keeping individual answers anonymous wherever the team can see them is a product covenant, not science.

  • One question, ~30 seconds

    What the evidence licenses

    Real workplace sessions of ~3 minutes showed the spacing effect; retrieval costs working memory

    Kim et al. 2019 · Zheng et al. 2023

    Where it is our call

    The 30-second size is untested. No study compares one question to several. This is the least-evidenced choice in the design.

  • Delivered in Slack, where the team already is

    What the evidence licenses

    Opt-in side channels die of non-adherence (daily SMS trial: median 34 answers in 11 months)

    Nelson et al. 2025 · Sigayret et al. 2026

    Where it is our call

    In-channel delivery is our bet on the adherence problem. It is an argument from a failure mode, not a measured success.

[07] Methods

Every paper, graded.

The grade column is ours. It says how much weight the design can carry, not whether we like the result.

Every paper reviewed, with its sample, its design and our grade.
PaperNDesignTaskGrade
Cepeda et al. 2006839 assessmentsQuantitative synthesisDistributed practice, verbal recallPeer-reviewed synthesis
Cepeda et al. 20081,354Experiment, gap x delayTrivia facts, up to 1 yearPeer-reviewed, one lab
Latimier et al. 202129 studies, 97 effectsMeta-analysisSpaced vs massed retrieval; schedulesPeer-reviewed meta-analysis
Rawson & Dunlosky 2011533Schedule experimentsKey concepts, 1-4 month delaysPeer-reviewed
Janes et al. 202048 + 22Course experimentsSuccessive relearning, real examsPeer-reviewed, small N
Yang et al. 202148,478 / 222 studiesMeta-analysisClassroom quizzing, stakes moderatorPeer-reviewed, bias-checked
Agarwal et al. 20141,408SurveyQuizzing and test anxietySelf-report perception
Kim et al. 201910,514 employeesObservational LMS dataWorkplace training questionsBig-N, not causal
Martinengo et al. 20243,371 / 23 studiesMeta-analysisSpaced digital education, cliniciansPeer-reviewed; quality caveats
Phillips et al. 20192,701 / 17 studiesSystematic reviewSpaced education, cliniciansMostly non-RCT
Larsen et al. 200940 analyzedRandomized crossoverResidents, 6-month examRCT, small N
Kerfoot et al. 2007537 residentsRCTSpaced email questionsDistal outcome null
Grad et al. 2021634 residentsCluster RCTCertification examNull result
Nelson et al. 2025293 residentsProspective cohortDaily SMS question, 11 monthsNull; adherence collapse
Zheng, Sun & Liu 202330Lab experimentRetrieval under WM loadSingle lab study, small N
[08] Limits

Where this page is weakest.

  • The exact dose is untested, still.

    No study compares one 30-second question a day to any alternative. We re-checked in 2026. Every bound in the assembly table is real; the assembled whole is a design, and this page should be read that way.

  • Adherence is the known failure mode.

    The closest published relative of our product died of non-adherence before cognition was tested, and unsupervised online learners produced nulls in the retrieval literature too. In-channel delivery is our answer; it is unproven.

  • The ridgeline comes from trivia facts.

    The gap-scaling map is one large study, one lab, one material type, and the printed optima are spline interpolations. The direction is corroborated by the 839-assessment synthesis; the exact percentages are softer than they look.

  • The headline spaced-retrieval effect needed a bias correction.

    Trim-and-fill pulled g = 1.02 down to 0.74, which means the raw literature overstates the effect. We print the corrected number, and the honest reading is 'large-ish, uncertain edges'.

  • Professional gains are proximal, not distal.

    The spaced-education meta shows knowledge gains at moderate certainty, but its authors flag bias risk in most included studies, and the two large resident trials that measured career-grade exams found nothing.

  • Daily is delivery, not the measured optimum.

    For long retention the measured optimal gaps are weeks, not days. Our scheduler spaces each question by days-to-weeks; the daily rhythm just carries whichever question is due. A reader skimming this page could conflate the two - do not.

  • The anxiety and brevity evidence is the weakest tier.

    Less-nervous is self-reported perception; the three-minute workplace sessions are observational vendor-platform data; the working-memory bound is 30 people. These rows support plausibility, not proof.

What we built on this finding

Atomic Reps runs the assembly above: one low-stakes retrieval question a day in Slack, with each question's return spaced by a plain, inspectable scheduling rule - no expanding-curve theater, because the evidence says even gaps do the job. The full case, and everything we are not claiming, is on the evidence hub.

Read the full case on the evidence hub
[09] Sources

The receipts.

The letter

One study at a time, from issue one.

The Retrieval is this page in instalments: one study worth knowing about, one idea worth a name, and one question you answer from memory in about thirty seconds. Everyone starts at issue one, so nothing in it assumes you read the last one.

Read issue one

An email address. Unsubscribe from any issue.

AtomicReps

Built on the evidence

Short, spaced practice beats cramming in the studies above. Daily is how we deliver it; the size of the dose is our call, not theirs.

Free for one channel.