The Dose. What the spacing research licenses.
You crammed the night before. What was left a month later?
Atomic Reps sells a 30-second daily retrieval rep, so this page argues our own dosing. No study tested exactly that dose. What follows shows which bound each study actually supplies, and marks every place where the assembly is our call rather than a finding.
- 01The best gap between reviews is not fixed - it scales with how long you need to remember. Measured optima ran from about 20-40% of a one-week goal down to 5-10% of a one-year goal. Practicing at the optimum improved recall 64% over no gap.
- 02Spacing out retrieval beats massing it: g = 0.74 after publication-bias correction, across 11 studies. The popular refinement - expanding the gaps over time - measured g = 0.034 against plain even spacing: nothing. The spacing does the work, not the schedule.
- 03For working professionals, spaced digital education now has a real meta-analysis: 23 studies, 19 of them randomized, knowledge gains around a third of a standard deviation, graded moderate certainty - with the authors' own quality warnings attached.
- 04Low stakes cost nothing measurable. The 222-study classroom meta-analysis found no reliable difference between high- and low-stakes quizzing, and in a 1,408-student survey, 72% said regular low-stakes quizzes made them less nervous. 6% said more.
- 05Nobody has tested exactly one 30-second question a day. The nearest attempt - a daily text-message question for 293 medical residents - found nothing, because most residents stopped answering. Adherence, not cognition, is where this dose can fail.
The gap is a dial, and it was mapped.
Start with the oldest and best-measured fact in this literature: when you review matters as much as whether you review. Cepeda and colleagues pooled 839 separate assessments across 317 experiments and found the same shape over and over. Repetitions spread over time beat the same repetitions bunched together, and the best gap grew as the retention goal grew.
One study then mapped the whole gap-by-delay grid. Over 1,300 people learned trivia facts, reviewed them after gaps from minutes to 3.5 months, then sat a test up to a year later. For every retention goal there was a sweet spot: review too soon and the second pass adds little, wait too long and the memory is gone before you return. Practicing at the measured optimum improved final recall by 64% over no gap at all (d = 1.1).
The sweet spot is a ratio, not a constant, and the ratio itself shrinks: roughly 20-40% of a one-week goal, down to 5-10% of a one-year goal. Want to remember something for five weeks? The interpolated optimum was about an 8-day gap. For a year, about 27 days. Kang's 2016 practice guide compressed this into the working rule most tools use - space reviews at about 10-20% of the retention interval.
It licenses spreading practice out and letting the gaps grow with the goal. It does not name 30 seconds, it does not name daily, and it was measured on trivia facts in a lab. Those are the first two places our design goes-beyond-the-data, and the assembly section below marks them.
Effect sizes on this page: Hedges' g, Cohen's d and SMD all measure a gap in standard deviations (0.2 small, 0.5 medium, 0.8 large); I2 is how much of the spread between studies is real disagreement rather than chance; trim-and-fill re-estimates an effect after allowing for studies that were probably never published; GRADE is the review's own rating of how far its evidence can be trusted. Redrawn from Cepeda, Vul, Rohrer, Wixted & Pashler 2008 (N=1,354, trivia facts, retention intervals of 7/35/70/350 days). Gaps shown are the paper's cubic-spline interpolated optima (3/8/12/27 days, i.e. 43%/23%/17%/8% of the interval); the raw within-study optima were 1/11/21/21 days. At the optimum vs a zero-day gap, final recall improved 64% (d = 1.1). One study, one lab, one material type.
Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T. & Rohrer, D., Psychological Bulletin 132(3) (2006)Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T. & Pashler, H., Psychological Science 19(11) (2008)
Space the recall itself. Skip the fancy curve.
The retrieval practice page showed that retrieving beats rereading. This literature asks the follow-up: does it matter when you retrieve? Latimier's meta-analysis of 29 studies says yes, a lot. Spacing retrieval attempts out beats massing the same attempts together at g = 0.74. That is the estimate after publication-bias correction pulled it down from 1.02, and the correction is worth naming: the raw literature flatters the effect.
The same meta-analysis demolishes a popular refinement. Expanding schedules - short gaps first, then progressively longer ones, the signature feature of several flashcard algorithms - measured g = 0.034 against plain even spacing. Nothing. The authors' own words: the results do not support the wide belief that intervals should be progressively increased. Spacing is the active ingredient; the fancy curve is branding.
The most practical version of this is called successive relearning: retrieve until correct, then come back days later and do it again. After 533 participants and more than 100,000 scored responses, Rawson and Dunlosky printed a prescription. Recall a concept correctly three times up front, then relearn it three times at widely spaced intervals. In a real upper-level course, students who did this every few days beat their own baseline material by a letter grade or more in both experiments (d = 0.54 and d = 1.10, about 10-16 points). Small classes, one course, but this is the closest published relative of a spaced daily rep.
The efficiency numbers matter for dosing. In the relearning experiments, once students had learned a concept, each spaced relearning pass took about a minute or two per concept. Durable memory does not come from hours. It comes from returns.
Drawn from Latimier, Peyre & Ramus 2021. The spaced-vs-massed estimate is the trim-and-fill corrected g = 0.74 (uncorrected g = 1.02, I2 = 51%); the expanding-vs-uniform comparison is g = 0.034, 95% CI [-0.10, 0.17], p = .62, I2 = 0%. With more than four exposures per item, expanding schedules trended better (g = 0.2, n.s.) - the authors call that tentative.
Latimier, A., Peyre, H. & Ramus, F., Educational Psychology Review 33 (2021)Rawson, K. A. & Dunlosky, J., J. Experimental Psychology: General 140(3) (2011)Janes, J. L., Dunlosky, J., Rawson, K. A. & Jasnow, A., Applied Cognitive Psychology 34(5) (2020)Carpenter, S. K., Pan, S. C. & Butler, A. C., Nature Reviews Psychology 1 (2022)
Two decades of pushed questions, nulls included.
The retrieval practice page was blunt that professional evidence for the technique is thin overall. For the specific delivery pattern we sell - short spaced questions pushed to working adults - the evidence is thicker, because medicine has been testing it for two decades under the name spaced education.
Martinengo's 2024 synthesis covers 23 studies of spaced digital education for health professionals, 19 of them randomized, 3,371 participants. Spaced beat massed online education on post-intervention knowledge by about a third of a standard deviation (SMD 0.32, CI 0.13 to 0.51), graded moderate certainty, and held a similar edge on retention. The authors attach their own warning - most included studies carry unclear or high risk of bias, and heterogeneity is substantial. An earlier Phillips review of 17 clinician studies found the same direction: knowledge moves, behavior sometimes, patient outcomes almost never measured.
The nulls from the retrieval practice page still stand, and one new one joins them. The 537-resident email trial moved its own quizzes but not the specialty exam. The 634-resident cluster trial was null on certification scores. And in 2025, the nearest published relative of our exact product - one multiple-choice question texted to 293 pediatric residents every weekday morning - found no exam effect for the plainest possible reason: residents stopped answering. Median 34 questions answered in eleven months; a fifth never answered one.
Read that last null carefully, because we did. It is not evidence that a daily question cannot work. Adherence collapsed before anyone could test the cognition. It is evidence that an opt-in side channel with no social context dies. That is the strongest argument in this literature for delivering the question where the team already is - and it is an argument about plumbing, not proof that our plumbing works.
Redrawn from Martinengo et al. 2024 (23 studies, N=3,371, 19 RCTs): knowledge SMD 0.32, 95% CI [0.13, 0.51], I2 = 66%, GRADE moderate; retention SMD 0.38 [0.10, 0.65], I2 = 0%; simulation-based surgical skills SMD 1.15 [0.34, 1.96], I2 = 74%, GRADE low. On the skills row the paper prints two pooled figures for the same two studies; we take the abstract's 1.15, the one carrying the GRADE rating, over the results text's 1.24 [0.84, 1.64]. The authors flag the low quality of included studies; the skills bar is the least certain and shown muted for that reason.
Martinengo, L. et al., Journal of Medical Internet Research 26 (2024)Phillips, J. L., Heneka, N., Bhattarai, P., Fraser, C. & Shaw, T., Medical Education 53(9) (2019)Nelson, A. et al., Cureus 17(10) (2025)
Low stakes are supported. Thirty seconds is a bet.
Two design choices remain: the rep is low-stakes, and it is one question. The stakes choice is the well-supported one. Across Yang's 222-study classroom meta-analysis, low-stakes quizzing measured g = 0.477 against high-stakes at 0.441 - no reliable difference, and the authors conclude that even low-stakes quizzes reliably promote learning. The anxiety side agrees. Agarwal's team surveyed 1,408 students after years of regular low-stakes clicker quizzes, and 72% said the quizzing made them less nervous about tests. Honesty requires the other rows of that table: 6% said more nervous, and it is a perception survey, not an anxiety instrument.
The brevity choice has weaker direct evidence, and we say so. No meta-analysis carries a quiz-length or session-duration moderator - we looked. What exists is circumstantial. In Kim's 10,514-employee workplace dataset, real training sessions already averaged about three minutes and two to three questions, and the spacing effect showed up anyway. Zheng's lab study of retrieval under working-memory load found the benefit only when the task left spare capacity, which argues for small bites (30 people, one study).
The word for tiny lessons is microlearning. The honest read of that literature: it is mostly marketing. The two serious reviews are scoping reviews of about 17 studies each, with one randomized trial between them and zero studies measuring the outcomes that matter most. We do not cite microlearning as evidence. The evidence is the spacing and retrieval literature above; micro is just the size our dose happens to be.
Yang, C., Luo, L., Vadillo, M. A., Yu, R. & Shanks, D. R., Psychological Bulletin 147(4) (2021)Agarwal, P. K., D'Antonio, L., Roediger, H. L., McDermott, K. B. & McDaniel, M. A., J. Applied Research in Memory and Cognition 3(3) (2014)Kim, A. S. N., Wong-Kee-You, A. M. B., Wiseheart, M. & Rosenbaum, R. S., Behavior Research Methods 51(4) (2019)Zheng, Y., Sun, P. & Liu, X. L., npj Science of Learning 8 (2023)
The design, bound by bound.
Here is the whole design, bound by bound. The left column is what a study actually licenses; the right column is the part we chose. If a row's right column is empty of evidence, that is the point of the table - you should know which parts are engineering judgment.
Practice is spaced, gaps grow with the goal
What the evidence licenses
Optimal gap scales with the retention interval; spaced beats massed with total time equal
Cepeda 2006, 2008 · Kang 2016 (10-20% working rule)
Where it is our call
Per-question spacing runs on a deterministic queue, not daily repeats of one item. The daily rhythm is delivery; the science is in the gaps between reps of the same question.
The rep is retrieval, spaced
What the evidence licenses
Spacing retrieval attempts beats massing them (g = 0.74, bias-corrected)
Latimier et al. 2021
Where it is our call
We use even, floor-bounded gaps, not an expanding curve - supported by the same meta-analysis's null on expanding schedules.
Come back until it sticks
What the evidence licenses
Retrieve to criterion, then relearn in later spaced sessions; letter-grade course gains
Rawson & Dunlosky 2011 · Janes et al. 2020
Where it is our call
Missed questions requeue sooner and resurface. We do not enforce the printed 3x3 criterion; our cadence is looser than the studied one.
Low stakes, never graded publicly
What the evidence licenses
No detected penalty vs high stakes; most students report less test anxiety
Yang et al. 2021 · Agarwal et al. 2014
Where it is our call
Keeping individual answers anonymous wherever the team can see them is a product covenant, not science.
One question, ~30 seconds
What the evidence licenses
Real workplace sessions of ~3 minutes showed the spacing effect; retrieval costs working memory
Kim et al. 2019 · Zheng et al. 2023
Where it is our call
The 30-second size is untested. No study compares one question to several. This is the least-evidenced choice in the design.
Delivered in Slack, where the team already is
What the evidence licenses
Opt-in side channels die of non-adherence (daily SMS trial: median 34 answers in 11 months)
Nelson et al. 2025 · Sigayret et al. 2026
Where it is our call
In-channel delivery is our bet on the adherence problem. It is an argument from a failure mode, not a measured success.
Every paper, graded.
The grade column is ours. It says how much weight the design can carry, not whether we like the result.
| Paper | N | Design | Task | Grade |
|---|---|---|---|---|
| Cepeda et al. 2006 | 839 assessments | Quantitative synthesis | Distributed practice, verbal recall | Peer-reviewed synthesis |
| Cepeda et al. 2008 | 1,354 | Experiment, gap x delay | Trivia facts, up to 1 year | Peer-reviewed, one lab |
| Latimier et al. 2021 | 29 studies, 97 effects | Meta-analysis | Spaced vs massed retrieval; schedules | Peer-reviewed meta-analysis |
| Rawson & Dunlosky 2011 | 533 | Schedule experiments | Key concepts, 1-4 month delays | Peer-reviewed |
| Janes et al. 2020 | 48 + 22 | Course experiments | Successive relearning, real exams | Peer-reviewed, small N |
| Yang et al. 2021 | 48,478 / 222 studies | Meta-analysis | Classroom quizzing, stakes moderator | Peer-reviewed, bias-checked |
| Agarwal et al. 2014 | 1,408 | Survey | Quizzing and test anxiety | Self-report perception |
| Kim et al. 2019 | 10,514 employees | Observational LMS data | Workplace training questions | Big-N, not causal |
| Martinengo et al. 2024 | 3,371 / 23 studies | Meta-analysis | Spaced digital education, clinicians | Peer-reviewed; quality caveats |
| Phillips et al. 2019 | 2,701 / 17 studies | Systematic review | Spaced education, clinicians | Mostly non-RCT |
| Larsen et al. 2009 | 40 analyzed | Randomized crossover | Residents, 6-month exam | RCT, small N |
| Kerfoot et al. 2007 | 537 residents | RCT | Spaced email questions | Distal outcome null |
| Grad et al. 2021 | 634 residents | Cluster RCT | Certification exam | Null result |
| Nelson et al. 2025 | 293 residents | Prospective cohort | Daily SMS question, 11 months | Null; adherence collapse |
| Zheng, Sun & Liu 2023 | 30 | Lab experiment | Retrieval under WM load | Single lab study, small N |
Where this page is weakest.
The exact dose is untested, still.
No study compares one 30-second question a day to any alternative. We re-checked in 2026. Every bound in the assembly table is real; the assembled whole is a design, and this page should be read that way.
Adherence is the known failure mode.
The closest published relative of our product died of non-adherence before cognition was tested, and unsupervised online learners produced nulls in the retrieval literature too. In-channel delivery is our answer; it is unproven.
The ridgeline comes from trivia facts.
The gap-scaling map is one large study, one lab, one material type, and the printed optima are spline interpolations. The direction is corroborated by the 839-assessment synthesis; the exact percentages are softer than they look.
The headline spaced-retrieval effect needed a bias correction.
Trim-and-fill pulled g = 1.02 down to 0.74, which means the raw literature overstates the effect. We print the corrected number, and the honest reading is 'large-ish, uncertain edges'.
Professional gains are proximal, not distal.
The spaced-education meta shows knowledge gains at moderate certainty, but its authors flag bias risk in most included studies, and the two large resident trials that measured career-grade exams found nothing.
Daily is delivery, not the measured optimum.
For long retention the measured optimal gaps are weeks, not days. Our scheduler spaces each question by days-to-weeks; the daily rhythm just carries whichever question is due. A reader skimming this page could conflate the two - do not.
The anxiety and brevity evidence is the weakest tier.
Less-nervous is self-reported perception; the three-minute workplace sessions are observational vendor-platform data; the working-memory bound is 30 people. These rows support plausibility, not proof.
Atomic Reps runs the assembly above: one low-stakes retrieval question a day in Slack, with each question's return spaced by a plain, inspectable scheduling rule - no expanding-curve theater, because the evidence says even gaps do the job. The full case, and everything we are not claiming, is on the evidence hub.
Read the full case on the evidence hubThe receipts.
The spacing law
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T. & Rohrer, D., Psychological Bulletin 132(3) (2006)
- Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T. & Pashler, H., Psychological Science 19(11) (2008)
- Kang, S. H. K., Policy Insights from the Behavioral and Brain Sciences 3(1) (2016)
- Carpenter, S. K., Pan, S. C. & Butler, A. C., Nature Reviews Psychology 1 (2022)
Spacing the retrieval
- Latimier, A., Peyre, H. & Ramus, F., Educational Psychology Review 33 (2021)
- Rawson, K. A. & Dunlosky, J., J. Experimental Psychology: General 140(3) (2011)
- Rawson, K. A. & Dunlosky, J., Educational Psychology Review 24 (the efficiency companion) (2012)
- Janes, J. L., Dunlosky, J., Rawson, K. A. & Jasnow, A., Applied Cognitive Psychology 34(5) (2020)
Professionals and delivery
- Martinengo, L. et al., Journal of Medical Internet Research 26 (2024)
- Phillips, J. L., Heneka, N., Bhattarai, P., Fraser, C. & Shaw, T., Medical Education 53(9) (2019)
- Larsen, D. P., Butler, A. C. & Roediger, H. L., Medical Education 43(12) (2009)
- Kerfoot, B. P. et al., Journal of Urology 177(4) (2007)
- Grad, R. et al., Advances in Health Sciences Education 26(3) (2021)
- Nelson, A. et al., Cureus 17(10) (2025)
Stakes, brevity, and bounds
- Yang, C., Luo, L., Vadillo, M. A., Yu, R. & Shanks, D. R., Psychological Bulletin 147(4) (2021)
- Agarwal, P. K., D'Antonio, L., Roediger, H. L., McDermott, K. B. & McDaniel, M. A., J. Applied Research in Memory and Cognition 3(3) (2014)
- Kim, A. S. N., Wong-Kee-You, A. M. B., Wiseheart, M. & Rosenbaum, R. S., Behavior Research Methods 51(4) (2019)
- Zheng, Y., Sun, P. & Liu, X. L., npj Science of Learning 8 (2023)
- De Gagne, J. C. et al., JMIR Medical Education 5(2), microlearning scoping review (2019)
- Taylor, A.-D. & Hung, W., Educational Technology Research and Development 70 (2022)
The letter
One study at a time, from issue one.
The Retrieval is this page in instalments: one study worth knowing about, one idea worth a name, and one question you answer from memory in about thirty seconds. Everyone starts at issue one, so nothing in it assumes you read the last one.
An email address. Unsubscribe from any issue.
AtomicReps30″Built on the evidence
Short, spaced practice beats cramming in the studies above. Daily is how we deliver it; the size of the dose is our call, not theirs.
Free for one channel.
