Skip to content

Spaced Repetition in Language Learning: What Studies Show

Flashcards spread on a desk representing a spaced vocabulary review schedule

Spaced repetition has one of the strongest evidence bases in learning science: a 2022 meta-analysis of second-language studies found medium-to-large retention gains from spacing out practice. But the exact schedule an app uses matters less than the marketing suggests. The same meta-analysis found equally-spaced and "smart" expanding schedules perform about the same. What matters more is matching your review interval to how long you actually need to remember the word.

What this post covers

  • Does spaced repetition actually work for language learning?

  • Does it matter which spacing schedule you use?

  • What matters more than the schedule shape?

  • Why can a review that feels successful still fail you later?

  • How should this change the way you review vocabulary?

  • What doesn't this research tell us?

Does Spaced Repetition Actually Work for Language Learning?

Yes, and the evidence for it in second-language learning specifically is unusually strong for an education claim. A 2022 meta-analysis in Language Learning pooled 98 effect sizes from 48 experiments and 3,411 learners, looking specifically at second-language vocabulary studies rather than general education research. Spaced practice, reviewing material with a gap between repetitions instead of cramming it all at once, produced a medium-to-large effect on retention. An effect size of roughly g = 0.76 on tests taken right after the study period, and a larger g = 1.15 on tests given after a delay.

The underlying idea isn't new. Hermann Ebbinghaus described the basic shape of this problem back in the 19th century with his forgetting curve, the observation that memory fades fastest right after learning and then more slowly over time, which is exactly what a review schedule is trying to counteract. The 2022 meta-analysis puts a modern, SLA-specific number on how large that effect actually is for people learning a second language, more than a century after Ebbinghaus first sketched the shape of the curve.

That second number is the more interesting one. Spacing doesn't just help retention in general; its advantage grows the longer you wait before testing whether the word actually stuck. Cramming can look competitive with spacing if you quiz someone immediately, but the gap widens substantially once real time has passed, which is the only kind of recall that matters for actually speaking a language later.

This research matches what most experienced teachers already know from watching students. What the meta-analysis adds is a number. Most people already agree that "spacing helps" as a general hypothesis. "Spacing produces roughly double the retention advantage on delayed tests compared to immediate ones" is something you can actually build a review schedule around.

Does It Matter Which Spacing Schedule You Use?

Less than most spaced repetition software imply. The same meta-analysis found that equally-spaced review schedules and "smart" expanding schedules produced statistically indistinguishable results. Plenty of flashcard apps market a proprietary spacing algorithm as the reason to choose them over a simpler tool, an algorithm that supposedly calculates the ideal moment to show you a card again, expanding the interval more precisely than a basic fixed schedule could. Kim and Webb's meta-analysis tested exactly this comparison, equal spacing (the same gap every time) against expanding spacing (gaps that grow with each successful review), and found no statistically significant difference between them.

That's a genuinely inconvenient finding for a lot of app marketing copy. Across the pool of second-language vocabulary studies analyzed, the specific choice between equal and expanding spacing wasn't the variable driving the results, which suggests scheduling algorithms are less central to why spacing works than the marketing around them implies.

Anki's default algorithm, SM-2, is a well-known example of an expanding schedule: it starts new cards at short intervals and lengthens the gap each time you answer correctly, shortening it again if you get something wrong. Quizlet and Memrise implement their own versions of the same basic idea. None of that's wrong. It's just not obviously the thing that explains why spaced repetition works, according to the evidence, even though it's usually the thing apps lead with.

What Matters More Than the Schedule Shape?

Matching the gap between reviews to how long you actually need to remember the word, rather than following one fixed interval for everything. A separate, earlier meta-analysis addressed this directly. Cepeda and colleagues (2006) reviewed 839 assessments across 317 experiments on verbal recall and found that the ideal gap between study sessions increases as the amount of time you need to retain the material increases. In plain terms, if you need to remember something for a test next week, a shorter review gap works best. If you need to remember it for years, the optimal gap between reviews is considerably longer than what most fixed schedules default to.

This research is about verbal recall broadly, not language acquisition specifically, and it's worth being upfront about that rather than implying otherwise. But the mechanism it describes, matching interval length to retention need, applies directly to a genuine practical problem in language learning. A vocabulary app has no way of knowing whether you're learning a word for a client call next Tuesday or building it into permanent vocabulary you'll use for years. It applies the same interval logic either way.

A teacher assembling materials by hand can make that distinction a student's app can't. A word that's only useful for an upcoming trip or presentation probably doesn't need the same long-term spacing treatment as core vocabulary the learner will use indefinitely. That distinction matters just as much when AI is doing the material generation: an AI tool can produce a fresh vocabulary list in seconds, but deciding how those words fit into a learner's actual review needs is a judgment call the tool doesn't make for you, one of the limits we cover in more depth in Chapter 1 of our AI guide for language teachers, on what AI is and isn't good at in a teaching context.

Why Can a Review That Feels Successful Still Fail You Later?

Because short review gaps and long review gaps look about equally effective on a quiz taken right away, even though they produce very different results once real time has passed. This is the practical sting in the Kim and Webb findings. Shorter spacing gaps performed about as well as longer gaps on immediate posttests, the kind of quiz you'd take moments after reviewing. On delayed posttests, the gap that had felt sufficiently effective in the short term underperformed noticeably.

Getting a flashcard right today tells you almost nothing about whether you'll remember that word in three weeks, especially if the interval before that review was short. The feeling of "I know this one" after a quick, frequent review is a poor predictor of durable memory. It's a comfortable feeling, and it's exactly the feeling a badly-spaced review schedule is good at producing.

How Should This Change the Way You Review Vocabulary?

Stretch your review intervals further than feels comfortable, and stop assuming a specific app's algorithm is doing more work than the basic principle of spacing itself. A few practical adjustments follow directly from the evidence above, and none of them require switching tools.

  • Don't judge a review schedule by how easy it feels. If every review feels effortless, the gaps are probably too short to build durable memory.
  • Match the interval to the reason you're learning the word. A trip-specific phrase can get a shorter review cycle; core vocabulary you'll use indefinitely deserves longer gaps than most default settings assume.
  • Whatever tool you use, plain flashcards, Anki, Quizlet, or a teacher-built list, trust the spacing principle more than the specific algorithm behind it. The evidence doesn't show a clear winner among scheduling methods.
  • Test yourself with a real delay occasionally, not just the next scheduled review, to get an honest read on what's actually sticking.

SRS apps are still useful tools, but their marketing claims about "smarter" algorithms deserve to be held a bit more loosely than the copy implies.

What Doesn't the Research on Spaced Repetition Tell Us?

It doesn't tell us much about grammar or speaking fluency, and most of the underlying studies are short, lab-based experiments rather than long-term classroom research. Both meta-analyses cited here focus on vocabulary and verbal recall specifically. Kim and Webb's review is explicitly about second-language vocabulary learning; Cepeda and colleagues studied general verbal recall tasks, not language acquisition at all. Neither directly tells you whether spacing helps someone develop speaking fluency, internalize grammar rules, or build listening comprehension the same way it helps them remember that "insurance paperwork" means what it means.

It's a reasonable bet that some version of the spacing effect applies more broadly. Skill development in general tends to benefit from distributed rather than massed practice. But betting on it and having direct evidence for it are different things, and this article is making the narrower, better-supported claim rather than the broader one.

Most of the underlying studies are also relatively short in duration and conducted under lab or lab-adjacent conditions, students studying word lists for an experiment, not for years of language use in the wild. That doesn't invalidate the findings, but it's a reason to treat the specific effect sizes as directional rather than as a promise about what will happen in any individual classroom or app.

Making Your Review Time Count

Most language teachers already assume spaced repetition works, and this research doesn't really challenge that assumption. Its strongest, best-supported claim is narrower. The specific shape of your review schedule probably matters less than app marketing suggests, and matching your review interval to why you're learning a word matters more than most tools account for.

If you want a broader look at how session length and frequency affect retention, separate from spacing specifically, we've written about that research too. Between the two, the practical takeaway is the same. Consistency and honest interval-matching beat whichever algorithm badge is on the app you're using.

How Does Edumo Handle Spaced Review of Vocabulary?

Edumo auto-generates flashcards directly from the vocabulary in a text a teacher creates or adapts, rather than from a generic pre-built deck, so what a learner reviews is tied to material they've actually encountered. It doesn't claim a proprietary scheduling breakthrough the evidence above doesn't support.

If you want to see how auto-generated flashcards fit into a lesson, take a look at Edumo.