Back to the notebook
Learning science · 8 min read

Spaced repetition, and the four buttons everyone presses wrong

The theory takes two minutes to explain and almost nobody gets the practice right. What FSRS does that fixed intervals cannot, why your grading is probably lying to the scheduler, and the three ways an app can ruin it anyway.

SW
Sebastian Walter
Founder, Lemnly
April 12, 2026
A leather-bound notebook beside a laptop, warm light from the side
Learning science
Spaced repetition, and the four buttons everyone presses wrong
The short version
  • Forgetting then recalling is the mechanism. Not forgetting would teach you nothing.
  • Fixed intervals (1 day, 3 days, 7 days…) treat every card the same. Your cards are not the same.
  • FSRS models each card’s difficulty and stability separately, and learns from your actual grading.
  • Which means your grading has to be honest. Most people’s isn’t, and it quietly wrecks the schedule.

Every learner has had this happen. A word you knew last Tuesday, gone. You looked it up, you used it in a sentence, you were certain. The brain recycled it anyway.

It is tempting to take that personally. It is not personal, and more to the point it is not a malfunction — it is the exact mechanism you are going to exploit.

The two-minute version

Ebbinghaus, 1885, memorising nonsense syllables and testing himself at intervals. The decline he found is roughly exponential: most of the loss in the first day, then it flattens.

The famous chart stops there, which is why most people take the wrong lesson from it. Here is the part that matters:

0%50%100%first readreview 1review 2time →
Each successful recall resets retention and flattens what follows. The gaps widen because the memory genuinely lasts longer, not because the app is going easy on you. Illustrative shape — real curves differ per person and per card, which is the whole reason schedulers model them individually.
Memory is not a recording. It is closer to a muscle, and you train it by making it strain just before it would have failed.

Why interval ladders fail

Plenty of flashcard apps still schedule on a fixed ladder — one day, three days, a week, a fortnight — and the problem with that is not subtle. It assumes cards are interchangeable.

They are conspicuously not. In Spanish, casa and acometer do not behave remotely alike in my head. A ladder over-rehearses the first and under-rehearses the second, so you spend your five minutes reviewing words you would know in your sleep while the ones you nearly had slip away.

It also assumes your forgetting curve matches everyone else’s. Sleep, age, time of day, how emotionally loaded a word is, whether you have a hook for it in a language you already speak — all of that moves the curve. A scheduler ignoring per-card and per-person signal is playing averages on data that is not average.

What FSRS does instead

FSRS — the Free Spaced Repetition Scheduler — does three things SM-2 and its descendants do not.

  • It separates difficulty from stability. Stability is how long the memory lasts; difficulty is how hard this particular card is for you. Both update after every review. The model works out which of your cards are sticky and which are slippery, rather than assuming.
  • It targets a retention probability, typically around 90%, and schedules to hit it. Aim higher and you review more; aim lower and you review less. The tradeoff becomes explicit instead of inherited from a heuristic somebody wrote in the eighties.
  • It was fitted to real review logs — an enormous pile of them, from the open-source SRS community — rather than derived from reasoning about how memory ought to work.

The published comparisons consistently show FSRS hitting a given retention with fewer reviews than SM-2, and the gap widening as your deck grows. I am deliberately not putting a percentage here. You will see specific numbers quoted around the internet, including by people selling things, and the honest answer is that it depends on your deck, your consistency and your grading.

The four buttons, and how you are probably using them

This is the part I care about most, because it is the part you control and the part almost everyone gets wrong. After each card: Again, Hard, Good, Easy.

What each one means
  • Again — you did not recall it. You could not have used the word cold.
  • Hard — you got there, but it cost you. You searched, you frowned, you nearly guessed.
  • Good — it came back without strain. This should be most of your cards on a normal day.
  • Easy — instant, effortless, no hesitation at all. Rare.
What people actually do
  • Pressing Good because Again feels like failure.
  • Pressing Easy on the third sighting of a word to feel like you are progressing.
  • Pressing Hard as a way of saying “I liked this card”.
  • Grading on how you feel about your progress rather than what just happened in your head.

The buttons are not a self-assessment and they are not a scorecard. They are the only channel through which the scheduler learns anything about you, and a consistently generous “Good” will stretch your intervals past what your memory supports. Then retention drops, you see more failures, and you conclude that spaced repetition does not work for you.

Consistency matters more than accuracy. A scheduler can work with a grader who is reliably a bit harsh. It cannot work with one who is moody.

Three ways an app ruins it anyway

A good scheduler underneath does not save an app that gets these wrong, and I have used several that do.

  1. Card glut. Adding every word you ever met guarantees you drown and quit. The curation is the work, not the collection. This is why Lemnly proposes only the words you do not already know — it checks your own history and a shared cache first — and why I would rather you added twelve good cards from a chapter than four hundred mediocre ones.
  2. Naked cards. One word on the front, one translation on the back, no context. That is a coin toss with extra steps. The sentence you first met the word in is what anchors the recall, which is why it gets attached automatically rather than being an optional field you would never fill in.
  3. Punishing people for having lives. A streak that shatters on one missed day trains users to lie to it, or to bulk-grade to protect it. Lemnly’s streak survives a missed day and does not shout at you about it, because the behaviour worth protecting is coming back after a gap — not never having one.
My take

The most damaging idea in this whole field is that more review is better. It is intuitive, it feels virtuous, and it is wrong in a specific and expensive way: reviewing a card before it is due teaches you almost nothing and costs you the willpower you needed tomorrow.

If you take one habit from this article, make it the boring one — stop when the queue is empty. The empty queue is not the app running out of things to show you. It is the app telling you that you are done.

Questions people ask about this

Is FSRS actually better than SM-2?
For most decks, yes — it reaches the same retention with fewer reviews, and the advantage grows as the deck grows. It is not magic and it will not rescue a deck full of bad cards. Anki ships FSRS now too, so this is no longer a reason to switch tools.
How many new cards a day should I add?
Fewer than you want to. Ten to fifteen is sustainable for most people alongside a life; twenty-five is a sprint you can hold for a few weeks. Remember every card you add today is a review you inherit for the next two years, so the question is not what you can absorb today but what you can service in March.
Should I study in both directions?
If you want to speak, yes. Recognising a word and producing it are separate skills with separate memories, and a recognition-only deck will let you believe you know things you cannot say. Reverse cards are slower and more irritating, which is roughly proportional to how much they are doing.
What happens if I miss a week?
Nothing catastrophic. You come back to a backlog, spread it over a few days rather than clearing it in one sitting, and expect more failures than usual for a while. Stability rebuilds much faster than it originally built. The people who quit are the ones who try to clear three hundred cards in one evening.
Can I just use the app without understanding any of this?
Yes, and most people should. The only part of this article you genuinely need is the section on the four buttons — grade honestly and consistently, and the scheduler handles the rest.

Thirty days, if you want to feel it

  • Days 1–2. Pick one language. Read one short thing. Tap the words you do not know. Do not build a system.
  • Days 3–14. Review daily. Five minutes. Add nothing new. Spend the fortnight calibrating your buttons — this is the whole exercise.
  • Day 15. Read something longer. Notice how many words you no longer need to tap.
  • Days 16–30. Daily review, plus ten or fifteen new words a week from whatever you are reading.

At day 30 you will not have a transformed vocabulary. You will have something more useful: a calibrated sense of what “Good” means, and a five-minute habit that costs you nothing to keep.

More on the curve itself in the forgetting curve piece, and on what all this vocabulary is actually for in the thresholds article.

A chapter tonight.
The words still there next week.

Free. Pick a story at your level, tap the words you don’t know, and let Lemnly bring them back at the right time.

iPhone & iPad · coming soonApp StoreSoonor open it in your browser

No credit card · Spanish, German and English today