Split View: 다시 읽지 말고 덮고 떠올려라 — 시험 효과의 원 논문과 간격 반복 알고리즘
다시 읽지 말고 덮고 떠올려라 — 시험 효과의 원 논문과 간격 반복 알고리즘
들어가며 — 가장 인기 있는 공부법이 가장 배신하는 공부법
시험 전날을 떠올려 보세요. 대부분의 사람이 하는 일은 같습니다. 교재와 노트를 다시 읽는 것. 형광펜을 긋고, 밑줄 친 부분을 또 읽고, 눈에 익은 페이지를 보며 안도합니다. "이제 알겠다."
그 안도감이 함정입니다. 다시 읽기가 만드는 것은 기억이 아니라 유창함 — 글자가 눈에 미끄러지듯 들어오는 익숙함 — 이고, 뇌는 그 유창함을 "안다"로 오역합니다. 논문으로 읽는 심리학 일곱 번째 편은 이 착각을 실험으로 폭로한 논문과, 그 대안을 코드 수준까지 파 봅니다. 시리즈에서 가장 실용적인 편이 될 것입니다.
2006년 실험 — 다시 읽기 대 스스로 시험 보기
헨리 뢰디거(Henry Roediger)와 제프리 카픽(Jeffrey Karpicke)의 2006년 논문(Psychological Science)의 설계는 교과서적입니다. 대학생들에게 짧은 지문을 학습시키되, 학습 방식을 달리했습니다.
- SSSS 조건: 지문을 네 번 반복해서 읽음 (순수 다시 읽기)
- SSST 조건: 세 번 읽고, 마지막엔 책을 덮고 기억나는 대로 써 봄 (시험 1회)
- STTT 조건: 한 번만 읽고, 세 번 연속 회상 시험 (시험 3회)
그리고 최종 시험을 두 시점으로 나눠 봤습니다. 5분 뒤, 그리고 일주일 뒤.
| 조건 | 5분 뒤 회상률 | 일주일 뒤 회상률 |
|---|---|---|
| SSSS (읽기 4회) | 약 83% | 약 40% |
| STTT (읽기 1회 + 시험 3회) | 약 71% | 약 61% |
5분 뒤 시험에서는 다시 읽기가 이깁니다. 그런데 일주일이 지나자 순위가 완전히 뒤집힙니다. 네 번 읽은 그룹은 절반 이상을 잃었지만, 한 번 읽고 세 번 인출한 그룹은 대부분을 지켰습니다. 읽기는 단기전의 기술이고, 인출은 장기전의 기술입니다.
더 잔인한 디테일이 있습니다. 학습 직후 "일주일 뒤 얼마나 기억할 것 같나"를 예측하게 했더니, 다시 읽기 그룹이 자신의 기억을 가장 높게 예측했습니다. 실제로는 가장 많이 잊을 그룹이 말입니다. 아는 느낌과 아는 것은 다른 시스템이고, 다시 읽기는 전자만 부풀립니다.
이것이 시험 효과(testing effect)입니다. 시험은 평가 도구이기 이전에 학습 도구라는 것. 기억에서 무언가를 꺼내는 행위 자체가 그 기억의 저장 강도를 높입니다. 이 효과는 이후 교실 현장 연구들과 메타분석에서 반복 확인된, 학습과학에서 가장 신뢰받는 발견 중 하나입니다.
두 번째 재료 — 간격 효과
시험 효과와 짝을 이루는 것이 간격 효과(spacing effect)입니다. 1885년 에빙하우스의 망각 곡선 이래 가장 오래 살아남은 기억 법칙로, 같은 총량의 복습이라도 몰아서(massed) 하는 것보다 간격을 두고(spaced) 하는 것이 장기 기억에 압도적으로 유리합니다. 니컬러스 세피다(Nicholas Cepeda) 연구팀의 2006년 메타분석(254개 연구, 1만 4천여 명)은 이를 정량화했고, 후속 연구는 실용적인 경험칙까지 제시했습니다. 복습 간격을 시험까지 남은 기간의 10~20% 정도로 잡으라는 것입니다(한 달 뒤 시험이면 3~6일 간격).
왜 간격이 효과적일까요. 유력한 설명은 "바람직한 어려움(desirable difficulty)"입니다. 복습 시점에 기억이 살짝 흐려져 있어야 인출이 노력을 요구하고, 노력이 든 인출일수록 기억을 강하게 재저장합니다. 어제 외운 단어를 오늘 복습하면 너무 쉽고, 한 달 뒤면 이미 사라져 있습니다. 잊힐 듯 말 듯한 최적의 순간이 존재하는 것입니다.
그렇다면 카드 수천 장 각각의 "최적의 순간"을 사람이 추적할 수 있을까요? 없습니다. 그래서 알고리즘이 등장합니다.
코드로 이해하기 — 안키의 심장, SM-2 알고리즘
간격 반복 소프트웨어(안키 등)의 원형은 1980년대 피오트르 워즈니악(Piotr Wozniak)의 SuperMemo에서 나온 SM-2 알고리즘입니다. 원리는 간단합니다. 카드마다 난이도 계수를 유지하고, 회상에 성공할 때마다 다음 복습 간격을 그 계수만큼 늘립니다. 실패하면 간격을 리셋합니다.
def sm2_review(quality, reps, interval, ease):
"""One SM-2 update after a review.
quality: 0-5 self-grade (>=3 means recalled)
reps: successful reviews in a row so far
interval: current interval in days
ease: difficulty factor, starts at 2.5
Returns (reps, interval, ease) for scheduling the next review.
"""
if quality < 3: # forgot -> relearn from scratch
return 0, 1, ease
# ease drifts with how hard the recall felt (min 1.3)
ease = max(1.3, ease + 0.1 - (5 - quality) * (0.08 + (5 - quality) * 0.02))
if reps == 0:
interval = 1 # first success: see it tomorrow
elif reps == 1:
interval = 6 # second success: ~a week out
else:
interval = round(interval * ease) # then multiply out
return reps + 1, interval, ease
# simulate one card always recalled with quality 4
reps, interval, ease = 0, 0, 2.5
schedule = []
for _ in range(7):
reps, interval, ease = sm2_review(4, reps, interval, ease)
schedule.append(interval)
print(schedule) # -> [1, 6, 15, 37, 90, 219, 533]
출력을 보세요. 1일, 6일, 15일, 37일, 90일... 복습 간격이 기하급수로 벌어집니다. 7번의 복습으로 한 장의 카드가 1년 반 뒤까지 스케줄되는 것입니다. 간격 효과(점점 벌어지는 최적 간격)와 시험 효과(매 복습이 인출 시험)를 코드 20줄이 동시에 구현하고 있습니다. 매일 새 카드 20장을 추가해도 일일 복습량이 폭발하지 않는 이유가 이 지수적 간격에 있습니다.
참고로 현대의 안키는 SM-2의 후손에 더해, 기억 확률을 명시적으로 모델링하는 FSRS 같은 최신 스케줄러도 제공합니다. 하지만 어떤 스케줄러든 심장은 같습니다. 꺼내 보기, 잊기 직전에.
우리가 가져갈 것 — 공부법의 교체 목록
이번 편의 실천은 명확합니다. 학습 루틴에서 다음 교체를 실행하는 것입니다.
- 다시 읽기 → 덮고 떠올리기. 챕터를 읽었으면 책을 덮고 백지에 요약해 봅니다. 막히는 지점이 정확히 복습할 지점입니다. 유창함의 안도감 대신 인출의 불편함을 선택하는 것 — 의식적 연습의 "컴포트존 밖" 원리가 기억에 적용된 형태입니다.
- 형광펜 → 질문 만들기. 밑줄은 미래의 다시 읽기를 예약하는 행위입니다. 대신 그 자리에서 스스로에게 낼 질문을 만들어 두세요. 좋은 질문 목록이 그대로 시험지가 됩니다.
- 벼락치기 → 간격 스케줄. 시험까지 4주라면 같은 8시간을 하루에 몰지 말고 3~5일 간격의 4회로 쪼갭니다. 총량이 같아도 결과가 다릅니다.
- 암기 과목 → 간격 반복 도구. 영단어, 자격증, 의학·법률 지식처럼 사실 밀도가 높은 영역은 안키류 도구의 홈그라운드입니다. 단, 카드는 남이 만든 것보다 직접 만든 것이 낫습니다. 카드 만들기 자체가 첫 인출 연습이기 때문입니다.
개발자에게도 그대로 적용됩니다. 새 언어의 문법, 자주 잊는 명령어, 시스템 설계 패턴 — "검색하면 되지"가 통하지 않는 면접과 장애 상황이라는 최종 시험이 있다는 점에서, 엔지니어의 공부도 결국 인출의 게임입니다.
원문 읽기 가이드
- 시험 효과: Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249-255.
- 간격 메타분석: Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354-380.
- 현장 적용 리뷰: Dunlosky, J., et al. (2013). Improving students' learning with effective learning techniques. Psychological Science in the Public Interest, 14(1), 4-58. — 10가지 공부법의 등급표. 다시 읽기와 형광펜이 최하 등급을 받는 장면이 백미입니다.
읽기 팁: 뢰디거와 카픽 논문은 그림 1(5분/2일/1주 시험의 조건별 막대그래프) 하나에 모든 것이 있습니다. 던로스키 리뷰는 표 4의 기법별 종합 등급만 봐도 공부법 책 열 권을 대체합니다. 다음 편은 몸 시리즈의 논문판, 수면 부채의 도스-반응 실험입니다.
Do Not Reread — Close the Book and Recall: The Original Testing-Effect Paper and the Spaced-Repetition Algorithm
Introduction — The Most Popular Study Method Is the One That Betrays You Most
Picture the night before an exam. Almost everyone does the same thing: rereading the textbook and their notes. They run a highlighter across the page, read the underlined parts again, and feel reassured by pages that already look familiar. "Now I know this."
That reassurance is the trap. What rereading builds is not memory but fluency — the ease of words sliding smoothly across your eyes — and the brain mistranslates that fluency as "I know this." This seventh installment of Psychology, Straight from the Papers digs into the paper that exposed this illusion experimentally, and into its alternative all the way down to the code. It will be the most practical installment in the series.
The 2006 Experiment — Rereading vs. Testing Yourself
The design of the 2006 paper by Henry Roediger and Jeffrey Karpicke (Psychological Science) is textbook-clean. College students studied short passages, but the way they studied differed.
- SSSS condition: read the passage four times (pure rereading)
- SSST condition: read three times, then close the book and write down everything recalled (one test)
- STTT condition: read once, then three consecutive recall tests (three tests)
And they split the final test across two time points: five minutes later, and one week later.
| Condition | Recall after 5 min | Recall after 1 week |
|---|---|---|
| SSSS (4 reads) | about 83% | about 40% |
| STTT (1 read + 3 tests) | about 71% | about 61% |
On the test five minutes later, rereading wins. But after a week the ranking flips completely. The group that read four times lost more than half, while the group that read once and retrieved three times kept most of it. Reading is a skill for the short game; retrieval is a skill for the long game.
There is a crueler detail. When students were asked right after studying to predict "how much will you remember a week from now," the rereading group predicted their own memory the highest — even though they were in fact the group that would forget the most. Feeling that you know and actually knowing are different systems, and rereading inflates only the former.
This is the testing effect. Before a test is an assessment tool, it is a learning tool. The very act of pulling something out of memory strengthens that memory's storage. The effect has been confirmed again and again in later classroom field studies and meta-analyses, and is one of the most trusted findings in the science of learning.
The Second Ingredient — The Spacing Effect
The partner of the testing effect is the spacing effect. It is the longest-surviving law of memory since Ebbinghaus's forgetting curve of 1885: for the same total amount of review, spacing it out (spaced) is overwhelmingly better for long-term memory than cramming it together (massed). The 2006 meta-analysis by Nicholas Cepeda's team (254 studies, more than 14,000 people) quantified this, and follow-up work even offered a practical rule of thumb: set your review interval to roughly 10-20% of the time left until the test (a test a month away means a 3-6 day interval).
Why is spacing effective? The leading explanation is "desirable difficulty." At the moment of review, the memory has to be slightly faded so that retrieval demands effort, and the more effortful the retrieval, the more strongly the memory is re-stored. Reviewing today a word you memorized yesterday is too easy; a month later it is already gone. There exists an optimal moment, right on the edge of being forgotten.
So can a person track the "optimal moment" for each of thousands of cards? No. That is where the algorithm comes in.
Understanding It Through Code — SM-2, the Heart of Anki
The prototype for spaced-repetition software (Anki and the like) is the SM-2 algorithm, which came out of Piotr Wozniak's SuperMemo in the 1980s. The principle is simple: keep a difficulty factor for each card, and every time you recall it successfully, stretch the next review interval by that factor. Fail, and the interval resets.
def sm2_review(quality, reps, interval, ease):
"""One SM-2 update after a review.
quality: 0-5 self-grade (>=3 means recalled)
reps: successful reviews in a row so far
interval: current interval in days
ease: difficulty factor, starts at 2.5
Returns (reps, interval, ease) for scheduling the next review.
"""
if quality < 3: # forgot -> relearn from scratch
return 0, 1, ease
# ease drifts with how hard the recall felt (min 1.3)
ease = max(1.3, ease + 0.1 - (5 - quality) * (0.08 + (5 - quality) * 0.02))
if reps == 0:
interval = 1 # first success: see it tomorrow
elif reps == 1:
interval = 6 # second success: ~a week out
else:
interval = round(interval * ease) # then multiply out
return reps + 1, interval, ease
# simulate one card always recalled with quality 4
reps, interval, ease = 0, 0, 2.5
schedule = []
for _ in range(7):
reps, interval, ease = sm2_review(4, reps, interval, ease)
schedule.append(interval)
print(schedule) # -> [1, 6, 15, 37, 90, 219, 533]
Look at the output. 1 day, 6 days, 15 days, 37 days, 90 days... the review interval widens exponentially. Seven reviews schedule a single card as far as a year and a half out. Twenty lines of code implement the spacing effect (an optimal interval that keeps widening) and the testing effect (every review is a retrieval test) at the same time. This exponential spacing is why adding 20 new cards a day does not make your daily review load explode.
For reference, modern Anki, on top of SM-2's descendants, also offers newer schedulers like FSRS that model recall probability explicitly. But whatever the scheduler, the heart is the same: retrieve it, right before you forget it.
What We Take Away — A Swap List for How You Study
The practice for this installment is clear: run the following swaps in your study routine.
- Rereading → close and recall. Once you have read a chapter, close the book and summarize it on a blank page. Wherever you get stuck is exactly what to review. It means choosing the discomfort of retrieval over the reassurance of fluency — the "outside the comfort zone" principle of deliberate practice applied to memory.
- Highlighting → writing questions. Underlining is scheduling a future reread. Instead, right there, write the questions you will ask yourself. A good list of questions becomes your exam paper as is.
- Cramming → a spaced schedule. If the test is four weeks away, do not pile the same 8 hours into one day; split it into four sessions at 3-5 day intervals. Same total, different result.
- Memory-heavy subjects → a spaced-repetition tool. Fact-dense areas like vocabulary, certifications, and medical or legal knowledge are the home turf of Anki-style tools. That said, cards you make yourself beat cards made by others — because making the card is itself the first retrieval practice.
It applies to developers just the same. The syntax of a new language, commands you keep forgetting, system-design patterns — because there is a final exam, the interview and the outage, where "I will just search for it" does not work, an engineer's studying is ultimately a game of retrieval too.
Reading Guide
- Testing effect: Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249-255.
- Spacing meta-analysis: Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354-380.
- Field-application review: Dunlosky, J., et al. (2013). Improving students' learning with effective learning techniques. Psychological Science in the Public Interest, 14(1), 4-58. — A ranking table of 10 study techniques. The highlight is the scene where rereading and highlighting receive the lowest grade.
Reading tip: the Roediger and Karpicke paper puts everything in Figure 1 (a bar chart by condition for the 5-minute, 2-day, and 1-week tests). For the Dunlosky review, the overall grade-by-technique in Table 4 alone replaces ten study-method books. The next installment is the paper edition of the body series, the dose-response experiment on sleep debt.