필사 모드: The Structure of Practice: Why Practice That Goes Well and Practice That Lasts Are Different
EnglishIntroduction
Most advice about practice is about quantity. How long, how often, how many hours.
But what has been relatively well verified in learning research is not quantity, it is arrangement. How you divide up the same total, what order you mix things in, what form you repeat them in. And this field contains one finding that is extremely inconvenient in practice.
The conditions that make things go best while you are practising and the conditions that leave the most behind afterwards are not the same.
That mismatch is the centre of this post. The point is that the reason self-designed practice mostly fails is not laziness but this mismatch. This post is not confined to a particular domain; it deals with learning in general.
1. What has been set out as effective
The broadest summary is the review Dunlosky and colleagues published in Psychological Science in the Public Interest in 2013. They examined ten learning techniques students actually use and assigned each a utility grade (Dunlosky et al., 2013, read 2026-08-16).
Only two received high utility: practice testing and distributed practice. The moderate grade went to elaborative interrogation, self-explanation, and interleaved practice. The low grade went to summarisation, highlighting, the keyword mnemonic, imagery use, and rereading.
The list worth noticing is the low-grade one. Highlighting and rereading are the methods students use most, and that is exactly why the authors put those two into the review. The fact that the most widely used methods received the lowest grade announces the subject of this post in advance.
2. Spacing
This is the difference between using the same time in one block and splitting it up.
The meta-analysis Cepeda and colleagues published in Psychological Bulletin in 2006 covered 184 articles, 317 experiments, and 839 measurements (Cepeda et al., 2006, read 2026-08-16).
The core conclusion of this study was not simply that spacing is good but that the optimal gap changes with the goal. According to the authors' analysis, the gap between study sessions and the interval to the final test work together, and as the retention interval lengthens, the gap that produces maximum retention lengthens as well.
Translated into practice, it comes out like this. If the test is tomorrow, a short gap is better; if you want to still use it in six months, a much longer gap is better. There is no single correct gap.
Let me mark one boundary. What this meta-analysis covered is verbal recall tasks. Nothing in this study licenses saying that the same curve applies intact to motor skills or complex problem solving.
3. Retrieval practice
This is pulling something out of memory instead of rereading it. Closing the book and saying what you just read; trying the problem before looking at the answer.
The review Serra and colleagues published in Behavioral Sciences in 2025 gathers the main meta-analytic figures in this field (Serra et al., 2025, read 2026-08-16).
Set out, they run as follows. In Roediger and Karpicke's 2006 summary, effect sizes were distributed between 0.31 and 1.26. Rowland's 2014 meta-analysis reported a moderate effect. g = 0.50. Pan and Rickard's 2018 analysis reported d = 0.40 even in transfer situations where the test format changed. Yang and colleagues' 2021 analysis reported g = 0.50 in actual classroom settings.
So this is not an effect that appears only in the laboratory, and some of it survives a change in test format. And while retrieval with feedback is generally better than retrieval without, even retrieval without feedback sometimes beats additional study — that is how the review sums it up.
There are boundary conditions too. The review Polack and Miller set out in 2022 lists the situations where retrieval practice does not work well (Polack & Miller, 2022, read 2026-08-16). When the error rate is high and there is no feedback; testing after a very short delay such as five minutes; when there was not enough initial learning in the first place for a retrievable memory to have formed; and when the gap is stretched so far that retrieval itself fails.
The last item matters practically. Retrieval practice is not a method to use from a standing start of knowing nothing. It works only when a minimum base is there.
4. Mixing practice and contextual interference
This is the difference between drilling one thing in a block and mixing several. On the motor-learning side it has long been studied under the name contextual interference.
The systematic review and meta-analysis Czyż and colleagues published in Scientific Reports in 2024 quantified this effect. 54 studies were included in the meta-analysis, and a total of 2,068 participants were analysed on delayed retention tests (Czyż et al., 2024, read 2026-08-16).
The overall retention effect was a standardised mean difference of 0.63 in the three-level model, 95 percent confidence interval 0.33 to 0.93. In the random-effects model it was 0.71, confidence interval 0.41 to 1.01. A moderate size.
But split that average up and the picture changes completely.
- Laboratory conditions: 0.92, confidence interval 0.48 to 1.36. A large effect.
- Field conditions: 0.23, confidence interval minus 0.16 to 0.62. Not significant.
- Under 18: 0.02. Effectively nothing.
- Adults: 0.63.
- Over 60: 1.45.
The authors explicitly noted that an effect that is large in the laboratory is nearly absent in the field. Heterogeneity is large as well. The overall heterogeneity index was 90 percent. And on methodological quality, only 3 of 59 studies received a high grade.
A similar tension exists in the cognitive domain. The study Do and Thomas published in the Journal of Intelligence in 2023 reconfirmed the benefit of interleaved practice in category learning while reporting that the participants did not recognise that benefit (Do & Thomas, 2023, read 2026-08-16).
There is material pointing the other way as well. The preprint Rowlandson and Simpson released in 2023 reports that two large classroom experiments on mathematical concept learning failed to detect an interleaving effect (confirmed in the same search result, read 2026-08-16). This is not yet a peer-reviewed paper and it is a single report, so it should be read with exactly that much weight. Still, the fact that material pointing in a different direction exists is worth recording.
Summed up, the evidence grade for interleaved practice is one notch below retrieval practice. That is consistent with Dunlosky's review giving interleaved practice a moderate grade.
5. Why self-designed practice so often fails
This is the centre of the post.
There is a concept called desirable difficulties. It refers to conditions that impede initial learning but promote long-term retention and transfer. Spacing, retrieval practice, and mixing all belong to it.
The review Binks wrote in the Journal of Evaluation in Clinical Practice in 2026 sets out this concept and attaches an important warning. What has sufficient evidence across several fields is formative assessment, interleaved practice, distributed practice, and productive-failure approaches — but feeling harder must not be uncritically equated with greater educational benefit (Binks, 2026, read 2026-08-16).
This warning cuts both ways. It cuts the illusion that comfortable practice is effective, and it cuts the illusion in the opposite direction that difficult practice is automatically effective.
And there is an explanation of why people cannot correct this for themselves. The piece de Bruin wrote in Medical Science Educator in 2023 sets out that the relationship between perceived effort and perceived learning determines the learner's choices (de Bruin, 2023, read 2026-08-16).
Put simply, it goes like this. People use performance at this present moment as the indicator of learning. Reread and the text flows easily, and when it flows easily you feel you know it well. Drill one thing in a block and each round goes better, and when it goes better you feel you are improving.
In both cases the feeling is accurate. Performance at that moment really is good. What is wrong is the part where that performance is used as an indicator of the future. Do and Thomas's result shows this point directly. The participants really did gain, and did not recognise the gain.
Let me draw the mismatch with a constructed example. Suppose two people study the same material for the same amount of time. One reads the material through carefully three times and underlines the important places. On the last pass almost every sentence is familiar, and when they finish the sense of having understood it well is strong. The other reads it once, closes it, writes down what they remember, and checks only the missing parts, repeating that three times. Because the blanks are exposed every time, the feeling at the end is worse than the first person's.
What matters here is that both people's feelings are accurate. Accessibility at the moment they finish studying really is higher for the first person. The divergence comes two weeks later. The effect sizes for retrieval practice above are measurements of exactly that two-weeks-later difference. And neither of the two can see two weeks ahead at the moment they finish studying. This is an example constructed for explanation and does not describe any particular experiment.
So this problem is not a problem of will. It is a problem of information. The only real-time signal a person has access to happens to be the misleading one. The solution, too, is not will but changing when the measurement is taken.
6. How to read effect sizes
Let me note briefly how the numbers in this post should be handled.
First, averages hide conditions. The overall contextual-interference effect of 0.63 is the laboratory's 0.92 and the field's 0.23 combined. You have to ask first which side your own situation is closer to.
Second, you have to look at whether the confidence interval includes zero. The field-condition interval was minus 0.16 to 0.62. That means the possibility of no effect has not been excluded.
Third, when heterogeneity is large, the meaning of the average shrinks. Heterogeneity of 90 percent suggests the studies may be measuring different things from one another.
Fourth, you have to look at quality grades. That only 3 of 59 studies were high quality means this whole literature has to be read carefully.
Apply those four and the conclusion that remains is this. Retrieval practice and spacing can be used with relative confidence. Mixing looks right in direction, but it is better not to be optimistic about the size.
7. If you are actually laying it out
I will write down only what follows directly from the above.
- For the same total time, split rather than block. The longer the target period, the longer the gap.
- Change rereading into retrieval. Close the material, pull it out first, then check.
- Attach a check to the retrieval. Especially at a stage where you still do not know much, retrieval without feedback is risky.
- Do not start with retrieval from a base of nothing. A minimum of initial learning comes first.
- Introduce mixing, but lower your expectations. Especially under field conditions.
- Do not use the feeling during practice as a performance indicator. Judge by trying it a few days later with no preparation.
The last item is the hardest to carry out and the most important. Change that one and the rest follow. Because once the habit of checking a few days later exists, the wrong methods expose themselves.
8. Evidence grade by claim
| Claim | Evidence grade | Source |
|---|---|---|
| Practice testing and distributed practice have high utility | Replicated | Dunlosky et al. 2013 |
| Rereading and highlighting have low utility | Replicated | Dunlosky et al. 2013 |
| The optimal gap lengthens with the retention interval | Replicated, verbal tasks only | Cepeda et al. 2006 |
| The retrieval practice effect is moderate in size | Replicated | The meta-analyses set out by Serra et al. 2025 |
| Retrieval weakens without basic learning and feedback | Replicated | Polack & Miller 2022 |
| Contextual interference aids retention | Conditionally verified | Czyż et al. 2024, clear only in the laboratory |
| Interleaved practice works in the classroom too | Contested | Do & Thomas 2023 versus Rowlandson & Simpson 2023 |
| Learners judge poorly which methods are effective for them | Replicated | Do & Thomas 2023, de Bruin 2023 |
| If it feels hard, the effect is large | No evidence | Binks 2026 warns against this explicitly |
What has not been verified
The largest gap is generalisation across domains. Most of the evidence in this post comes from verbal recall, category learning, and simple motor tasks. There is no evidence that it applies at the same magnitude to complex professional skills or to creative work.
Second, literature quality. That only 3 of 59 studies in the contextual-interference meta-analysis received a strong quality grade is a serious limitation in itself. The authors stated that they did not perform a formal publication-bias assessment such as a funnel plot or Egger's test.
Third, the field effect of interleaved practice is an ongoing dispute. I will not say which side is right. I will point out only that popular material recommending interleaved practice strongly mostly does not mention this dispute.
Fourth, this post has not dealt with motivation. Even the most effective practice structure is worth nothing if you do not keep going. And the factors that make you keep going are of a different kind from the variables handled here.
Closing
What can be said with confidence about the structure of practice comes to three sentences.
Splitting it up is better. Pulling it out instead of rereading is better. And the feeling during practice cannot be used as an indicator of learning.
The third supports the other two. Because as long as you trust the feeling, you keep going back to the comfortable side.
The same conclusion comes out on the experience side too. Fourteen Things Table Tennis Taught Me has an item saying that playing only for fun will not make you better. In research language, desirable difficulties; in experience language, the place where you lose. Only, the sentence from the experience side slides easily into "if it is hard it is good," and as we saw above that is an extension with no evidence behind it.
There is also a post that takes the deliberate practice debate inside the narrow frame of language learning. What the Research Shows: Interpreter Training, Cognitive Ability, and the Deliberate Practice Debate looks at the same dispute in the special case of interpreter training. That one handles domain-specific problems; this one handled only the structure that does not care which domain you are in.
Related posts:
- Previous post: Experts Do Not Know More, They See Differently — expertise as perception.
- Next post: What It Means to Be Smart — what intelligence measurement says and does not say.
- Deliberate Practice Path Tool — a tool that makes gaps and repetition cycles visible.
- Memory Lab — a tool for doing retrieval practice directly.
References
Below are the sources checked directly on 2026-08-16 while writing this post. The links are the addresses actually read, and some are public API responses from Europe PMC.
- Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving Students Learning With Effective Learning Techniques. Psychological Science in the Public Interest
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin
- Serra, M. J., Kaminske, A. N., Nebel, C., & Coppola, K. M. (2025). The Use of Retrieval Practice in the Health Professions: A State-of-the-Art Review. Behavioral Sciences, 15(7), 974
- Polack, C. W., & Miller, R. R. (2022). Testing Improves Performance as Well as Assesses Learning. Journal of Experimental Psychology: Animal Learning and Cognition
- Czyż, S. H., Wójcik, A. M., Solarská, P., & Kiper, P. (2024). High contextual interference improves retention in motor learning: systematic review and meta-analysis. Scientific Reports
- Do, L. A., & Thomas, A. K. (2023). The Underappreciated Benefits of Interleaving for Category Learning. Journal of Intelligence / Rowlandson, P., & Simpson, A. (2023). Interleaving in mathematical category learning. PsyArXiv preprint
- Binks, S. (2026). Why Desirable Difficulties Work: Review and Caveats. Journal of Evaluation in Clinical Practice / de Bruin, A. B. H. (2023). Supporting Students to Engage with Desirable Difficulties. Medical Science Educator
현재 단락 (1/88)
Most advice about practice is about quantity. How long, how often, how many hours.