Introduction — The Graph You Have Seen Is Not in the Paper
There is a picture you have almost certainly seen. Experience on the horizontal axis, confidence on the vertical. The curve rockets upward through the beginner zone, hits the "Peak of Mount Stupid," plunges into the "Valley of Despair," and only much later settles onto a gentle plateau. The caption always carries the same name: the Dunning-Kruger effect.
That curve is not in the original paper. There is no peak, no valley, no plateau. The picture was manufactured somewhere on the internet, given the paper as a name tag, and set loose. The figure actually printed in the 1999 paper looks far more boring. And the distance between the boring original and the flashy copy is the whole of this installment.
Following the rules of Psychology, Straight from the Papers, we start with what the original paper actually did. Then we follow the sharpest criticism ever leveled at the effect — that meaningless random numbers produce the same graph — and count what survives.
What the 1999 Paper Actually Did — Cornell Undergraduates, Humor, Logic, Grammar
The 1999 Journal of Personality and Social Psychology paper by Justin Kruger and David Dunning consists of four studies. Every participant was an undergraduate in a Cornell psychology class.
- Study 1 (humor): participants rated how funny 30 jokes were, and the ratings of 8 professional comedians served as the answer key. 65 participants.
- Study 2 (logical reasoning): 20 LSAT-style logic items. 45 participants.
- Study 3 (grammar): standardized English grammar items. 84 participants.
- Study 4 (retraining): after taking the logic task, some participants were given a short course in logic and then rated themselves again. 140 participants.
The heart of the procedure is the question asked after the test: "In percentile terms, where do you think your ability and your score today rank among your Cornell classmates?" Participants were then divided into quartiles by their actual scores, and the average self-assessment and the average actual score were plotted side by side for each quartile. That figure is the signature chart of the paper. The horizontal axis is not experience but performance quartile, there are two lines, and there is no peak.
The Real Result Is a Little Different from "Incompetent People Are Overconfident"
Transferring the approximate values from the grammar task (Study 3) into a table gives this.
| Performance quartile | Actual score (percentile) | Estimate of own ability (percentile) | Direction |
|---|---|---|---|
| Bottom 25 percent | about 10 | about 67 | Large overestimate |
| 2nd quartile | about 32 | about 66 | Overestimate |
| 3rd quartile | about 62 | about 71 | Slight overestimate |
| Top 25 percent | about 89 | about 72 | Underestimate |
Two things stand out here. First, the self-assessment line is almost flat across the quartiles. While actual scores stretch from the 10th percentile to the 89th, nearly eightfold, the self-estimates merely drift between the mid-60s and the low 70s. What the data say, in other words, is much closer to almost everyone rates themselves a little above average than to "only incompetent people overestimate themselves."
Second, there is a fact that popular summaries almost never mention. The top 25 percent underestimated themselves by roughly 15 to 20 percentile points. Kruger and Dunning address this in the paper themselves, adding the explanation that people who do well found the task easy and therefore assume everyone else did about as well (false consensus). That is why the two lines in the paper cross at the right-hand edge.
At this point the popular version is already wrong in two places: the shape of the graph, and the conclusion about the top performers.
The Sharpest Objection — Random Numbers Draw the Same Graph
The real problem comes next. Generate self-assessment scores and actual scores as random numbers with no relationship to each other, split them into quartiles by actual score, and plot the average of each group. Astonishingly, you get a crossing figure that looks exactly like the one in the paper. Two forces are overlapping.
The first is regression to the mean. The bottom quartile is not a gathering of "people with low ability" alone; it also contains "people who had a bad day." Look at those same people through a different measurement, their self-assessment, and they drift back toward the average. The top quartile drifts back in the opposite direction. The moment you cut the quartiles by actual score, the crossing of the two lines is booked in advance.
The second is the better-than-average effect combined with measurement error. People tend to cluster their self-assessments a little above the midpoint, and the correlation with actual ability is not especially high. That flattens the self-assessment line in every quartile. Overlay a flat line and a steep line and you get an X.
The first systematic statement of this objection was the 2002 paper by Joachim Krueger and Ross Mueller. They reported that correcting for measurement reliability and controlling for regression effects shrinks the Dunning-Kruger pattern sharply. Edward Nuhfer and colleagues then nailed the argument down in pictures in their 2016 and 2017 Numeracy papers. Pure random-number simulations produced graphs indistinguishable from the one printed in textbooks, and when more than 1,000 real responses were drawn as individual scatter points instead of quartile averages, a completely different picture appeared: most people know their own ability fairly accurately, and severe overestimators are a minority. The conclusion of the Nuhfer team is scathing. What created the illusion was not the human mind but the convention for drawing the graph.
The 2020 Reanalysis — 929 People, and the Residue That Remains
The cleanest test of everything above is the 2020 Intelligence paper by Gilles Gignac and Marcin Zajenkowski. The sample was 929 people. They measured objective intelligence test scores alongside self-assessed intelligence, then abandoned quartile-average graphs entirely and ran two tests on individual-level data: one a nonlinear (quadratic) regression, the other a test for heteroscedasticity. If the Dunning-Kruger hypothesis were true, self-assessment error would have to be systematically larger in the low-ability range.
Here is the result. The correlation between self-assessment and actual intelligence was about 0.28, a clear positive relationship, and the nonlinear term was effectively meaningless. Only the heteroscedasticity test left a faint trace: the spread of self-assessment error is a little wider at lower ability. The title of the paper states the conclusion outright. It is mostly a statistical artefact.
One more thing on top. The 2006 study by Katherine Burson, Richard Larrick and Joshua Klayman showed that manipulating task difficulty flips the pattern. On easy tasks the bottom performers overestimate themselves badly, but on hard tasks it is the top performers who underestimate themselves more severely. What sets the direction of the error was not human incompetence but the difficulty of the task.
The Irony of the Internet — an Effect Turned into a Weapon
Online, the uses of this effect converge on almost exactly one thing: putting someone else down. "Textbook Dunning-Kruger."
Whoever writes that sentence is doing three things at once. They are citing a graph that does not exist in the original paper, they are not testing the accuracy of their own judgment, and they are missing that the actual conclusion of the paper is about a blind spot everyone has rather than about "a certain type of person." Dunning himself has said repeatedly, in essays and interviews since, that the effect is a tool for self-examination, not a label to pin on other people.
The point that hundreds of accumulated citations do not make a claim sturdy was already made once in the ego depletion installment. Dunning-Kruger is a slightly different case. The literature was not inflated. Rather, a summary written by people who never read the paper spread far more widely than the original.
So What to Discard and What to Keep
What to discard: that curve with its peak and its valley. And the character sketch that says "the less competent you are, the more confident you feel." The first has no source, and the second is in large part a picture manufactured by the drawing convention of quartile averages. Pinning this name on a particular person or group is a usage that not even the original paper supports.
What to keep 1 — self-assessment is a function of structure, not of ability. As the Burson team showed, changing the difficulty changes the direction of the error, and the more immediate and specific the feedback, the more accurate self-assessment becomes. So the job in practice is not to comment on how confident someone is but to shorten the feedback loop. That is also why the real conditions for deliberate practice counts immediate feedback among the core requirements.
What to keep 2 — Study 4 of the original paper still stands. When bottom-quartile participants were given a short course in logic, not only did their scores rise, the accuracy of their self-assessments rose along with them. This design is not entirely free of regression effects either, but at least it is a different kind of evidence from the illusion of the quartile graph. The ability to notice what you do not know arrives as a by-product of training.
What to keep 3 — be suspicious, on principle, of any chart that splits a group by extreme values and then compares averages. The internal report claiming that the bottom 20 percent of performers improved after training, the dashboard measuring gains against the worst month on record: all of it sits on the same trap. The most practical legacy of the Dunning-Kruger debate lies not in psychology but on this statistical side.
A Guide to Reading the Originals
- Original paper: Kruger, J., & Dunning, D. (1999). Unskilled and unaware of it: How difficulties in recognizing one's own incompetence lead to inflated self-assessments. Journal of Personality and Social Psychology, 77(6), 1121-1134.
- Statistical objection: Krueger, J., & Mueller, R. A. (2002). Unskilled, unaware, or both? The better-than-average heuristic and statistical regression predict errors in estimates of own performance. Journal of Personality and Social Psychology, 82(2), 180-188.
- Difficulty effect: Burson, K. A., Larrick, R. P., & Klayman, J. (2006). Skilled or unskilled, but still unaware of it: How perceptions of difficulty drive miscalibration in relative comparisons. Journal of Personality and Social Psychology, 90(1), 60-77.
- Random-number simulation: Nuhfer, E., Fleisher, S., Cogan, C., Wirth, K., & Gaze, E. (2017). How random noise and a graphical convention subverted behavioral scientists' explanations of self-assessment data: Numeracy underlies better alternatives. Numeracy, 10(1), Article 4.
- Reanalysis: Gignac, G. E., & Zajenkowski, M. (2020). The Dunning-Kruger effect is (mostly) a statistical artefact: Valid approaches to testing the hypothesis with individual differences data. Intelligence, 80, 101449.
Reading tip: with the 1999 paper, look at Figure 1 (the humor task) and Figure 3 (the grammar task) before you read the text. How flat the self-assessment line is, and how the two lines invert at the right-hand edge, register in about 30 seconds. Then set the random-number simulation figure from Nuhfer 2017 beside it and the whole debate falls into place. For Gignac and Zajenkowski 2020, the highlight is the figure that presents individual scatter plots instead of quartile averages.
현재 단락 (1/39)
There is a picture you have almost certainly seen. Experience on the horizontal axis, confidence on ...