Split View: 똑똑하다는 것의 정의: 무엇이 측정되고 무엇이 측정되지 않는가
똑똑하다는 것의 정의: 무엇이 측정되고 무엇이 측정되지 않는가
들어가며
똑똑하다는 말은 하루에도 여러 번 쓰이지만, 그 안에는 최소한 세 가지 다른 것이 섞여 있습니다.
하나는 심리측정학이 다루는 통계적 규칙성입니다. 둘째는 그 규칙성을 설명하려는 이론 모형입니다. 셋째는 일상에서 사람을 평가할 때 쓰는 사회적 용법입니다. 이 셋은 서로 다르고, 섞으면 대화가 성립하지 않습니다.
이 글의 목표는 어느 쪽이 옳은지를 정하는 것이 아니라 지형을 그리는 것입니다. 여기에는 여전히 열려 있는 논쟁이 있고, 저는 그것을 닫지 않겠습니다.
1. 시작하기 전에 분명히 해둘 것
이 글은 검사와 모형에 대한 글입니다. 사람에 대한 글이 아닙니다. 그래서 세 가지를 먼저 못 박아두겠습니다.
첫째, 측정된 능력은 사람의 가치가 아닙니다. 인지 검사 점수는 특정 시점에 특정 조건에서 특정 종류의 과제를 수행한 기록입니다. 그 기록으로 한 사람의 가치나 존엄을 말할 수 없습니다. 이것은 예의상 붙이는 말이 아니라 측정이라는 행위의 논리적 한계입니다. 어떤 도구든 자기가 재도록 만들어진 것만 잽니다.
둘째, 이 글은 어떤 집단과 어떤 집단을 비교하지 않습니다. 그런 비교는 이 글의 주제가 아니며, 여기서 다루는 어떤 자료도 그런 결론의 근거로 쓰일 수 없습니다.
셋째, 사람을 줄 세우는 데 이 글의 내용을 쓸 수 없습니다. 아래에서 보겠지만, 측정 도구들의 예측력은 그런 용도로 쓰기에 한참 부족합니다.
2. 양의 상관 다발이라는 관찰
지능 연구의 출발점은 이론이 아니라 자료의 패턴입니다. 서로 성격이 다른 인지 과제들의 점수가 대체로 양의 상관을 보인다는 관찰입니다. 어휘 검사와 도형 추론처럼 겉으로는 관련 없어 보이는 과제들 사이에도 양의 상관이 나타납니다.
McGrew와 동료들이 2023년 Journal of Intelligence에 발표한 논문은 이 지점을 정확하게 짚습니다. 저자들은 일반지능이 주류 지능 연구의 일차적 사실이 아니며, 일차적 사실은 양의 상관 다발이고 일반요인은 그 사실에 대한 하나의 해석이라고 씁니다 (McGrew 외, 2023, 2026-08-16 읽음).
이 구분이 이 글에서 가장 중요한 한 줄입니다. 상관 다발이 존재한다는 것은 반복 관찰된 사실입니다. 그 다발의 배후에 하나의 원인이 있다는 것은 해석입니다. 앞의 것은 튼튼하고, 뒤의 것은 논쟁 중입니다.
3. CHC라는 주류 모형
심리측정학에서 가장 널리 쓰이는 구조 모형은 CHC입니다. Cattell, Horn, Carroll 세 사람의 작업이 통합된 이름입니다.
이 모형은 능력을 층으로 나눕니다. 맨 위에 일반요인, 그 아래에 유동추론, 결정지능, 시각처리, 청각처리, 작업기억, 인출 유창성, 처리속도 같은 넓은 능력들, 그 아래에 훨씬 많은 좁은 능력들이 놓입니다.
McGrew와 동료들의 2023년 연구는 심리측정 네트워크 분석이라는 다른 방법으로 같은 자료를 봤을 때, 지배적인 일반요인을 전제하지 않고도 일곱 개의 넓은 CHC 차원이 확인되었다고 보고합니다. 즉 넓은 능력들의 구분 자체는 방법을 바꿔도 남았습니다.
여기서 CHC가 이론적 진리라는 뜻은 아닙니다. 검사 도구들이 CHC 틀에 맞춰 만들어졌기 때문에 자료가 그 틀에 맞게 나오는 부분도 있습니다. 모형과 측정 도구가 서로를 강화하는 구조는 이 분야가 오래 안고 있는 문제입니다.
4. 일반요인의 해석은 왜 논쟁 중인가
McGrew와 동료들은 이 논쟁을 두 진영으로 정리합니다. 한쪽은 위계 모형에서 넓은 CHC 점수들이 여전히 해석 가치를 갖는다고 보고, 다른 쪽은 넓은 능력들이 전체 IQ를 넘어 제공하는 정보가 미미하다고 봅니다. 저자들은 네트워크 분석을 이 교착을 넘어서는 대안으로 제시합니다.
같은 논문이 짚는 또 하나의 구분이 심리측정적 일반요인과 이론적 일반요인입니다. 앞의 것은 요인분석으로 뽑아낸 통계량이고, 뒤의 것은 그 뒤에 어떤 생물학적 기제가 있다는 주장입니다. 저자들은 공통원인 모형이 이 둘을 뒤섞어 이론적 혼란을 만든다고 비판합니다.
일상 대화에서 벌어지는 혼란도 정확히 여기입니다. "머리가 좋다"는 말이 검사 점수를 가리키는지, 그 점수 뒤에 있다고 가정된 무언가를 가리키는지가 구분되지 않습니다. 그리고 대부분의 경우 후자로 쓰입니다.
5. 예측력은 양쪽 방향으로 과장된다
여기서 정직하게 밝혀둘 것이 있습니다. 저는 이번에 인지 능력 검사와 학업이나 소득 같은 결과 사이의 구체적인 상관계수를 제시하는 자료를 확인하지 못했습니다. 그래서 그 수치는 이 글에 쓰지 않겠습니다. 기억나는 숫자를 적는 것은 이 시리즈에서 하지 않기로 한 일입니다.
대신 확인한 것이 하나 있습니다. 예측력 추정치 자체가 과대추정되어 왔다는 방법론적 발견입니다.
Sackett과 동료들이 2022년 Journal of Applied Psychology에 발표한 논문은 인사 선발 분야의 메타분석들이 범위 제한을 보정하는 방식을 재검토했습니다 (Sackett 외, 2022, 2026-08-16 읽음).
저자들은 흔히 쓰여 온 다섯 가지 보정 방식을 검토한 뒤, 각각이 상당한 과대보정을 낳는 문제를 갖고 있으며 그 결과 많은 선발 도구의 타당도가 실질적으로 과대추정되어 왔다고 결론지었습니다. 재계산한 추정치에서 순위가 높던 도구들은 대체로 여전히 높았지만, 평균 타당도 값은 0.10에서 0.20 정도 낮아졌습니다. 그리고 구조화 면접이 가장 높은 순위로 올라왔습니다.
이 발견이 왜 중요한지는 두 방향으로 읽어야 합니다.
한 방향은 과장을 낮추는 것입니다. 널리 인용되던 수치들이 방법론적 이유로 부풀려져 있었다면, 그 수치에 기대어 세운 주장들도 함께 약해집니다.
다른 방향은 반대 과장도 경계하는 것입니다. 저자들의 결론은 선발 도구가 쓸모없다는 것이 아니라, 예측 관계가 이전에 생각했던 것보다 상당히 낮다는 것이었습니다. 낮아진 것과 사라진 것은 다릅니다.
그래서 이 항목의 정확한 요약은 이렇습니다. 측정된 능력의 예측력은 흔히 인용되는 것보다 작고, 동시에 0은 아닙니다. 양쪽 과장이 다 흔합니다.
6. 다중지능은 왜 널리 가르쳐지고 왜 근거가 약한가
다중지능 이론은 한국의 교육 현장과 자기계발 담론에 깊이 들어와 있습니다. 언어 지능, 논리수학 지능, 신체운동 지능처럼 여러 종류의 지능이 독립적으로 존재한다는 틀입니다.
먼저 이 이론이 왜 매력적인지를 인정하고 넘어가야 합니다. 하나의 척도로 사람을 줄 세우는 방식에 대한 반발에서 나왔고, 각자 다른 강점이 있다는 메시지는 그 자체로 좋은 메시지입니다. 이 틀을 배운 사람들이 어리석어서 배운 것이 아닙니다.
문제는 증거입니다. Waterhouse가 2023년 Frontiers in Psychology에 쓴 논문은 이 이론을 신경신화로 규정하고 세 가지 공백을 지적합니다 (Waterhouse, 2023, 2026-08-16 읽음).
첫째, 각 지능을 재는 표준 측정 도구가 존재하지 않습니다. 둘째, 요인분석 결과는 지능들이 독립적이라기보다 서로 상관되어 있음을 보여주며, 이는 이론의 핵심 주장과 어긋납니다. 셋째, 각 지능에 대응하는 신경 상관물이 확인되지 않았습니다.
같은 논문이 보급 정도를 보여주는 조사들을 인용합니다. 미국 예비교사의 90퍼센트가 다중지능 전략을 수업에 쓸 계획이라고 답했고, 퀘벡 교사의 94퍼센트가 실제로 쓰고 있다고 답했습니다. 반면 캐나다, 미국, 영국 교사 중 다중지능을 신경신화로 인식한 비율은 25퍼센트였습니다.
이 격차가 이 항목의 요점입니다. 널리 믿어진다는 것과 검증되었다는 것은 완전히 다른 문제입니다.
7. 학습 유형이라는 더 흔한 신화
같은 구조가 더 크게 나타나는 것이 학습 유형입니다. 시각형, 청각형, 체험형으로 나누고 그에 맞춰 가르치면 더 잘 배운다는 주장입니다.
Brown이 2023년 Frontiers in Education에 쓴 리뷰는 이 믿음의 보급률과 증거 상태를 함께 정리합니다 (Brown, 2023, 2026-08-16 읽음).
보급률부터 보겠습니다. Dekker와 동료들의 2012년 조사에서 영국 교사의 93퍼센트, 네덜란드 교사의 96퍼센트가 이 가설을 지지했습니다. Newton과 Salvi의 2020년 조사에서는 89.1퍼센트였습니다. Hughes와 동료들의 2020년 조사에서 호주 교사의 79퍼센트 이상이 지지했습니다.
증거 상태는 이렇습니다. Pashler와 동료들이 2008년에 지적한 대로, 이 가설을 제대로 검정하려면 교차 상호작용을 보여야 합니다. 즉 시각형에게는 시각 자료가, 청각형에게는 청각 자료가 더 나아야 하고, 두 선이 교차해야 합니다. Brown의 정리에 따르면 검토된 연구의 75퍼센트가 그 교차 상호작용을 산출하지 못했습니다. 이후의 여러 연구에서도 필수적인 교차 상호작용은 나타나지 않았습니다.
즉 학습 유형에 맞춘 수업이 효과적이라는 주장은 대중적으로 널리 퍼져 있지만 근거가 뒷받침하지 않는 주장으로 분류됩니다.
여기서 조심할 것이 하나 있습니다. 사람마다 선호하는 학습 방식이 있다는 것은 별개의 이야기입니다. 선호는 실제로 존재합니다. 무너지는 것은 선호의 존재가 아니라, 선호에 맞추면 더 잘 배운다는 연결 고리입니다.
8. 전문성은 영역에 붙어 있다
이 시리즈의 앞 두 글에서 본 내용이 여기에도 걸립니다.
Sala와 동료들의 2019년 2차 메타분석은 인지훈련의 효과가 훈련한 과제와 그 근처를 넘어서지 않는다는 결론을 내렸습니다. 위약 효과와 출판 편향을 통제하면 원전이 효과 크기와 참분산이 0이 되었습니다 (Sala 외, 2019, 2026-08-16 읽음).
그리고 체스 연구가 보여준 것처럼, 전문가의 우위는 그 영역의 패턴이 살아 있을 때 나타납니다 (Bartlett 외, 2013, 2026-08-16 읽음).
이 둘을 합치면 실용적으로 중요한 결론이 나옵니다. 어떤 사람이 한 영역에서 대단히 뛰어나다는 사실은 다른 영역에 대해 알려주는 것이 매우 적습니다. 이것은 겸손을 권하는 말이 아니라 자료가 말하는 바입니다.
9. 사회적으로 쓰이는 똑똑하다
마지막으로, 실제 대화에서 이 단어가 어떻게 쓰이는지를 보겠습니다. 이 절은 연구 결과가 아니라 관찰입니다. 그렇게 읽어주시기 바랍니다.
일상에서 누군가를 똑똑하다고 말할 때, 실제로 가리키는 것은 대체로 이런 것들입니다.
- 말이 빠르고 설명이 매끄럽다.
- 이 자리에서 통용되는 어휘와 참조를 알고 있다.
- 자신 있게 말한다.
- 나와 비슷한 방식으로 생각한다.
이 목록에는 위에서 다룬 어떤 측정 개념도 없습니다. 유창함은 능력의 신호일 수 있지만 유창함 자체가 능력은 아니고, 자신감은 능력과 독립적으로 변동합니다. 그리고 마지막 항목은 능력에 대한 판단이 아니라 유사성에 대한 판단입니다.
이 관찰이 실용적으로 중요한 이유는, 사람들이 이 사회적 용법을 심리측정적 개념과 같은 것으로 다루기 때문입니다. 회의에서 말이 매끄러운 사람이 실제로 판단이 좋은지는 별개의 문제이고, 그것을 확인하려면 결과를 봐야 합니다.
10. 주장별 증거 등급
| 주장 | 증거 등급 | 근거 |
|---|---|---|
| 인지 과제 점수들 사이에 양의 상관 다발이 있다 | 반복 검증됨 | McGrew 외 2023 |
| 그 다발의 배후에 단일 원인이 있다 | 논쟁 중 | McGrew 외 2023이 두 진영을 정리 |
| CHC의 넓은 능력 구분은 방법을 바꿔도 남는다 | 반복 검증됨 | McGrew 외 2023 네트워크 분석 |
| 넓은 능력 점수가 전체 IQ 이상의 정보를 준다 | 논쟁 중 | McGrew 외 2023이 인용한 대립 진영 |
| 선발 도구의 타당도가 과대추정되어 왔다 | 방법론적 재분석 | Sackett 외 2022 |
| 다중지능 이론에 표준 측정 도구가 없다 | 비판 문헌의 정리 | Waterhouse 2023 |
| 다중지능이 교육 현장에 널리 쓰인다 | 조사 자료 | Waterhouse 2023이 인용한 설문들 |
| 학습 유형에 맞춘 수업이 효과적이다 | 대중적이지만 근거 부족 | Brown 2023, Pashler 외 2008 인용 |
| 전문성은 영역을 넘어 일반화되지 않는다 | 반복 검증됨 | Sala 외 2019, Bartlett 외 2013 |
| 일상적 똑똑함 판단이 측정 개념과 일치한다 | 근거 없음 | 이 글의 관찰, 검증 자료 없음 |
무엇이 검증되지 않았나
첫째, 앞에서 밝힌 대로 저는 이 글에 인지 능력과 실제 성과 사이의 구체적 상관계수를 쓰지 않았습니다. 신뢰할 만한 형태로 확인하지 못했기 때문입니다. 이 글에서 가장 큰 공백입니다.
둘째, 일반요인의 해석 논쟁은 열려 있습니다. 저는 어느 쪽으로도 결론짓지 않았고, 그것이 현재 상태를 가장 정확하게 반영한다고 생각합니다. 다만 이 논쟁이 열려 있다는 것과 아무것도 모른다는 것은 다릅니다. 상관 다발의 존재는 논쟁 대상이 아닙니다.
셋째, 다중지능에 대한 판단은 비판 문헌 한 편에 크게 기대고 있습니다. Waterhouse는 이 이론의 오랜 비판자이고, 저는 반대 진영의 자료를 이번에 직접 읽지 못했습니다. 그 사실을 감안해서 읽어야 합니다. 다만 표준 측정 도구가 없다는 지적은 확인하거나 반박하기가 비교적 쉬운 종류의 주장입니다.
넷째, 학습 유형에 대해서는 최근 몇 년 사이 반론 성격의 메타분석도 발표되고 있다는 것을 검색 과정에서 봤습니다. 다만 그 논문들을 직접 읽지 못했으므로 내용을 요약하지 않겠습니다. 이 항목의 결론이 앞으로 조정될 가능성은 열어둡니다.
다섯째, 이 글은 능력의 발달, 환경의 영향, 교육의 효과를 다루지 않았습니다. 구조 모형에 대한 글과 발달에 대한 글은 다른 글입니다.
마치며
정리하면 이렇습니다.
서로 다른 인지 과제 점수들이 양의 상관을 이룬다는 관찰은 견고합니다. 그 관찰을 무엇으로 설명할지는 아직 논쟁 중입니다. 그 논쟁을 대중적으로 대체해 온 다중지능과 학습 유형은 널리 퍼져 있지만 근거가 약합니다. 그리고 어떤 측정치든 예측력은 흔히 인용되는 것보다 작습니다.
이 네 문장을 합치면 실용적인 결론이 하나 나옵니다. 자기 자신이든 남이든, 사람을 일반 능력의 크기로 요약하려는 시도는 자료가 지지하지 않습니다. 확인 가능한 것은 훨씬 좁고 구체적인 것들입니다. 무엇을 할 수 있는지, 어떤 조건에서 그렇게 되는지, 무엇이 아직 안 되는지.
그리고 처음에 적은 문장을 다시 적겠습니다. 측정된 능력은 사람의 가치가 아닙니다. 이 글에 나온 어떤 수치도 그런 용도로 쓰일 수 없습니다.
이어지는 글:
- 이전 글: 연습의 구조 — 간격, 인출, 섞기.
- 다음 글: 전문가가 되는 길 — 지금까지의 내용을 실행 가능한 형태로.
- 사고 실험실 — 판단의 편향을 직접 겪어보는 도구.
- 논리 실험실 — 논증 구조를 분해해보는 도구.
참고 자료
아래는 이 글을 쓰면서 2026-08-16에 직접 확인한 자료입니다. 링크는 실제로 읽은 주소이며, 일부는 Europe PMC와 Semantic Scholar의 공개 API 응답입니다.
- McGrew, K. S., Schneider, W. J., Decker, S. L., & Bulut, O. (2023). A Psychometric Network Analysis of CHC Intelligence Measures. Journal of Intelligence
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection. Journal of Applied Psychology
- Waterhouse, L. (2023). Why multiple intelligences theory is a neuromyth. Frontiers in Psychology
- Brown, S. B. R. E. (2023). The persistence of matching teaching and learning styles. Frontiers in Education
- Sala, G., Aksayli, N. D., Tatlidil, K. S., Tatsumi, T., Gondo, Y., & Gobet, F. (2019). Near and Far Transfer in Cognitive Training: A Second-Order Meta-Analysis. Collabra: Psychology
- Bartlett, J. C., Boggan, A. L., & Krawczyk, D. C. (2013). Expertise and processing distorted structure in chess. Frontiers in Human Neuroscience
What It Means to Be Smart: What Gets Measured and What Does Not
Introduction
The word smart gets used several times a day, but at least three different things are mixed inside it.
One is the statistical regularity that psychometrics deals with. The second is the theoretical model that tries to explain that regularity. The third is the social usage we reach for when we size someone up in daily life. These three are different, and mixing them means the conversation does not hold together.
The goal of this post is not to decide which side is right but to draw the terrain. There are debates here that remain open, and I am not going to close them.
1. What to make clear before starting
This post is about tests and models. It is not about people. So let me nail down three things first.
First, measured ability is not the worth of a person. A score on a cognitive test is a record of performing a particular kind of task under particular conditions at a particular moment. You cannot speak of a person's worth or dignity from that record. This is not a line added out of politeness but a logical limit of the act of measuring. Any instrument measures only what it was built to measure.
Second, this post does not compare one group with another group. Comparisons of that kind are not the subject of this post, and none of the material handled here can be used as grounds for a conclusion of that kind.
Third, the content of this post cannot be used to rank people. As we will see below, the predictive power of these measuring instruments falls far short of that use.
2. The observation called the positive manifold
The starting point of intelligence research is not a theory but a pattern in the data. It is the observation that scores on cognitive tasks of quite different character tend to correlate positively. A positive correlation shows up even between tasks that look unrelated on the surface, such as a vocabulary test and figural reasoning.
The paper McGrew and colleagues published in Journal of Intelligence in 2023 puts a finger on exactly this point. The authors write that general intelligence is not the primary fact of mainstream intelligence research, that the primary fact is the positive manifold, and that the general factor is one interpretation of that fact (McGrew et al., 2023, read 2026-08-16).
This distinction is the most important line in this post. That a positive manifold exists is a repeatedly observed fact. That there is a single cause behind the manifold is an interpretation. The former is sturdy, and the latter is under debate.
3. CHC, the mainstream model
The structural model most widely used in psychometrics is CHC. The name comes from the integrated work of three people, Cattell, Horn, and Carroll.
The model divides ability into strata. At the top sits the general factor. Below it are broad abilities such as fluid reasoning, crystallised intelligence, visual processing, auditory processing, working memory, retrieval fluency, and processing speed. Below those sit far more numerous narrow abilities.
The 2023 study by McGrew and colleagues reports that when the same data were examined with a different method, psychometric network analysis, seven broad CHC dimensions were confirmed without presupposing a dominant general factor. That is, the distinction among the broad abilities survived the change of method.
This does not mean CHC is theoretical truth. Part of the reason the data come out fitting the CHC frame is that the test instruments were built to fit that frame. The structure in which a model and its measuring instruments reinforce each other is a problem this field has carried for a long time.
4. Why the interpretation of the general factor is still contested
McGrew and colleagues sort this debate into two camps. One side holds that in a hierarchical model the broad CHC scores still carry interpretive value. The other side holds that the information broad abilities provide beyond full-scale IQ is negligible. The authors put network analysis forward as an alternative that gets past this deadlock.
Another distinction the same paper points out is the one between a psychometric general factor and a theoretical general factor. The former is a statistic extracted by factor analysis, and the latter is the claim that some biological mechanism sits behind it. The authors criticise the common-cause model for mixing the two and producing theoretical confusion.
The confusion that happens in everyday conversation sits at exactly this spot. When someone says another person has a good head, whether that points at a test score or at something assumed to lie behind the score goes unmarked. And in most cases it is used in the second sense.
5. Predictive power is exaggerated in both directions
There is something to state honestly here. This time I could not confirm a source giving concrete correlation coefficients between cognitive ability tests and outcomes such as schooling or income. So I will not put those figures in this post. Writing down a number from memory is one of the things this series has decided not to do.
Instead there is one thing I did confirm. It is a methodological finding, that the estimates of predictive power have themselves been overestimated.
The paper Sackett and colleagues published in Journal of Applied Psychology in 2022 re-examined the way meta-analyses in personnel selection correct for range restriction (Sackett et al., 2022, read 2026-08-16).
The authors reviewed five commonly used correction methods and concluded that each carries a problem producing substantial overcorrection, and that the validity of many selection instruments has in consequence been materially overestimated. In the recalculated estimates, the instruments that ranked high generally still ranked high, but the mean validity values fell by something on the order of 0.10 to 0.20. And the structured interview came up to the highest rank.
Why this finding matters has to be read in two directions.
One direction is to bring the exaggeration down. If figures that were widely cited were inflated for methodological reasons, the claims built on top of those figures weaken along with them.
The other direction is to guard against the opposite exaggeration. The authors' conclusion was not that selection instruments are useless but that the predictive relationships are considerably lower than had been thought. Lowered and gone are not the same thing.
So the accurate summary of this item is this. The predictive power of measured ability is smaller than what is commonly cited, and at the same time it is not zero. Both exaggerations are common.
6. Why multiple intelligences is widely taught and why its evidence is weak
Multiple intelligences theory has gone deep into Korean education and into self-improvement discourse. It is a frame in which several kinds of intelligence exist independently, such as linguistic intelligence, logical-mathematical intelligence, and bodily-kinesthetic intelligence.
We should first acknowledge why the theory is attractive and then move on. It came out of a reaction against ranking people on a single scale, and the message that each person has different strengths is a good message in itself. The people who learned this frame did not learn it out of foolishness.
The problem is the evidence. The paper Waterhouse wrote in Frontiers in Psychology in 2023 defines the theory as a neuromyth and points to three gaps (Waterhouse, 2023, read 2026-08-16).
First, no standard instrument exists for measuring each intelligence. Second, factor analysis results show the intelligences to be correlated with one another rather than independent, which runs against the core claim of the theory. Third, no neural correlate corresponding to each intelligence has been confirmed.
The same paper cites surveys showing how far the theory has spread. In the United States, 90 percent of pre-service teachers answered that they planned to use multiple intelligences strategies in their teaching, and 94 percent of teachers in Quebec answered that they actually use them. By contrast, the proportion of teachers in Canada, the United States, and the United Kingdom who recognised multiple intelligences as a neuromyth was 25 percent.
This gap is the point of this item. Being widely believed and having been verified are completely different matters.
7. Learning styles, the more common myth
The same structure appears on a larger scale in learning styles. This is the claim that people divide into visual, auditory, and kinesthetic types, and that teaching matched to the type makes them learn better.
The review Brown wrote in Frontiers in Education in 2023 sets out the prevalence of this belief together with the state of the evidence (Brown, 2023, read 2026-08-16).
Take the prevalence first. In the 2012 survey by Dekker and colleagues, 93 percent of teachers in the United Kingdom and 96 percent of teachers in the Netherlands endorsed the hypothesis. In the 2020 survey by Newton and Salvi it was 89.1 percent. In the 2020 survey by Hughes and colleagues, more than 79 percent of Australian teachers endorsed it.
The state of the evidence is this. As Pashler and colleagues pointed out in 2008, testing this hypothesis properly requires showing a crossover interaction. That is, visual material should be better for visual types and auditory material better for auditory types, and the two lines should cross. By Brown's account, 75 percent of the studies reviewed failed to produce that crossover interaction. In the various studies since then, the essential crossover interaction has not appeared either.
So the claim that teaching matched to learning styles is effective is classified as a claim that is widely spread in popular belief but not supported by the evidence.
There is one thing to be careful about here. That people have a preferred way of learning is a separate story. Preferences really do exist. What collapses is not the existence of the preference but the link that says matching the preference makes for better learning.
8. Expertise is attached to its domain
What we saw in the first two posts of this series catches here as well.
The second-order meta-analysis by Sala and colleagues in 2019 concluded that the effects of cognitive training do not go beyond the trained task and its neighbourhood. Once placebo effects and publication bias were controlled for, the far-transfer effect size and the true variance became zero (Sala et al., 2019, read 2026-08-16).
And as chess research showed, the expert's advantage appears when the patterns of that domain are alive (Bartlett et al., 2013, read 2026-08-16).
Put the two together and a practically important conclusion comes out. The fact that someone is extremely good in one domain tells you very little about another domain. This is not a plea for humility but what the data say.
9. Smart as the word is used socially
Finally, let us look at how the word actually gets used in real conversation. This section is an observation, not a research finding. Please read it that way.
When we call someone smart in daily life, what we are actually pointing at is generally something like this.
- They talk fast and explain smoothly.
- They know the vocabulary and the references that circulate in this room.
- They speak with confidence.
- They think in a way similar to mine.
None of the measurement concepts handled above appears in this list. Fluency can be a signal of ability, but fluency itself is not ability, and confidence fluctuates independently of ability. And the last item is not a judgement about ability but a judgement about similarity.
The reason this observation matters practically is that people treat this social usage as the same thing as the psychometric concept. Whether the person who speaks smoothly in a meeting actually has good judgement is a separate question, and checking that means looking at outcomes.
10. Evidence grade by claim
| Claim | Evidence grade | Source |
|---|---|---|
| There is a positive manifold among scores on cognitive tasks | Replicated | McGrew et al. 2023 |
| A single cause sits behind that manifold | Under debate | McGrew et al. 2023 sets out the two camps |
| The CHC distinction among broad abilities survives a change of method | Replicated | McGrew et al. 2023 network analysis |
| Broad ability scores give information beyond full-scale IQ | Under debate | The opposing camp cited by McGrew et al. 2023 |
| The validity of selection instruments has been overestimated | Methodological reanalysis | Sackett et al. 2022 |
| Multiple intelligences theory has no standard measuring instrument | Summary of the critical literature | Waterhouse 2023 |
| Multiple intelligences is widely used in education | Survey data | The surveys cited by Waterhouse 2023 |
| Teaching matched to learning styles is effective | Popular but insufficiently supported | Brown 2023, citing Pashler et al. 2008 |
| Expertise does not generalise beyond its domain | Replicated | Sala et al. 2019, Bartlett et al. 2013 |
| Everyday judgements of smartness match the measurement concepts | No evidence | Observation in this post, no verifying data |
What has not been verified
First, as stated above, I did not put concrete correlation coefficients between cognitive ability and actual performance into this post, because I could not confirm them in a trustworthy form. This is the largest gap in the post.
Second, the debate over how to interpret the general factor is open. I have not concluded in either direction, and I think that reflects the current state most accurately. But this debate being open and nothing being known are different things. The existence of the positive manifold is not what is in dispute.
Third, the judgement about multiple intelligences leans heavily on a single piece of critical literature. Waterhouse is a long-standing critic of the theory, and this time I did not read material from the opposing camp directly. That should be taken into account in reading it. Still, the point that no standard measuring instrument exists is the kind of claim that is comparatively easy to confirm or to refute.
Fourth, on learning styles, I saw in the course of searching that meta-analyses of a rebutting character have also been published over the last few years. But I did not read those papers directly, so I will not summarise their contents. I leave open the possibility that the conclusion of this item will be adjusted in future.
Fifth, this post did not deal with the development of ability, the influence of environment, or the effects of education. A post about structural models and a post about development are different posts.
Closing
To sum up.
The observation that scores on different cognitive tasks form positive correlations is solid. What to explain that observation with is still under debate. Multiple intelligences and learning styles, which have popularly stood in for that debate, are widespread but weakly evidenced. And whatever the measure, its predictive power is smaller than what is commonly cited.
Put the four sentences together and one practical conclusion comes out. The attempt to summarise a person, yourself or anyone else, by the size of a general ability is not supported by the data. What can be checked is far narrower and more concrete. What someone can do, under what conditions they can do it, and what they still cannot do.
And let me write again the sentence I wrote at the beginning. Measured ability is not the worth of a person. No figure in this post can be used for that purpose.
Related posts:
- Previous post: The Structure of Practice — spacing, retrieval, mixing.
- Next post: The Road to Becoming an Expert — everything so far, in an executable form.
- Thinking Lab — a tool for going through biases of judgement yourself.
- Logic Lab — a tool for taking argument structures apart.
References
Below are the sources checked directly on 2026-08-16 while writing this post. The links are the addresses actually read, and some are public API responses from Europe PMC and Semantic Scholar.
- McGrew, K. S., Schneider, W. J., Decker, S. L., & Bulut, O. (2023). A Psychometric Network Analysis of CHC Intelligence Measures. Journal of Intelligence
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection. Journal of Applied Psychology
- Waterhouse, L. (2023). Why multiple intelligences theory is a neuromyth. Frontiers in Psychology
- Brown, S. B. R. E. (2023). The persistence of matching teaching and learning styles. Frontiers in Education
- Sala, G., Aksayli, N. D., Tatlidil, K. S., Tatsumi, T., Gondo, Y., & Gobet, F. (2019). Near and Far Transfer in Cognitive Training: A Second-Order Meta-Analysis. Collabra: Psychology
- Bartlett, J. C., Boggan, A. L., & Krawczyk, D. C. (2013). Expertise and processing distorted structure in chess. Frontiers in Human Neuroscience