Skip to content
Published on

What It Means to Be Smart: What Gets Measured and What Does Not

Share
Authors

Introduction

The word smart gets used several times a day, but at least three different things are mixed inside it.

One is the statistical regularity that psychometrics deals with. The second is the theoretical model that tries to explain that regularity. The third is the social usage we reach for when we size someone up in daily life. These three are different, and mixing them means the conversation does not hold together.

The goal of this post is not to decide which side is right but to draw the terrain. There are debates here that remain open, and I am not going to close them.

1. What to make clear before starting

This post is about tests and models. It is not about people. So let me nail down three things first.

First, measured ability is not the worth of a person. A score on a cognitive test is a record of performing a particular kind of task under particular conditions at a particular moment. You cannot speak of a person's worth or dignity from that record. This is not a line added out of politeness but a logical limit of the act of measuring. Any instrument measures only what it was built to measure.

Second, this post does not compare one group with another group. Comparisons of that kind are not the subject of this post, and none of the material handled here can be used as grounds for a conclusion of that kind.

Third, the content of this post cannot be used to rank people. As we will see below, the predictive power of these measuring instruments falls far short of that use.

2. The observation called the positive manifold

The starting point of intelligence research is not a theory but a pattern in the data. It is the observation that scores on cognitive tasks of quite different character tend to correlate positively. A positive correlation shows up even between tasks that look unrelated on the surface, such as a vocabulary test and figural reasoning.

The paper McGrew and colleagues published in Journal of Intelligence in 2023 puts a finger on exactly this point. The authors write that general intelligence is not the primary fact of mainstream intelligence research, that the primary fact is the positive manifold, and that the general factor is one interpretation of that fact (McGrew et al., 2023, read 2026-08-16).

This distinction is the most important line in this post. That a positive manifold exists is a repeatedly observed fact. That there is a single cause behind the manifold is an interpretation. The former is sturdy, and the latter is under debate.

3. CHC, the mainstream model

The structural model most widely used in psychometrics is CHC. The name comes from the integrated work of three people, Cattell, Horn, and Carroll.

The model divides ability into strata. At the top sits the general factor. Below it are broad abilities such as fluid reasoning, crystallised intelligence, visual processing, auditory processing, working memory, retrieval fluency, and processing speed. Below those sit far more numerous narrow abilities.

The 2023 study by McGrew and colleagues reports that when the same data were examined with a different method, psychometric network analysis, seven broad CHC dimensions were confirmed without presupposing a dominant general factor. That is, the distinction among the broad abilities survived the change of method.

This does not mean CHC is theoretical truth. Part of the reason the data come out fitting the CHC frame is that the test instruments were built to fit that frame. The structure in which a model and its measuring instruments reinforce each other is a problem this field has carried for a long time.

4. Why the interpretation of the general factor is still contested

McGrew and colleagues sort this debate into two camps. One side holds that in a hierarchical model the broad CHC scores still carry interpretive value. The other side holds that the information broad abilities provide beyond full-scale IQ is negligible. The authors put network analysis forward as an alternative that gets past this deadlock.

Another distinction the same paper points out is the one between a psychometric general factor and a theoretical general factor. The former is a statistic extracted by factor analysis, and the latter is the claim that some biological mechanism sits behind it. The authors criticise the common-cause model for mixing the two and producing theoretical confusion.

The confusion that happens in everyday conversation sits at exactly this spot. When someone says another person has a good head, whether that points at a test score or at something assumed to lie behind the score goes unmarked. And in most cases it is used in the second sense.

5. Predictive power is exaggerated in both directions

There is something to state honestly here. This time I could not confirm a source giving concrete correlation coefficients between cognitive ability tests and outcomes such as schooling or income. So I will not put those figures in this post. Writing down a number from memory is one of the things this series has decided not to do.

Instead there is one thing I did confirm. It is a methodological finding, that the estimates of predictive power have themselves been overestimated.

The paper Sackett and colleagues published in Journal of Applied Psychology in 2022 re-examined the way meta-analyses in personnel selection correct for range restriction (Sackett et al., 2022, read 2026-08-16).

The authors reviewed five commonly used correction methods and concluded that each carries a problem producing substantial overcorrection, and that the validity of many selection instruments has in consequence been materially overestimated. In the recalculated estimates, the instruments that ranked high generally still ranked high, but the mean validity values fell by something on the order of 0.10 to 0.20. And the structured interview came up to the highest rank.

Why this finding matters has to be read in two directions.

One direction is to bring the exaggeration down. If figures that were widely cited were inflated for methodological reasons, the claims built on top of those figures weaken along with them.

The other direction is to guard against the opposite exaggeration. The authors' conclusion was not that selection instruments are useless but that the predictive relationships are considerably lower than had been thought. Lowered and gone are not the same thing.

So the accurate summary of this item is this. The predictive power of measured ability is smaller than what is commonly cited, and at the same time it is not zero. Both exaggerations are common.

6. Why multiple intelligences is widely taught and why its evidence is weak

Multiple intelligences theory has gone deep into Korean education and into self-improvement discourse. It is a frame in which several kinds of intelligence exist independently, such as linguistic intelligence, logical-mathematical intelligence, and bodily-kinesthetic intelligence.

We should first acknowledge why the theory is attractive and then move on. It came out of a reaction against ranking people on a single scale, and the message that each person has different strengths is a good message in itself. The people who learned this frame did not learn it out of foolishness.

The problem is the evidence. The paper Waterhouse wrote in Frontiers in Psychology in 2023 defines the theory as a neuromyth and points to three gaps (Waterhouse, 2023, read 2026-08-16).

First, no standard instrument exists for measuring each intelligence. Second, factor analysis results show the intelligences to be correlated with one another rather than independent, which runs against the core claim of the theory. Third, no neural correlate corresponding to each intelligence has been confirmed.

The same paper cites surveys showing how far the theory has spread. In the United States, 90 percent of pre-service teachers answered that they planned to use multiple intelligences strategies in their teaching, and 94 percent of teachers in Quebec answered that they actually use them. By contrast, the proportion of teachers in Canada, the United States, and the United Kingdom who recognised multiple intelligences as a neuromyth was 25 percent.

This gap is the point of this item. Being widely believed and having been verified are completely different matters.

7. Learning styles, the more common myth

The same structure appears on a larger scale in learning styles. This is the claim that people divide into visual, auditory, and kinesthetic types, and that teaching matched to the type makes them learn better.

The review Brown wrote in Frontiers in Education in 2023 sets out the prevalence of this belief together with the state of the evidence (Brown, 2023, read 2026-08-16).

Take the prevalence first. In the 2012 survey by Dekker and colleagues, 93 percent of teachers in the United Kingdom and 96 percent of teachers in the Netherlands endorsed the hypothesis. In the 2020 survey by Newton and Salvi it was 89.1 percent. In the 2020 survey by Hughes and colleagues, more than 79 percent of Australian teachers endorsed it.

The state of the evidence is this. As Pashler and colleagues pointed out in 2008, testing this hypothesis properly requires showing a crossover interaction. That is, visual material should be better for visual types and auditory material better for auditory types, and the two lines should cross. By Brown's account, 75 percent of the studies reviewed failed to produce that crossover interaction. In the various studies since then, the essential crossover interaction has not appeared either.

So the claim that teaching matched to learning styles is effective is classified as a claim that is widely spread in popular belief but not supported by the evidence.

There is one thing to be careful about here. That people have a preferred way of learning is a separate story. Preferences really do exist. What collapses is not the existence of the preference but the link that says matching the preference makes for better learning.

8. Expertise is attached to its domain

What we saw in the first two posts of this series catches here as well.

The second-order meta-analysis by Sala and colleagues in 2019 concluded that the effects of cognitive training do not go beyond the trained task and its neighbourhood. Once placebo effects and publication bias were controlled for, the far-transfer effect size and the true variance became zero (Sala et al., 2019, read 2026-08-16).

And as chess research showed, the expert's advantage appears when the patterns of that domain are alive (Bartlett et al., 2013, read 2026-08-16).

Put the two together and a practically important conclusion comes out. The fact that someone is extremely good in one domain tells you very little about another domain. This is not a plea for humility but what the data say.

9. Smart as the word is used socially

Finally, let us look at how the word actually gets used in real conversation. This section is an observation, not a research finding. Please read it that way.

When we call someone smart in daily life, what we are actually pointing at is generally something like this.

  • They talk fast and explain smoothly.
  • They know the vocabulary and the references that circulate in this room.
  • They speak with confidence.
  • They think in a way similar to mine.

None of the measurement concepts handled above appears in this list. Fluency can be a signal of ability, but fluency itself is not ability, and confidence fluctuates independently of ability. And the last item is not a judgement about ability but a judgement about similarity.

The reason this observation matters practically is that people treat this social usage as the same thing as the psychometric concept. Whether the person who speaks smoothly in a meeting actually has good judgement is a separate question, and checking that means looking at outcomes.

10. Evidence grade by claim

ClaimEvidence gradeSource
There is a positive manifold among scores on cognitive tasksReplicatedMcGrew et al. 2023
A single cause sits behind that manifoldUnder debateMcGrew et al. 2023 sets out the two camps
The CHC distinction among broad abilities survives a change of methodReplicatedMcGrew et al. 2023 network analysis
Broad ability scores give information beyond full-scale IQUnder debateThe opposing camp cited by McGrew et al. 2023
The validity of selection instruments has been overestimatedMethodological reanalysisSackett et al. 2022
Multiple intelligences theory has no standard measuring instrumentSummary of the critical literatureWaterhouse 2023
Multiple intelligences is widely used in educationSurvey dataThe surveys cited by Waterhouse 2023
Teaching matched to learning styles is effectivePopular but insufficiently supportedBrown 2023, citing Pashler et al. 2008
Expertise does not generalise beyond its domainReplicatedSala et al. 2019, Bartlett et al. 2013
Everyday judgements of smartness match the measurement conceptsNo evidenceObservation in this post, no verifying data

What has not been verified

First, as stated above, I did not put concrete correlation coefficients between cognitive ability and actual performance into this post, because I could not confirm them in a trustworthy form. This is the largest gap in the post.

Second, the debate over how to interpret the general factor is open. I have not concluded in either direction, and I think that reflects the current state most accurately. But this debate being open and nothing being known are different things. The existence of the positive manifold is not what is in dispute.

Third, the judgement about multiple intelligences leans heavily on a single piece of critical literature. Waterhouse is a long-standing critic of the theory, and this time I did not read material from the opposing camp directly. That should be taken into account in reading it. Still, the point that no standard measuring instrument exists is the kind of claim that is comparatively easy to confirm or to refute.

Fourth, on learning styles, I saw in the course of searching that meta-analyses of a rebutting character have also been published over the last few years. But I did not read those papers directly, so I will not summarise their contents. I leave open the possibility that the conclusion of this item will be adjusted in future.

Fifth, this post did not deal with the development of ability, the influence of environment, or the effects of education. A post about structural models and a post about development are different posts.

Closing

To sum up.

The observation that scores on different cognitive tasks form positive correlations is solid. What to explain that observation with is still under debate. Multiple intelligences and learning styles, which have popularly stood in for that debate, are widespread but weakly evidenced. And whatever the measure, its predictive power is smaller than what is commonly cited.

Put the four sentences together and one practical conclusion comes out. The attempt to summarise a person, yourself or anyone else, by the size of a general ability is not supported by the data. What can be checked is far narrower and more concrete. What someone can do, under what conditions they can do it, and what they still cannot do.

And let me write again the sentence I wrote at the beginning. Measured ability is not the worth of a person. No figure in this post can be used for that purpose.

Related posts:

References

Below are the sources checked directly on 2026-08-16 while writing this post. The links are the addresses actually read, and some are public API responses from Europe PMC and Semantic Scholar.