Skip to content
Published on

What the Research Shows: Interpreter Training, Cognitive Ability, and the Deliberate Practice Debate

Share
Authors

Introduction

The previous post looked at what interpreters train. The question in this one is different. Does that training actually change a person's cognitive ability?

The answer you see most often online is that interpreters have outstanding working memory. That sentence is about half right. The problem is which half. That interpreters score higher than a comparison group on a working memory test, and that interpreter training raises working memory, are completely different claims. The first has been observed in several studies; the second is not settled.

This post is about that gap. And it will show, rather than hide, the places where the literature in this field contradicts itself.

1. Working memory: there is a meta-analysis, but I read only the abstract

The most frequently cited meta-analysis is the one Mellinger and Hanson published in 2019 in the journal Interpreting, volume 21, issue 2 (pages 165 to 195). Using random-effects models, it pooled the difference between professional interpreters and comparison groups, and the correlation between working memory capacity and simultaneous interpreting performance. It examined as moderators the difference between auditory and visual stimuli, and the difference between storage tasks and processing tasks.

According to the abstract, professional interpreters had significantly higher working memory capacity than comparison groups, and a positive correlation appeared between working memory capacity and simultaneous interpreting quality.

Here is my honest disclosure. The full text of this paper sits behind a paywall, so I could not read it. I was unable to check exactly how many studies were included, what the effect sizes were, or how publication bias was tested. So I use this meta-analysis only as directional information: that a working memory difference between interpreters and comparison groups has been reported. I will not cite magnitudes.

And even directional information already calls for caution. Correlation is not causation. People with good working memory may have become interpreters, or interpreting may have improved their working memory. That distinction is the axis of this entire post.

2. Executive function: two reviews reached different conclusions

More interesting than working memory is executive function. The standard frame divides it into three parts: inhibition, shifting, and updating.

There are two reviews on this, and their conclusions differ.

The first is the systematic review that Nour and colleagues published in 2020 in Interpreting, volume 22, issue 2. It examined 17 studies. According to the abstract, it found no evidence of an interpreter advantage in inhibition; in updating, the interpreter group scored higher than comparison groups, but longitudinal studies showed no overall improvement among trainees; and in shifting, scores improved as a result of training. This paper is also paywalled, so I read only the abstract.

The second is the systematic review and meta-analysis that Hu and Fan published in 2021 in Forum for Linguistic Studies, volume 3, issue 1. This one is open access. Searching both international and Chinese-language databases, it pooled 98 tasks from 29 primary studies. The results were substantial evidence of an interpreter advantage in shifting (g of 0.68 across 7 effects based on the Wisconsin Card Sorting Test, and g of -0.32 across 8 effects based on switching cost), no effect in inhibition (g of 0.13 across 6 Stroop task effects), and mixed results in working memory updating. For updating, within-group training effects were significant with g between 0.58 and 0.71, but between-group comparisons were not significant, with g between -0.03 and 0.18.

Laid out as a table, here is where the two reviews overlap and where they part.

Sub-abilityNour et al. (2020, 17 studies, abstract only)Hu and Fan (2021, 29 studies and 98 tasks, full text)
InhibitionNo evidence of an advantageNo effect (g = 0.13)
ShiftingImproved as a result of trainingSubstantial advantage (g = 0.68)
UpdatingAdvantage in group comparison, no improvement longitudinallySignificant within group, not significant between groups

Of the three sub-abilities, the two reviews reach the same conclusion on inhibition, and point the same direction on shifting. The point of divergence is updating, and in fact both reviews report that in updating the result flips depending on the design. What this literature agrees on, then, is not that interpreters have generally good cognitive ability, but something far narrower.

3. Why they diverge (1): cross-sectional and longitudinal designs

The problem both reviews single out is design.

A cross-sectional design compares a group of interpreters and a group of non-interpreters at a single point in time. When a difference appears in that design, two interpretations remain available. Either the training changed them, or people who were already like that entered the profession.

A longitudinal design tracks the same people before and after training. Only that design lets you speak of a training effect. And Nour and colleagues wrote that in updating ability, longitudinal studies tend to show no improvement. In Hu and Fan's meta-analysis too, updating was significant only in within-group comparisons. A within-group comparison means comparing the same group before and after with no control group, so it cannot rule out simple test-repetition effects or maturation effects.

To sum up, the interpreter advantage that appears in updating may still be a result of selection rather than of training, while in shifting the evidence for a training effect is relatively better. That is the most that can be said right now.

4. Why they diverge (2): does selection come first, or training?

The fact established in the previous post becomes precisely the problem here. The AIIC document on best practices for conference interpreting training programmes states that applicants must pass an aptitude test before admission. Geneva FTI likewise requires passing the faculty's oral entrance examination.

So the population of interpreters is a population filtered twice. Once at the entrance examination, and again at the graduation assessment. If you compare such a group with ordinary university students and call the resulting difference a training effect, you may be measuring the effect of the examination rather than of the training.

It is not that researchers are unaware of this. Jouravlev and colleagues, in their 2021 polyglot brain-imaging study in Cerebral Cortex, volume 31, issue 1, wrote themselves that their cross-sectional design could not establish causation. They could not distinguish whether the reduced activation they observed was the result of extensive experience and training, or a trait that draws certain people to that kind of training in the first place. They added that longitudinal studies and studies of genetic bases are needed.

Popular pieces introducing a paper often erase the limits the paper itself set down. They will not be erased here.

5. What brain imaging research can and cannot say

Polyglot brain research is the most frequently cited strand on this topic. Two papers here.

The study by Jouravlev, Mineroff, Blank and Fedorenko, published in 2021 in Cerebral Cortex, volume 31, issue 1 (pages 62 to 76), covered 17 polyglots. Nine of them were so-called hyperpolyglots handling between 10 and 55 languages, and the overall range was 5 to 55 languages. They were compared with 17 monolingual speakers matched on age, sex, handedness and intelligence, and data from 217 non-polyglots was also consulted separately. The result was that when processing English, their native language, activation of the left-hemisphere language network was lower and its extent smaller.

The follow-up study by Malik-Moraleda and colleagues, published in 2024 in the same journal, volume 34, issue 3 (article number bhae049), investigated 34 polyglots (including 16 hyperpolyglots) with precision fMRI. The mean was 14.6 languages, the median 10, and the range 5 to 54. The definition of polyglot was by self-report: someone reporting a minimum proficiency of at least 4 out of 20 in five or more languages, and an upper-intermediate proficiency of at least 10 in one of their non-native languages. The main results were that native, non-native and unknown languages all activated the left-hemisphere language network, that the response was stronger the higher the self-reported proficiency, and that the native language was an exception, showing a similar or lower response.

The limits stated by the authors of the second paper are carried over here as written. Proficiency was self-reported rather than objectively measured; comprehension of the experimental materials was not directly verified; proficiency correlates with age of acquisition, which can introduce a confound into the interpretation; and a sample of 34 is small for individual-differences research.

What these studies say is the observation that the language network of polyglots responds differently. What they do not say is that learning many languages produces such a brain. Making the second statement requires a longitudinal design, and within what I was able to check, no such study exists.

6. Sample size: the default in this field is between ten and forty people

Gathering the samples of the studies cited so far into one place shows the reality of this field.

StudySampleDesign
Gile (1999)10 professional interpretersRepeated interpreting of the same speech
Jouravlev et al. (2021)17 polyglots, 17 controlsCross-sectional fMRI
Hikmat & Hasan (2022)36 studentsPre-post comparison
Amos et al. (2023)36 per experimentEye-tracking
Malik-Moraleda et al. (2024)34 polyglotsCross-sectional precision fMRI
Yang et al. (2025)15 professionals, 30 traineesLag measurement

Two things can be read out of this table.

First, small samples are not in themselves researcher negligence. Assembling fifteen professional conference interpreters with eight or more years of experience is a different order of difficulty from assembling three hundred undergraduates. Yang and colleagues also wrote that the difficulty of recruiting expert participants is a shared problem in interpreting research.

Second, that said, effect sizes drawn from small samples swing widely. So when reading this kind of study it is safer to take the direction and leave the numbers. The reason I have written the numbers in this post is not so that you will believe them, but so that you can see alongside them the sample they came from.

A third thing to add is within-person variability. Daniel Gile, in a 1999 study in HERMES, volume 12, issue 23 (pages 153 to 172), had 10 professional interpreters interpret the same speech simultaneously. Errors and omissions did not cluster in particular passages, a substantial share appeared in only a few participants, and when the same person interpreted the same speech again, new errors appeared in passages they had got right the first time. Gile explains this with the tightrope hypothesis. Because interpreters work close to the saturation point of their processing capacity, they break down at small fluctuations regardless of the difficulty of the source text. This is a theory and a model, and the experiment above is an observation consistent with that model.

7. The deliberate practice debate: original, meta-analysis, replication

The academic root of the story known as the ten thousand hours is the paper Ericsson, Krampe and Tesch-Römer published in 1993 in Psychological Review. It divided violin students into three groups and compared accumulated practice time, and concluded that individual differences in eventual performance are largely explained by past and present amounts of practice.

The sample of that original study was 30 people. Ten in each group. I confirmed this number in the body of the replication paper introduced below.

Then came two shocks.

The first is the meta-analysis Macnamara, Hambrick and Oswald published in 2014 in Psychological Science, volume 25, issue 8 (pages 1608 to 1618). The proportion of performance variance accounted for by deliberate practice was 26 percent in games, 21 percent in music, 18 percent in sports, 4 percent in education, and less than 1 percent in the professions. The authors concluded that deliberate practice matters, but not as much as has been claimed.

The second is the replication study Macnamara and Maitra published in 2019 in Royal Society Open Science, volume 6, article number 190327. It was run as pre-registered with outcomes withheld, and used a double-blind procedure. The sample was 39 violinists (13 in each group). The results were as follows. The core finding, that accumulated deliberate practice corresponds to level of skill, did not replicate. The variance explained by practice alone was about 48 percent in the original but 26 percent in the replication, a level similar to the 23 percent average from the music-domain meta-analysis. Accumulated practice hours up to age 18 came to 8224 hours for the top group and 9844 hours for the next group, reversing the order.

8. The rebuttal from Ericsson's side, also set down here

Introduce only one side and this post too ends up a biased summary. In 2020 Ericsson published a response to Macnamara and Hambrick in Psychological Research.

His point is that the critics defined deliberate practice far too broadly. He holds that all five criteria from the 1993 paper must be met for something to count as deliberate practice: a well-defined task with a clear goal that the participant understands; the participant being able to perform that task themselves; immediate and actionable feedback on performance; the opportunity to perform similar tasks repeatedly; and individualised instruction and guidance from a teacher.

He takes issue with the 2014 meta-analysis defining deliberate practice as engagement in structured activities created specifically to improve performance, and calls that structured practice rather than deliberate practice. That definition takes in attending lectures, studying alone, and coach-led group activities, and these lack the essential requirements, individualised teacher instruction among them.

How should this rebuttal be read? Here is how I see it. Ericsson's point is logically sound. That an effect measured with a broad definition is small does not mean the effect measured with a narrow definition is also small. At the same time this rebuttal comes at a price. The narrower the definition, the less data satisfies its conditions, and the harder the theory becomes to disconfirm. If immediate feedback from an individual teacher is an essential requirement, then most adult language learners are outside the scope of the theory to begin with.

This debate is not over. So citing the ten thousand hours as though it were evidence in a piece about language learning is inaccurate at this point in time.

What has not been verified

Here is what this post could not confirm.

First, the proposition that interpreter training improves general cognitive ability has not been verified. That the evidence for a training effect is relatively better in shifting ability is the extent of what can be said, and on inhibition the two reviews converge on there being no effect.

Second, I could not find evidence that the cognitive effects of interpreter training transfer to tasks outside interpreting. A change in executive function test scores and a change in everyday performance are separate matters.

Third, I read only the abstract of Mellinger and Hanson's meta-analysis, so I can say nothing about effect sizes or publication bias. The same goes for the review by Nour and colleagues.

Fourth, whether the differences observed in polyglot brain research are the result or the cause of language learning has not been distinguished. The authors wrote as much themselves.

Fifth, I am not in a position to adjudicate which side of the deliberate practice debate is right. I have only carried over both sets of claims and their grounds.

Sixth, none of the research in this post involved Korean speakers. That results may differ with a different language pair is a limit the cited studies noted themselves.

Closing

The conclusion of this post is not satisfying. That is because there is no satisfying conclusion.

Still, something remains. A substantial part of the cognitive characteristics observed in the interpreter population may be the result of selection, and what is relatively well supported as a result of training covers a narrow area. And even the most famous story of all, that amount of practice explains skill, saw its magnitude halved in the course of replication.

Is this bad news for a learner? I think the opposite. That it is not determined by amount of practice alone also means there is room to change the kind and the conditions of practice. How to use that room is the subject of the fourth post.

If you want to break a learning path into stages and manage it, see /tools/skill-path; if you want to attach repeated practice to a game format, see /tools/language-quest.

References

Every link was opened and checked directly on 16 August 2026. For paywalled literature, each entry notes that only the abstract was read.