Introduction — Nobody Feels Guilty Using a Calculator
Almost no adult beats themselves up for being bad at mental arithmetic. Barely anyone worries about their memory because they can't remember a single phone number. Ask someone to find their way somewhere new without a map and they'll struggle, but nobody calls that a decline in ability.
And yet, after sending off a report that AI drafted, most people feel a little uneasy. Paste in an entire block of code wholesale, watch the tests pass, and relief arrives along with something faintly unsettling.
This asymmetry is the real question. Why does one kind of delegation feel like nothing at all, while another feels uncomfortable? Is it just unfamiliarity with something new, or is that discomfort pointing at something accurate?
The short answer: automation's effect on ability isn't uniform. And that lack of uniformity is exactly what makes this interesting. Some abilities can be handed off with nothing happening at all. Some get measurably eroded. And most of today's knowledge work currently sits in a middle zone where the answer isn't in yet.
What's Fine to Lose — Where Offloading Is a Pure Win
First, something worth admitting honestly: most of the history of cognitive offloading is a history of success.
When writing was invented, Socrates worried in the Phaedrus that it would ruin memory. He was about half right. We really can't memorize entire epics anymore. And in exchange, civilization as a whole broke free of the bottleneck of individual memory capacity. Not a bad trade.
The calculator is the same. In exchange for losing the ability to multiply three-digit numbers in your head, people got to spend their time on what to multiply instead of the multiplication itself. Same goes for spell-checkers, same goes for phone books.
Pull out the pattern and it looks like this. Cases where offloading is safe generally satisfy three conditions. First, the result can be independently verified. Whether a calculator got it wrong is caught by a rough mental estimate. Second, the ability isn't raw material for some other ability. Memorizing a phone number isn't an input into any other kind of thinking. Third, a situation without the tool realistically never comes up.
Break any one of these three conditions and the story changes. And there are real cases where they do break.
GPS as a Blurry Boundary Line
Wayfinding is the first case where condition three, above, starts to wobble. Batteries die. Signals drop.
A 2020 study by Dahmani and Bohbot in Scientific Reports gets cited often. Fifty drivers had their lifetime GPS usage and various spatial memory abilities measured, and the more GPS a person had used over their lifetime, the worse their spatial memory scored when they had to navigate without it. Thirteen of them were followed up three years later, and the ones who had used GPS more in the interim showed a larger drop in hippocampus-dependent spatial memory.
Don't stop here, though. The study's own researchers spell out its limits.
The main result is a correlational study. The arrow of causation could run the other way — someone with a naturally poor sense of direction may have simply come to rely on GPS more, and that's arguably the more natural explanation. The longitudinal piece narrows that possibility down somewhat, but only 13 people came back for the three-year follow-up. Thirteen. That's not enough to wash out individual noise.
So the accurate summary is: not "GPS ruins your brain," but "there's a signal that wayfinding ability changes along with usage, and that signal is still small, and the direction of causation hasn't been settled." The first sentence is a headline. The second sentence is the research.
Aviation — The Best-Documented Case, and an Ending That Runs Against the Conventional Wisdom
The field that has studied automation dependence longest and most seriously is aviation. Here, eroded skill can kill people.
The conventional wisdom goes like this: "pilots today lean so hard on autopilot they've lost the ability to hand-fly a plane." A 2013 report from an FAA working group on flight path management systems actually did flag degraded manual flying skill and situational awareness as major concerns.
But the research that actually put this conventional wisdom to the test found something else. A 2014 study in Human Factors by Casner and colleagues, "Manual Flying Skills in the Era of Automated Cockpits," had 16 airline pilots systematically vary automation levels in a Boeing 747-400 simulator, flying both normal and abnormal scenarios, and probed what they were thinking about throughout the flights.
The result split two ways.
Instrument scanning and stick-and-rudder skills mostly held up fine — even though the pilots themselves reported "I don't get to hand-fly much these days." Skills baked into muscle memory and sensorimotor loops held up better than expected.
By contrast, the cognitive tasks needed for manual flight had clearly deteriorated. Keeping track in your head of where you're actually headed, recalling the next procedure, translating an ATC instruction into a route plan — errors clustered here, and the same mistakes kept recurring. The researchers' summary is compact: cognitive skills decay faster than psychomotor skills.
Why does this matter? Conventional wisdom said "hands get rusty." What had actually gone rusty was the mind — not the ability to grip the yoke, but the ability to hold the whole situation in your head while gripping it.
The limits here are worth stating plainly too. Sixteen participants, in a simulator, reflecting a specific aircraft type and a specific era's training practices. Transplanting a result from pilots — a highly selected, continuously drilled population — onto knowledge workers in general is a large leap on its own. But one lesson from this study does seem likely to transfer: assuming you already know what's going to erode is usually wrong.
Automation Bias — What We Do When the Tool Is Wrong
Separate from skill decay, automation carries a more immediate problem: we fail to catch it when the tool is wrong.
Parasuraman and Manzey's 2010 review in Human Factors on automation complacency and bias is the standard reference point here. Three findings stand out.
Automation bias produces both omission errors and commission errors. Missing something because the system failed to flag it is an omission error; carrying out a bad suggestion the system made anyway is a commission error. The second is scarier, because it includes cases where the person actually knew the right answer and reversed it anyway.
It shows up in both novices and experts. This isn't a problem experience solves.
Training and instruction barely block it. Even after being explicitly told "this system isn't perfect, always verify," a substantial share of the bias remains. And it gets worse under multitasking load. When there's a lot else going on, attention drifts away from whatever already looks like it's working fine.
There are concrete numbers too. In an empirical study of a prescription decision-support system by Goddard and colleagues, when clinicians received bad advice, 5.2 percent of the time they changed a correct answer to a wrong one. Five percent sounds small. It isn't, when the denominator is millions of prescriptions a day.
One caution worth flagging here: most of these studies artificially create conditions where bad advice is actually mixed in, to make it measurable. The real-world error rate could differ, and lab tasks carry less real-world weight than the actual thing. Even so, the direction is consistent across multiple domains.
Delegating Boredom, Delegating Judgment
Now the central distinction of this piece. I split delegation into two kinds.
Delegating boredom means handing off the labor of execution after what needs to be done has already been decided. Reformat this table. Fill in the tests for this function. Pull the action items out of these meeting notes. There's a correct answer, you can check whether it was hit, and checking costs less than doing it yourself.
Delegating judgment means handing off the decision about what needs to be done at all. Is this problem worth solving? Is this the right metric? Is the question we're currently trying to answer even the right question to begin with?
The difference between the two isn't difficulty. It's reversibility.
Picture one concrete scene. A ticket comes in saying the dashboard is slow. Hand it to AI and in thirty seconds you get an excellent query optimization plan, complete with index suggestions, and response time really does drop by half. It looks like perfect delegation.
Except nobody has actually looked at that dashboard in two months. What actually needed to happen was deleting the screen.
This is what happens when you delegate judgment. A brilliant answer arrives for the wrong question, and because the answer is brilliant, nobody goes back to look at the question. The quality of the answer is what blocks the verification. If the result had come out badly, you'd have gone back to the start; because it came out well, you move on instead.
This is exactly why judgment matters. Judgment is the place where you notice a question was wrong in the first place. And that noticing usually happens inside the process of producing the answer — the way you open the query yourself and stumble on "wait, why does this table's recent access log look like this." Skip that process, and the chance to notice disappears along with it. What gets lost isn't ability. It's the chance encounter.
| Delegating boredom | Delegating judgment | |
|---|---|---|
| What gets handed off | How to do it | What to do |
| Verifying the answer | Usually possible | Only becomes visible later |
| When it's wrong | Just redo it | The whole direction has to be reversed |
| Do it repeatedly | You gain time | Your eye for problems gets less developed |
Don't read this table as a fixed rule. Real work doesn't fall cleanly into either box, and the same piece of work moves between the two depending on the situation. This isn't a classification chart — it's a single question: which of the two is what I'm actually handing off right now?
What This Generation Is Handing Off — and the Evidence That Isn't There Yet
Writing, coding, analysis. These are what's currently being delegated at massive scale. Which box these fall into is exactly the question, and the honest answer is they straddle both. Drafting is closer to boredom; deciding what to say is judgment. And in real work, the two don't separate cleanly. Anyone who has had the experience of figuring out what they think while writing already knows this.
An honest summary of what research currently exists looks like this.
A 2025 CHI paper from Microsoft and Carnegie Mellon researchers surveyed 319 knowledge workers and collected 936 real usage examples. Result: higher trust in generative AI was associated with less reported critical thinking, and higher confidence in one's own ability was associated with more. The limits are substantial. It's self-report, it's cross-sectional, and what it measured was perception, not performance. "The degree to which I feel I thought critically" and "the degree to which I actually thought critically" are not the same variable.
MIT Media Lab's 2025 EEG study, popularly known as "Your Brain on ChatGPT," reported differences in brain connectivity during essay writing depending on tool-use condition, and got a lot of attention. This one calls for particular caution. It's an arXiv preprint that hasn't gone through peer review, with 54 participants, only 18 of whom made it to the fourth session. The authors themselves disclosed that they chose to release it before review, and formal commentary on the EEG methodology and reproducibility has since been raised. Reading a difference in brain connectivity as equivalent to a decline in ability is a much heavier claim than this data can support.
And here's a genuinely good cautionary case. The "Google effect" study by Sparrow and colleagues, published in Science in 2011, is a classic that's been cited for fifteen years running. But the study's famous Experiment 1 — a Stroop task showing that difficult questions prime computer-related words — failed to replicate in a 2020 precision replication. With 89 participants, a design that incorporated input from the original authors, the Bayes factor said the data were roughly five times more consistent with the null model. This paper also notes that the same effect failed to replicate in an earlier large-scale replication project.
An important qualifier: this replication targeted only the priming portion. The other result from the same paper — that people remember where information is stored better than they remember the content itself, the transactive-memory finding — was not tested by this replication. So "the Google effect is false" is also the wrong summary. The accurate summary is: this widely cited finding splits into at least two parts, one of which has now failed to replicate twice, and the other of which hasn't yet been directly tested.
This is exactly how every piece of research on AI and cognition needs to be read right now. This literature shares precisely the same vulnerabilities as the replication crisis in social psychology: small samples, self-report, an appealing narrative, and a distribution structure where confirmation spreads far faster than disconfirmation.
And the most honest sentence is this: we do not yet have any evidence about the long-term effects of losing exactly the abilities this generation is currently handing off. Conversational AI has been part of daily work for a few years now, and this question demands longitudinal research measured in decades. That research has only just begun. Anyone making a strong claim right now — on either side — is speaking from intuition, not data. Myself included.
This Piece Was Written With That Tool
This blog is written with a lot of help from AI. So was this post. Writing a piece about whether automation erodes skill, using automation to write it, without disclosing that — doesn't hold together.
Honestly, the line between what I delegated in this piece and what I didn't wasn't nearly as clean as the table above. I handed off finding papers, checking abstracts, cross-checking numbers. What I didn't hand off was things like "where does the failed replication of the Google effect belong in this argument" or "does the aviation study's reversal go before or after the conventional wisdom." I tried delegating those too. What came back wasn't bad. I didn't use it anyway. That judgment call was, itself, the thing this piece is actually about.
And while writing this, I actually got caught by it once. In an early draft, I nearly repeated the conventional wisdom that "pilots are losing the ability to hand-fly a plane" as-is. It was only after opening the original paper that I found the actual result runs closer to the opposite. This wasn't the tool failing — it was me, almost skipping the verification step. I'm leaving this in as proof that the distinction above isn't just an abstraction.
Closing — Not a Rule, But a Posture
I don't want to hand you a rule. A sentence like "never delegate judgment" can't be kept, and doesn't need to be. Some judgment can be handed off with nothing happening at all, and more importantly, the moment you mistake today's limits for permanent ones, that rule goes stale within a few years.
Instead, I'll suggest one posture: pause one more time when the result is good.
When a result is bad, we go back and check on our own. What's dangerous is when the answer comes back brilliant. That's exactly the moment you stop examining the question, exactly the moment the dashboard nobody looks at doesn't get deleted, exactly the moment a clinician overturns the answer they originally had right.
So in front of a result that came out well, I try adding just one sentence: "if this is the answer, what was the question again?" It takes about ten seconds. Most of the time, nothing happens. Once in a while, those ten seconds save two months of work. Exactly how much to delegate, and how much, will keep changing — today's tools, and today's limits, will likely look different in a few years. But the habit of holding onto the question instead of the answer will stay useful for as long as making good decisions itself doesn't disappear, no matter how the tools change.
현재 단락 (1/59)
Almost no adult beats themselves up for being bad at mental arithmetic. Barely anyone worries about ...