The headlines say AI is eroding our minds. The studies behind them are real — but they may be measuring the wrong thing, at the wrong moment. Here is a hypothesis they don't test.
A 2025 study of 666 people found a strong negative correlation between how much someone uses AI tools and how they score on critical-thinking tests — around r = −0.68, with the effect strongest in the youngest users. A separate brain-imaging experiment found weaker executive-control activity when people wrote essays with AI help than when they wrote unaided. The mechanism named in both: cognitive offloading — letting the tool do the thinking instead of doing it yourself.
Taken at face value, the verdict writes itself: the smarter the tool, the lazier the mind. But two things should make us pause before we accept it.
In the Phaedrus, Plato has King Thamus warn that the invention of writing will make people forgetful: they will trust external marks instead of their own memory, and gain "the appearance of wisdom, not its reality." The very first "this will make us stupid" panic was not about computers. It was about the alphabet.
He was right about the small thing and wrong about the big one. Writing did weaken raw memory — but it freed the mind for more complex thought and let knowledge accumulate across generations. The printing press drew the same fear; so did the calculator, so did Google. Each externalised some mental load, and each time the freed-up capacity moved up, not away.
Here is the uncomfortable part of the panic studies. Their critical-thinking tests were built and validated in a pre-AI world — they measure how good you are at exactly the thing the AI now does for you. If a new skill is being built in its place — precise prompting, judgement, steering, checking a synthesis rather than producing it — there is no test for that yet, because we don't know how to measure it. So what gets recorded as "decline" may just be a shift into a dimension the instruments can't see.
Working memory is finite. Cognitive load theory (Sweller) says that when you remove the routine burden, more of it is available for higher-order work. That part is solid. But there is a catch the panic studies miss: the freed capacity does not switch to higher thinking instantly. First the brain needs offline time to reorganise — to consolidate, connect, and abstract what it just took in.
This is the whole argument in one word, and it is a word you probably said to yourself already: the higher layer has to be learned. Freed capacity is a door, not a staircase. It opens the possibility; it does not climb for you.
There is a hidden assumption we have been carrying: that AI is a net relief from the start. At first, it isn't. Especially with multi-AI — running several models in parallel, comparing their outputs, weighing them, deciding — the tool doesn't remove the old task. It hands you a new one you have no trained routine for: more parallel sources, faster judgement, a bigger synthesis to hold in your head at once. Today's minds simply aren't practised at processing information at this density, or deciding at this speed. That is not laziness. That is a load we were never trained to carry — yet.
A concrete example makes the load visible. The EQUORA Institute's Trilith Method™ runs three models as a deliberate stack — Claude for synthesis, ChatGPT for monitoring, Gemini for cross-validation — and, crucially, sets them against each other: adversarial multi-model verification, with a human as the deciding judge. The hard part is not reading three answers. It is holding three conflicting outputs at once and adjudicating between them — the maximum version of the synthesis load. No prior schema exists for that. It is exhausting at first for exactly the reason it is powerful: it forces the higher layer to do work it has never been trained for.
This is the shape of every learning curve: it dips before it climbs. And it quietly resolves a puzzle from the decline studies — that the youngest, heaviest users score worst. If the curve goes down first, the ones who plunged in earliest, fastest, at the highest volume are simply deepest in the dip. Not dumber — just furthest along the painful early stretch. The open question is whether they climb out.
The dip is not guaranteed to turn upward, though. A schema only builds if the difficulty sits in the learner's reach — hard enough to stretch, not so hard it collapses into giving up. Multi-AI's speed and volume land in that zone for some (a learning curve) and above it for others (overwhelm, and a retreat into offloading). Same tool, two outcomes — again.
There is a network for exactly this. The »default mode network« — the medial prefrontal cortex, posterior cingulate, and hippocampus — is most active precisely when you are not focused on an external task: in the shower, on a walk, staring out a window. Far from idling, it is doing some of the brain's most important work: after learning, the hippocampus replays what you took in, triggers this network, and that replay predicts what you'll remember later. Crucially, this spontaneous activity enhances abstraction and generalisation — it turns raw facts into patterns. That is the raw material of higher thought, and it is built during downtime.
Put the two mechanisms together and the usual conclusion flips. AI removes the lower load — good. The freed capacity then needs idle time to reorganise into something higher — also good, and this is where the coffee-break, the walk, the staring-into-space actually earn their keep. The threat was never the tool. The threat is that we refill the freed time immediately with external input — the endless scroll — and never let the »Default Mode Network« do its work.
Which reframes even the "wasted" break. Real idle — the shower, the unfilled pause — is when the rearrangement happens. The scroll is not idle: it fills the exact gap the network needs with fresh external stimulus, and the DMN switches off. So the honest distinction is not coffee vs. work. It is empty time vs. filled time.
Nobody has run that study — AI is too new, and we have no validated instrument for the "conductor" skill or for the idle-time payoff. That is not a weakness in the hypothesis. That is a gap, and gaps are where research starts. Three mechanisms stack into one model: AI first overloads (the curve dips as the schema builds); real idle time lets the freed capacity reorganise upward; and offloading is the escape hatch for those who don't make it through the dip. The prediction that follows: the outcome is not set by how much AI you use — the thing current studies measure — but by whether you survive the schema-building dip, protect the idle time, and resist the retreat into offloading. That is the study waiting to be run.
Help build the study →