Advertisement
Advertisement
Advertisement
22 August 2026ยท7 min readยทBy Victor Holm

Are We Thinking Correctly About AI Intelligence?

Melanie Mitchell explores how developmental psychology can assess AI's alien intelligence.

Are We Thinking Correctly About AI Intelligence?

AI intelligence is being measured all wrong, and Melanie Mitchell thinks she knows why.

She's watched the field lurch from one benchmark to the next, celebrating each new high score as if it meant something real, and she's spent years doing it from her post at the Santa Fe Institute. But in a recent conversation with mathematician Steven Strogatz, she argued that our entire framework for evaluating machine cognition is built on shaky ground. The tools we use, the assumptions we make, they were designed for a different kind of question entirely. It's all wrong. We've been measuring the wrong things.

The Alien Mind Problem

Here's the core tension. Large language models are trained on human text, human images, human code. They absorb our collective output and regurgitate it in statistically plausible patterns. Yet the way these systems actually process information bears almost no resemblance to human cognition.

Mitchell calls this "alien intelligence." Not because AI comes from another planet, but because it operates through mechanisms so fundamentally different from our own that we can't easily map our intuitions onto them. Babies are also alien intelligences in this sense, and so are dogs and crows and dolphins. The difference is that we've spent a century developing rigorous methods to study those minds.

Those methods, she argues, are exactly what AI research needs right now.

What Psychology Knows That AI Doesn't

Developmental psychologists don't just ask whether a baby can solve a puzzle. They probe how the baby approaches it, what errors it makes, how it recovers from failure, and they do so with a level of granularity that turns a simple play session into a rich dataset of cognitive struggle and adaptation. Comparative psychologists studying animal cognition do the same thing. But they design careful experiments that isolate specific abilities and rule out simpler explanations, often comparing species side by side to see which mental tools are shared and which are uniquely human. It's a meticulous craft. So the question isn't just "can they?" It's "how, and why, and what does that reveal about the mind's architecture?

AI benchmarking, by contrast, has become a kind of arms race. Throw a test at a model, watch it score well, declare victory. It's a hollow win. Nobody asks whether the model is actually reasoning or just pattern-matching on training data, and nobody checks whether it can transfer its apparent skills to a slightly different context, so we're left with scores that don't mean much beyond the test itself.

That last point has haunted the field for years. It hadn't learned to play. It had learned to execute a fixed routine. And while modern models are vastly more flexible, the question of whether they've truly transcended that limitation remains open.

The Math Breakthrough Nobody Can Explain

Recent events have made this question more urgent. Recent events have made this question more urgent, including recent AI-assisted breakthroughs in mathematics.

A computer circuit board with a brain on it

By any reasonable standard, that looks like creativity. But Mitchell urges caution.

The model trained on enormous amounts of mathematical text, including textbooks and lectures. It had access to essentially the entire corpus of human mathematical knowledge. That's a staggering thought. When you've ingested everything, the notion of "transfer" becomes murky, because the boundary between recalling a stored pattern and truly creating something new starts to blur in ways we can't easily define. So is combining two known fields a creative leap? Or just a sophisticated search through a vast space of possibilities? Hard to say.

It generated hundreds of thousands of candidates. Most were garbage. Occasionally, one was interesting. But a human had to sort through the junk to find the gems.

The modern AI might be doing something similar, just at a vastly larger scale. We can't see how many wrong paths it explored, how many dead ends it hit before finding the solution. That invisibility is the problem.

Six Principles for Better Measurement

Mitchell has proposed six principles for assessing AI cognitive capacity, drawing directly on methods from psychology. The details matter less than the underlying shift in perspective. She wants AI researchers to stop treating models as black boxes that either pass or fail, and start treating them as subjects to be studied with the same rigor we apply to human and animal minds.

The stakes are practical, not just philosophical. So how much can we truly trust AI to perform the tasks that genuinely matter, the ones we depend on daily without a second thought? We've got to ask how closely we need to supervise it. That's the real question. But what happens when we deploy these systems in domains where mistakes carry real consequences, where a single error could ripple through lives, finances, or safety in ways we can't easily reverse? Don't underestimate that risk.

Market Context: According to the Stanford AI Index Report, documented AI safety incidents surged by 56.4% from 2023 to 2024, causing real harm including financial losses, legal consequences, safety risks, and in some cases, loss of life.

Those questions become answerable only when we understand what the machines are actually doing.

The Cautionary Tale of a Counting Horse

Mitchell invoked a famous episode from the early 1900s. A horse named Clever Hans appeared to do arithmetic, tapping out answers to math problems with his hoof. The scientific community was amazed. Then psychologists discovered the horse was actually reading subtle cues from his human handlers. He wasn't doing math at all. He was responding to unconscious body language.

The lesson endures. When an intelligent-seeming system performs impressively, we need to ask what it's really responding to. With AI, the equivalent of Clever Hans's cues might be statistical regularities in training data that we don't recognize as shortcuts. The performance looks like understanding. It might just be sophisticated mimicry.

That distinction determines everything about how we use these tools.

Mitchell's broader point is that AI research has drifted away from its roots in cognitive science. But the abandonment went too far. In the rush to scale up, researchers lost the tools for asking what intelligence actually is.

The machines are changing faster than our ability to understand them. We can't keep up. But that's not a reason to panic, and it's not a reason to throw our hands up in defeat; it's a reason to slow down, take a breath, and adopt better methods that give us a fighting chance. So let's be honest with ourselves. The science of AI needs to catch up with the engineering of AI, and we've got a lot of ground to cover before we're truly in step.

If it doesn't, we risk being fooled again and again, celebrating achievements that we fundamentally misunderstand.

Frequently Asked Questions

What is Melanie Mitchell's main criticism of how AI intelligence is currently measured?

Melanie Mitchell argues that AI intelligence is being measured all wrong because the field celebrates each new benchmark high score as if it meant something real. She contends that the tools and assumptions used are built on shaky ground and were designed for a different kind of question entirely, so we've been measuring the wrong things.

Why does Mitchell call AI 'alien intelligence'?

Mitchell calls AI 'alien intelligence' not because it comes from another planet, but because it operates through mechanisms so fundamentally different from human cognition that we can't easily map our intuitions onto them. This term also applies to babies, dogs, crows, and dolphins, but we have developed rigorous methods to study those minds, which AI research lacks.

How do developmental psychologists study baby cognition compared to AI benchmarking?

Developmental psychologists don't just ask whether a baby can solve a puzzle; they probe how the baby approaches it, what errors it makes, and how it recovers from failure, turning a play session into a rich dataset of cognitive struggle and adaptation. In contrast, AI benchmarking has become an arms race where a model is thrown a test, scores well, and victory is declared without asking whether it's reasoning or pattern-matching.

What lesson does the story of Clever Hans the counting horse illustrate for AI?

The story of Clever Hans, a horse that appeared to do arithmetic but was actually reading subtle cues from human handlers, illustrates that we need to ask what an intelligent-seeming system is really responding to. In AI, the equivalent of Hans's cues might be statistical regularities in training data that we don't recognize as shortcuts, so impressive performance might be sophisticated mimicry rather than understanding.

According to the article, why is the recent AI-assisted breakthrough in mathematics viewed with caution?

The article cautions that the AI model trained on enormous amounts of mathematical text, including textbooks and lectures, so it had access to essentially the entire corpus of human mathematical knowledge. This blurs the boundary between recalling stored patterns and creating something new, making it unclear whether combining known fields is a creative leap or a sophisticated search through a vast space of possibilities.

Victor Holm
Written by
Science Correspondent

Victor Holm reports on science and discovery, with a particular interest in physics, biology and the questions that drive research forward. He looks for the wonder in how the world works.

๐Ÿ’ฌ Comments (0)

Sign in to leave a comment.

No comments yet. Be the first!

Advertisement