Why AI can pass the bar exam but cannot load your dishwasher

Why AI can pass the bar exam but cannot load your dishwasher

Everything that feels hard to a person turns out to be easy to build. Everything that feels effortless turns out to be brutally hard. That flip is not a curiosity. It is the most useful thing you can know about what to hand a machine.

A computer can beat the best chess player who has ever lived. It can beat the best Go player, and Go is the game people spent decades saying machines would never crack. It can pass the bar exam. It can write a decent sonnet before you have finished asking for one.

That same computer, wired into the best robot body money can buy, cannot walk into your kitchen and load the dishwasher.

Loading a dishwasher is not an intellectual achievement. A three year old can do a rough version of it. Nobody has ever put it on a CV. So why can a machine that demolishes grandmasters not manage it?

The same eight jobs, scored twice

One bar for how hard a person finds it, one for how hard a machine finds it.

Hard for a personThe jobHard for a machine

Press the button and watch the order invert. The pattern is not noise, and the explanation for it took 600 million years to write.

The name for that flip is Moravec's paradox, after the roboticist who pointed out in the 1980s that high level reasoning needs surprisingly little computation while ordinary seeing and moving needs an enormous amount.

Naming a thing is not explaining it, though, and the explanation is the good part. I found it in a book called A Brief History of Intelligence by Max Bennett, and it has changed how I think about what to hand a machine and what to keep. This is the first of three posts pulling out the parts that matter if you work with AI rather than build it.

Intelligence was never designed. It piled up.

Bennett's argument is that intelligence did not arrive in one go, and was never designed at all. It accumulated. Five times in 600 million years, brain design took a real leap, and each leap solved a problem the one before it could not handle.

The important part is what happened to the old layers. Nothing was thrown away. Every leap got bolted on top of what was already running, and the thing underneath carried on running. Evolution is a hoarder.

Drag 600 million years

The layers stack up and stay up, which is precisely the point.

🌊
No brains yet
700 million years ago

Five words, in order: steering, reinforcing, simulating, mentalizing, speaking. Hold those and you have the whole argument.

Steering arrived about 600 million years ago with the first worms, and it sorted the world into good and bad. Reinforcing arrived with the first fish, and it added trial and error. Simulating arrived with the first mammals, and it added imagination. Mentalizing arrived with the first primates, and it added the ability to model another mind. Speaking arrived with us, and it let one brain install its contents in another.

Treat the dates as approximate. Bennett does too. The order is what matters.

Which means the newest layer was the easy one to copy

Now put that next to the dishwasher.

Reasoning and language sit at the top of the pile, and they are new. Language is roughly two million years old, which against 600 million is a coat of paint. Acting sensibly in a messy physical world sits at the bottom, and it has been tuned for hundreds of millions of years by the blunt method of everything that got it wrong being eaten.

So of course the top layer was the easy one to copy. It is shallow, and better still, it is written down. Every argument, explanation, instruction and story humans have ever produced is sitting there in text, which happens to be exactly the training material a language model needs.

The old layers were never written down anywhere. There is no manual for picking up an unfamiliar mug, no essay explaining how to tell a wobbly chair from a solid one. That knowledge was compiled into wetware across a timescale nobody can really picture, and there is no copy of it to train on.

That is Moravec's paradox, no longer just an observation but an explanation.

The oldest layer, and why it still matters to you

Stay with the bottom layer for a minute, because it is the one that keeps surprising me.

Before it, animals like jellyfish had loose nets of nerves and bodies that were the same in every direction. A jellyfish has no front. It pulses and drifts and catches whatever floats into it.

Then came a body with a left and a right, a front and a back, and a definite forward. That sounds like a minor engineering change and it is not. The moment an animal has a forward, it has one question to answer, over and over, all day, forever. Go toward this, or get away from it?

That single question grew the first brain. Not a mind, nothing like a mind. A small hub of cells at the front end, where incoming signals get stamped with one of two labels. Good, or bad.

Neuroscientists call that stamp valence, and for the first hundred million years or so it was the entire intellectual repertoire of animal life. The worm has no concept of poison, or danger, or death. It has a circuit that says: this smell, minus, turn around.

Feed the worm, or poison it

Click the pond to drop a smell. Two sensors, one rule, and no idea what any of it means.

valence0.00

That plus and minus never left. Your recoil at a foul smell and your pull toward coffee are direct descendants of this circuit, running underneath everything else you do.

The worm also got the seed of all learning that follows: linking a cue to a consequence. Pair a smell with something that makes it sick and it learns to avoid that smell. Simple enough, except that wiring everything to everything would drown the animal in coincidence, so evolution built rules around it. Two events only link if they land within about a second of each other. A loud cue drowns out a faint one predicting the same thing. Anything you meet constantly gets filed as background. And if you already predict something reliably, a new cue arriving alongside gets ignored as redundant.

Those four rules were solved, roughly, in a worm. All four are still live problems in machine learning.

A brain is a movement organ

One more thing about that layer, because Bennett makes a point of it and it is easy to skip past.

Nothing about having a brain is inevitable. Plenty of successful living things never bothered with one. Plants do fine. Fungi do fine. A sea squirt actually has a small brain as a swimming larva, finds a rock, attaches itself, and then digests its own brain, because once you have stopped moving you do not need one.

That is the clue about what brains are actually for. A brain is not a thinking organ that happens to sit inside a moving animal. It is a movement organ. It exists to answer where to go next, and every clever thing it later learned to do was built on top of that job.

Hold onto that when you get to the AI comparisons, because it cuts at something. A system with no body never has to answer the question brains were invented to answer.

What to do with this

Here is the practical bit, and the reason I am writing three of these rather than one.

If you can work out which layer a job sits on, you have a decent guess at how well a machine will handle it before you try. That turns out to be a better guide than any benchmark, because a benchmark tells you how a model did on somebody else's task, and the layer tells you why.

Drafting, summarising, explaining, translating, arguing a position: all top layer, all written down somewhere, all things to hand over early and without much anxiety. Anything needing a settled sense of what is actually good, or a working model of how physical things behave, or judgment about a situation nobody has ever written down: those are the bottom layers, and that is where your hand stays on the wheel.

The next two posts go through the middle of the stack, which is where it stops being a nature documentary and starts being useful.

Next: what a worm and a fish can tell you about the hardest part of setting an AI loose, which is saying what good actually looks like.

Share