
Two ideas that change what you expect from a machine
The first is that reward is a forecast rather than a feeling. The second is that imagination is a cost saving device. Between them they explain most of what is good and most of what is missing in the AI you use every day.
Last time the story got as far as the first brain: a hub at the front end of a worm that stamped the world with a plus or a minus, and nothing else. Powerful, and completely blind.
Then the ocean got dangerous. Around 500 million years ago predators appeared and an arms race started, and fixed reflexes stopped being good enough. If your whole repertoire is hardwired, a predator only has to solve you once.
Layer two: try it and see
The fix was trial and error. Do something, see how it goes, and if it worked, become more likely to do it again. Thousands of small adjustments across a lifetime, tuning an animal toward whatever actually pays off in its actual surroundings rather than the surroundings its ancestors had.
Vertebrates run this in a lump deep in the middle of the brain called the basal ganglia, and the architecture has a name, one that engineers landed on independently about half a billion years later. It is called actor critic. One part picks the move. Another part grades how things are going. The grade is the signal that gradually adjusts what the first part chooses next time.
So far, so tidy. What comes next is the best thing in the book.
Dopamine is not the pleasure chemical
You have heard that it is. Dopamine hit, dopamine detox, dopamine as the little reward chemical that makes you feel good. It is wrong, and the real story is far more useful.
Dopamine is a forecast. It is about expecting good things, not about receiving them.
Here is the experiment that pinned it down, run by Wolfram Schultz on monkeys. Record from dopamine neurons, give a monkey a squirt of juice, and dopamine spikes. Fine, that looks like the pleasure story. Then start playing a tone a moment before the juice arrives, over and over, until the monkey learns the tone means juice is coming.
Read stage two again, because it is the whole thing. Once the juice is predictable, the juice itself produces no dopamine response at all. The response has jumped backwards in time to the earliest reliable signal that something good is coming.
And stage three is the part that settles it. Play the tone, withhold the juice, and dopamine does not merely stay flat. It dips below the baseline at exactly the moment the juice should have arrived. That dip is disappointment. Not a metaphor for disappointment, the actual physical event.
So dopamine is not tracking reward. It is tracking the error in the prediction of reward. Better than expected, spike. Exactly as expected, nothing. Worse than expected, dip.
In machine learning this is called temporal difference learning, and Richard Sutton formalised it in the 1980s, before anyone knew the brain was already doing it. The insight is that you should not wait for the final outcome to learn. You learn from every change in your prediction along the way. That solves a real problem: if you win a chess game in 40 moves, which move deserves the credit? Waiting until the end and rewarding all 40 equally is hopeless.
The proof point is a backgammon program from 1992 called TD Gammon. It learned by playing itself, using nothing but temporal difference learning, and reached world class strength. Then it started playing opening moves that human champions had examined and rejected. The champions went back, looked again, and concluded the machine was right. Human opening theory got rewritten because of it.
What that means when you set an AI loose
Once you see reward as a forecast rather than a feeling, a lot of things line up.
Addiction stops being mysterious: a substance hijacks the forecasting system directly, so wanting goes through the roof while liking stays flat. People describe desperately wanting something they no longer enjoy at all. Dopamine was always the wanting. Curiosity stops needing a separate explanation too, because surprise itself triggers dopamine, which means novelty pays and exploration is built into the same machinery.
But here is the bit that matters for your work.
Every one of these systems, the fish and the machine alike, is optimising against a definition of good that something else supplied. In the fish, the hypothalamus supplies it: hunger, cold, thirst, the basic business of staying alive. In a machine, you supply it. That is the actual job when you hand something over, and it is much harder than writing a prompt.
It is why "write me a report" gets you a report that technically answers and lands wrong, while "here is what a good one looks like, here are three we were happy with, here is what we would never say" gets you something usable. You are not describing the task. You are building the forecast it optimises against.
It is also why the far end of handing things over is the part to go slowly on. A system that improves its own work week to week is only as good as its definition of good. Get that wrong and it does not fail loudly. It gets confidently worse, quickly, in the direction you pointed it.
Layer three: imagination as a cost saving device
Trial and error works, but look at the bill. To learn a cliff is bad, you have to fall off it. Reinforcement learning requires you to take the action and take the consequence, and in a dangerous world that is an expensive way to get an education.
About 200 million years ago the first mammals found the shortcut, and it is the leap that starts to feel like a mind. Run the action in your head first.
The machinery is the neocortex, the wrinkly outer layer mammals have and most other animals do not. It builds an internal model of the world and runs that model forward in time. Instead of learning only by doing, an animal can learn before doing. It can picture the fall, feel the internal warning, and not step off the edge.
Bennett dwells on one fact about the neocortex that I find as striking as he does: it looks nearly identical wherever you cut it. The same repeating six layered patch handles seeing, hearing, touch and planning. That hints it runs one core operation on whatever you feed it. Point it at the eyes and you get sight. Point it at goals and you get planning.
He is careful to call that a hypothesis rather than a settled fact. But you can see immediately why it has AI researchers so worked up, because one general algorithm scaled by repetition is a very appealing thing to try to copy.
You do not see reality. You see your best guess.
And if the brain builds a simulation, then the world you are experiencing right now is that simulation. You do not perceive reality directly. You perceive your brain's best guess about it, continuously corrected by your senses rather than delivered by them.
The AI echo is unmistakable. AlphaZero reached superhuman strength at Go and chess by simulating future move sequences and evaluating them, not by reflex and not by pattern matching alone. That is the machine version of the rat imagining both arms of the maze.
What is worth noticing is how recent that is. Really good look ahead planning in software is a couple of decades old at best. Your ancestors have been imagining the future since the dinosaurs were walking around above them.
What to take from this one
Two things, and they pull in opposite directions, which is the useful part.
The first is that a machine given a good definition of what counts as better will grind its way to answers no human would have found, and it will do it in places where the moves are cheap and the score is clear. That is the fish layer, and it really is superhuman when the conditions suit it.
The second is that the whole thing rests on a definition of good you wrote. The machine has no hypothalamus. Nothing underneath is correcting you, unasked, when the target drifts away from what you actually wanted, so that job stays yours, permanently, and it is the one to spend your time on.
Next, and last: the two newest layers, why today's AI is made almost entirely of the very top one, and how to tell which layer your own problem sits on.