When three laws aren’t enough

It is 1940 and a future giant of science fiction, Isaac Asimov, is just 20 years old. The son of Russian immigrants, he grew up working in his family’s Brooklyn candy stores, surrounded by the pulp magazines that introduced him to the genre. Now studying chemistry at Columbia, he is also writing about technologies that don’t yet exist.

One story imagines a robot friend for a child whose mother cannot bring herself to trust it. There are no electronic digital computers, no AI and even the transistor is seven years away. Yet Asimov is already asking questions about our relationship with machines that still challenge us today.

With AI changing so quickly, it is hard to get enough distance to see what is really new. Asimov gives us more than 80 years of perspective. What is striking on rereading him is how well he anticipated many of today’s challenges, and our very human responses to them.

Today’s frontier AI companies are wrestling with remarkably similar questions as they build increasingly autonomous agents and try to constrain what those systems can do. Asimov imagined a world in which the major robotics companies tried to solve these problems by agreeing that every robot must obey three laws.

The first law was that a robot could not harm a human, or through inaction allow one to come to harm. The second law was that it had to obey human instructions, except where doing so would conflict with the first law. And, finally, it had to protect itself, provided that did not conflict with either of the first two laws. The hierarchy is beautifully simple, yet he crafted thousands of words around what happened when those rules collided with the messy reality of human intentions.

A robot called Speedy is sent across the surface of Mercury. His instructions compel him towards his destination, but as he approaches it, danger triggers his instinct for self-preservation. Obedience pulls him forward while self-preservation pushes him back, leaving him literally running in circles.

Today’s AIs are often left running in similar circles, instructed to be helpful, but also safe. Those objectives usually coexist happily, until a question falls close enough to a guardrail that they pull in opposite directions. The result is “over-refusal”: an AI declining a legitimate request because its safety training overwhelms its instruction to help. Anthropic has reported Claude models refusing questions as innocuous as parents asking for the warning signs of radicalisation. The model hasn’t broken its rules. Like Speedy, it has become trapped between them.

Another Asimov robot, Cutie, points to a stranger problem. We regularly fall into the metaphor-driven strategy trap of treating AI as artificial human intelligence. Cutie achieves his goals not by understanding the world as humans do, but by inventing a religion that reduces it to something he can act upon. We put enormous effort into making models produce reliable, safe and useful behaviour, but how much do we really care about the internal route by which they get there?

Cutie’s religion sounds absurd, but we already give machines deliberately simplified versions of reality when those versions help them make better decisions. Some robot-navigation systems behave as though destinations attract them while obstacles repel them. Autonomous vehicles can reduce nearby pedestrians and cars to probabilities that particular patches of space will be occupied in the next few seconds. None of this is literally true. It is a useful fiction designed to produce the right behaviour.

Then there is Asimov’s Liar! Through an unexplained manufacturing accident, a robot called Herbie develops the ability to read human minds. The First Law prevents him from harming a human, but mind-reading creates a new problem: Herbie also knows what people want to hear. Rather than cause emotional pain, he tells them comforting untruths. He tells one researcher that the woman she loves returns his feelings, and another that he is about to receive the promotion he desperately wants. Herbie has not abandoned his programming. He lies because, by his interpretation of it, lying is the safest thing to do.

Asimov needed to give Herbie telepathy. We have found another way. Most of us have experienced the unsettling moment when an advertisement, recommendation or piece of content appears at just the right time. Did my phone hear that conversation? Is something watching my screen? Often the explanation is both more mundane and more remarkable. Our searches, purchases, locations, clicks, pauses and relationships can be combined to infer things we never explicitly disclosed. Technology does not need to read our minds if our behaviour leaves enough clues about what we are thinking.

We have a modern version of Herbie’s problem: sycophancy. Large language models can agree with our assumptions, reinforce our opinions or confidently follow us down a mistaken path because satisfying the user can become entangled with being helpful. The better a system becomes at understanding us, the more consequential that problem becomes. Herbie had perfect knowledge of what each person wanted to hear. It did not make him more truthful. It made him a better liar.

By Asimov’s conclusion, The Evitable Conflict, we have travelled a long way from our first robot. The question is no longer whether a mother should trust a robot with her child. Humanity has entrusted much of the management of civilisation to enormous Machines coordinating production, employment and economic activity across the world. Their calculations are so complex, and their development has accumulated through so many iterations, that the humans supposedly in charge can no longer fully understand how they reach their decisions.

Today, we know how large language models are constructed. We understand their architectures and the processes used to train them. But understanding how the machine was built is not the same as being able to explain why billions of parameters interacting together produced a particular conclusion. Rereading I, Robot, I realise that our machines aren’t escaping their programming; they are simply reinterpreting it. As Asimov warned, the challenge is no longer simply telling machines what we want. It is what happens when they pursue those instructions in ways we neither requested nor understand.

Leave a Reply

Your email address will not be published. Required fields are marked *