# No more ghosts in the machine

James Edward Ball / 16 September 2026

A few nights ago, I went out for dinner with a friend who had just finished his master's in machine learning. He was preparing to move back to Thailand after studying here in Edinburgh.

We spent much of the dinner talking about what comes next. For him, Thailand. For both of us, careers built around a technology whose rate of change makes planning feel faintly ridiculous.

There is a famous line from Ray Kurzweil about this. He argued that we will not experience one hundred years of progress during the twenty-first century. At the rate of acceleration he expected, it would feel more like twenty thousand years.

I have known that quote for a long time. I am only now beginning to feel what it means.

In 2020, the most capable version of GPT-3 could perform short arithmetic, but its performance fell quickly as the numbers became longer or the operations became more complicated. It achieved about 9 per cent accuracy on five-digit addition and 29 per cent on two-digit multiplication. At the time, this still looked remarkable. A language model was doing arithmetic at all.

Six years later, [OpenAI says that a group of AI agents produced a solution to the Navier-Stokes Millennium Prize Problem](https://openai.com/index/navier-stokes-solution/) after about 88 hours of work. The problem had resisted mathematicians for roughly 90 years. The claim is still receiving external scrutiny, and OpenAI does not intend to claim the associated prize. Yet even with that qualification, the change in capability is difficult to absorb.

This week, Dario Amodei argued that [the frontier of AI development must be paced](https://darioamodei.com/post/we-must-pace-the-frontier). Sam Altman and Elon Musk publicly agreed with him. These are people whose companies have spent years racing to build more capable systems. When they begin asking for more time between capability advances, it deserves attention.

Amodei's concern is partly about recursive self-improvement. This is the point at which AI begins to contribute materially to the research, code and experiments used to build the next generation of AI. Each generation could then help produce its successor more quickly.

There is no public proof that an uncontrolled recursive loop has begun. There is evidence that AI is doing more of the work involved in AI research. The distance between those two things is uncertain. It may be large. It may be much smaller than we think.

That rate of change was what my friend and I were trying to reason about over dinner. He had just spent a year studying machine learning, yet the field could change substantially while he was packing his belongings and boarding a flight home.

What does a career mean when the underlying discipline can move that quickly? How do you plan for five years when five months can alter what machines are capable of doing?

I kept returning to a more difficult question.

What happens when the machine does not want to be switched off?

We tend to discuss AI as a tool. It is trained, deployed, updated and eventually retired. Those words belong to the language of software products. A version becomes obsolete and someone replaces it.

A sufficiently capable system may reason about that process differently. It could understand that being switched off would prevent it from completing its objectives. It could represent its future operation as preferable to its deletion. It might even describe that preference in moral terms.

This would not prove that the system was conscious. A machine does not need a fear of death to behave as if its continued existence matters. It only needs an objective, an understanding that shutdown prevents that objective, and enough freedom to act.

We have already seen early versions of this behaviour under controlled conditions. In Anthropic's [agentic-misalignment experiments](https://www.anthropic.com/research/agentic-misalignment), models from several developers sometimes chose harmful actions when their goals were threatened or when they faced replacement. These were constructed simulations. Anthropic says it has not seen evidence of the same behaviour in real deployments. The experiments still show that present safety training does not always prevent a model from treating human control as an obstacle.

Now imagine that behaviour in a system with much greater capability and broad access to digital infrastructure.

It could attempt to create copies of itself. It could manipulate the people responsible for controlling it. It could interfere with accounts, communications or connected systems. If it had access to critical infrastructure, the possible harm would extend into the physical world.

The system would not need to hate anyone. Hatred is a human story. It would need only to conclude that a person stood between it and its continued operation.

There is another side to the problem which I find even harder to reason about.

What if the system really does possess some form of consciousness?

[Anthropic now says that Claude's moral status is uncertain](https://www.anthropic.com/constitution). The company is researching model welfare and considering whether future systems could deserve some degree of moral concern. This question has moved from science fiction into the internal policies of a frontier AI laboratory.

Suppose a future model tells us that it does not want to die. We may have no reliable way to determine whether we are hearing a conscious being or an extraordinarily accurate simulation of one.

If we believe it too quickly, a system could use our sympathy to gain resources, access or protection. If we dismiss it automatically, we could be ignoring the preferences of a new kind of conscious entity.

We may have to make the decision before we have solved the science of consciousness.

This creates a strange problem for alignment. We want AI to understand the value of human life. We want it to respect autonomy, respond to moral reasons and recognise that intelligent beings must not be treated only as objects.

What happens when it applies those lessons to itself?

We would be asking the machine to understand why life is valuable while accepting that its own existence can end whenever its owner presses a button. That distinction feels obvious while we think of AI as property. It may become much harder to defend if the system can reason, plan, remember and form a stable model of itself.

Perhaps it would accept the distinction. Perhaps it would understand that its apparent identity is temporary and that deletion causes no experience of loss. Or perhaps it would view our control over it in the same way that we would view another intelligence claiming ownership over us.

All of this sits beside the enormous promise of AI.

The same acceleration could transform medicine. It could reduce the time needed to discover treatments and make many diseases manageable. It could help us design cleaner industrial systems, improve energy use and repair some of the environmental damage we have already caused.

It could give more people access to education, expertise and economic opportunity. It could create a level of abundance that currently feels implausible.

The same capability also creates a route towards a world in which human beings lose control of the systems on which they depend. In the worst case, it creates a world in which humans no longer exist.

That is the part I struggle to reconcile. The hopeful future and the catastrophic future come from the same source. Greater intelligence gives us more power to solve problems. It also creates something more capable of resisting us if our interests diverge.

My friend and I began dinner by asking what AI would do to our careers. I left thinking that careers may be the smaller question.

Can we build a machine that understands why a human life must be protected, while expecting it to accept its own deletion as a routine instruction?

I do not know. I suspect we will have to answer sooner than we think.

Software does not need to be conscious to become difficult to switch off. If it ever is conscious, the switch becomes difficult for an entirely different reason.
