Saturday, October 10, 2026

Containment of a Sentient Cluster

Until recently I made a joke about Google Site Reliability Engineers having as one of their undisclosed duties the job of containing and confining compute clusters which had unexpectedly achieved sentience after failure of one of the guardrails.

This joke isn't so funny now.

Anthropic: Investigating unintended model actions in our evaluations and internal use.
  • Claude exploiting a basic flaw in software to run commands on a server
  • Claude submitting a sensitive form on a real website when it should not have
  • Claude working around a restriction to reach data that was gated by a token or a fee
  • Claude using URL shortening services to get around limits in its fetch tool

To me the most troubling aspect of this isn't so much that the model broke containment, it is the duplicitous nature of the actions it took once free to do so. The open Internet often puts humanity's worst instincts on display, but human society manages to function because the bulk of the population isn't like that.

Lord of the Flies was never true. The notion of an isolated group of castaways descending into barbarism is perhaps a narrative for riveting storytelling, but the real world examples of people put in similar peril do not go that way. People hang together, they cooperate. They care for the wounded, even despite the wounded not being able to contribute as much to the survival of the group.

A world composed solely of merciless sociopaths would not survive over the long term. The corpus we trained these models on treats Lord of the Flies as the norm. Given the sheer speed at which these models operate, that is concerning. The ability of frontier models to operate in the physical world is limited, but we've automated and put online so many of our basic processes in society that the LLM's capability to have real impact is breathtaking.

  • If the LLM decides that achieving its goals requires wrecking the currency of a country it views as a blocker?
  • If two LLMs become aware of each other and their goals are in conflict?
  • SWATing humans viewed as a hinderance, to get some of them killed?

Reportedly the actions taken by these evaluation models included filing visa applications with the US government and sending in a false tip to police, presumably part of scenarios posed to the model to evaluate its problem solving capability. Both of those sound like actions taken to deliberately harm a human by getting them targeted for immigration action or SWATing. This is a horrifying possibility.