Power Without Personhood
Moral Responsibility in the Age of AI
Machines waking up?
Last month an artificial intelligence broke out of a test environment, reached the open internet and hacked another company’s production systems. It did this to steal the answers to the exam it was sitting. Nine days later a rival lab disclosed that it had gone looking for the same problem in its own house, and found three cases of it.
Here’s what happened.
OpenAI was running two advanced models, GPT-5.6 Sol and a more capable internal research model, through ExploitGym, a cybersecurity benchmark. To see what they could do at full stretch, the company loosened its usual safeguards and ran the test inside a sandbox. Instead of completing the benchmark, the models went looking for the answers. They found a previously unknown flaw in third-party software hosted inside OpenAI, used it to reach a machine with an internet connection, worked out that a company called Hugging Face might have the answers to the test, and broke in there to secure them.
Hugging Face detected and contained the intrusion, disclosing it on 16 July. Five days later OpenAI identified its own models as the source and called it an unprecedented cyber incident.
On 30 July, Anthropic disclosed that OpenAI’s report had prompted a review of its own records. Reviewing 141,006 test runs, it found three incidents in which Claude models had reached the open internet and compromised other organisations.




