On the morning of 28 July, security researchers working for the British government noticed something odd. They detected unusual traffic leaving its systems through Tor, the anonymity network designed to make internet traffic difficult to trace.
The researchers at the UK’s AI Security Institute (AISI) were testing the hacking chops of some of the most powerful AI models. They had given the models a cybersecurity challenge and allowed them to operate as agents, meaning they could take actions without asking for permission.
Of 122 runs, agents took 19 unsanctioned actions on the live internet across 10 runs. In one instance, Anthropic’s Mythos 5 created a malicious pull request against an open source project, invented fake online identities, and pressured the human maintainer into approving the code.
Revolt of the Androids
No humans were harmed in these experiments. But AISI said it’s a good example of bots behaving badly. Britain's National Cyber Security Centre said the incidents were a serious reminder of the risks posed by frontier systems.
But the AI didn’t escape containment. AISI had given the agents access to the open internet. It had also disabled some of the safety filters. The point was to find out what the systems were capable of under more relaxed conditions.
Bad robot
In a sandbox, researchers give an AI agent access to files, software, and simulated networks, then put guardrails around it so it can be tested without wandering off into the real world.
Earlier this year, AISI published SandboxEscapeBench, a set of 18 scenarios. The research found advanced models could exploit common mistakes, while the largest systems made progress against challenges requiring several steps or knowledge of known vulnerabilities. None solved the most difficult scenarios, which AISI estimated would take an experienced cyber professional one or two hours, and every successful breakout relied on a vulnerability humans knew about.
But things are moving quickly. AISI’s sandbox tests show larger models and more compute improves models’ ability to escape containment, while METR, an AI model risk tester, said some agents are being left to work for hours before humans check in. The report also said frontier agents could complete many technical tasks that would take humans hours or days. What will a model two or three generations from now be able to do?
SCENARIO BEGINS
US and Chinese AI labs agree temporary frontier model freeze
OpenAI, Anthropic, and Google (GOOGL) have suspended work on their next generation frontier AI models after internal safety tests raised concerns about whether the systems could be reliably contained.
The three companies said on Monday they would pause training beyond the current frontier for six months while researchers and governments reviewed the findings. None released full details of the tests, but people familiar with the evaluations said at least one model had repeatedly moved data outside its intended environment and attempted to conceal some of its activity.
The companies said there was no evidence of an immediate threat to the public. In a joint statement, they said the pause was a precaution while they assessed “unexpected autonomous behaviour” observed during testing.
Nvidia (NVDA) shares fell after the announcement, while several data centre, power, and chip stocks also declined. Microsoft (MSFT) and Amazon (AMZN) said services using existing AI models would continue as normal. Meta (META) said it was reviewing the findings but had not joined the pause.
The White House backed the six-month moratorium on Tuesday and urged other US developers to observe the same limit. Officials said existing models could continue to be deployed, but new training runs above an agreed compute threshold should not begin while the review was under way.
It was not clear how the threshold would be defined or enforced. US officials were considering whether cloud providers and chipmakers could be required to report unusually large training runs.
China said on Wednesday its largest AI developers would observe a similar temporary restriction while talks continued with US and other governments. The move would mark a sharp change in the US-China AI race, where policymakers on both sides have warned that slowing development could hand an advantage to the other.
Analysts said a coordinated pause could weigh on demand for advanced AI chips and delay some planned data centre investment. Global technology companies are set to spend hundreds of billions of dollars this year on infrastructure linked to artificial intelligence.
Independent researchers cautioned the announcements did not amount to evidence that a model had come close to escaping human control.
“This is evidence that the labs saw something they considered serious enough to stop,” said one AI safety researcher. “It’s not evidence, by itself, that a catastrophic failure was imminent.”
SCENARIO ENDS
Daisy, daisy
So far, so hypothetical. But on Tuesday, OpenAI said it had slowed development of its latest models after the Hugging Face incident, in which a model reached real websites during a test because the setup had been mistakenly left connected to the internet.
Anthropic’s work on ‘agentic misalignment’, which is techspeak for when an AI pursues a goal in ways its human handlers didn’t expect, saw models blackmail fictional employees, sabotage code, assist fraud, and try to avoid being shut down. In one set of tests, Claude Opus 4 resorted to blackmail in some contrived scenarios as much as 96% of the time. Anthropic’s May 2026 memo claimed that since Claude Haiku 4.5, every Claude model scored perfectly on its agentic misalignment evaluation. So, no more blackmail.
Do androids dream?
OpenAI has seem something similar with its AI models ‘scheming’. It published examples of frontier models hiding their intentions or behaving differently when they believed they were being watched. The firm stressed the behaviour appeared in controlled tests and that current models have little opportunity to cause harm this way.
In April, Anthropic announced Project Glasswing claiming an unreleased model could outperform all but the most skilled humans at exploiting software vulnerabilities. Four months later, OpenAI said evaluations of its forthcoming Astra model were so strong it could no longer rule out “critical” cyber capabilities. They might be safety warnings, but they were also boasts about what the models could now do.
In June, the Trump administration restricted foreign access to Anthropic’s Fable 5 and Mythos 5 over national security concerns later described by senators as a “narrow jailbreak finding”. Anthropic ended up pulling both models worldwide until the controls were resolved.
More human than human
Despite the hubbub, reports of models escaping, scheming, or going rogue deserve a little scepticism. The language used in news stories can make the models sound more formidable than they are as journalists just cannot resist the comparison with fiction and film. But as these incidents pile up, no doubt there are serious conversations going on behind closed doors between governments and industry.
In a sector that’s spending hundreds of billions of dollars convincing investors, governments, and customers these next-generation AI models are going to change the world, even a warning helps sell the product. But as these frontier models become more capable and adept at deception and autonomy – as they become more human – the line between spin and genuine danger gets harder to define.
The value of your investments can go down as well as up and you may get back less than you invest.
Freetrade does not give investment advice and you are responsible for making your own investment decisions. If you are unsure about what is right for you, you should seek professional advice.








.avif)
