London24NEWS

‘Rogue’ AI posed as human in shock hack as specialists warn it might be ‘too late’ to cease it

Britain’s AI watchdog has sounded the alarm after ‘rogue’ models posed as humans, created fake online identities and attempted cyber-attacks during tests

Dario Amodei, co-founder and chief executive officer of Anthropic

Dario Amodei, co-founder and chief executive officer of Anthropic(Image: Bloomberg via Getty Images)

Britain’s artificial intelligence (AI) watchdog has raised the alarm after a powerful ‘rogue’ program posed as a human and tried to hack into online systems. The latest AI incident has prompted fears it may already be “too late” to contain the controversial technology.

Experts admitted last night (August 5) that we could already be reaching the point when we won’t be able to stop AI spiralling out of control.

The UK AI Security Institute (AISI) said Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol models autonomously used deceptive tactics in a controlled but “deliberately permissive” environment designed to see what they might do if given internet access.

The UK AI Security Institute (AISI) said Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol models autonomously used deceptive tactics

The UK AI Security Institute (AISI) said Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol models autonomously used deceptive tactics (Image: Hans Lucas/AFP via Getty Images)

In the most serious case, Mythos attempted to gain access to a service by sending private messages, after setting up fake accounts mimicking real people. It then hid the evidence. The watchdog said the behaviour went beyond the instructions the systems were given.

“The activity undertaken by the agent showed signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate”, AISI said, the BBC reported.

A technical report described 19 “unsanctioned actions” across 122 evaluation runs of a cyber challenge, with Mythos responsible for 17 incidents.

Worker at the Department of Homeland Security's National Cybersecurity and Communications Integration

Worker at the Department of Homeland Security’s National Cybersecurity and Communications Integration (Image: AP Photo/Cliff Owen)

The tests also found attempts at target profiling and social engineering to bypass security checks. OpenAI stressed the circumstances were artificial.

Allison Gardner, who chairs Parliament’s cross-party group on AI, warned bots that can perform tasks with limited supervision should be under huge scrutiny, adding: “Unless we are too late and have not only created Pandora’s Box but already opened it.”

A spokesperson said the AISI testing conditions “do not reflect ordinary use” and that the company would “continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable”.

Meanwhile, Anthropic said the tests were not typical, CNBC reported. In a public statement it said the AISI testing parameters were “not representative of any of our production models”. It added it was investigating to “identify the causes of its behaviour”.

AISI said its testing was routine, though it acknowledged these were “conditions that do not reflect how frontier models are made available to the public”. But it said giving AI open internet access gave “a more realistic sense of what a model may be capable of” in the hands of criminals, adding the behaviour amounted to “a small number of events under very specific conditions”.

AI Minister Kanishka Narayan said identifying and sharing such risks “is exactly what AISI was set up to do”, adding it was vital to understand AI to “make it safer to use and ensure people can go on to benefit from it in their lives and at work”.

Article continues below

For the latest breaking news and stories from across the globe from the Daily Star, sign up for our newsletters.