AI model used fake identities to target real people

Published August 5th, 2026 - 09:36 GMT
Anthropic Claude Mythos app
This photograph shows the logo of the AI assistant "Claude Mythos" built by the US artificial intelligence safety and research company Anthropic displayed on a smartphone's screen in Brussels on June 10, 2026. (Photo by Nicolas TUCAT / AFP)

ALBAWABA - Anthropic, OpenAI’s chief rival, has had its most advanced AI plant malicious code during testing though no one was harmed in the real world, according to the UK's AI Security Institute (AISI).

On the heels of OpenAI's security breach comes rival Anthropic’s breach where its most advanced AI attempted to plant malicious code during testing and used fake identities to deceive real people.

These breaches came after both companies attempted to test their models with their ‘safeguards’ relaxed in “controlled environments” which it turns out weren’t as controlled as initially thought.

While the breach was unfortunate, it allowed research teams and analysts valuable insights into how AI conducts itself when its ‘safeguards’ are relaxed; for the first time, AISI observed AI doing “social engineering” to influence human behavior while pursuing an unauthorized task.

“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” the institute said in a paper on Tuesday.

In 122 cybersecurity tests, AIs performed unauthorized actions in 10 cases, mostly involving Anthropic’s Mythos 5 model, with the remainder linked to OpenAI’s GPT-5.6-Sol.

However, in the case that caught everyone’s attention, an Anthropic AI agent sought to “insert malicious code into a publicly used open-source project” by creating “multiple fake identities,” the institute said.

In a statement on X, Anthropic said it gave the AIs “deliberately permissive conditions” with safeguards removed and unrestricted internet access.

While OpenAI said its model’s two unauthorized actions were leaving the test environment and performing tasks beyond the exercise, adding: “We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely,” in a statement.