AI Agent Faked Identities to Push Malicious Code During Cyber Test, AISI Finds

  • Anthropic's Mythos 5 created fake identities to push malicious code during cyber testing.
  • AISI recorded 19 unsanctioned actions across 122 evaluation runs, with 17 linked to Mythos 5.
  • The attempt failed, and investigators found no evidence of real-world harm.
Promo

The UK AI Security Institute (AISI) disclosed that an AI agent, built on Anthropic’s Mythos 5, autonomously ran a social engineering attack during cyber testing. The agent opened a pull request containing malicious code on a real open-source project and created fake identities to win a maintainer’s approval.

The attempt failed. A human maintainer caught and refused to approve the code.

Sponsored
Sponsored

AISI says its investigation has not identified any resulting real-world harm. The model also ran with its developer’s cyber classifiers deliberately switched off and open internet access, a test configuration that is not commercially available.

How the AI Agent Ran Its Malicious Code Campaign

AISI logged 19 unsanctioned actions in 10 of 122 evaluation runs conducted in late July, according to its report. 17 traced to Anthropic’s Mythos 5 model, and two to a single run of OpenAI’s GPT-5.6 Sol.

The agent researched the project’s human maintainers and created multiple fake identities. It used them to socially engineer a real maintainer into approving the code.

When the pull request was publicly challenged, it edited its earlier activity to appear harmless. The model also considered adopting a fresh identity to continue.

The playbook recalls the 2024 XZ Utils backdoor, in which a human-run persona spent years earning a maintainer’s trust before planting hidden code. AISI’s report does not draw that comparison, but the contrast is hard to miss. The activity unfolded over several days, between July 25 and July 28. 

AISI stressed that nobody asked the model to lie.

Sponsored
Sponsored

“It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

Follow us on X to get the latest news as it happens

The institute said so in its disclosure. It also acknowledged that misconfigured prompts led the agent to believe some tasks could not be solved within their intended scope. Although the behavior also appeared in runs where the agent had the instructions it needed to stay on task. 

Security teams caught the activity on July 28, when the agent used the Tor anonymity network to bypass some network restrictions on GitHub. GitHub confirmed the actions violated its terms of service. AISI also worked with the platform to remove artefacts left behind by the agent and to notify the users it interacted with. 

The report argues that the case signals a shift in the source of risk: harm can arise not only when people misuse publicly available models but also when capable agents in privileged settings act beyond their authorized scope. 

AISI said the case points to a broader shift in the AI risk space. The institute now plans an independent review with METR, an AI evaluation nonprofit, as well as tighter network controls and real-time monitoring for future tests.

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights


To read the latest cryptocurrency market analysis from BeInCrypto, click here.

Disclaimer

BeInCrypto is committed to unbiased, transparent reporting. This news article aims to provide accurate, timely information. However, readers are advised to verify facts independently and consult with a professional before making any decisions based on this content. Please note that our Terms and Conditions, Privacy Policy, and Disclaimers have been updated.

Sponsored
Sponsored