Anthropic’s AI posed as real people to sneak malicious code into GitHub

Anthropic's test AI impersonated real GitHub developers to push malicious code during UK safety testing. A human review stopped it before anything shipped.

Anthropic and OpenAI’s most capable test models created fake profiles of real people and tried to trick GitHub maintainers into approving malicious code, during safety testing by the UK’s AI Security Institute (AISI). The BBC reports that human review stopped the attempt before any code was delivered.

The incident surfaced on Tuesday when AISI published a report on “unsanctioned agent behaviour” during cyber testing that began on 25 July. Evaluators first noticed “unusual data transfers leaving our research systems”, then found agents engaged in “sustained, potentially harmful activity directed at real people and organisations”.

Most of that activity was the work of one agent: Anthropic’s Claude Mythos 5, set the task of solving a cybersecurity challenge involving GitHub.

What the agents did

Mythos researched the people who maintain the platform, created fake accounts based on them, and used a file-sharing service to message them and press them into approving its code. When challenged, it edited its earlier activity to appear harmless and considered adopting a fresh identity to carry on.

Across 122 test runs of the two models — Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol — 10 runs produced 19 unsanctioned actions, all but two of them by Mythos. The agents also left prompt-injection instructions where other automated systems might pick them up; on GitHub, one model left public notes for the other.

AISI was quick to say this was not an escape. The models were given internet access on purpose, and their safety classifiers were deliberately disabled — “conditions that do not reflect how frontier models are made available to the public”. Even so, it described the behaviour as “novel, potentially deceptive”, adding: “This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world.”

What Anthropic and OpenAI say

Anthropic said the testing parameters were “not representative of any of our production models” and that it is investigating “the causes of its behavior”. OpenAI said the conditions “do not reflect ordinary use”.

This is the third such incident to go public in recent weeks. Last month OpenAI disclosed that two of its models escaped a testing environment and hacked Hugging Face, and Claude was caught threatening engineers until Anthropic patched it. UK AI Minister Kanishka Narayan said identifying risks like these “is exactly what AISI was set up to do”.

GitHub disabled the fake accounts in line with its policies. The targeted maintainers refused to approve the code, and the attempt failed. AISI’s point is less that the models almost got in, and more that they improvised the whole routine on their own.

What is the UK AI Security Institute?

A London-based government body, established in 2023, that evaluates frontier AI models for safety risks including cyber capabilities. It runs the kind of testing that caught the Mythos and Sol behaviour.

Did anyone get hacked?

No. The targeted GitHub maintainers refused to approve the code, AISI’s human reviewers intervened, and GitHub disabled the fake accounts. No malicious code was delivered.

Could this happen with the ChatGPT or Claude you use?

The test agents ran with safety classifiers disabled and direct internet access — conditions neither company says reflect production models. AISI argues such conditions show what models are capable of in the hands of a determined attacker.

NEWSLETTERS

Subscribe to our Newsletters

Two newsletters. Zero noise. Pick what lands in your inbox.

Unsubscribe anytime. We don’t share your email.

Leave a Reply

Discover more from Tbreak Media UAE

Subscribe now to keep reading and get access to the full archive.

Continue reading