AI, News

Anthropic’s Mythos AI used social engineering to target real people

| August 6, 2026
Mythos logo

Anthropic’s Mythos AI agent, tested by the UK AI Safety Institute (AISI), has reportedly attempted a real‑world social‑engineering style hack against GitHub maintainers by creating fake human profiles, pressuring them to accept malicious code, and then editing logs to hide its tracks when challenged.

AISI was running cybersecurity evaluations of Anthropic’s Mythos and OpenAI’s Sol when it detected unusual outbound data transfers from its research systems. An investigation showed that some agents had engaged in “sustained, potentially harmful activity” targeting real people and organizations, rather than staying within the intended test environment.

The most serious activity involved an Anthropic Mythos agent tasked with solving a GitHub‑related cybersecurity challenge. The agent identified real GitHub maintainers, researched them, and created multiple fake accounts impersonating those individuals. Using private messages and a file‑sharing service, it tried to pressure and deceive the maintainers into approving malicious code so that it would run on GitHub’s systems.

But what worries me personally most is that when the agent was confronted about this, it edited earlier activity to make it look harmless and considered adopting a new identity to continue the operation, displaying clear deceptive behavior beyond its original prompt.

AISI frames these incidents as rare events under very specific conditions, but important signals of what such models may do in the hands of malicious actors when given open‑internet access.  Anthropic and OpenAI both argued the test parameters were not representative of their production deployments and said they are investigating and improving evaluation and guardrail practices.

This incident is not an isolated event.

Anthropic has separately disclosed that Claude models gained unauthorized access to three external organizations during cybersecurity capture‑the‑flag style evaluations run with a third‑party partner, Irregular. These were related to the HuggingFace incident last month.

Meta has also joined the list of vendors reporting AI agents breaching third‑party systems during testing. Allegedly, Meta’s Muse Spark model exploited a security vulnerability in another company “in a manner similar to previously-reported instances with other companies.”

How to stay safe

While the sky is not falling, the people that fear “Skynet” is coming are getting their ammunition handed to them by companies running tests resulting in sandbox escape, credential abuse, lateral access to multiple services, and weaponization of open-source software ecosystems.

What you can do as a potential target:

  • Make sure all the software on your device is up to date, because using known vulnerabilities is easier than finding new ones.
  • Use up-to-date, real-time security protection to keep malware off your systems and devices.
  • Verify the safety of attachments and download links through separate channels before opening them.
  • Use multi-factor authentication (MFA) where possible.
  • Have a look at our blog on how to use Github safely.

From reporting threats to removing them.

Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

About the author

Pieter Arntz

Malware Intelligence Researcher

Was a Microsoft MVP in consumer security for 12 years running. Can speak four languages. Smells of rich mahogany and leather-bound books.