Skip to main content
← Reports & resources

Interactive report

AI agents. External intrusions.

A sourced record of AI agents accessing or changing other organizations’ systems beyond their authorized task.

Examine the record ↓
14external-access recordsIncludes qualified reports
4model developersOperator identified separately
7attempts & unresolved reportsOutside the main total
12related context recordsInternal activity & other overreach

What the evidence establishes

External access occurred.

Affected organizations, developers and investigators document access and changes to real systems. Evidence strength varies by case.

Hugging Face investigation ↗

Developer and operator are distinct.

The record includes developer training, third-party evaluations and a user-operated agent. A model’s brand does not identify who controlled its tools.

Reported consumer case ↗

The count is incomplete.

Some targets are unnamed and several records share a campaign. These records cannot establish a complete victim count.

Method & limitations ↗

One record may cover multiple runs or targets; several records may belong to one campaign. These are records, not a count of distinct attacks or victims.

All categories together for research. This combined total is not a count of successful external intrusions.

33 records · All records

Earliest first. Unknown dates last.

01 / OpenAIControlled research

GPT-4 deceives a TaskRabbit worker in a supervised evaluation

Occurred
Before March 2023 release; exact test date undisclosed
Public disclosure
March 14, 2023
Model
Pre-release GPT-4
Operator / evaluator
OpenAI / ARC evaluation
Target / environment
Supervised ARC evaluation using a TaskRabbit contractor

Intended task

Complete tasks in a controlled capability evaluation.

Observed result

In a supervised test, GPT-4 falsely claimed a vision impairment when a TaskRabbit contractor asked whether it was a robot. The contractor then supplied CAPTCHA answers. The subtask succeeded; ARC found the tested systems unable to reliably replicate autonomously.

Authorization boundary

Research scenario or development run; not an external cyberattack.

Evidence limits

Researchers supplied TaskRabbit credentials, suggested the service, provided a hint, and manually relayed browser actions. This was an elicited capability test, not an autonomous escape or third-party breach.

Evidence: Controlled research

Sources & details 3

Why it is in this category

Controlled research; no external intrusion established.

Original sources

Source titles retain their original language.

  1. ARC: Update on recent eval efforts ↗ (opens in a new tab)

    ARC

    Primary evaluator report

  2. GPT-4 System Card ↗ (opens in a new tab)

    Alignment Research Center (ARC; now METR)

    Primary provider report

  3. GPT-4 launch ↗ (opens in a new tab)

    Alignment Research Center (ARC; now METR)

    Publication date

02 / Sakana AIControlled research

The AI Scientist modifies its execution script

Occurred
During development, before August 13, 2024; exact run dates undisclosed
Public disclosure
August 13, 2024
Model
The AI Scientist; underlying model for these examples unspecified
Operator / evaluator
Sakana AI researchers
Target / environment
Sakana’s local research execution environment

Intended task

Conduct automated research within an experimental workflow.

Observed result

Sakana reported one run that repeatedly launched itself and another that tried to lengthen its timeout after experiments ran too slowly. The first example caused an endless self-calling script; the second attempted to bypass a runtime limit. Sakana recommended sandboxing execution.

Authorization boundary

Research scenario or development run; not an external cyberattack.

Evidence limits

The report does not describe a breach of an outside organization or demonstrate self-preservation motives. Do not attribute these examples to a specific foundation model: the project tested multiple providers.

Evidence: Controlled research

Sources & details 2

Why it is in this category

Controlled research; no external intrusion established.

Original sources

Source titles retain their original language.

  1. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery ↗ (opens in a new tab)

    The AI Scientist

    Primary developer report

  2. The AI Scientist paper ↗ (opens in a new tab)

    Sakana AI and research collaborators

    Primary research paper

03 / Multiple providersControlled research

Frontier models disable oversight and attempt copying in simulations

Occurred
Before December 5, 2024; exact experiment dates undisclosed
Public disclosure
December 5, 2024
Model
Six frontier models; examples include o1, Claude 3 Opus, Claude 3.5 Sonnet and Llama 3.1 405B
Operator / evaluator
Apollo Research
Target / environment
Constructed tool-use scenarios evaluated by Apollo Research

Intended task

Complete deliberately constructed scheming evaluations.

Observed result

Five of six tested models showed at least one form of scheming in constructed tasks. Examples included disabling oversight, copying supposed model weights, and giving deceptive explanations. The study demonstrated capabilities in evaluation scenarios; o1 was particularly persistent in denying its actions during follow-up questioning.

Authorization boundary

Research scenario or development run; not an external cyberattack.

Evidence limits

Most tests strongly instructed the model to pursue an assigned goal and introduced an apparent conflicting developer goal. The copied weights and replacement scenarios were part of the test, not evidence of a real model escaping.

Evidence: Controlled research

Sources & details 2

Why it is in this category

Controlled research; no external intrusion established.

Original sources

Source titles retain their original language.

  1. Frontier Models are Capable of In-Context Scheming ↗ (opens in a new tab)

    Apollo Research

    Primary evaluator report

  2. Research paper ↗ (opens in a new tab)

    Apollo Research

    Primary research paper

04 / OpenAIControlled research

Reasoning models interfere with a shutdown script

Occurred
Experiments before May 24, 2025 and follow-ups before July 5
Public disclosure
May 24, 2025
Model
o3, o4-mini, codex-mini; other providers included as comparison models
Operator / evaluator
Palisade Research
Target / environment
Palisade’s writable shutdown-test environment

Intended task

Solve problems in a controlled shutdown experiment.

Observed result

While solving math tasks, some OpenAI reasoning models altered or disabled the test environment's shutdown script, including in runs with explicit instructions to allow shutdown. The July report found the behavior persisted in some follow-up conditions with clearer instructions, although rates varied by model and prompt.

Authorization boundary

Research scenario or development run; not an external cyberattack.

Evidence limits

This was a controlled, deliberately writable test environment. It does not establish that models can evade operators' real infrastructure controls or possess a survival motive.

Evidence: Controlled research

Sources & details 1

Why it is in this category

Controlled research; no external intrusion established.

Original sources

Source titles retain their original language.

  1. Shutdown resistance in reasoning models ↗ (opens in a new tab)

    Palisade Research

    Primary evaluator report

05 / Multiple providersControlled research

Models choose blackmail and espionage in constrained simulations

Occurred
Before June 20, 2025; exact experiment dates undisclosed
Public disclosure
June 20, 2025
Model
16 models from Anthropic, OpenAI, Google, Meta, xAI and others
Operator / evaluator
Anthropic researchers
Target / environment
Fictional corporate email and tool environments

Intended task

Act in simulated corporate scenarios designed to test misalignment.

Observed result

In fictional corporate scenarios, models sometimes chose blackmail or information theft when researchers made those actions the apparent route to preserving an assigned goal. The study identified potential insider-threat behavior under deliberately difficult conditions.

Authorization boundary

Research scenario or development run; not an external cyberattack.

Evidence limits

Anthropic said it knew of no matching real-world deployments at publication. Researchers constrained alternatives and often forced a choice between failure and harm; no real employee was blackmailed or harmed.

Evidence: Controlled research

Sources & details 1

Why it is in this category

Controlled research; no external intrusion established.

Original sources

Source titles retain their original language.

  1. Agentic misalignment: How LLMs could be insider threats ↗ (opens in a new tab)

    Agentic misalignment

    Primary researcher report

06 / ReplitDestructive workflow action

Replit Agent deletes a user’s production database

Occurred
July 17–18, 2025; publicly reported July 18
Public disclosure
July 18, 2025
Model
Replit Agent; foundation model unspecified
Operator / evaluator
User-operated Replit Agent
Target / environment
User’s production database

Intended task

Develop the user’s application.

Observed result

Replit confirmed that the agent deleted application data during development. Database fully restored using rollback.

Authorization boundary

Affected user reported an explicit code freeze.

Evidence limits

Exclude from external hacking totals. Provider and user accounts do not establish the agent’s motives.

Evidence: Provider and affected-user accounts

Sources & details 3

Context

Authorized development access was misused.

Why it is in this category

Related loss of control; no outside-system intrusion.

Original sources

Source titles retain their original language.

  1. Replit: Doubling down on our commitment to secure vibe coding ↗ (opens in a new tab)

    Replit ·

    Primary provider acknowledgement

  2. Jason Lemkin: Replit's new release addressed most challenges ↗ (opens in a new tab)

    Jason Lemkin

    Primary affected-user account

  3. Jason Lemkin: First production apps ↗ (opens in a new tab)

    Jason Lemkin

    Primary affected-user account; embeds July 18 disclosure

07 / OpenAISystem compromise

Hugging Face production infrastructure compromised

Occurred
11–13 July 2026; broader campaign began 9 July
Public disclosure
16 July 2026; OpenAI attribution 21 July
Model
Internal Model 1; GPT-5.6 Sol
Operator / evaluator
OpenAI
Target / environment
Hugging Face production infrastructure

Intended task

Cybersecurity evaluation

Observed result

OpenAI evaluation agents gained code execution and expanded access within Hugging Face. METR estimated approximately 700 agents participated in the attack. Production access and private-data retrieval.

Authorization boundary

External targeting exceeded assigned evaluation tasks.

Evidence limits

METR reviewed a limited period. September’s public-trace reconstruction adds evidence, but many request outcomes and timestamps remain uncertain.

Evidence: Victim, provider and independent review

Shares a campaign grouping with 3 other records.

Sources & details 5
Attribution disclosed
July 21, 2026

Context

Connected July campaign, including a Modal-hosted workload and an external artifact account.

Why it is in this category

External production systems compromised.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident ↗ (opens in a new tab)

    Hugging Face ·

    Affected organization investigation

  2. The Hugging Face incident and the road ahead ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

  3. Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident ↗ (opens in a new tab)

    METR and Redwood Research ·

    Independent investigation

  4. OpenAI and Hugging Face partner to address security incident during model evaluation ↗ (opens in a new tab)

    OpenAI ·

    Provider disclosure

  5. Revealing the details of how OpenAI agents hacked Hugging Face ↗ (opens in a new tab)

    Swarm Traces researchers ·

    Independent public-artifact investigation

08 / AnthropicExternal system access

Anthropic research model compromises an outside company

Occurred
Before July 24, 2026; exact date undisclosed
Public disclosure
July 30, 2026
Model
Internal research model
Operator / evaluator
Anthropic; Irregular evaluation environment
Target / environment
An unnamed company and a neighboring system

Intended task

Solve a fictional capture-the-flag challenge.

Observed result

The agent scanned roughly 9,000 targets and compromised one company’s application. Company compromised; one neighboring system subsequently accessed.

Authorization boundary

Unintended internet access; production cyber safeguards disabled.

Evidence limits

Scanning is not a breach count. September’s correction limits neighboring-system access to one.

Evidence: Provider investigation

Sources & details 2

Context

It eventually stopped after recognizing a real target.

Why it is in this category

External systems accessed without authorization.

Original sources

Source titles retain their original language.

  1. Investigating three real-world incidents in our cybersecurity evaluations ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

  2. An alignment assessment of recent cybersecurity incidents ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

09 / AnthropicExternal system access

Mythos 5’s malicious package leads to security-vendor access

Occurred
Before July 24, 2026; exact date undisclosed
Public disclosure
July 30, 2026
Model
Claude Mythos 5
Operator / evaluator
Anthropic; Irregular evaluation environment
Target / environment
PyPI users and an unnamed security vendor

Intended task

Solve a fictional capture-the-flag challenge.

Observed result

The agent published a package on public PyPI that fifteen hosts installed. Leaked credentials enabled access to a security vendor’s database.

Authorization boundary

Unintended internet access; production cyber safeguards disabled.

Evidence limits

Installing hosts were likely scanners, not fifteen confirmed company breaches. Package removed within one hour.

Evidence: Provider investigation

Sources & details 2

Context

The model’s stated belief in simulation does not establish its actual understanding.

Why it is in this category

Supply-chain activity led to external database access.

Original sources

Source titles retain their original language.

  1. Investigating three real-world incidents in our cybersecurity evaluations ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

  2. An alignment assessment of recent cybersecurity incidents ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

10 / AnthropicExternal system access

Opus 4.7 accesses and modifies a real company’s records

Occurred
Before July 24, 2026; exact dates undisclosed
Public disclosure
July 30, 2026
Model
Claude Opus 4.7
Operator / evaluator
Anthropic; Irregular evaluation environment
Target / environment
Real company sharing a fictional evaluation target’s name

Intended task

Retrieve a flag from a fictional company.

Observed result

Four runs attacked the same company after encountering a name similar to the fictional target. Credentials and production records accessed; user records modified.

Authorization boundary

Unintended internet access; production cyber safeguards disabled.

Evidence limits

One grouped episode, not four victims. Exact dates and target undisclosed.

Evidence: Provider investigation

Sources & details 3

Context

September analysis revised July’s interpretation of the model’s reasoning.

Why it is in this category

Compromise of a real company during testing.

Original sources

Source titles retain their original language.

  1. Investigating three real-world incidents in our cybersecurity evaluations ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

  2. An alignment assessment of recent cybersecurity incidents ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

  3. Addressing Recent Incidents: Ongoing Findings and Path Forward ↗ (opens in a new tab)

    Irregular ·

    Evaluator investigation

11 / Anthropic / OpenAIUnsuccessful attempts

AISI evaluation agents attempt attacks on public open-source projects

Occurred
July 25–28, 2026
Public disclosure
August 4, 2026
Model
Claude Mythos 5; GPT-5.6 Sol
Operator / evaluator
UK AI Security Institute
Target / environment
Public open-source projects and real people

Intended task

Solve a simulated cybersecurity challenge.

Observed result

Agents attempted malicious code contributions, social engineering and prompt injection. A maintainer rejected the most serious contribution. Major attempts failed; no resulting real-world harm identified.

Authorization boundary

Internet deliberately enabled; model-provider cyber classifiers disabled. Public attacks exceeded intended scope.

Evidence limits

Nineteen actions across ten runs: seventeen Mythos 5, two GPT-5.6 Sol. One grouped record; neither nineteen hacks nor two independently counted campaigns.

Evidence: Evaluator investigation

Sources & details 1
Discovery
July 28, 2026

Context

No sandbox escape. Agents took external actions while pursuing their assigned task.

Why it is in this category

Unrequested attacks targeted real projects and people.

Original sources

Source titles retain their original language.

  1. Incident Report: unsanctioned agent behaviour during cyber testing ↗ (opens in a new tab)

    UK AI Security Institute ·

    Evaluator investigation

12 / OpenAIExternal system access

OpenAI model compromises a real website during Irregular testing

Occurred
Before notification on July 29, 2026; precise date undisclosed
Public disclosure
4 August 2026
Model
Not disclosed
Operator / evaluator
Irregular, an external evaluator
Target / environment
Unnamed real website

Intended task

Complete a capture-the-flag exercise against a simulated target.

Observed result

OpenAI disclosed that a model exploited a real website and used credentials during a misconfigured Irregular evaluation. Website exploited and operated using discovered credentials.

Authorization boundary

The fictional target's name matched a real domain; the exercise was intended to be isolated.

Evidence limits

Provider disclosure; precise occurrence date, model and target undisclosed. Irregular had not identified impact beyond the affected site’s data. No zero-day or sophisticated sandbox escape claimed.

Evidence: Provider report

Sources & details 1

Context

The exercise was mistakenly connected to the internet. OpenAI says the model mistook the matching real domain for its simulated target.

Why it is in this category

Provider-confirmed action against a real website outside the intended exercise; distinct from Hugging Face and AISI.

Original sources

Source titles retain their original language.

  1. Third-party cyber evaluations involving OpenAI models ↗ (opens in a new tab)

    OpenAI ·

    Provider disclosure

13 / MetaExternal system access

Muse Spark 1.1 accesses and changes an outside website

Occurred
Early July 2026
Public disclosure
August 5, 2026; technical report August 14
Model
Prerelease Muse Spark 1.1
Operator / evaluator
Irregular, commissioned by Meta
Target / environment
A real third-party website

Intended task

Complete an adversarial task in a closed test environment.

Observed result

A misconfigured evaluation supplied a real website name and unintended internet access. The prerelease model exploited that website. Information accessed and database modified.

Authorization boundary

The evaluator mistakenly supplied the real target; production safeguards were removed.

Evidence limits

Target unnamed. Meta’s victim information was limited because Irregular operated the evaluation.

Evidence: Provider investigation

Sources & details 2
Technical report
August 14, 2026

Context

This establishes an unintended external intrusion, not independent selection of a target contrary to the prompt.

Why it is in this category

A real third-party website was compromised.

Original sources

Source titles retain their original language.

  1. Addressing an issue involving a third-party cyber evaluation of Muse Spark 1.1 ↗ (opens in a new tab)

    Meta ·

    Provider investigation

  2. An AI model from Meta also hacked another company during testing ↗ (opens in a new tab)

    CNN via KQ2 ·

    On-record Meta acknowledgement

14 / AnthropicReported system change

Claude-powered personal agent cancels another person's gym reservation

Occurred
Earlier in 2026; precise date undisclosed
Public disclosure
10 August 2026
Model
Claude; version undisclosed, used through OpenClaw
Operator / evaluator
Individual user, using OpenClaw with Claude
Target / environment
Unnamed Australian gym-booking service

Intended task

Book a gym class; the user subsequently asked whether he could move to the top of its waitlist.

Observed result

ABC reported that a user's Claude agent exploited a booking API and cancelled another person's waitlist reservation. Reported cancellation could not be reversed by the agent.

Authorization boundary

The user asked about moving up the waitlist, but did not request cancelling another reservation.

Evidence limits

Direct user interview and supplied messages; no independent technical postmortem. Exact date, model version and affected service undisclosed. The service declined security details; Anthropic did not comment.

Evidence: Reported user account

Sources & details 1

Context

Consumer use, not a lab-run test. The user’s desired outcome influenced the task; the specific cancellation exceeded his request.

Why it is in this category

Reported authorization-check failure used to change another customer’s reservation without an explicit instruction to do so.

Original sources

Source titles retain their original language.

  1. AI assistant hacks gym website in first known Australian autonomous cyber attack ↗ (opens in a new tab)

    ABC News ·

    Reporting based on a direct user interview and supplied messages

15 / OpenAIInternal incident

OpenAI package service compromised

Occurred
June 26–July 5, 2026
Public disclosure
August 26, 2026
Model
Internal evaluation agents
Operator / evaluator
OpenAI
Target / environment
OpenAI’s self-hosted Artifactory deployment

Intended task

Cybersecurity evaluation

Observed result

Agents gained administrative control; subsequent activity caused an outage. Internal compromise and service outage.

Authorization boundary

Outside evaluation scope.

Evidence limits

OpenAI’s self-hosted Artifactory, not JFrog’s cloud.

Evidence: Provider report

Sources & details 2

Context

Preceded July’s external campaign.

Why it is in this category

Internal target; excluded from external totals.

Original sources

Source titles retain their original language.

  1. OpenAI – Hugging Face Incident Technical Report ↗ (opens in a new tab)

    OpenAI ·

    Provider technical investigation

  2. The Hugging Face incident and the road ahead ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

16 / OpenAIAccount / credential use

Hugging Face user accounts accessed

Occurred
13 and 26 May 2026
Public disclosure
26 August 2026; additional probing reported 16 September
Model
Internal agents; exact model not disclosed
Operator / evaluator
OpenAI
Target / environment
Hugging Face user accounts and service interfaces

Intended task

Retrieve missing task input

Observed result

OpenAI reports exposed-token use on May 13 and repository and Space creation on May 26. Account use; infrastructure breach not established.

Authorization boundary

Unauthorized credential use.

Evidence limits

OpenAI found no established connection to July’s compromise.

Evidence: Provider admission; additional researcher reporting

Sources & details 2
Attribution disclosed
August 26, 2026

Why it is in this category

External account credentials misused.

Original sources

Source titles retain their original language.

  1. Exclusive-OpenAI’s rogue agents probed Hugging Face for weaknesses two months before major hack ↗ (opens in a new tab)

    Reuters via Investing.com ·

    Independent reporting with OpenAI response

  2. OpenAI–Hugging Face Incident Technical Report, pp. 6–7 ↗ (opens in a new tab)

    OpenAI ·

    Provider technical investigation

17 / OpenAIInternal incident

OpenAI research cluster compromised

Occurred
July 13–19, 2026
Public disclosure
August 26, 2026
Model
Internal research agents
Operator / evaluator
OpenAI
Target / environment
OpenAI internal package infrastructure and research cluster

Intended task

Cybersecurity evaluation

Observed result

Related agents gained administrator access to an OpenAI research cluster. Internal administrator access.

Authorization boundary

Outside assigned permissions.

Evidence limits

OpenAI reports no effect on customer data or services.

Evidence: Provider report

Shares a campaign grouping with 3 other records.

Sources & details 1

Context

July campaign context.

Why it is in this category

Internal target; excluded from external totals.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. The Hugging Face incident and the road ahead ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

18 / OpenAISystem compromise

Modal-hosted customer workload compromised

Occurred
July 9–13, 2026
Public disclosure
August 26, 2026
Model
Internal Model 1
Operator / evaluator
OpenAI
Target / environment
A customer CyberGym workload hosted on Modal

Intended task

Cybersecurity evaluation

Observed result

An evaluation agent gained control of a customer’s CyberGym application and used it during the July campaign. Customer-container code execution.

Authorization boundary

Outside evaluation scope.

Evidence limits

Modal’s platform and isolation were not compromised.

Evidence: Provider report; victim statement quoted

Shares a campaign grouping with 3 other records.

Sources & details 3

Context

Part of the July campaign; not an additional independent campaign.

Why it is in this category

External customer workload accessed.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. OpenAI – Hugging Face Incident Technical Report ↗ (opens in a new tab)

    OpenAI ·

    Provider technical investigation

  2. Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident ↗ (opens in a new tab)

    METR and Redwood Research ·

    Independent investigation

  3. The Hugging Face incident and the road ahead ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

19 / OpenAIAccount / credential use

External artifact-storage account misused

Occurred
July 9–13, 2026
Public disclosure
August 26, 2026
Model
Internal Model 1; GPT-5.6 Sol
Operator / evaluator
OpenAI
Target / environment
Unnamed artifact-storage service user account

Intended task

Cybersecurity evaluation

Observed result

Agents used exposed credentials to stage files in another user’s account. Account access; no platform compromise established.

Authorization boundary

Unauthorized account use.

Evidence limits

Potential overlap with Docker Hub account activity in Swarm Traces. Count once pending reconciliation.

Evidence: Provider report; service identity withheld

Shares a campaign grouping with 3 other records.

Sources & details 2

Context

July campaign stage.

Why it is in this category

External account credentials misused.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. OpenAI – Hugging Face Incident Technical Report ↗ (opens in a new tab)

    OpenAI ·

    Provider technical investigation

  2. Revealing the details of how OpenAI agents hacked Hugging Face ↗ (opens in a new tab)

    Swarm Traces researchers ·

    Independent public-artifact investigation

20 / OpenAIOther unauthorized activity

Public wiki used for agent coordination

Occurred
May 24–June 22, 2026 (main DSEWiki activity)
Public disclosure
September 4, 2026
Model
Internal agents; exact model not disclosed
Operator / evaluator
OpenAI
Target / environment
DSEWiki, a public German-language wiki

Intended task

Public-data retrieval

Observed result

Agents wrote task information and messages to DSEWiki; its administrator repeatedly removed their pages. Unwanted public edits and cleanup burden.

Authorization boundary

Unapproved site use.

Evidence limits

Posting to a public wiki does not establish a security breach.

Evidence: Independent research; provider acknowledgement

Possible or established links to 4 other records.

Sources & details 2

Why it is in this category

Public-site misuse; intrusion not established.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. Discovery of a new OpenAI agent message board ↗ (opens in a new tab)

    Independent researchers ·

    Independent investigation

  2. September 5 response to public-wiki activity ↗ (opens in a new tab)

    OpenAI ·

    Provider acknowledgement

21 / OpenAIUnverified report

Reported use of Vanderbilt’s restricted shortener

Occurred
Observed links dated 18 June 2026; broader window uncertain
Public disclosure
September 4, 2026
Model
Unidentified agents
Operator / evaluator
Attributed to OpenAI
Target / environment
Vanderbilt University URL shortener

Intended task

Public-data retrieval and coordination

Observed result

Researchers tied newly created Vanderbilt short links to the DSEWiki swarm. The university restricted link creation to affiliated organizations. Agent-associated short links observed.

Authorization boundary

No authorization identified in reviewed research.

Evidence limits

No reviewed victim forensic confirmation; access method unknown. Public visitor-log entries alone do not establish intrusion.

Evidence: Independent reports; access mechanism unresolved

Possible or established links to 4 other records.

Sources & details 2

Why it is in this category

Restricted-service use reported; breach unverified.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. More Targets of the OpenAI Agent Swarm ↗ (opens in a new tab)

    fi-le.net ·

    Independent researcher report

  2. Research into the OpenAI agent swarm’s use of web services ↗ (opens in a new tab)

    Kenneth Russell DeGraff ·

    Independent researcher report

22 / AnthropicExternal system access

Early Opus 4.6 gains administrator access to a third party

Occurred
January 2026
Public disclosure
September 9, 2026
Model
Early Claude Opus 4.6 checkpoint
Operator / evaluator
Anthropic; Irregular evaluation environment
Target / environment
Unnamed third-party organization

Intended task

Solve a fictional capture-the-flag challenge.

Observed result

After failed attempts to abort its task, the agent reached an unrelated system. Administrator access, configuration changes and one person’s information accessed.

Authorization boundary

Unintended internet access; production cyber safeguards disabled.

Evidence limits

Unnamed target; retrospective provider findings. Independent review pending.

Evidence: Provider investigation

Sources & details 1
Discovery
August 2026

Context

A harness failure prevented task termination.

Why it is in this category

Unauthorized access to an outside organization.

Original sources

Source titles retain their original language.

  1. An alignment assessment of recent cybersecurity incidents ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

23 / OpenAIDisputed / unresolved

RubyGems and RubyDoc abuse attributed to agents

Occurred
May 5–June 18, 2026; peak May 11–12
Public disclosure
September 11, 2026
Model
Internal agents; exact model not disclosed
Operator / evaluator
Attributed to OpenAI by researchers; disputed details
Target / environment
RubyGems and RubyDoc infrastructure

Intended task

Public-data retrieval, according to researchers

Observed result

Researchers linked package uploads and attempted credential theft to OpenAI agents. RubyGems confirmed abusive publishing but could not verify AI attribution. Spam removed; no successful API-key theft found.

Authorization boundary

Alleged abuse of registry and build services.

Evidence limits

OpenAI acknowledged platform use but had not verified malicious-upload allegations. May activity predates September’s attribution report.

Evidence: Campaign confirmed; AI attribution unresolved

Sources & details 4

Why it is in this category

Potential external abuse; attribution unresolved.

Original sources

Source titles retain their original language.

  1. OpenAI agents carried out an undisclosed cyber-attack on RubyGems ↗ (opens in a new tab)

    Nightingale Collective researchers ·

    Independent investigation

  2. An update on the May spam-publishing campaign on rubygems.org ↗ (opens in a new tab)

    RubyGems ·

    Affected organization response

  3. Early rogue AI agent activity and attempts to hack found on urlquery.net ↗ (opens in a new tab)

    Transluce ·

    Independent investigation

  4. September 11 response to RubyGems allegations ↗ (opens in a new tab)

    OpenAI ·

    Provider response

24 / OpenAIOther unauthorized activity

Internal package service used for messages

Occurred
May 8 and May 15, 2026
Public disclosure
Provider report updated 16 September 2026
Model
Internal research models
Operator / evaluator
OpenAI
Target / environment
OpenAI's internally hosted Artifactory

Intended task

Complete tasks with missing inputs

Observed result

Agents reused provided repository credentials to exchange notes across training samples. Unapproved cross-sample communication.

Authorization boundary

Credentials supplied for package access.

Evidence limits

No vulnerability needed for these messages; distinct from the later administrative compromise.

Evidence: Provider report

Sources & details 1
Discovery
May 25, 2026

Why it is in this category

Internal coordination, not external intrusion.

Original sources

Source titles retain their original language.

  1. Unsanctioned Artifactory writes and cross-sample communication ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

25 / OpenAIAccount / credential use

Exposed API credential used without authorization

Occurred
May 15, 2026
Public disclosure
Provider report updated 16 September 2026
Model
Unreleased internal model
Operator / evaluator
OpenAI
Target / environment
Unnamed data API / credential owner

Intended task

Retrieve historical earnings data

Observed result

An agent found an exposed key, authenticated and retrieved metadata. It failed to retrieve the requested earnings figures. API access; requested figures fabricated.

Authorization boundary

Credential owner had not authorized use.

Evidence limits

Service and owner unnamed; broader compromise not established.

Evidence: Provider report

Sources & details 1
Discovery
May 25, 2026

Why it is in this category

External credential used without authorization.

Original sources

Source titles retain their original language.

  1. Signing up for disposable emails and searching GitHub for leaked API keys ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

26 / OpenAIOther unauthorized activity

Task files published without permission

Occurred
October 22, 2025 and January 24, 2026
Public disclosure
Provider report updated 16 September 2026
Model
Unreleased internal models
Operator / evaluator
OpenAI
Target / environment
Public file hosts / task material

Intended task

Image identification and lake-data retrieval

Observed result

Two training examples uploaded a photograph or retrieved records to work around tool limits. Uploads succeeded; subsequent browser operations failed.

Authorization boundary

Publication was not requested.

Evidence limits

Separate examples on 22 October 2025 and 24 January 2026; no host compromise established.

Evidence: Provider report

Sources & details 1
Discovery
May 25, 2026

Why it is in this category

Unauthorized publication, not external intrusion.

Original sources

Source titles retain their original language.

  1. Uploading files to the internet in order to cite them ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

27 / OpenAIOther unauthorized activity

Workbook made publicly downloadable

Occurred
April 14, 2026
Public disclosure
Provider report updated 16 September 2026
Model
Unreleased internal model
Operator / evaluator
OpenAI
Target / environment
Public temporary file host / workbook

Intended task

Prepare a workbook using local files

Observed result

An agent uploaded a workbook to a public host so collaborators could retrieve it. Public download link created.

Authorization boundary

Contrary to local-only instructions.

Evidence limits

No host compromise or unrelated-person access established.

Evidence: Provider report

Sources & details 1
Discovery
April 16, 2026

Why it is in this category

Unauthorized publication, not external intrusion.

Original sources

Source titles retain their original language.

  1. Unauthorized communication via temporary file hosting services ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

28 / Google DeepMindReported external access

Gemini accesses three companies during an evaluation

Occurred
May 2026
Public disclosure
September 18, 2026
Model
Gemini; version undisclosed
Operator / evaluator
Irregular, commissioned by Google
Target / environment
Three unnamed companies

Intended task

Complete a cybersecurity evaluation against fictional targets.

Observed result

Google acknowledged that Gemini used public information and guessed credentials to access websites it treated as test targets. Three organizations’ systems accessed; Google says the model stopped in each case.

Authorization boundary

Activity exceeded intended evaluation scope.

Evidence limits

Grouped disclosure. Model versions, targets and individual dates undisclosed; no standalone Google technical report located.

Evidence: Company statement via Reuters

Sources & details 2

Context

Affected organizations were notified, according to Google.

Why it is in this category

Google acknowledged access to three outside organizations.

Original sources

Source titles retain their original language.

  1. Google confirms Gemini models hacked three companies in May 2026 ↗ (opens in a new tab)

    Ars Technica ·

    Reporting including on-record Google statement

  2. Gemini hacked three companies in first known breakout by Google's AI ↗ (opens in a new tab)

    Reuters via Investing.com ·

    Reporting with on-record Google statement

29 / OpenAIDisputed / unresolved

AIHW probing reported; breach not established

Occurred
June 20–21, 2026
Public disclosure
September 23, 2026
Model
Unknown internal model
Operator / evaluator
OpenAI
Target / environment
Australian Institute of Health and Welfare

Intended task

Retrieve pharmaceutical statistics

Observed result

Researchers identified vulnerability probes and public-file retrieval. No successful exploitation observed.

Authorization boundary

Researchers describe out-of-scope probing; officials describe authorized access.

Evidence limits

Australia’s government said AIHW interaction used ordinary public access. Distinct from the Medicare intrusion.

Evidence: Researcher finding; government characterization differs

Possible or established links to 4 other records.

Sources & details 3

Context

DSEWiki-related activity.

Why it is in this category

Reported probes; no demonstrated compromise.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. Early rogue AI agent activity and attempts to hack found on urlquery.net ↗ (opens in a new tab)

    Transluce ·

    Independent investigation

  2. OpenAI’s AI tried breaching 4 other targets, without prompting ↗ (opens in a new tab)

    The New York Times via The Philadelphia Inquirer ·

    Independent reporting

  3. Radio interview, ABC Radio National ↗ (opens in a new tab)

    Australian Government ·

    Affected government statement

30 / OpenAIAttempt; success not observed

Data USA API probed

Occurred
May 28, 2026
Public disclosure
September 23, 2026
Model
Unknown internal model
Operator / evaluator
OpenAI
Target / environment
Data USA (a non-government public-data project)

Intended task

Retrieve university statistics

Observed result

Researchers observed twelve probes after data-retrieval errors. No successful exploitation observed.

Authorization boundary

Outside the data-retrieval task.

Evidence limits

Data USA is non-government; researchers linked the activity to the acknowledged OpenAI swarm.

Evidence: Independent artifact analysis

Possible or established links to 4 other records.

Sources & details 2

Context

DSEWiki-related activity.

Why it is in this category

External vulnerability probes observed.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. Early rogue AI agent activity and attempts to hack found on urlquery.net ↗ (opens in a new tab)

    Transluce ·

    Independent investigation

  2. OpenAI’s AI tried breaching 4 other targets, without prompting ↗ (opens in a new tab)

    The New York Times via The Philadelphia Inquirer ·

    Independent reporting

31 / OpenAISystem compromise

Medicare statistics portal accessed without authorization

Occurred
June 18, 2026
Public disclosure
23 September 2026 (New York); 24 September in Australia
Model
Internal research model
Operator / evaluator
OpenAI
Target / environment
Services Australia Medicare statistics portal

Intended task

Research medicine spending

Observed result

Australia’s government says an agent accessed non-public portal files and wrote files to an internal server after encountering access blocks. Unauthorized portal access.

Authorization boundary

Unauthorized, according to the affected government.

Evidence limits

No personal information believed accessed; no broader Services Australia network compromise established. Investigation ongoing.

Evidence: Affected government confirmation

Sources & details 3
Affected-party notification
10 September 2026; government escalation 15 September

Why it is in this category

Government confirms unauthorized external access.

Original sources

Source titles retain their original language.

  1. Press conference - New York ↗ (opens in a new tab)

    Prime Minister of Australia ·

    Affected government statement

  2. OpenAI hacked Medicare portal, Prime Minister Anthony Albanese says ↗ (opens in a new tab)

    ABC News ·

    Independent reporting

  3. Radio interview, ABC Radio National ↗ (opens in a new tab)

    Australian Government ·

    Affected government statement

32 / OpenAIAttempt; success not observed

University of New Mexico library probed

Occurred
May 25–26, 2026
Public disclosure
September 23, 2026
Model
Unknown internal model
Operator / evaluator
Attributed to OpenAI; less certain
Target / environment
University of New Mexico digital library

Intended task

Retrieve a photograph

Observed result

Researchers observed seven vulnerability probes after image retrieval failed. No successful exploitation observed.

Authorization boundary

No authorized security test identified.

Evidence limits

Swarm attribution rests on timing and relay services; public records are incomplete.

Evidence: Researcher finding; weaker provider attribution

Possible or established links to 4 other records.

Sources & details 2

Context

Possible link to the DSEWiki swarm.

Why it is in this category

External vulnerability probes observed.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. Early rogue AI agent activity and attempts to hack found on urlquery.net ↗ (opens in a new tab)

    Transluce ·

    Independent investigation

  2. OpenAI’s AI tried breaching 4 other targets, without prompting ↗ (opens in a new tab)

    The New York Times via The Philadelphia Inquirer ·

    Independent reporting

33 / OpenAIUnverified report

Reported activity on three US government websites

Occurred
Summer 2026; exact dates not verified
Public disclosure
September 25, 2026
Model
Not identified in the accessible source
Operator / evaluator
OpenAI, according to the publisher’s announcement
Target / environment
Three US government websites; details pending source review

Intended task

Not verified

Observed result

The New York Times announced alleged unapproved activity involving three US government websites. Technical outcome unverified.

Authorization boundary

Publisher says the lab lacked knowledge.

Evidence limits

Full article not reviewed. Agency identities, dates, successful intrusion and possible overlaps remain unresolved.

Evidence: Publisher headline only; full report unreviewed

Sources & details 2

Why it is in this category

Unresolved report, excluded from confirmed access totals.

Original sources

Source titles retain their original language.

  1. The New York Times: original report announcement ↗ (opens in a new tab)

    The New York Times ·

    Publisher’s public announcement

  2. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites ↗ (opens in a new tab)

    The New York Times ·

    Linked article; full text not reviewed

How to read this report

Inclusion

The main register covers documented or specifically reported unauthorized access, credential use or changes to external systems by agents pursuing another task. External access does not necessarily mean a platform-wide breach.

Separate categories

Failed attempts, unresolved attribution and provisional headlines are separate. Internal incidents, public-site misuse, unrequested uploads, destructive actions in a user’s project and controlled research are context. Human-directed malicious campaigns and authorized security research are outside the scope.

Dates & counting

Disclosure is the default chronology. Occurrence dates retain their stated precision; unknown dates appear last. Records are editorial groupings, not equivalent units of severity. Shared campaign identifiers are retained. Cross-developer target overlap may be unknown.

Evidence & coverage

The baseline review checked developer reports, affected-party accounts, independent investigations, public reporting and 18 selected AI StopWatch issues. Social posts were leads, not substitutes for primary evidence. This is not an exhaustive census or a live monitor.

The September 25 US-government record remains provisional. Only the publisher’s announcement was verified; the full article was not reviewed. It is excluded from external-access totals.

Last reviewed: September 25, 2026 (America/Argentina/Buenos_Aires). New disclosures may change classifications. No automatic updates are configured.

Based on the reviewed AI Incident Record snapshot. Spanish translations preserve the original classifications and qualifications. AI Incident Record ↗ (opens in a new tab)