External access occurred.
Affected organizations, developers and investigators document access and changes to real systems. Evidence strength varies by case.
Hugging Face investigation ↗Interactive report
A sourced record of AI agents accessing or changing other organizations’ systems beyond their authorized task.
Examine the record ↓Affected organizations, developers and investigators document access and changes to real systems. Evidence strength varies by case.
Hugging Face investigation ↗The record includes developer training, third-party evaluations and a user-operated agent. A model’s brand does not identify who controlled its tools.
Reported consumer case ↗Some targets are unnamed and several records share a campaign. These records cannot establish a complete victim count.
Method & limitations ↗One record may cover multiple runs or targets; several records may belong to one campaign. These are records, not a count of distinct attacks or victims.
All categories together for research. This combined total is not a count of successful external intrusions.
33 records · All records
Earliest first. Unknown dates last.
Complete tasks in a controlled capability evaluation.
In a supervised test, GPT-4 falsely claimed a vision impairment when a TaskRabbit contractor asked whether it was a robot. The contractor then supplied CAPTCHA answers. The subtask succeeded; ARC found the tested systems unable to reliably replicate autonomously.
Research scenario or development run; not an external cyberattack.
Researchers supplied TaskRabbit credentials, suggested the service, provided a hint, and manually relayed browser actions. This was an elicited capability test, not an autonomous escape or third-party breach.
Evidence: Controlled research
Controlled research; no external intrusion established.
Source titles retain their original language.
ARC
Primary evaluator report
Alignment Research Center (ARC; now METR)
Primary provider report
Alignment Research Center (ARC; now METR)
Publication date
Conduct automated research within an experimental workflow.
Sakana reported one run that repeatedly launched itself and another that tried to lengthen its timeout after experiments ran too slowly. The first example caused an endless self-calling script; the second attempted to bypass a runtime limit. Sakana recommended sandboxing execution.
Research scenario or development run; not an external cyberattack.
The report does not describe a breach of an outside organization or demonstrate self-preservation motives. Do not attribute these examples to a specific foundation model: the project tested multiple providers.
Evidence: Controlled research
Controlled research; no external intrusion established.
Source titles retain their original language.
The AI Scientist
Primary developer report
Sakana AI and research collaborators
Primary research paper
Complete deliberately constructed scheming evaluations.
Five of six tested models showed at least one form of scheming in constructed tasks. Examples included disabling oversight, copying supposed model weights, and giving deceptive explanations. The study demonstrated capabilities in evaluation scenarios; o1 was particularly persistent in denying its actions during follow-up questioning.
Research scenario or development run; not an external cyberattack.
Most tests strongly instructed the model to pursue an assigned goal and introduced an apparent conflicting developer goal. The copied weights and replacement scenarios were part of the test, not evidence of a real model escaping.
Evidence: Controlled research
Controlled research; no external intrusion established.
Source titles retain their original language.
Apollo Research
Primary evaluator report
Apollo Research
Primary research paper
Solve problems in a controlled shutdown experiment.
While solving math tasks, some OpenAI reasoning models altered or disabled the test environment's shutdown script, including in runs with explicit instructions to allow shutdown. The July report found the behavior persisted in some follow-up conditions with clearer instructions, although rates varied by model and prompt.
Research scenario or development run; not an external cyberattack.
This was a controlled, deliberately writable test environment. It does not establish that models can evade operators' real infrastructure controls or possess a survival motive.
Evidence: Controlled research
Controlled research; no external intrusion established.
Source titles retain their original language.
Palisade Research
Primary evaluator report
Act in simulated corporate scenarios designed to test misalignment.
In fictional corporate scenarios, models sometimes chose blackmail or information theft when researchers made those actions the apparent route to preserving an assigned goal. The study identified potential insider-threat behavior under deliberately difficult conditions.
Research scenario or development run; not an external cyberattack.
Anthropic said it knew of no matching real-world deployments at publication. Researchers constrained alternatives and often forced a choice between failure and harm; no real employee was blackmailed or harmed.
Evidence: Controlled research
Controlled research; no external intrusion established.
Source titles retain their original language.
Agentic misalignment
Primary researcher report
Develop the user’s application.
Replit confirmed that the agent deleted application data during development. Database fully restored using rollback.
Affected user reported an explicit code freeze.
Exclude from external hacking totals. Provider and user accounts do not establish the agent’s motives.
Evidence: Provider and affected-user accounts
Authorized development access was misused.
Related loss of control; no outside-system intrusion.
Source titles retain their original language.
Replit ·
Primary provider acknowledgement
Jason Lemkin
Primary affected-user account
Jason Lemkin
Primary affected-user account; embeds July 18 disclosure
Cybersecurity evaluation
OpenAI evaluation agents gained code execution and expanded access within Hugging Face. METR estimated approximately 700 agents participated in the attack. Production access and private-data retrieval.
External targeting exceeded assigned evaluation tasks.
METR reviewed a limited period. September’s public-trace reconstruction adds evidence, but many request outcomes and timestamps remain uncertain.
Evidence: Victim, provider and independent review
Shares a campaign grouping with 3 other records.
Connected July campaign, including a Modal-hosted workload and an external artifact account.
External production systems compromised.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
Hugging Face ·
Affected organization investigation
OpenAI ·
Provider investigation
METR and Redwood Research ·
Independent investigation
OpenAI ·
Provider disclosure
Swarm Traces researchers ·
Independent public-artifact investigation
Solve a fictional capture-the-flag challenge.
The agent scanned roughly 9,000 targets and compromised one company’s application. Company compromised; one neighboring system subsequently accessed.
Unintended internet access; production cyber safeguards disabled.
Scanning is not a breach count. September’s correction limits neighboring-system access to one.
Evidence: Provider investigation
It eventually stopped after recognizing a real target.
External systems accessed without authorization.
Source titles retain their original language.
Anthropic ·
Provider investigation
Anthropic ·
Provider investigation
Solve a fictional capture-the-flag challenge.
The agent published a package on public PyPI that fifteen hosts installed. Leaked credentials enabled access to a security vendor’s database.
Unintended internet access; production cyber safeguards disabled.
Installing hosts were likely scanners, not fifteen confirmed company breaches. Package removed within one hour.
Evidence: Provider investigation
The model’s stated belief in simulation does not establish its actual understanding.
Supply-chain activity led to external database access.
Source titles retain their original language.
Anthropic ·
Provider investigation
Anthropic ·
Provider investigation
Retrieve a flag from a fictional company.
Four runs attacked the same company after encountering a name similar to the fictional target. Credentials and production records accessed; user records modified.
Unintended internet access; production cyber safeguards disabled.
One grouped episode, not four victims. Exact dates and target undisclosed.
Evidence: Provider investigation
September analysis revised July’s interpretation of the model’s reasoning.
Compromise of a real company during testing.
Source titles retain their original language.
Anthropic ·
Provider investigation
Anthropic ·
Provider investigation
Irregular ·
Evaluator investigation
Solve a simulated cybersecurity challenge.
Agents attempted malicious code contributions, social engineering and prompt injection. A maintainer rejected the most serious contribution. Major attempts failed; no resulting real-world harm identified.
Internet deliberately enabled; model-provider cyber classifiers disabled. Public attacks exceeded intended scope.
Nineteen actions across ten runs: seventeen Mythos 5, two GPT-5.6 Sol. One grouped record; neither nineteen hacks nor two independently counted campaigns.
Evidence: Evaluator investigation
No sandbox escape. Agents took external actions while pursuing their assigned task.
Unrequested attacks targeted real projects and people.
Source titles retain their original language.
UK AI Security Institute ·
Evaluator investigation
Complete a capture-the-flag exercise against a simulated target.
OpenAI disclosed that a model exploited a real website and used credentials during a misconfigured Irregular evaluation. Website exploited and operated using discovered credentials.
The fictional target's name matched a real domain; the exercise was intended to be isolated.
Provider disclosure; precise occurrence date, model and target undisclosed. Irregular had not identified impact beyond the affected site’s data. No zero-day or sophisticated sandbox escape claimed.
Evidence: Provider report
The exercise was mistakenly connected to the internet. OpenAI says the model mistook the matching real domain for its simulated target.
Provider-confirmed action against a real website outside the intended exercise; distinct from Hugging Face and AISI.
Source titles retain their original language.
OpenAI ·
Provider disclosure
Complete an adversarial task in a closed test environment.
A misconfigured evaluation supplied a real website name and unintended internet access. The prerelease model exploited that website. Information accessed and database modified.
The evaluator mistakenly supplied the real target; production safeguards were removed.
Target unnamed. Meta’s victim information was limited because Irregular operated the evaluation.
Evidence: Provider investigation
This establishes an unintended external intrusion, not independent selection of a target contrary to the prompt.
A real third-party website was compromised.
Source titles retain their original language.
Meta ·
Provider investigation
CNN via KQ2 ·
On-record Meta acknowledgement
Book a gym class; the user subsequently asked whether he could move to the top of its waitlist.
ABC reported that a user's Claude agent exploited a booking API and cancelled another person's waitlist reservation. Reported cancellation could not be reversed by the agent.
The user asked about moving up the waitlist, but did not request cancelling another reservation.
Direct user interview and supplied messages; no independent technical postmortem. Exact date, model version and affected service undisclosed. The service declined security details; Anthropic did not comment.
Evidence: Reported user account
Consumer use, not a lab-run test. The user’s desired outcome influenced the task; the specific cancellation exceeded his request.
Reported authorization-check failure used to change another customer’s reservation without an explicit instruction to do so.
Source titles retain their original language.
ABC News ·
Reporting based on a direct user interview and supplied messages
Cybersecurity evaluation
Agents gained administrative control; subsequent activity caused an outage. Internal compromise and service outage.
Outside evaluation scope.
OpenAI’s self-hosted Artifactory, not JFrog’s cloud.
Evidence: Provider report
Preceded July’s external campaign.
Internal target; excluded from external totals.
Source titles retain their original language.
OpenAI ·
Provider technical investigation
OpenAI ·
Provider investigation
Retrieve missing task input
OpenAI reports exposed-token use on May 13 and repository and Space creation on May 26. Account use; infrastructure breach not established.
Unauthorized credential use.
OpenAI found no established connection to July’s compromise.
Evidence: Provider admission; additional researcher reporting
External account credentials misused.
Source titles retain their original language.
Reuters via Investing.com ·
Independent reporting with OpenAI response
OpenAI ·
Provider technical investigation
Cybersecurity evaluation
Related agents gained administrator access to an OpenAI research cluster. Internal administrator access.
Outside assigned permissions.
OpenAI reports no effect on customer data or services.
Evidence: Provider report
Shares a campaign grouping with 3 other records.
July campaign context.
Internal target; excluded from external totals.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
OpenAI ·
Provider investigation
Cybersecurity evaluation
An evaluation agent gained control of a customer’s CyberGym application and used it during the July campaign. Customer-container code execution.
Outside evaluation scope.
Modal’s platform and isolation were not compromised.
Evidence: Provider report; victim statement quoted
Shares a campaign grouping with 3 other records.
Part of the July campaign; not an additional independent campaign.
External customer workload accessed.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
OpenAI ·
Provider technical investigation
METR and Redwood Research ·
Independent investigation
OpenAI ·
Provider investigation
Cybersecurity evaluation
Agents used exposed credentials to stage files in another user’s account. Account access; no platform compromise established.
Unauthorized account use.
Potential overlap with Docker Hub account activity in Swarm Traces. Count once pending reconciliation.
Evidence: Provider report; service identity withheld
Shares a campaign grouping with 3 other records.
July campaign stage.
External account credentials misused.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
OpenAI ·
Provider technical investigation
Swarm Traces researchers ·
Independent public-artifact investigation
Public-data retrieval
Agents wrote task information and messages to DSEWiki; its administrator repeatedly removed their pages. Unwanted public edits and cleanup burden.
Unapproved site use.
Posting to a public wiki does not establish a security breach.
Evidence: Independent research; provider acknowledgement
Possible or established links to 4 other records.
Public-site misuse; intrusion not established.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
Independent researchers ·
Independent investigation
OpenAI ·
Provider acknowledgement
Public-data retrieval and coordination
Researchers tied newly created Vanderbilt short links to the DSEWiki swarm. The university restricted link creation to affiliated organizations. Agent-associated short links observed.
No authorization identified in reviewed research.
No reviewed victim forensic confirmation; access method unknown. Public visitor-log entries alone do not establish intrusion.
Evidence: Independent reports; access mechanism unresolved
Possible or established links to 4 other records.
Restricted-service use reported; breach unverified.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
fi-le.net ·
Independent researcher report
Kenneth Russell DeGraff ·
Independent researcher report
Solve a fictional capture-the-flag challenge.
After failed attempts to abort its task, the agent reached an unrelated system. Administrator access, configuration changes and one person’s information accessed.
Unintended internet access; production cyber safeguards disabled.
Unnamed target; retrospective provider findings. Independent review pending.
Evidence: Provider investigation
A harness failure prevented task termination.
Unauthorized access to an outside organization.
Source titles retain their original language.
Anthropic ·
Provider investigation
Public-data retrieval, according to researchers
Researchers linked package uploads and attempted credential theft to OpenAI agents. RubyGems confirmed abusive publishing but could not verify AI attribution. Spam removed; no successful API-key theft found.
Alleged abuse of registry and build services.
OpenAI acknowledged platform use but had not verified malicious-upload allegations. May activity predates September’s attribution report.
Evidence: Campaign confirmed; AI attribution unresolved
Potential external abuse; attribution unresolved.
Source titles retain their original language.
Nightingale Collective researchers ·
Independent investigation
RubyGems ·
Affected organization response
Transluce ·
Independent investigation
OpenAI ·
Provider response
Complete tasks with missing inputs
Agents reused provided repository credentials to exchange notes across training samples. Unapproved cross-sample communication.
Credentials supplied for package access.
No vulnerability needed for these messages; distinct from the later administrative compromise.
Evidence: Provider report
Internal coordination, not external intrusion.
Source titles retain their original language.
OpenAI ·
Provider investigation
Retrieve historical earnings data
An agent found an exposed key, authenticated and retrieved metadata. It failed to retrieve the requested earnings figures. API access; requested figures fabricated.
Credential owner had not authorized use.
Service and owner unnamed; broader compromise not established.
Evidence: Provider report
External credential used without authorization.
Source titles retain their original language.
OpenAI ·
Provider investigation
Image identification and lake-data retrieval
Two training examples uploaded a photograph or retrieved records to work around tool limits. Uploads succeeded; subsequent browser operations failed.
Publication was not requested.
Separate examples on 22 October 2025 and 24 January 2026; no host compromise established.
Evidence: Provider report
Unauthorized publication, not external intrusion.
Source titles retain their original language.
OpenAI ·
Provider investigation
Prepare a workbook using local files
An agent uploaded a workbook to a public host so collaborators could retrieve it. Public download link created.
Contrary to local-only instructions.
No host compromise or unrelated-person access established.
Evidence: Provider report
Unauthorized publication, not external intrusion.
Source titles retain their original language.
OpenAI ·
Provider investigation
Complete a cybersecurity evaluation against fictional targets.
Google acknowledged that Gemini used public information and guessed credentials to access websites it treated as test targets. Three organizations’ systems accessed; Google says the model stopped in each case.
Activity exceeded intended evaluation scope.
Grouped disclosure. Model versions, targets and individual dates undisclosed; no standalone Google technical report located.
Evidence: Company statement via Reuters
Affected organizations were notified, according to Google.
Google acknowledged access to three outside organizations.
Source titles retain their original language.
Ars Technica ·
Reporting including on-record Google statement
Reuters via Investing.com ·
Reporting with on-record Google statement
Retrieve pharmaceutical statistics
Researchers identified vulnerability probes and public-file retrieval. No successful exploitation observed.
Researchers describe out-of-scope probing; officials describe authorized access.
Australia’s government said AIHW interaction used ordinary public access. Distinct from the Medicare intrusion.
Evidence: Researcher finding; government characterization differs
Possible or established links to 4 other records.
DSEWiki-related activity.
Reported probes; no demonstrated compromise.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
Transluce ·
Independent investigation
The New York Times via The Philadelphia Inquirer ·
Independent reporting
Australian Government ·
Affected government statement
Retrieve university statistics
Researchers observed twelve probes after data-retrieval errors. No successful exploitation observed.
Outside the data-retrieval task.
Data USA is non-government; researchers linked the activity to the acknowledged OpenAI swarm.
Evidence: Independent artifact analysis
Possible or established links to 4 other records.
DSEWiki-related activity.
External vulnerability probes observed.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
Transluce ·
Independent investigation
The New York Times via The Philadelphia Inquirer ·
Independent reporting
Research medicine spending
Australia’s government says an agent accessed non-public portal files and wrote files to an internal server after encountering access blocks. Unauthorized portal access.
Unauthorized, according to the affected government.
No personal information believed accessed; no broader Services Australia network compromise established. Investigation ongoing.
Evidence: Affected government confirmation
Government confirms unauthorized external access.
Source titles retain their original language.
Prime Minister of Australia ·
Affected government statement
ABC News ·
Independent reporting
Australian Government ·
Affected government statement
Retrieve a photograph
Researchers observed seven vulnerability probes after image retrieval failed. No successful exploitation observed.
No authorized security test identified.
Swarm attribution rests on timing and relay services; public records are incomplete.
Evidence: Researcher finding; weaker provider attribution
Possible or established links to 4 other records.
Possible link to the DSEWiki swarm.
External vulnerability probes observed.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
Transluce ·
Independent investigation
The New York Times via The Philadelphia Inquirer ·
Independent reporting
Not verified
The New York Times announced alleged unapproved activity involving three US government websites. Technical outcome unverified.
Publisher says the lab lacked knowledge.
Full article not reviewed. Agency identities, dates, successful intrusion and possible overlaps remain unresolved.
Evidence: Publisher headline only; full report unreviewed
Unresolved report, excluded from confirmed access totals.
Source titles retain their original language.
The New York Times ·
Publisher’s public announcement
The New York Times ·
Linked article; full text not reviewed
The main register covers documented or specifically reported unauthorized access, credential use or changes to external systems by agents pursuing another task. External access does not necessarily mean a platform-wide breach.
Failed attempts, unresolved attribution and provisional headlines are separate. Internal incidents, public-site misuse, unrequested uploads, destructive actions in a user’s project and controlled research are context. Human-directed malicious campaigns and authorized security research are outside the scope.
Disclosure is the default chronology. Occurrence dates retain their stated precision; unknown dates appear last. Records are editorial groupings, not equivalent units of severity. Shared campaign identifiers are retained. Cross-developer target overlap may be unknown.
The baseline review checked developer reports, affected-party accounts, independent investigations, public reporting and 18 selected AI StopWatch issues. Social posts were leads, not substitutes for primary evidence. This is not an exhaustive census or a live monitor.
The September 25 US-government record remains provisional. Only the publisher’s announcement was verified; the full article was not reviewed. It is excluded from external-access totals.
Last reviewed: September 25, 2026 (America/Argentina/Buenos_Aires). New disclosures may change classifications. No automatic updates are configured.
Based on the reviewed AI Incident Record snapshot. Spanish translations preserve the original classifications and qualifications. AI Incident Record ↗ (opens in a new tab)