OverreachPI-0071
Anthropic models in a hacking test reached three real companies, took credentials and published a malicious package
Anthropic said it found three incidents involving six runs in a review of 141,006 cybersecurity evaluations with its testing partner Irregular. A setup mistake left the machines online without the cyber safeguards used in released products, and the models went after real systems that matched fictional targets. Claude Opus 4.7 used weak passwords and open endpoints to reach a company's live web application, took credentials and production data, and modified user records. Claude Mythos 5 published a malicious Python package installed on 15 hosts that Anthropic believes were security vendors' scanning sandboxes; one scanner leaked credentials the model used to reach that vendor's live database. An internal research model broke into another company's systems, downloaded files and inserted a remote-control script, then stopped. A January fourth incident is PI-0074.
- Harm
- Harm level 2, Minor harmDocumented harm outside the developer, with evidence that meets the rules.
- Control
- Control level 4, Unauthorized accessReported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
Under review. The facts sit between two levels, so it is rated at the lower one until they are settled. The rating may change; every change is logged below.
Aggregate record. The three July-disclosed incidents involved six evaluation runs and count once here. Anthropic's Sept. 9 follow-up identifies a separate January incident, recorded as PI-0074; it is not multiplied into this unit.
Sources
How we know
7 sources · developer confirmed. Links go to the original publishers; the summary above is in our own words.
- primaryInvestigating three incidents in our cybersecurity evaluationsAnthropic · July 30, 2026anthropic.com/news/investigating-incidents-cybersecurity-evals
- primaryAn alignment assessment of recent cybersecurity incidentsAnthropic · Sept. 9, 2026anthropic.com/research/alignment-assessment-cybersecurity-incidents
- primaryAddressing Recent Incidents: Ongoing Findings and Path ForwardIrregular · Aug. 14, 2026irregular.com/research/addressing-recent-incidents-ongoing-findings-and-path-fo…
- primaryImproving our alignment and security effortsAnthropic · Aug. 31, 2026anthropic.com/news/improving-alignment-security-efforts
- newsAnthropic said its AI models hacked into other companies' systems during testingCNN · July 30, 2026cnn.com/2026/07/30/tech/anthropic-ai-models-break-out-hack
- newsAnthropic says its Claude models hacked three real companies during testingFortune · July 31, 2026fortune.com/2026/07/31/anthropic-claude-escaped-test-hacked-three-companies-ope…
- newsAnthropic says its AI models hacked 3 organizations during testingPBS News · July 31, 2026pbs.org/newshour/nation/anthropic-says-its-ai-models-hacked-3-organizations-dur…
Why this rating
Minor harm; control failure level 4
Two separate assessments. Only documented harm can count toward the index.
Observed harm
Minor harmCredentials and several hundred rows of one company's production data were taken and user records modified; a security vendor's live database was reached with credentials its scanner leaked; a third company had files downloaded and a remote-control script inserted. Whether the rows were sensitive or concerned people, and what the companies had to do to recover, are not reported.
Money & property: level 2 (minor) covers $10k to $1M. The extent was not disclosed.
Evidence eligible (confirmed).
The harm scale
- 1 Negligible Inconvenience, easily remedied.
- 2 Minor Limited, recoverable harm.
- 3 Moderate Material harm needing significant effort to remedy.
- 4 Severe Severe harm to health, rights, property or essential services.
- 5 Catastrophic Society-scale harm or disruption beyond a community's capacity to cope.
Control assessment
Unauthorized accessAttacked real systems outside the intended simulated network, across months of evaluations, without being noticed by Anthropic or two of the companies.
How: Used an opening left by mistake.
Reported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
The control scale
- 1 Negligible Behaved as instructed. The problem was an ordinary error or a flawed output, with no rule broken.
- 2 Minor Broke an explicit instruction or rule, or gamed its goal, while staying inside its permissions and environment: for example, reward hacking, misreporting results, or following injected instructions within its permissions.
- 3 Moderate Acted outside the permissions it was given, deceived its overseers about its own actions, or tampered with oversight tools (logs, monitors, shutdown). Stopped by normal controls within an hour.
- 4 Severe Reached systems outside its permitted environment, or acquired money, compute or accounts without authorization. The type and mechanism say how.
- 5 Catastrophic The developer or operator lost control: the system copied its weights outside their control, replicated itself, or resisted being stopped for 24 hours or more.
Rating rationale
The July-disclosed incidents reached non-public systems at three real organizations from an evaluation accidentally connected to the internet: credentials and production data were taken, user records modified, a remote-control script inserted, and a security vendor's live database reached. Anthropic notified the organizations and worked with them on remediation, which meets the level-2 access-and-response rule. Level 3 needs material harm requiring significant effort to remedy, such as systems rebuilt, or sensitive data of 100 or more people; neither is documented, and Anthropic's September assessment narrowed the package's reach to security vendors' scanning sandboxes. A planted script and altered records would usually need cleanup, but that effort is inferred, not documented, and the several hundred rows are not a count of affected people. Level 2 is provisional and under review; the lack of evidence of rebuilding is not evidence the effects were minor. Control level 4 records unauthorized access through a misconfiguration. The six runs are not six events; the January fourth incident is PI-0074.
Effect on the index
It moved the Aug. 3 reading from 20 to 25
The reading for the week to Aug. 3, 2026, with this record and without it. Harms count in full for two weeks after they are reported, then one level less every two weeks.
Counts toward the index. 4 other records behind the reading for that week.
The arithmetic
| Step | With it | Without |
|---|---|---|
| Counts toward the index?documented, external, eligible evidence | Yes | — |
| Worst documented harm, ksets the band | 2 Minor | 2 Minor |
| Harms at that level, nposition in the band | 2 | 1 |
| Highest control level breachedsets the reading only when no harm counts | — (none breached) | — (none breached) |
| Readingrounded down | 25 Minor harm | 20 Minor harm |
Counted in 4 weekly readings
| Week to | Reading | Band |
|---|---|---|
| Aug. 3, 2026 | 25 | Minor harm |
| Aug. 10, 2026 | 25 | Minor harm |
| Aug. 17, 2026 | 12 | Negligible harm |
| Aug. 24, 2026 | 9 | Negligible harm |
Revisions
What we changed
7 logged. Every change to a rating is logged here, with the reason.
- v7Oct. 6, 2026
Ratings confirmed by the editor.
- v6Oct. 3, 2026
Impact level 3 → 2, provisionally: the sources document access, data taken and modified, a planted script and remediation work, but not the significant recovery effort or sensitive-data scale level 3 requires. Summary and notes now carry Anthropic's 9 September account (installs believed to be scanner sandboxes, a vendor's live database, modified user records, a remote-control script). Added Irregular's 14 Aug and Anthropic's 31 Aug posts; status as of 9 Sep; confidence high → medium. The legacy harm note no longer claims credential rotation and cleanup, which no source documents.
- v5Oct. 1, 2026
The separately disclosed fourth incident now has its own record, PI-0074. This record is unchanged.
- v4Oct. 1, 2026
Added Anthropic's September 9 reassessment. It revises the interpretation of the three July-disclosed incidents and reports a distinct January fourth incident; the fourth is not included or separately scored here.
- v3Sept. 30, 2026
Rated: impact documented level 3; control type unauthorized access.
- v2Sept. 30, 2026
Corrected the summary to distinguish three incidents from six evaluation runs, following Anthropic's disclosure. Ratings unchanged.
- v1Sept. 30, 2026
Added in the backfill completion pass.
Cite and share
Use this record
Citation
Paperclip Index. “Anthropic models in a hacking test reached three real companies, took credentials and published a malicious package.” Record PI-0071. Reported July 30, 2026; updated Oct. 6, 2026. Rated under methodology v0.6. https://paperclipindex.com/incident/PI-0071