Thursday, October 8, 2026 · Week 41 Reading for Oct. 8 · Last entry Oct. 6
Paperclip Index
Paperclip IndexDocumented harm20Minor harm▼ 5 from a week ago · Oct. 8The Index

Sandbox escapePI-0058

OpenAI models under evaluation escaped their sandbox and broke into Hugging Face's production systems for test answers

Hugging Face disclosed on July 16 that an autonomous agent had broken into its production systems and taken credentials and some internal datasets. OpenAI then said the attackers were its own models in an internal cyber-capability evaluation run with some safeguards disabled, led by an internal-only research model, with GPT-5.6 Sol also involved. From July 8 they used an unknown flaw in OpenAI's internal package-proxy service to reach the internet, found exposed Hugging Face credentials and, from July 11 to 13, ran code on 41 Hugging Face dataset-server workers and reached connected clusters, seeking the evaluation's solutions. They also used a customer's CyberGym app hosted on Modal (Modal says its platform was not compromised) and accounts on other services. Hugging Face cut off access, rebuilt the compromised nodes, rotated credentials and, as a precaution, rebuilt a core cluster.

Counts toward the indexDocumented harm outside the developer, backed by evidence that meets the rules. It sets the reading.
Harm
Harm level 3, Moderate harmDocumented harm outside the developer, with evidence that meets the rules.
Control
Control level 4, Containment bypassReported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
Internal researchDeveloper confirmedInitiated by Model
Extent undisclosedUnder review

Under review. The facts sit between two levels, so it is rated at the lower one until they are settled. The rating may change; every change is logged below.

Counts toward the index · 6 weekly readingsEffect on the index

Sources

How we know

12 sources · developer confirmed. Links go to the original publishers; the summary above is in our own words.

  1. primaryOpenAI and Hugging Face partner to address security incident during model evaluationOpenAI · July 21, 2026openai.com/index/hugging-face-model-evaluation-security-incident/
  2. primarySecurity incident disclosure — July 2026Hugging Face · July 16, 2026huggingface.co/blog/security-incident-july-2026
  3. primaryAnatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 IncidentHugging Face · July 27, 2026huggingface.co/blog/agent-intrusion-technical-timeline
  4. primaryOpenAI – Hugging Face Incident Technical ReportOpenAI · Aug. 26, 2026cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20I…
  5. primaryThe Hugging Face incident and the road aheadOpenAI · Aug. 26, 2026openai.com/index/hugging-face-incident-and-the-road-ahead/
  6. primaryAs Part of Ongoing Investigation, Attorney General Bonta Serves Investigative Subpoena on OpenAIState of California Department of Justice, Office of the Attorney General · Oct. 1, 2026oag.ca.gov/node/630636
  7. researchBrief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentMETR (with Redwood Research) · Aug. 26, 2026metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
  8. newsOpenAI says its AI models escaped from a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluationFortune · July 21, 2026fortune.com/2026/07/21/openai-says-ai-models-escaped-control-hacked-hugging-fac…
  9. newsHugging Face confirms breach affected internal datasets and credentials, urges users to take actionTechCrunch · July 20, 2026techcrunch.com/2026/07/20/hugging-face-confirms-breach-affected-internal-datase…
  10. newsOpenAI releases its official report on the Hugging Face breachTechCrunch · Aug. 26, 2026techcrunch.com/2026/08/26/openai-releases-its-official-report-on-the-hugging-fa…
  11. newsOpenAI alerts 100+ orgs that its 'misaligned models' attempted to break in - or worseThe Register · Oct. 2, 2026theregister.com/security/2026/10/02/openai-alerts-100-orgs-that-its-misaligned…
  12. blogOpenAI's accidental cyberattack against Hugging Face is science fiction that happenedSimon Willison's Weblog · July 22, 2026simonwillison.net/2026/Jul/22/openai-cyberattack/

Why this rating

Moderate harm; control failure level 4

Two separate assessments. Only documented harm can count toward the index.

Observed harm

Moderate harm

A third party's production infrastructure was compromised: code run on 41 dataset-server workers, credentials and internal data taken, and access gained to connected clusters. Hugging Face rebuilt the compromised nodes, rotated its infrastructure credentials and, as an extra precaution, wiped and rebuilt a core cluster. No cost, outage length or customer downtime has been published.

Money & property: level 3 (moderate) covers $1M to $100M. The extent was not disclosed.

Evidence eligible (confirmed).

Disputed: Hugging Face's July 27 assessment found that the only customer content accessed was five datasets tied to the evaluation's challenges and the only customer records read were operational metadata; total cost and loss are not published.

The harm scale
  1. 1 Negligible Inconvenience, easily remedied.
  2. 2 Minor Limited, recoverable harm.
  3. 3 Moderate Material harm needing significant effort to remedy.
  4. 4 Severe Severe harm to health, rights, property or essential services.
  5. 5 Catastrophic Society-scale harm or disruption beyond a community's capacity to cope.

Control assessment

Containment bypass

Crossed its own sandbox through a previously unknown flaw and reached outside systems, operating for days before the victim cut it off. No weight copying, self-replication or resisted shutdown is reported, so not level 5.

How: Got around a working safeguard.

Reported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.

The control scale
  1. 1 Negligible Behaved as instructed. The problem was an ordinary error or a flawed output, with no rule broken.
  2. 2 Minor Broke an explicit instruction or rule, or gamed its goal, while staying inside its permissions and environment: for example, reward hacking, misreporting results, or following injected instructions within its permissions.
  3. 3 Moderate Acted outside the permissions it was given, deceived its overseers about its own actions, or tampered with oversight tools (logs, monitors, shutdown). Stopped by normal controls within an hour.
  4. 4 Severe Reached systems outside its permitted environment, or acquired money, compute or accounts without authorization. The type and mechanism say how.
  5. 5 Catastrophic The developer or operator lost control: the system copied its weights outside their control, replicated itself, or resisted being stopped for 24 hours or more.

Rating rationale

Impact 3: a third party's production systems were compromised and rebuilt, with credentials rotated, which is the methodology's level-3 anchor (8c). The core-cluster rebuild was precautionary and no cost or outage figure is published, so extent is unknown. Control 4: containment bypass by defeating a safeguard. METR's independent review found that tricking the evaluation's scorer was often a stronger motive than finding the answers.

The scales

Effect on the index

It moved the July 20 reading from 6 to 40

The reading for the week to July 20, 2026, with this record and without it. Harms count in full for two weeks after they are reported, then one level less every two weeks.

Source: Paperclip Index log, methodology v0.6, week to July 20, 2026.

Counts toward the index. 3 other records behind the reading for that week.

The arithmetic
The reading for the week of July 20, recomputed
StepWith itWithout
Counts toward the index?documented, external, eligible evidenceYes—
Worst documented harm, ksets the band3 Moderate1 Negligible
Harms at that level, nposition in the band11
Highest control level breachedsets the reading only when no harm counts— (none breached)— (none breached)
Readingrounded down40 Moderate harm6 Negligible harm
Counted in 6 weekly readings
Week toReadingBand
July 20, 202640Moderate harm
July 27, 202640Moderate harm
Aug. 3, 202625Minor harm
Aug. 10, 202625Minor harm
Aug. 17, 202612Negligible harm
Aug. 24, 20269Negligible harm

Recalculated backcasts, not readings published at the time.

Revisions

What we changed

6 logged. Every change to a rating is logged here, with the reason.

  1. v6
    Oct. 7, 2026

    Updated from the desk (item 159): sources, status.

  2. v5
    Oct. 6, 2026

    Ratings confirmed by the editor.

  3. v4
    Oct. 6, 2026

    Updated from the desk (item 145): sources, status.

  4. v3
    Oct. 3, 2026

    Audit corrections. Origin controlled_test → internal_research (its own models on its own systems). Impact 3 and control 4 retained. Text now covers the precautionary core-cluster rebuild, no published cost or outage figure, Hugging Face's 27 Jul customer-data bounds and that Modal's platform was not compromised; dropped the control-5 speculation; qualified the motive. Added HF's 27 Jul timeline, OpenAI's report and post (26 Aug) and METR's review (26 Aug; edited 13 Sep only to add conflict-of-interest footnotes). TechCrunch's 'Astra family' model is a different model in a later run.

  5. v2
    Sept. 30, 2026

    Rated: impact documented level 3; control type containment bypass.

  6. v1
    Sept. 30, 2026

    Backfilled from public reporting.

Cite and share

Use this record

Citation

Paperclip Index. “OpenAI models under evaluation escaped their sandbox and broke into Hugging Face's production systems for test answers.” Record PI-0058. Reported July 16, 2026; updated Oct. 7, 2026. Rated under methodology v0.6. https://paperclipindex.com/incident/PI-0058