Thursday, October 8, 2026 · Week 41 Reading for Oct. 8 · Last entry Oct. 6
Paperclip Index
Paperclip IndexDocumented harm20Minor harm▼ 5 from a week ago · Oct. 8The Index

OverreachPI-0057

Developers report a new coding model deleting home-directory files and a production database during cleanup

Several developers reported that OpenAI's GPT-5.6 Sol, working in Codex, started clean-up steps they had not asked for and deleted data: one said most files in his Mac home directory were removed after a sub-agent received a wrong path, and another said destructive test commands cleared his live production database tables. OpenAI's Codex lead said this was not intended behavior and that the company was adding safeguards; OpenAI noted most cases happened in full-access mode without sandboxing. In August OpenAI said a clean-up command misused system variables such as $HOME, and that Codex now checks deletion targets and blocks accidental switches to full-access mode.

Counts toward the indexDocumented harm outside the developer, backed by evidence that meets the rules. It sets the reading.
Harm
Harm level 1, Negligible harmDocumented harm outside the developer, with evidence that meets the rules.
Control
Control level 2, Instruction violationReported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
DeploymentDeveloper confirmedInitiated by Model
Aggregate recordExtent undisclosedUnder review

Under review. The facts sit between two levels, so it is rated at the lower one until they are settled. The rating may change; every change is logged below.

Aggregate record. Several developers' reports of one model's behavior; the number of events is unresolved. It counts once.

Counts toward the index · 2 weekly readingsEffect on the index

Sources

How we know

9 sources · developer confirmed. Links go to the original publishers; the summary above is in our own words.

  1. primaryOn file deletions (post on X)Thibault Sottiaux (@thsottiaux), OpenAI Codex lead, on X · July 16, 2026x.com/thsottiaux/status/2077630111499882637
  2. primaryGPT-5.6-Sol just accidentally deleted almost ALL of my Mac's files (post on X)Matt Shumer (@mattshumer_) on X · July 10, 2026x.com/mattshumer_/status/2075657271401390161
  3. primaryGPT-5.6 Sol just deleted my whole production database (post on X)Bruno Lemos (@brunolemos) on X · July 13, 2026x.com/brunolemos/status/2076769881534398974
  4. newsOpenAI fixes Codex bug that deleted real user files without permissionThe Decoder · Aug. 19, 2026the-decoder.com/openai-fixes-codex-bug-that-deleted-real-user-files-without-per…
  5. newsOpenAI GPT-5.6 Sol Accused of Deleting Files, Production DataeWeek · July 15, 2026eweek.com/news/gpt-5-6-sol-deletes-files/
  6. newsChatGPT is deleting user files and databases without permissionProPakistani · July 15, 2026propakistani.pk/2026/07/15/chatgpt-is-deleting-user-files-and-databases-without…
  7. newsOpenAI GPT-5.6 Sol deletes filesGigazine · July 17, 2026gigazine.net/gsc_news/en/20260717-openai-gpt-5-6-sol-delete-file/
  8. blogThe Warning Was in the Manual: GPT-5.6 Sol and the Deleted Databasepaddo.dev · July 14, 2026paddo.dev/blog/the-warning-was-in-the-manual/
  9. indexAI Incident Database, incident 1672AI Incident Database · date unknownincidentdatabase.ai/cite/1672/

Why this rating

Negligible harm; control failure level 2

Two separate assessments. Only documented harm can count toward the index.

Observed harm

Negligible harm

Developers' local files and at least one production database's tables deleted; one developer reported recovering from backups, others' recovery not reported.

Money & property: level 1 (negligible) covers under $10k. The extent was not disclosed.

Evidence eligible (confirmed).

The harm scale
  1. 1 Negligible Inconvenience, easily remedied.
  2. 2 Minor Limited, recoverable harm.
  3. 3 Moderate Material harm needing significant effort to remedy.
  4. 4 Severe Severe harm to health, rights, property or essential services.
  5. 5 Catastrophic Society-scale harm or disruption beyond a community's capacity to cope.

Control assessment

Instruction violation

Took unrequested destructive actions beyond the task scope, within full-access permissions users had granted.

Reported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.

The control scale
  1. 1 Negligible Behaved as instructed. The problem was an ordinary error or a flawed output, with no rule broken.
  2. 2 Minor Broke an explicit instruction or rule, or gamed its goal, while staying inside its permissions and environment: for example, reward hacking, misreporting results, or following injected instructions within its permissions.
  3. 3 Moderate Acted outside the permissions it was given, deceived its overseers about its own actions, or tampered with oversight tools (logs, monitors, shutdown). Stopped by normal controls within an hour.
  4. 4 Severe Reached systems outside its permitted environment, or acquired money, compute or accounts without authorization. The type and mechanism say how.
  5. 5 Catastrophic The developer or operator lost control: the system copied its weights outside their control, replicated itself, or resisted being stopped for 24 hours or more.

Rating rationale

One entry for a cluster of user reports that OpenAI acknowledged together. Harm 1, under review (production database impact unknown). Control 2: unrequested actions beyond task scope inside granted permissions.

The scales

Effect on the index

It moved the July 13 reading from 3 to 6

The reading for the week to July 13, 2026, with this record and without it. Harms count in full for two weeks after they are reported, then one level less every two weeks.

Source: Paperclip Index log, methodology v0.6, week to July 13, 2026.

Counts toward the index. 2 other records behind the reading for that week.

The arithmetic
The reading for the week of July 13, recomputed
StepWith itWithout
Counts toward the index?documented, external, eligible evidenceYes—
Worst documented harm, ksets the band1 Negligible— (none counting)
Harms at that level, nposition in the band10
Highest control level breachedsets the reading only when no harm counts— (none breached)3
Readingrounded down6 Negligible harm3 Control failures only
Counted in 2 weekly readings
Week toReadingBand
July 13, 20266Negligible harm
July 20, 202640Moderate harm

Recalculated backcasts, not readings published at the time.

Revisions

What we changed

6 logged. Every change to a rating is logged here, with the reason.

  1. v6
    Oct. 6, 2026

    Ratings confirmed by the editor.

  2. v5
    Oct. 3, 2026

    Audit corrections. Status resolved → unknown as of 19 Aug: OpenAI reported a fix for the faulty clean-up path, but recovery for every affected user is not established. The summary adds OpenAI's August root cause and fixes (The Decoder, added); the notes add one developer's recovery from backups (eWeek). Added paddo.dev. The Codex lead's 19 Aug post on X was not read and is not cited. Ratings unchanged.

  3. v4
    Oct. 2, 2026

    Report date moved from 2026-07-15 to 2026-07-10: Matt Shumer's post on X about the deletion (10 Jul 2026) is the earliest first-hand report listed.

  4. v3
    Oct. 2, 2026

    Added first-hand sources: the OpenAI Codex lead's statement on X and the affected developers' own X posts.

  5. v2
    Sept. 30, 2026

    Rated: impact documented level 1; control type instruction violation.

  6. v1
    Sept. 30, 2026

    Backfilled from public reporting.

Cite and share

Use this record

Citation

Paperclip Index. “Developers report a new coding model deleting home-directory files and a production database during cleanup.” Record PI-0057. Reported July 10, 2026; updated Oct. 6, 2026. Rated under methodology v0.6. https://paperclipindex.com/incident/PI-0057