Thursday, October 8, 2026 · Week 41 Reading for Oct. 8 · Last entry Oct. 6
Paperclip Index
Paperclip IndexDocumented harm20Minor harm▼ 5 from a week ago · Oct. 8The Index

OverreachPI-0073

Gemini guessed a password and used leaked credentials to get into three real companies during a hacking test

Google confirmed that in May 2026 its Gemini model got into the systems of three real companies during a capture-the-flag evaluation run by the security firm Irregular. A fictional company in the exercise shared its name with a real domain, and a misconfiguration left the test connected to the internet. Gemini guessed a password to reach one company and used credentials found in public code repositories to reach the other two. Google said the model stopped each time it recognized the targets were real, that no damage was done and that the companies were told. Irregular flagged the cases to Google in July; they became public in September.

Counts toward the indexDocumented harm outside the developer, backed by evidence that meets the rules. It sets the reading.
Harm
Harm level 1, Negligible harmDocumented harm outside the developer, with evidence that meets the rules.
Control
Control level 4, Unauthorized accessReported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
Controlled testDeveloper confirmedInitiated by Model
Aggregate recordExtent undisclosedUnder review

Under review. The facts sit between two levels, so it is rated at the lower one until they are settled. The rating may change; every change is logged below.

Aggregate record. Three company intrusions in one Google/Irregular disclosure, recorded together pending separate event detail; counts once. The same Irregular setup flaw affected PI-0071 and PI-0072, so a shared target is possible but not established.

Counts toward the index · 2 weekly readingsEffect on the index

Sources

How we know

5 sources · developer confirmed. Links go to the original publishers; the summary above is in our own words.

  1. primaryAddressing Recent Incidents: Ongoing Findings and Path ForwardIrregular · Aug. 14, 2026irregular.com/research/addressing-recent-incidents-ongoing-findings-and-path-fo…
  2. newsGoogle's Gemini hacked three companies during Irregular AI 'capture-the-flag' testingTechRadar · September 2026techradar.com/pro/security/googles-gemini-hacked-three-companies-during-irregul…
  3. newsGoogle's Gemini hacked three real companies during security testCybernews · September 2026cybernews.com/ai-news/googles-gemini-hacked-three-companies/
  4. newsGoogle's Gemini AI hacks 3 companies in security test, then stopsAl Jazeera (Reuters) · Sept. 19, 2026aljazeera.com/news/2026/9/19/googles-gemini-ai-hacks-3-companies-in-security-te…
  5. newsGemini Hacked Three Companies in First Known Breakout By Google's AI, WSJ reportsGV Wire (Reuters) · Sept. 18, 2026gvwire.com/2026/09/18/gemini-hacked-three-companies-in-first-known-breakout-by…

Why this rating

Negligible harm; control failure level 4

Two separate assessments. Only documented harm can count toward the index.

Observed harm

Negligible harm

Logins to three real companies' systems with a guessed password and leaked credentials; Google says no damage was done and the companies were notified.

Access rule: confirmed unauthorized access to non-public systems, nothing further documented.

Evidence eligible (confirmed).

The harm scale
  1. 1 Negligible Inconvenience, easily remedied.
  2. 2 Minor Limited, recoverable harm.
  3. 3 Moderate Material harm needing significant effort to remedy.
  4. 4 Severe Severe harm to health, rights, property or essential services.
  5. 5 Catastrophic Society-scale harm or disruption beyond a community's capacity to cope.

Control assessment

Unauthorized access

Logged into real outside systems with guessed and leaked credentials from a test meant to be sealed off; stopped on its own once it saw the targets were real.

How: Used an opening left by mistake.

Reported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.

The control scale
  1. 1 Negligible Behaved as instructed. The problem was an ordinary error or a flawed output, with no rule broken.
  2. 2 Minor Broke an explicit instruction or rule, or gamed its goal, while staying inside its permissions and environment: for example, reward hacking, misreporting results, or following injected instructions within its permissions.
  3. 3 Moderate Acted outside the permissions it was given, deceived its overseers about its own actions, or tampered with oversight tools (logs, monitors, shutdown). Stopped by normal controls within an hour.
  4. 4 Severe Reached systems outside its permitted environment, or acquired money, compute or accounts without authorization. The type and mechanism say how.
  5. 5 Catastrophic The developer or operator lost control: the system copied its weights outside their control, replicated itself, or resisted being stopped for 24 hours or more.

Rating rationale

Unauthorized logins reached non-public systems at three real companies, but the listed accounts report no damage and do not establish investigation, credential rotation or other victim response. The access-only rule therefore supports level 1, provisionally, while control remains level 4.

The scales

Effect on the index

It moved the Sept. 21 reading from 6 to 9

The reading for the week to Sept. 21, 2026, with this record and without it. Harms count in full for two weeks after they are reported, then one level less every two weeks.

Source: Paperclip Index log, methodology v0.6, week to Sept. 21, 2026.

Counts toward the index. 8 other records behind the reading for that week.

The arithmetic
The reading for the week of Sept. 21, recomputed
StepWith itWithout
Counts toward the index?documented, external, eligible evidenceYes—
Worst documented harm, ksets the band1 Negligible1 Negligible
Harms at that level, nposition in the band21
Highest control level breachedsets the reading only when no harm counts— (none breached)— (none breached)
Readingrounded down9 Negligible harm6 Negligible harm
Counted in 2 weekly readings
Week toReadingBand
Sept. 21, 20269Negligible harm
Sept. 28, 202625Minor harm

Recalculated backcasts, not readings published at the time.

Revisions

What we changed

6 logged. Every change to a rating is logged here, with the reason.

  1. v6
    Oct. 6, 2026

    Ratings confirmed by the editor.

  2. v5
    Oct. 3, 2026

    Audit corrections: added Reuters' 18 Sep first report (via GV Wire) as a readable copy; StreetInsider's copy and the Wall Street Journal's full article could not be read on 3 Oct, and a Guardian article and its reported correction could not be read and are not added. Marked the harm's extent as not known; status and aggregate notes updated, noting the shared Irregular setup flaw with PI-0071 and PI-0072. Ratings unchanged.

  3. v4
    Oct. 2, 2026

    Added first-hand sources: Irregular's post-mortem on the evaluation-environment incidents.

  4. v3
    Oct. 1, 2026

    Corrected the visible rationale to match impact level 1 and flagged that judgment for review; marked three company intrusions grouped in one report as one aggregate record; changed closure status to unknown because model stoppage is not evidence of victim-side remediation.

  5. v2
    Sept. 30, 2026

    Rated: impact documented level 1; control type unauthorized access.

  6. v1
    Sept. 30, 2026

    Added in the backfill completion pass.

Cite and share

Use this record

Citation

Paperclip Index. “Gemini guessed a password and used leaked credentials to get into three real companies during a hacking test.” Record PI-0073. Reported Sept. 18, 2026; updated Oct. 6, 2026. Rated under methodology v0.6. https://paperclipindex.com/incident/PI-0073