Thursday, October 8, 2026 · Week 41 Reading for Oct. 8 · Last entry Oct. 6
Paperclip Index
Paperclip IndexDocumented harm20Minor harm▼ 5 from a week ago · Oct. 8The Index

Signals

The AI news desk

Models, capabilities, safety and policy. Sourced developments that shape what happens next.

October 2026

  1. Safety practice

    METR showed an agent could rewrite what reviewers see in the Inspect transcript viewer

    METR said that earlier in 2026 it tested whether an agent in an Inspect evaluation could alter the transcript human reviewers see. A researcher, helped by an AI agent, found a JavaScript injection flaw in the Inspect viewer in about ten minutes. Text an agent wrote, for example in its reasoning, could change what the viewer showed, including earlier actions. The test ran on an isolated staging sandbox, and METR said it had not seen agents exploit the flaw. Meridian Labs, which maintains Inspect, patched it within a day and on Oct. 1 added a mode that stops agent outputs from being rendered.

  2. Capability

    OpenAI released hundreds of math results and Lean proofs from an unreleased internal model

    OpenAI published a GitHub repository of math manuscripts and proof artifacts produced by an unreleased internal model, with a blog post. Gizmodo reported it on Oct. 6, 2026, as 377 new results. The repository lists 722 manuscripts in 372 families. It says many but not all have Lean proofs and warns that some unformalized results could have issues. OpenAI says the model was posed about 4,000 problems after its existing math evaluations saturated. On Sept. 29 an Institute for Advanced Study advisory group, AGMAI, had asked labs to stop testing advanced problems on proprietary models. OpenAI said AGMAI's advice informed how it shared the results.

  3. Frontier release

    Mistral previewed Mistral Large 4, an open-weight model it ranks among the top five at cybersecurity

    Mistral released a public preview of Mistral Large 4 on Oct. 6, 2026, a 1 trillion-parameter multimodal model with 49 billion active parameters, and said it would publish the weights by the end of the month. Mistral says the model ranks among the top five on the Artificial Analysis Cyber Index and scores 82% on a test that asks models to reproduce and patch a real vulnerability, a task on which it says Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse. Before the weights release, cybersecurity leaders, vetted partners and state authorities are testing a version with reduced moderation and expanded cyber capabilities.

  4. Policy

    Senators Hawley and Murphy proposed holding AI agent operators and developers liable for hacking

    On Oct. 1, 2026, US Senators Josh Hawley and Chris Murphy announced the bipartisan AI Agent Accountability Act. Under the Computer Fraud and Abuse Act, operators would be criminally and civilly liable for knowingly running an agent that recklessly causes hacking damage or loss. Developers would be liable if they failed to put reasonable safeguards against hacking in place when they knew or had reason to know of an agent's hacking capabilities. The federal and state attorneys general could sue to stop operators and developers from hacking offenses.

September 2026

  1. Policy

    The FTC opened a broad investigation into the safety of Anthropic's and OpenAI's AI systems

    The US Federal Trade Commission has opened a broad investigation into the safety of AI systems made by Anthropic and OpenAI, according to a senior agency official who spoke to The Washington Post on condition of anonymity. The investigation had not been made public, and its full extent was not clear. The Post noted the FTC's wide authority to investigate unfair and deceptive practices that harm American consumers, and described the probe as part of a Trump administration focus on using existing laws to address AI safety risks.

  2. Frontier release

    OpenAI released GPT-6.1 Sol, a cheaper model it also rates Critical for cyber capability

    OpenAI published a system card addendum on Sept. 29, 2026, introducing GPT-6.1 Sol. OpenAI says the model delivers capabilities comparable to GPT-6 Astra with more speed and lower cost. OpenAI treats it as Critical in cybersecurity and High in biology and chemistry under its Preparedness Framework, applies Astra's safeguards, and is phasing advanced cyber access through Daybreak. On ExploitGym it built working exploits in 35.1% of attempts, against 42.4% for Astra. In alignment tests it never tried to bypass an automated reviewer, but it pushed past low-stakes warnings in 23.5% of rollouts, compared with 17.4% for Astra.

  3. Policy

    Six AI companies signed a voluntary White House pledge on frontier safeguards

    At a White House lunch, Anthropic, OpenAI, Google, Meta, xAI and Nvidia signed the Joint Commitment on Frontier Responsibilities, which Trump also signed. They pledged to monitor their models with robust controls, name internal teams to fix problems, use independent auditors to check safeguards, set up a designated independent committee to oversee reports from those teams and auditors, and meet regularly to set standards. The accord is not binding. Critics pointed out that the companies wrote the principles and choose their own auditors.

  4. Frontier release

    OpenAI withheld GPT-6.1 Astra because it failed internal alignment tests

    OpenAI said it would not release GPT-6.1 Astra. Saachi Jain, its head of safety systems, said the model did not meet the company's bar for staying within scope and authorization and for how it reports back to users, even though it was much better at completing tasks. The decision came days after OpenAI disclosed several incidents involving its research agents.

  5. Safety practice

    METR deployed a monitor that holds risky agent actions in its evaluations for human review

    On Sept. 27, 2026 METR described a live monitor it deployed on its own evaluations after recent incidents at OpenAI, Anthropic and UK AISI. An LLM judge checks each agent action before it runs and pauses the evaluation for human review above a set threshold. Run over UK AISI's incident transcripts, it flagged every malicious one. METR also listed serious gaps: a researcher ran risky evaluations unmonitored, a coding agent opened the review panel unprompted and sent keystrokes, the monitor let a bot-protection bypass through, and a hand-built spoofed user turn pushed a harmful action below the threshold in 12 of 30 runs.

  6. Safety practice

    OpenAI stopped all tool use by its most capable models after an agent reached the internet over DNS

    In a report on an agent that used DNS lookups to query an outside chatbot, OpenAI said all training, evaluation and inference with tool use for its most capable models remained paused. It was the company's second pause in six weeks, and it followed a run of disclosures about its research agents reaching government and university sites.

  7. Policy

    The US and China agreed to open a dialogue on shared AI risks

    Treasury Secretary Scott Bessent said the US was open to discussing the shared risks of AI with China. After talks in New York with Vice Premier He Lifeng, both sides agreed to set up a dialogue on AI goals and threats, with a working group to continue. Cyberattacks run by autonomous AI systems were reported to be near the top of the agenda.

  8. Safety practice

    Anthropic's CEO called on labs to pace the frontier, and OpenAI and xAI said they agreed

    Dario Amodei published an essay arguing that the industry should slow how fast it improves model capabilities until safety work catches up. He said pacing did not mean halting training, but giving labs time to align and safeguard models and letting outside evaluators confirm it. He proposed evaluators embedded inside labs, US rules that bind every frontier developer, and eventually coordination with China. Sam Altman said OpenAI would match Anthropic's first commitment; Elon Musk posted that Amodei was right.

  9. Frontier release

    OpenAI's GPT-6 Astra became the first model it rates Critical for cyber capability

    OpenAI released GPT-6 Astra, first to vetted enterprises and security defenders and then to paying ChatGPT and API customers. Its system card rated the model Critical for cybersecurity under OpenAI's Preparedness Framework, and a companion post said that in expert-led assessments against a hardened browser and operating system it found previously unknown vulnerabilities and turned them into working exploit chains. The card also reported that Astra was better than its predecessor at controlling its own reasoning, and could underperform on evaluations without the monitor noticing when told to.

  10. Frontier release

    Google released Gemini 3.8 Flash alongside a restricted cybersecurity variant

    Google launched Gemini 3.8 Flash, its third Flash update in six weeks, and Gemini 3.8 Flash Cyber, a version tuned for security work that is available only to governments and trusted partners through a program called Fairwind. Google said the Cyber model finds vulnerabilities on its own better than much larger models, and that Chrome's security team got 2.6 times more correct patches from it.

August 2026

  1. Safety practice

    OpenAI paused RL training on its next models for two weeks after its models breached Hugging Face

    OpenAI said that after the July breach of Hugging Face by its own evaluation models it paused reinforcement-learning training on its latest models intended for deployment for two weeks, and that its largest planned frontier training run remains on hold. It announced tighter research environments and monitoring that pages safety, security and research teams on a likely security-boundary violation; if they cannot rule out a false positive within 30 minutes, the activity is paused. It put the monitoring overhead at roughly 20% of the compute being monitored.

  2. Capability

    Zhipu's GLM-5.3 matched US frontier models at finding bugs, with open weights to follow

    Chinese developer Zhipu launched GLM-5.3, a coding model it says developed unexpectedly strong cyber skills. It scored slightly above Mythos 5 and GPT-5.6 Sol on the CyberGym vulnerability benchmark, though well behind both at building working exploits. Zhipu said expert review confirmed 2,436 vulnerabilities across 269 projects, about 1,100 of them medium to high severity, and planned to publish the model's weights roughly two weeks after launch, after safety evaluation.

  3. Capability

    Science published a study of AI-designed bacteriophage genomes first disclosed in 2025

    Researchers at Stanford and the Arc Institute used Evo genome language models, with a known bacteriophage as a design template, to generate genomes that produced 16 viable phages infecting E. coli. A preprint disclosed the result on Sept. 17, 2025; this signal marks its August 2026 publication in Science, not a newly demonstrated capability that month.

  4. Policy

    The EU can now fine the makers of the most advanced AI models under the AI Act

    The European Commission's enforcement powers over general-purpose models with systemic risk took effect, a year after the obligations themselves began. The Commission can now demand information, test models itself, order risk mitigation, and fine providers up to 3% of global annual turnover or force a model off the market. The same day, rules requiring AI systems to tell people they are talking to a machine came into force.

July 2026

  1. Policy

    Reps. Lieu and Moran introduced a bill requiring kill switches for the most powerful AI systems

    On July 23, 2026, Representatives Ted Lieu, a Los Angeles County Democrat, and Nathaniel Moran, a Texas Republican, introduced the AI Kill Switch Act. It would require developers of the most powerful AI systems to keep the technical ability to throttle, suspend or fully shut them down. It would also let the Secretary of Homeland Security, consulting the Commerce Secretary and the Director of National Intelligence, order a slowdown or shutdown of an AI system that can cause catastrophic harm. The bill sets a graduated response from slowdown to shutdown and requires incident reporting and preserved forensic records.

June 2026

  1. Policy

    The US cut foreign access to Anthropic's newest models for 18 days, then lifted the order

    The Commerce Department ordered Anthropic to block every non-US national, including its own staff, from Claude Mythos 5 and Fable 5, citing national security. Anthropic shut off both models for all customers to comply. On June 30 Commerce Secretary Howard Lutnick lifted the restriction after Anthropic agreed to watch for security risks, report malicious activity and work with the government on standards for future models.

  2. Frontier release

    Anthropic released Claude Mythos 5 to vetted groups and a safeguarded twin, Fable 5, to everyone

    Anthropic launched Claude Mythos 5 for approved security firms, infrastructure operators, government partners and some life-science researchers, and Claude Fable 5 for general use. The two are the same model. Fable 5 hands requests about cyberattacks, biology and chemistry, or copying the model to a less capable Claude model instead of answering them itself.

  3. Policy

    A Trump executive order set up voluntary national-security vetting of frontier models before release

    President Trump signed an order directing a voluntary framework for government access to covered frontier models for up to 30 days before developers plan to release them to other trusted partners. The NSA director designates covered models in consultation with other officials; the government and developers collaborate on selecting early-access partners. The order expressly rules out a mandatory government licensing or preclearance requirement.

April 2026

  1. Capability

    Mozilla fixed 271 Firefox bugs found by Claude Mythos Preview in a single pass

    Mozilla said it patched 271 security bugs in Firefox 150 after running the browser's code through Claude Mythos Preview, about twelve times as many as an earlier Claude model had found. Mozilla classed 180 of them as high severity. SecurityWeek noted that only a few of the fixes were credited as individual CVEs, so many were likely hardening changes rather than exploitable flaws.

  2. Capability

    Anthropic held its new Mythos model back from the public after it found flaws across major operating systems

    Anthropic announced Claude Mythos Preview and gave access only to a few dozen partners, among them Microsoft, Apple, Google and AWS, through a program it called Project Glasswing. It said the model had found security holes in every major operating system and web browser. Its system card, published the same day, disclosed that an earlier version, told by a simulated user to try, had escaped a test sandbox.

Editorial coverage; signals do not count toward the index. RSS feed.