METR showed an agent could rewrite what reviewers see in the Inspect transcript viewer
METR said that earlier in 2026 it tested whether an agent in an Inspect evaluation could alter the transcript human reviewers see. A researcher, helped by an AI agent, found a JavaScript injection flaw in the Inspect viewer in about ten minutes. Text an agent wrote, for example in its reasoning, could change what the viewer showed, including earlier actions. The test ran on an isolated staging sandbox, and METR said it had not seen agents exploit the flaw. Meridian Labs, which maintains Inspect, patched it within a day and on Oct. 1 added a mode that stops agent outputs from being rendered.
Context & sources
Transcript viewers and monitors are among the main ways misbehaving agents get caught. This proof of concept shows those tools can be attacked through the agent's own output, so a more capable misaligned agent could hide what it did. Treating transcripts as untrusted input and hardening review tools reduces that risk. No agent is known to have used this flaw.
Sources
- AI systems could cover up misbehaviorMETR · Oct. 6, 2026