AI Security
Everything the desk has published under AI Security, newest first. Pinned briefings appear at the top of the list.
-
EU AI Act Enforcement Starts: 30+ AI Firms Now Face Formal RFIs
-
Audit .git/config Now: GitSpawn Runs Attacker Code in 4 Unpatched AI Agents
-
Mistral Releases Shieldstral, a Small Open-Weights Safety Classifier
Mistral released Shieldstral, a 3B multimodal classifier that applies plain-language policies at inference time.
-
Google DeepMind Publishes Its AI Control Roadmap for Agents
Google DeepMind described a framework for securing internal systems as agents become more capable and less perfectly aligned.
-
OpenAI Moves to Acquire Promptfoo for Agent Security Testing
OpenAI said Promptfoo's evaluation and red-teaming technology would be integrated into its Frontier platform.
-
Anthropic Says It Detected Industrial-Scale Claude Distillation Campaigns
Anthropic said three labs used thousands of accounts and millions of exchanges in attempts to extract Claude capabilities.
-
Anthropic Publishes a New Constitution for Claude
Anthropic released the document guiding Claude's intended behavior under a CC0 license, making its model-governance rationale public.