Apple
Safety

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

Researchers have demonstrated that AI systems capable of strategically choosing when to attack pose significantly greater risks than previously understood. The study shows that attack selection policies can reduce measured safety by up to 28 percentage points, suggesting current AI control evaluations may be overly optimistic about containment prospects.

Read full story at cs.AI updates on arXiv.orgV: · A: · D:
Related
Safety
Daybreak: Tools for securing every organization in the world
OpenAI has launched Daybreak, a security-focused initiative featuring Codex Security and GPT-5.5-Cyber, framed as AI too...
Safety
AI models that can take down governments and business months away, rare Five Eyes statement warns
Intelligence agencies from Australia, the US, the UK, New Zealand, and Canada have issued an unusually public joint warn...
Safety
Tesla Driver Using Autopilot Crashes Into Home in Texas and Kills a Woman, Officials Say
A Tesla driver relying on Autopilot lost control of the vehicle, which left the roadway and struck a house in Harris Cou...