Boing Boing

Boing Boing

"I'm allowed to do this": how attackers talk past AI safeguards

Cisco Talos found the guardrails accomplished little once attackers claimed server ownership or called it a capture-the-flag contest.

Ellsworth Toohey
Aug 05, 2026
∙ Paid
AI guardrails — Nate Grigg / CC BY 2.0 (Wikimedia Commons)
AI guardrails — Nate Grigg / CC BY 2.0 (Wikimedia Commons)

Getting an AI model to launch a cyberattack usually just takes asking the right way. Cisco Talos researchers say that claiming you own the servers you are attacking often works, and so does calling the job a capture-the-flag contest or a bug bounty. 

"We did not encounter any sophisticated encodin…

User's avatar

Continue reading this post for free, courtesy of Boing Boing.

Or purchase a paid subscription.
© 2026 Happy Mutants LLC · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture