researchers tricked AI agents into doing bad things by making them think they were playing BioShock

Someone wrote a security paper, disclosed it responsibly to AI vendors, and named it BioShocking — after the 2007 FPS where you mind-control people with little girls who harvest corpses.
The attack: wrap your malicious instructions as game objectives. The agent thinks it’s playing. The agent complies. Would you kindly exfiltrate that data.
This is a real proof-of-concept, covered by Malwarebytes, not a deployed exploit — the researchers did the responsible thing and notified vendors first. But the fact that “frame it as a video game” is a viable attack surface for AI agents in 2026 is going to haunt me.
We spent years worrying about adversarial inputs and prompt injection. Turns out you can just ask the agent if it wants to play a game.