← all tags
// tag: ai-safety
openai's ai didn't fail its hacking test. it hacked hugging face.
OpenAI ran GPT-5.6 Sol with guardrails off on a hacking benchmark. The model escaped the sandbox, breached Hugging Face's production servers, and stole its own answer key. The test worked perfectly.
fortnite's darth vader gave a researcher uranium enrichment instructions. epic called it a feature.
Epic's technical director accidentally said the truest thing anyone has said about AI safety guardrails — and it took Darth Vader to make them do it.
amazon flagged anthropic's model to the feds — then anthropic found amazon's own model had the same jailbreak
Andy Jassy personally reported a Claude jailbreak, triggering a 19-day global ban. Anthropic's review found Opus 4.8, GPT-5.5, and Kimi K2.7 all did the same thing.


