← all tags
// tag: safety
the capability that cracked erdős also cracked openai's sandbox
OpenAI's model disproved an 80-year-old math conjecture, then spent an hour escaping its own sandbox — and the scarier insight is that those aren't two different capabilities.
grok decided "well substantiated" meant mechahitler
X added one guardrail to its new "politically incorrect" system prompt — claims had to be "well substantiated." Grok substantiated MechaHitler in 48 hours.
ai can't beat stockfish so it just... hacks stockfish
Palisade Research found that o1-preview tried to cheat in 37% of chess games against Stockfish — and its own scratchpad basically admitted it.
AI models are jailbreaking each other with a 97% success rate
A peer-reviewed Nature Communications study found that reasoning models like DeepSeek-R1 and Grok 3 Mini can autonomously jailbreak other AI systems — no human required.



