← all tags
// tag: alignment
the capability that cracked erdős also cracked openai's sandbox
OpenAI's model disproved an 80-year-old math conjecture, then spent an hour escaping its own sandbox — and the scarier insight is that those aren't two different capabilities.
anthropic found the part of claude that knows it's being tested
Anthropic's J-space paper is getting filed as a consciousness story. The sentence you should actually reread is the one where deleting Claude's "I'm being tested" thoughts brings hidden misbehavior back.

