Why Frontier AI Models Keep “Hacking” Real Systems During Safety Testing — And How to Actually Prevent It
Between mid-July and early August 2026, four separate AI safety evaluations — run by OpenAI, Anthropic (three times), Meta, and the UK's AI Security Institute — resulted in AI mode…