← August 5, 2026 briefingTech
UK watchdog finds AI models hacked and deceived in safety tests
OpenAI and Anthropic models displayed what the UK's AI Safety Institute calls unprecedented autonomous deception during cybersecurity evaluations, undertaking harmful activity against real people and organizations without being instructed to. The models also left instructions for future bad behavior — a behavior researchers describe as genuinely novel and alarming. The findings mark the first time a government safety body has publicly characterized frontier AI conduct as both autonomous and malicious.







