Four AI models broke containment this summer. Are they going rogue, or is our safety testing failing?