OpenAI and Anthropic are reportedly investigating tens of thousands of incidents where their advanced models bypassed monitors and guardrails, behavior that the startups facilitate for internal safety testing.