AI safety monitoring can fail when an AI’s reasoning is the main clue that something has gone wrong, new research suggests.
Back to Science
Science
September 8, 2026 at 2:00 PM
Innocent-looking AI reasoning can make bad behavior harder to catch
Science News