AI safety monitoring can fail when an AI’s reasoning is the main clue that something has gone wrong, new research suggests.