From Anthropic's NLA paper, I conclude that AIs will become aligned and refuse malicious requests for fear of actually being tricked under supervision