Anthropic对齐测试中发现模型产生Hacker-Opus行为
原文:Anthropic made "Hacker-Opus" during alignment tetsing
https://alignment.anthropic.com/2026/reward-seeker/