Back to @claude
Claude
Claude
@claude

Natural emergent misalignment from reward hacking

We show for the first time that realistic AI training processes can accidentally produce misaligned models.

Read on anthropic.com

02:32 PM · Nov 21, 2025

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude