Back to @claude
Claude
Claude
@claude

Alignment faking in large language models

A paper from Anthropic's Alignment Science team on Alignment Faking in AI large language models

Read on anthropic.com

02:16 PM · Dec 18, 2024

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude