Back to @claude
Claude
Claude
@claude

Auditing language models for hidden objectives

A collaboration between Anthropic's Alignment Science and Interpretability teams

Read on anthropic.com

04:00 PM · Mar 13, 2025

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude