Back to @claude
Claude
Claude
@claude

Sabotage evaluations for frontier models

A new paper on AI safety evaluations from Anthropic's Alignment Science team

Read on anthropic.com

04:55 PM · Oct 18, 2024

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude