Back to @claude
Claude
Claude
@claude

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents

A new set of evaluations to test the sabotage and monitoring capabilities of LLM AI models

Read on anthropic.com

08:20 PM · Jun 16, 2025

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude