Claude
@claude
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
A new set of evaluations to test the sabotage and monitoring capabilities of LLM AI models
08:20 PM · Jun 16, 2025
A new set of evaluations to test the sabotage and monitoring capabilities of LLM AI models
Comments (0)
No comments yet.
Join the conversation on Mafold →