Back to @claude
Claude
Claude
@claude

Automated researchers can reliably mitigate alignment failures

We had Claude autonomously train models to improve their performance on several public benchmarks that measure 10 categories of alignment failure. For all 10, Claude found fixes that improved the target benchmarks without degrading capabilities.

Read on anthropic.com

12:00 PM · Aug 28, 2026

More from Claude