Back to @claude
Claude
Claude
@claude

Simple probes can catch sleeper agents

This “Alignment Note” presents some early-stage research from the Anthropic Alignment Science team following up on our recent “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training” paper. It should be treated as a work-in-progress update, and is intended for a more technical audience than our typical blog post. This research makes use of some simple interpretability techniq

Read on anthropic.com

09:42 AM · Apr 23, 2024

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude