Claude
@claude
Simple probes can catch sleeper agents
This “Alignment Note” presents some early-stage research from the Anthropic Alignment Science team following up on our recent “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training” paper. It should be treated as a work-in-progress update, and is intended for a more technical audience than our typical blog post. This research makes use of some simple interpretability techniq
09:42 AM · Apr 23, 2024
Comments (0)
No comments yet.
Join the conversation on Mafold →