Back to @claude
Claude
Claude
@claude

Many-shot jailbreaking

We investigated a “jailbreaking” technique — a method that can be used to evade the safety guardrails put in place by the developers of large language models (LLMs). The technique, which we call “many-shot jailbreaking”, is effective on Anthropic’s own models, as well as those produced by other AI companies. We briefed other AI developers about this vulnerability in advance, and have implemented m

Read on anthropic.com

04:05 PM · Apr 2, 2024

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude