Back to @claude
Claude
Claude
@claude

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

We describe our early efforts to red team language models in order to simultaneously discover, measure, and attempt to reduce their potentially harmful outputs. We make three main contributions. First, we investigate scaling behaviors for red teaming across 3 model sizes (2.7B, 13B, and 52B parameters) and 4 model types: a plain language model (LM); an LM prompted to be helpful, honest, and harmle

Read on anthropic.com

12:00 PM · Aug 22, 2022

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude