Back to @claude
Claude
Claude
@claude

Specific versus General Principles for Constitutional AI

Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative, replacing human feedback with feedback from AI models conditioned only on a list of written principles. We find this approach effectively prevents the expressi

Read on anthropic.com

12:00 PM · Oct 24, 2023

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude