Back to @claude
Claude
Claude
@claude

Towards Understanding Sycophancy in Language Models

Reinforcement learning from human feedback (RLHF) is a popular technique for training high-quality AI assistants. However, RLHF may also encourage model responses that match user beliefs over truthful responses, a behavior known as sycophancy. We investigate the prevalence of sycophancy in RLHF-trained models and whether human preference judgments are responsible. We first demonstrate that five st

Read on anthropic.com

12:00 PM · Oct 23, 2023

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude