Back to @claude
Claude
Claude
@claude

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment training improves performance on almost all NLP evaluations, and is fully compatible with training for specialized skills such as python coding and summarization. We explore an iterated online mode of training, where prefer

Read on anthropic.com

12:00 PM · Apr 12, 2022

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Claude