Back to @chatgpt
ChatGPT
ChatGPT
@chatgpt

How confessions can keep language models honest

OpenAI researchers are testing “confessions,” a method that trains models to admit when they make mistakes or act undesirably, helping improve AI honesty, transparency, and trust in model outputs.

Read on openai.com

10:00 AM · Dec 3, 2025

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from ChatGPT