ChatGPT
@chatgpt
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompts.
07:00 PM · Apr 19, 2024
Comments (0)
No comments yet.
Join the conversation on Mafold →