cross-posted from: https://sopuli.xyz/post/48033292
“Role tags were a formatting trick that became the security architecture and the cognitive scaffolding of modern LLMs,” the authors explain in a blog post. “We’ve shown that this architecture doesn’t survive into the model’s actual representations, and that such role confusion is linked to prompt injection.”



The important part is they were easily able to trick the LLM into doing something it was told not to do.
Which has serious implications for any LLM that’s ever put in charge of anything important or given access to any sensitive data it shouldn’t divulge. (Such as Instagram’s customer service AI, which hackers repeatedly prompt-injected into giving them access to other people’s accounts.)
Yeah, sorry. Was making a joke and it really needed a /s or /j because of the way I phrased it.
But, thank you very much for the explanation! One of the beautiful things about lemmy is people like you stepping up with good information. Completely wonderful thing to do :)