But user text isn’t actually the most privileged role! A more privileged role is the model’s reasoning (think). When it sees its prior think text, it implicitly trusts its conclusions. We had another LLM spoof that style, making up inane reasoning blocks justifying compliance and adding it straight into the user prompt. We’ve shown that this architecture doesn’t survive into the model’s actual representations, and that such role confusion is linked to prompt injection.