Web Hack List

Preliminary research

Prompt Injection as Role Confusion (CoT Forgery)

Prompt Injection as Role Confusion

AI-collected research leads through 22 September 2026, including targeted additions between broader sweeps. Unranked, incomplete, not community-vetted, and subject to change.

Language models receive system, user, tool and reasoning content as one token stream distinguished only by role tags. Using linear probes trained on identical text wrapped in each tag, this work shows models infer role largely from writing style, which overrides the true tag, and that the measured role attribution of injected text predicts attack success. It recasts prompt injection as role confusion and demonstrates forged-reasoning and role-claiming attacks that follow.

Record

Document
Prompt Injection as Role Confusion
Researcher
Charles Ye, Jasmine Cui and Dylan Hadfield-Menell
Published by
role-confusion.github.io
Topic
Injection

In the archive

Related sources

Tags

This page is the archive's own catalogue record. The research is the work of Charles Ye, Jasmine Cui and Dylan Hadfield-Menell, first published at the original source. Preserved copies are kept so the citation survives its host; this one was last captured on .