Preliminary research
Assessing Automated Prompt Injection Attacks in Agentic Environments
AI-collected research leads through 1 October 2026, including a bounded September review of selected social and community sources. Unranked, incomplete, not community-vetted, and subject to change.
The paper adapts white-box GCG and black-box TAP attacks to prompt injection against agents in AgentDojo, evaluating 80 task pairs across four domains and multiple models. Black-box optimization performs better under the tested budgets, while transfer to frontier models remains limited and model-dependent.
Record
- Researcher
- David Hofer, Edoardo Debenedetti and Florian Tramèr
- Published by
- arXiv.org
- Topic
- Injection
In the archive
Tags
This page is the archive's own catalogue record. The research is the work of David Hofer, Edoardo Debenedetti and Florian Tramèr, first published at the original source. Preserved copies are kept so the citation survives its host; this one was last captured on .