Preliminary research
Cache Me, Catch You: Exploiting LLM Caching Layers in vLLM, GPTCache & Friends
AI-collected research leads through 22 September 2026, including targeted additions between broader sweeps. Unranked, incomplete, not community-vetted, and subject to change.
Research on LLM serving caches (vLLM, GPTCache and peers) where the layer deciding whether two requests are 'the same' is fooled. All three cache types rest on serialize-key-reuse, so an attacker crafts colliding inputs: hash-colliding padding poisons a shared system-prompt or block-wise KV entry so a malicious block is skipped; near-neighbour embeddings make semantic and RAG caches serve a planted answer; byte-only image hashing makes moderation reuse a benign verdict. Three CVEs resulted.
Record
- Researcher
- Xiangfan Wu, Lingyun Ying, Haipeng Qu, Guoqiang Chen and Yacong Gu
- Published by
- i.blackhat.com
- Format
- Whitepaper
- Topic
- AI
In the archive
Related sources
- Cache Me, Catch You: Cache Related Security Threats in LLM Serving Frameworks
- Code Repository
- Kv Cache Collision advisory Advisory
- Image Hash Collision (tobytes) advisory Advisory
- PNG tRNS Transparency Bypass advisory Advisory
Tags
This page is the archive's own catalogue record. The research is the work of Xiangfan Wu, Lingyun Ying, Haipeng Qu, Guoqiang Chen and Yacong Gu, first published at the original source. Preserved copies are kept so the citation survives its host; this one was last captured on .