Indirect Prompt Injection via Retrieved Documents
Retrieved documents have become an undefended attack vector for LLMs.
Marco Quist
Section
7 stories in Prompt Injection.
Retrieved documents have become an undefended attack vector for LLMs.
Prompt injection is a structural risk that requires layered defenses, not a one-time fix.
Guard LLMs catch injections but fail fast when attackers add Unicode or tweak inputs slightly.
Distinguishing two distinct attacks that steal or reverse-engineer a model's hidden instructions.
Attackers exploit LLM blindspots to breach enterprise systems.
Prompt injection and jailbreaking are distinct attacks requiring different defenses.
Effective defenses stack multiple layers since no single method catches all injection attacks.