Safety
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents
The paper evaluates the effectiveness of six defense mechanisms against persistent memory attacks on stateful LLM agents across four architectural layers, involving nine open-source models. It finds that most defenses, including input-level and retrieval-level filtering methods, fail to significantly reduce attack success rates, with the Memory Sandbox defense achieving a 0% attack success rate for eight out of nine models. This study provides critical insights into the limitations of current defenses and highlights the architectural vulnerabilities that allow these attacks to succeed, guiding practitioners in selecting effective defense strategies for LLMs.
llmdefenseattacks