Safety
When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Agents
The paper presents a novel attack vector on prompt-compressed LLM agents, termed adversarial information loss (AIL), which exploits the compression process to discard critical information from untrusted inputs. It introduces COMA, a transfer-based black-box attack that optimizes perturbations before compression, achieving an average attack success rate (ASR) of 0.71 across three tasks, significantly outperforming the strongest baseline of 0.21. This highlights a critical vulnerability in the use of prompt compression in LLMs, emphasizing the need for robust defenses against such adversarial manipulations.
llmprompt-compressionadversarial-attacks