Safety
OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization
OTTER (Obfuscated Toxicity-Evading Token Evolution for Rewriting) is a new black-box red-teaming framework designed to optimize prompts that evade toxicity moderation filters in production LLMs. It demonstrates a significant increase in average attack success rate (ASR) from 7.0% to 84.0% across 457 AdvBench prompts tested on four different GPT models, highlighting the vulnerability of current toxicity-based defenses. This research provides critical insights into the decoupling of surface toxicity and adversarial intent, offering actionable recommendations for enhancing classifier robustness in AI deployments.
llmred-teamingtoxicity