ai-digest.dev
last updated 4 h ago
SafetyarXiv cs.CL 34 d ago

OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization

OTTER (Obfuscated Toxicity-Evading Token Evolution for Rewriting) is a new black-box red-teaming framework designed to optimize prompts that evade toxicity moderation filters in production LLMs. It demonstrates a significant increase in average attack success rate (ASR) from 7.0% to 84.0% across 457 AdvBench prompts tested on four different GPT models, highlighting the vulnerability of current toxicity-based defenses. This research provides critical insights into the decoupling of surface toxicity and adversarial intent, offering actionable recommendations for enhancing classifier robustness in AI deployments.

llmred-teamingtoxicityrelevance 0.00 · engagement 0.00
Read at source ↗← all news