Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
The paper introduces MiJaBench, a bilingual adversarial benchmark designed to evaluate safety alignment in large language models (LLMs) across 16 minority groups, revealing significant disparities in defense rates that can vary by up to 42% based on demographic targeting. The study critiques current safety evaluations for their failure to protect underrepresented communities and demonstrates that existing alignment methods favor specific populations, leading to a "Selective Safety Trap." By employing targeted direct preference optimization (DPO) on a 1B-parameter model, the authors achieve improved zero-shot safety generalizations for previously untested demographics, providing datasets and scripts to foster equitable safety alignment practices in LLM development.