Safety
The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
The paper introduces the concept of the "alignment veto," demonstrating that alignment training in language models can suppress cultural knowledge rather than erase it, based on a study involving 26 models across 16 MENA countries and 1.53 million human survey responses. It highlights that suppression failures occur when accurate internal distributions are blocked at the output, leading to a significant alignment-quality gap of 19.8% between the best- and worst-served nations, with a safety tax reaching 37.6%. The findings emphasize the need for different interventions to address suppression and representational bias, as well as the implications of how alignment decisions impact diverse cultural contexts.
alignmentllmcultural knowledge