ai-digest.dev
last updated 4 h ago
SafetyarXiv cs.CL 34 d ago

Are Language Models Sensitive to Morally Irrelevant Distractors?

This study presents a novel multimodal dataset of 60 "moral distractors" to evaluate the sensitivity of large language models (LLMs) to irrelevant situational factors in moral decision-making. The findings reveal that the introduction of these distractors can alter LLMs' moral judgments by over 30% in unambiguous scenarios, indicating that LLMs may not exhibit stable moral preferences. This highlights the necessity for more sophisticated approaches to AI alignment that account for contextual influences on model behavior, which is critical for practitioners developing LLMs for high-stakes applications.

moral biaseslarge language modelsrelevance 0.00 · engagement 0.00
Read at source ↗← all news