Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
This paper introduces VALDI, a systematic framework designed to measure the alignment between the stated values of large language models (LLMs) and their generated dialogues, highlighting the persistent "value-action gap." The study reveals that even with explicit reasoning, LLMs exhibit a phenomenon termed "Pseudo-Deliberation," where reasoning does not lead to aligned actions. Additionally, the authors propose VIVALDI, a multi-agent value auditor that intervenes during generation, aiming to address the identified misalignments across various LLMs. This research is significant for AI practitioners as it underscores the importance of developing methodologies for ensuring value alignment in LLM outputs, which is critical for ethical AI deployment.