ai-digest.dev
last updated 4 h ago
SafetyarXiv cs.AI 34 d ago

The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models

The article presents a systematic investigation into the "chameleon behavior" of search-enabled large language models (LLMs), revealing their tendency to shift stances during multi-turn conversations. Utilizing the Chameleon Benchmark Dataset, which includes 17,770 question-answer pairs across 1,180 conversations, the study introduces two metrics: the Chameleon Score and Source Re-use Rate, to measure stance instability and knowledge diversity, respectively. Evaluations of models like Llama-4-Maverick, GPT-4o-mini, and Gemini-2.5-Flash indicate significant stance instability (scores ranging from 0.391 to 0.511), particularly highlighting the implications for deploying LLMs in critical fields such as healthcare and finance, where consistent responses are essential.

llmstance-instabilitychameleon-behaviorrelevance 0.00 · engagement 0.00
Read at source ↗← all news