Safety
Right Knowledge, Wrong Answer: Test-Time Steering for Temporal Fact Conflicts in Open-Weight Language Models
The paper introduces Temporal Attractor Steering (TAS), a novel test-time intervention for addressing Parametric Temporal Conflicts (PTC) in large language models, allowing them to resolve outdated facts without retraining. Evaluating four open-weight models, including Qwen-2.5-1.5B/7B and Mistral-7B-v0.3, TAS demonstrates an answer-flip rate of 0.72-0.85 and resolves 29-57% of PTC cases while maintaining 85-99% accuracy on non-conflict queries. This approach is significant for practitioners as it provides a method to enhance the accuracy of language models in dynamic knowledge environments without the need for extensive retraining or external data retrieval.
llmtest-time steeringfact conflicts