ai-digest.dev
last updated 4 h ago
MultimodalarXiv cs.AI 34 d ago

Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models

The article presents FOCUS, a training-free and architecture-agnostic method designed to mitigate cross-image information leakage in Large Vision-Language Models (LVLMs) when processing multi-image inputs. By masking all but one image with random noise, FOCUS enables the model to concentrate on a single clear image, leading to improved performance on various multi-image benchmarks and demonstrating generalization to video understanding. This approach is significant for practitioners as it enhances multi-image reasoning capabilities without requiring additional training or changes to the model architecture.

vision-language-modelsinformation-leakagerelevance 0.00 · engagement 0.00
Read at source ↗← all news