ai-digest.dev
last updated 4 h ago
MultimodalarXiv cs.CL 34 d ago

MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos

The MMOU benchmark has been introduced to evaluate multimodal understanding and reasoning in long, complex videos, consisting of 20,000 questions and 11,877 curated videos across diverse domains. It assesses 13 fundamental skill categories requiring integration of visual, audio, and textual signals, with evaluations revealing significant performance gaps: the best closed-source model achieves 64.2% accuracy, while the top open-source model only reaches 46.8%. This benchmark underscores the limitations of current multimodal models in handling omni-modal reasoning over extended content, providing insights into failure modes that practitioners can address in future model development.

benchmarkreasoningvideosrelevance 0.00 · engagement 0.00
Read at source ↗← all news