Training
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
Oracle-RLAIF is a new fine-tuning framework for multi-modal video models that enhances video comprehension by employing reinforcement learning from ranking feedback instead of traditional supervised fine-tuning. The framework introduces an Oracle ranker that ranks model responses, replacing the costly human feedback process, and utilizes a novel rank-based loss function, $GRPO_{rank}$, for optimizing ordinal feedback. This approach demonstrates improved performance over existing video-language models across various benchmarks, offering a more cost-effective and flexible method for aligning large-scale models.
fine-tuningvideo-models