Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Zhang, Ruohong; Gui, Liangke; Sun, Zhiqing; Feng, Yihao; Xu, Keyang; Zhang, Yuanhan; Fu, Di; Li, Chunyuan; Hauptmann, Alexander; Bisk, Yonatan; Yang, Yiming

Computer Science > Computer Vision and Pattern Recognition

arXiv:2404.01258 (cs)

[Submitted on 1 Apr 2024 (v1), last revised 2 Apr 2024 (this version, v2)]

Title:Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Authors:Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander Hauptmann, Yonatan Bisk, Yiming Yang

View PDF HTML (experimental)

Abstract:Preference modeling techniques, such as direct preference optimization (DPO), has shown effective in enhancing the generalization abilities of large language model (LLM). However, in tasks involving video instruction-following, providing informative feedback, especially for detecting hallucinations in generated responses, remains a significant challenge. Previous studies have explored using large large multimodal models (LMMs) as reward models to guide preference modeling, but their ability to accurately assess the factuality of generated responses compared to corresponding videos has not been conclusively established. This paper introduces a novel framework that utilizes detailed video captions as a proxy of video content, enabling language models to incorporate this information as supporting evidence for scoring video Question Answering (QA) predictions. Our approach demonstrates robust alignment with OpenAI GPT-4V model's reward mechanism, which directly takes video frames as input. Furthermore, we show that applying this tailored reward through DPO significantly improves the performance of video LMMs on video QA tasks.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2404.01258 [cs.CV]
	(or arXiv:2404.01258v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2404.01258

Submission history

From: Ruohong Zhang [view email]
[v1] Mon, 1 Apr 2024 17:28:16 UTC (3,439 KB)
[v2] Tue, 2 Apr 2024 12:47:49 UTC (3,427 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators