VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

Jipeng Zhang; Kehao Miao; Renjie Pi; Zhaowei Wang; Runtao Liu; Rui Pan; Tong Zhang

arXiv:2506.13888·cs.CL·June 18, 2025

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

Jipeng Zhang, Kehao Miao, Renjie Pi, Zhaowei Wang, Runtao Liu, Rui Pan, Tong Zhang

PDF

Open Access

TL;DR

VL-GenRM introduces an iterative training framework using vision experts and rationales to improve vision-language reward models, effectively addressing biases and hallucinations for better model alignment.

Contribution

The paper presents a novel iterative training method that leverages vision experts, Chain-of-Thought rationales, and rejection sampling to enhance VL-RMs and mitigate hallucinations.

Findings

01

Improved hallucination detection accuracy.

02

Enhanced multimodal reasoning capabilities.

03

Superior performance on VL-RM benchmarks.

Abstract

Reinforcement Fine-Tuning (RFT) with verifiable rewards has advanced large language models but remains underexplored for Vision-Language (VL) models. The Vision-Language Reward Model (VL-RM) is key to aligning VL models by providing structured feedback, yet training effective VL-RMs faces two major challenges. First, the bootstrapping dilemma arises as high-quality training data depends on already strong VL models, creating a cycle where self-generated supervision reinforces existing biases. Second, modality bias and negative example amplification occur when VL models hallucinate incorrect visual attributes, leading to flawed preference data that further misguides training. To address these issues, we propose an iterative training framework leveraging vision experts, Chain-of-Thought (CoT) rationales, and Margin-based Rejection Sampling. Our approach refines preference datasets,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Natural Language Processing Techniques · Topic Modeling