ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

Jindi Lv; Hao Li; Jie Li; Yifei Nie; Fankun Kong; Yang Wang; Xiaofeng Wang; Zheng Zhu; Chaojun Ni; Qiuping Deng; Hengtao Li; Jiancheng Lv; Guan Huang

arXiv:2604.08168·cs.RO·April 10, 2026

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

Jindi Lv, Hao Li, Jie Li, Yifei Nie, Fankun Kong, Yang Wang, Xiaofeng Wang, Zheng Zhu, Chaojun Ni, Qiuping Deng, Hengtao Li, Jiancheng Lv, Guan Huang

PDF

1 Repo

TL;DR

ViVa introduces a novel video-generative value model for robot reinforcement learning that leverages pretrained video generators to improve value estimation and generalize across tasks.

Contribution

It repurposes a pretrained video generator for value estimation, capturing temporal dynamics and improving real-world robot task performance.

Findings

01

ViVa improves value estimation accuracy in real-world box assembly tasks.

02

Qualitative analysis shows ViVa produces more reliable signals reflecting task progress.

03

ViVa generalizes to novel objects by leveraging spatiotemporal priors from video corpora.

Abstract

Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due to partial observability and delayed feedback. Reinforcement learning addresses this via value functions, which assess task progress and guide policy improvement. However, existing value models built on vision-language models (VLMs) struggle to capture temporal dynamics, undermining reliable value estimation in long-horizon tasks. In this paper, we propose ViVa, a video-generative value model that repurposes a pretrained video generator for value estimation. Taking the current observation and robot proprioception as input, ViVa jointly predicts future proprioception and a scalar value for the current state. By leveraging the spatiotemporal priors of a pretrained video generator, our approach grounds value estimation in anticipated…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

gigaai-research/ViVa
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.