WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories

Yisu Zhang; Chenjie Cao; Tengfei Wang; Xuhui Zuo; Junta Wu; Jianke Zhu; Chunchao Guo

arXiv:2603.02049·cs.CV·March 3, 2026

WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories

Yisu Zhang, Chenjie Cao, Tengfei Wang, Xuhui Zuo, Junta Wu, Jianke Zhu, Chunchao Guo

PDF

Open Access

TL;DR

WorldStereo introduces a novel framework that combines camera-guided video generation with 3D scene reconstruction using geometric memory modules, enabling multi-view consistency and high-quality 3D outputs from diffusion models.

Contribution

The paper presents a new approach that bridges video generation and 3D reconstruction through geometric memories, improving camera control and multi-view consistency without joint training.

Findings

01

Enables precise camera control and consistent multi-view video generation.

02

Achieves high-fidelity 3D scene reconstruction from diffusion-based videos.

03

Demonstrates effectiveness across diverse scene generation tasks.

Abstract

Recent advances in foundational Video Diffusion Models (VDMs) have yielded significant progress. Yet, despite the remarkable visual quality of generated videos, reconstructing consistent 3D scenes from these outputs remains challenging, due to limited camera controllability and inconsistent generated content when viewed from distinct camera trajectories. In this paper, we propose WorldStereo, a novel framework that bridges camera-guided video generation and 3D reconstruction via two dedicated geometric memory modules. Formally, the global-geometric memory enables precise camera control while injecting coarse structural priors through incrementally updated point clouds. Moreover, the spatial-stereo memory constrains the model's attention receptive fields with 3D correspondence to focus on fine-grained details from the memory bank. These components enable WorldStereo to generate…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Vision and Imaging · 3D Shape Modeling and Analysis · Generative Adversarial Networks and Image Synthesis