FrozenRecon: Pose-free 3D Scene Reconstruction with Frozen Depth Models
Guangkai Xu, Wei Yin, Hao Chen, Chunhua Shen, Kai Cheng, Feng Zhao

TL;DR
This paper introduces a test-time optimization method that leverages pre-trained affine-invariant depth models to achieve robust, pose-free 3D scene reconstruction across diverse datasets, with minimal parameter tuning per video.
Contribution
It proposes a novel approach that freezes large-scale depth models and optimizes only a few parameters for consistent 3D reconstruction without requiring camera poses.
Findings
Achieves state-of-the-art results on five zero-shot datasets.
Ensures inter-frame depth consistency in challenging scenes.
Operates with only dozens of parameters per frame.
Abstract
3D scene reconstruction is a long-standing vision task. Existing approaches can be categorized into geometry-based and learning-based methods. The former leverages multi-view geometry but can face catastrophic failures due to the reliance on accurate pixel correspondence across views. The latter was proffered to mitigate these issues by learning 2D or 3D representation directly. However, without a large-scale video or 3D training data, it can hardly generalize to diverse real-world scenarios due to the presence of tens of millions or even billions of optimization parameters in the deep network. Recently, robust monocular depth estimation models trained with large-scale datasets have been proven to possess weak 3D geometry prior, but they are insufficient for reconstruction due to the unknown camera parameters, the affine-invariant property, and inter-frame inconsistency. Here, we…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsAdvanced Vision and Imaging · Optical measurement and interference techniques · Image Processing Techniques and Applications
