DORSal: Diffusion for Object-centric Representations of Scenes et al

Allan Jabri; Sjoerd van Steenkiste; Emiel Hoogeboom; Mehdi S. M.; Sajjadi; Thomas Kipf

arXiv:2306.08068·cs.CV·May 6, 2024·1 cites

DORSal: Diffusion for Object-centric Representations of Scenes et al

Allan Jabri, Sjoerd van Steenkiste, Emiel Hoogeboom, Mehdi S. M., Sajjadi, Thomas Kipf

PDF

Open Access

TL;DR

DORSal introduces a diffusion-based approach for high-fidelity 3D scene rendering that maintains object-level editing capabilities, outperforming existing methods on synthetic and real-world datasets.

Contribution

This paper adapts a diffusion architecture for 3D scene generation conditioned on object-centric representations, enabling scalable, high-quality rendering with editing features.

Findings

01

Achieves high-quality novel view synthesis in complex scenes

02

Enables object-level scene editing in 3D representations

03

Outperforms existing approaches on synthetic and real-world datasets

Abstract

Recent progress in 3D scene understanding enables scalable learning of representations across large datasets of diverse scenes. As a consequence, generalization to unseen scenes and objects, rendering novel views from just a single or a handful of input images, and controllable scene generation that supports editing, is now possible. However, training jointly on a large number of scenes typically compromises rendering quality when compared to single-scene optimized models such as NeRFs. In this paper, we leverage recent progress in diffusion models to equip 3D scene representation learning models with the ability to render high-fidelity novel views, while retaining benefits such as object-level scene editing to a large degree. In particular, we propose DORSal, which adapts a video diffusion architecture for 3D scene generation conditioned on frozen object-centric slot-based…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsGenerative Adversarial Networks and Image Synthesis · Domain Adaptation and Few-Shot Learning · Model Reduction and Neural Networks

MethodsDiffusion