Lotus: Diffusion-based Visual Foundation Model for High-quality Dense   Prediction

Jing He; Haodong Li; Wei Yin; Yixun Liang; Leheng Li; Kaiqiang Zhou,; Hongbo Zhang; Bingbing Liu; Ying-Cong Chen

arXiv:2409.18124·cs.CV·January 22, 2025·2 cites

Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction

Jing He, Haodong Li, Wei Yin, Yixun Liang, Leheng Li, Kaiqiang Zhou,, Hongbo Zhang, Bingbing Liu, Ying-Cong Chen

PDF

Open Access 1 Repo 10 Models

TL;DR

Lotus is a diffusion-based model tailored for dense prediction tasks that predicts annotations directly, simplifies the diffusion process, and achieves state-of-the-art zero-shot performance with high efficiency.

Contribution

The paper introduces Lotus, a novel diffusion-based dense prediction model that predicts annotations directly and reformulates the diffusion process into a single step for improved speed and accuracy.

Findings

01

Achieves state-of-the-art zero-shot depth and normal estimation.

02

Significantly faster inference compared to existing diffusion methods.

03

Improves practical applications like 3D reconstruction.

Abstract

Leveraging the visual priors of pre-trained text-to-image diffusion models offers a promising solution to enhance zero-shot generalization in dense prediction tasks. However, existing methods often uncritically use the original diffusion formulation, which may not be optimal due to the fundamental differences between dense prediction and image generation. In this paper, we provide a systemic analysis of the diffusion formulation for the dense prediction, focusing on both quality and efficiency. And we find that the original parameterization type for image generation, which learns to predict noise, is harmful for dense prediction; the multi-step noising/denoising diffusion process is also unnecessary and challenging to optimize. Based on these insights, we introduce Lotus, a diffusion-based visual foundation model with a simple yet effective adaptation protocol for dense prediction.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

envision-research/lotus
pytorch

Models

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsImage Retrieval and Classification Techniques · Video Analysis and Summarization · Machine Learning and Data Classification

MethodsDiffusion