3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding

Yiping Chen; Jinpeng Li; Wenyu Ke; Yang Luo; Jie Ouyang; Zhongjie He; Li Liu; Hongchao Fan; Hao Wu

arXiv:2603.23447·cs.CV·March 25, 2026

3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding

Yiping Chen, Jinpeng Li, Wenyu Ke, Yang Luo, Jie Ouyang, Zhongjie He, Li Liu, Hongchao Fan, Hao Wu

PDF

Open Access

TL;DR

3DCity-LLM introduces a multi-modality large language model tailored for 3D city-scale perception, utilizing a new large dataset and a coarse-to-fine encoding strategy to enhance urban scene understanding.

Contribution

The paper presents a novel framework and dataset for 3D city-scale vision-language tasks, advancing large language models' capabilities in urban environment perception.

Findings

01

Outperforms existing state-of-the-art methods on benchmarks.

02

Provides a large, high-quality dataset with diverse urban scenarios.

03

Demonstrates improved spatial reasoning and urban understanding.

Abstract

While multi-modality large language models excel in object-centric or indoor scenarios, scaling them to 3D city-scale environments remains a formidable challenge. To bridge this gap, we propose 3DCity-LLM, a unified framework designed for 3D city-scale vision-language perception and understanding. 3DCity-LLM employs a coarse-to-fine feature encoding strategy comprising three parallel branches for target object, inter-object relationship, and global scene. To facilitate large-scale training, we introduce 3DCity-LLM-1.2M dataset that comprises approximately 1.2 million high-quality samples across seven representative task categories, ranging from fine-grained object analysis to multi-faceted scene planning. This strictly quality-controlled dataset integrates explicit 3D numerical information and diverse user-oriented simulations, enriching the question-answering diversity and realism of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Advanced Neural Network Applications · Robotics and Sensor-Based Localization