Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling

Yuran Wang; Bohan Zeng; Chengzhuo Tong; Wenxuan Liu; Yang Shi; Xiaochen Ma; Hao Liang; Yuanxing Zhang; Wentao Zhang

arXiv:2512.12675·cs.CV·April 14, 2026

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling

Yuran Wang, Bohan Zeng, Chengzhuo Tong, Wenxuan Liu, Yang Shi, Xiaochen Ma, Hao Liang, Yuanxing Zhang, Wentao Zhang

PDF

1 Repo 1 Models 2 Datasets

TL;DR

Scone is a unified model that improves subject-driven image generation by effectively combining composition and distinction, enabling accurate multi-subject rendering in complex scenes.

Contribution

The paper introduces Scone, a novel understanding-generation framework that jointly models composition and distinction, along with a new benchmark SconeEval for evaluation.

Findings

01

Scone outperforms existing models in composition tasks.

02

Scone effectively distinguishes multiple subjects in complex images.

03

The approach enhances semantic alignment and subject identity preservation.

Abstract

Subject-driven image generation has advanced from single- to multi-subject composition, while neglecting distinction, the ability to distinguish and generate the correct subject when inputs contain multiple candidates. This limitation restricts effectiveness in complex, realistic visual settings. We propose Scone, a unified understanding-generation method that integrates composition and distinction. Scone enables the understanding expert to act as a semantic bridge, conveying semantic information and guiding the generation expert to preserve subject identity while minimizing interference. A two-stage training scheme first learns composition, then enhances distinction through semantic alignment and attention-based masking. We also introduce SconeEval, a benchmark for evaluating both composition and distinction across diverse scenarios. Experiments demonstrate that Scone outperforms…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Ryann-Ran/Scone
github

Models

🤗
Ryann829/Scone
model· 5 dl· ♡ 4
5 dl♡ 4

Datasets

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.