T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Dongzhi Jiang; Ziyu Guo; Renrui Zhang; Zhuofan Zong; Hao Li; Le Zhuo; Shilin Yan; Pheng-Ann Heng; Hongsheng Li

arXiv:2505.00703·cs.CV·July 2, 2025

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Dongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong, Hao Li, Le Zhuo, Shilin Yan, Pheng-Ann Heng, Hongsheng Li

PDF

2 Repos 2 Models

TL;DR

T2I-R1 introduces a novel reasoning-enhanced text-to-image generation model that employs bi-level chain-of-thought reasoning and reinforcement learning to improve image quality and coherence, surpassing existing models on multiple benchmarks.

Contribution

The paper proposes a new bi-level CoT reasoning framework with reinforcement learning for text-to-image generation, integrating semantic and token-level reasoning to enhance performance.

Findings

01

Achieved 13% improvement on T2I-CompBench

02

Achieved 19% improvement on WISE benchmark

03

Surpassed state-of-the-art model FLUX

Abstract

Recent advancements in large language models have demonstrated how chain-of-thought (CoT) and reinforcement learning (RL) can improve performance. However, applying such reasoning strategies to the visual generation domain remains largely unexplored. In this paper, we present T2I-R1, a novel reasoning-enhanced text-to-image generation model, powered by RL with a bi-level CoT reasoning process. Specifically, we identify two levels of CoT that can be utilized to enhance different stages of generation: (1) the semantic-level CoT for high-level planning of the prompt and (2) the token-level CoT for low-level pixel processing during patch-by-patch generation. To better coordinate these two levels of CoT, we introduce BiCoT-GRPO with an ensemble of generation rewards, which seamlessly optimizes both generation CoTs within the same training step. By applying our reasoning strategies to the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Models

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.