IGD: Instructional Graphic Design with Multimodal Layer Generation

Yadong Qu; Shancheng Fang; Yuxin Wang; Xiaorui Wang; Zhineng Chen; Hongtao Xie; Yongdong Zhang

arXiv:2507.09910·cs.CV·July 15, 2025

IGD: Instructional Graphic Design with Multimodal Layer Generation

Yadong Qu, Shancheng Fang, Yuxin Wang, Xiaorui Wang, Zhineng Chen, Hongtao Xie, Yongdong Zhang

PDF

TL;DR

IGD introduces a novel multimodal layer generation approach for graphic design that uses natural language instructions, combining parametric rendering and diffusion models to produce editable, high-quality design files efficiently.

Contribution

The paper presents a new paradigm for graphic design automation using multimodal understanding, parametric rendering, and diffusion models, enabling scalable and editable design generation.

Findings

01

Outperforms existing methods in design quality and flexibility.

02

Supports end-to-end training for complex graphic tasks.

03

Provides a standardized format for multi-scenario design files.

Abstract

Graphic design visually conveys information and data by creating and combining text, images and graphics. Two-stage methods that rely primarily on layout generation lack creativity and intelligence, making graphic design still labor-intensive. Existing diffusion-based methods generate non-editable graphic design files at image level with poor legibility in visual text rendering, which prevents them from achieving satisfactory and practical automated graphic design. In this paper, we propose Instructional Graphic Designer (IGD) to swiftly generate multimodal layers with editable flexibility with only natural language instructions. IGD adopts a new paradigm that leverages parametric rendering and image asset generation. First, we develop a design platform and establish a standardized format for multi-scenario design files, thus laying the foundation for scaling up data. Second, IGD…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.