SLayR: Scene Layout Generation with Rectified Flow
Cameron Braunstein, Hevra Petekkaya, Jan Eric Lenssen, Mariya Toneva,, Eddy Ilg

TL;DR
SLayR is a transformer-based model that generates diverse and plausible scene layouts from ambiguous prompts, outperforming baselines and setting new state-of-the-art results in layout plausibility and variety.
Contribution
The paper introduces SLayR, a novel rectified flow-based transformer model for text-to-layout generation, with a new benchmark and evaluation procedure.
Findings
SLayR surpasses existing baselines in layout plausibility and variety.
SLayR achieves state-of-the-art performance with fewer parameters.
The new benchmark effectively evaluates layout generation quality.
Abstract
We introduce SLayR, Scene Layout Generation with Rectified flow, a novel transformer-based model for text-to-layout generation which can then be paired with existing layout-to-image models to produce images. SLayR addresses a domain in which current text-to-image pipelines struggle: generating scene layouts that are of significant variety and plausibility, when the given prompt is ambiguous and does not provide constraints on the scene. SLayR surpasses existing baselines including LLMs in unconstrained generation, and can generate layouts from an open caption set. To accurately evaluate the layout generation, we introduce a new benchmark suite, including numerical metrics and a carefully designed repeatable human-evaluation procedure that assesses the plausibility and variety of generated images. We show that our method sets a new state of the art for achieving both at the same time,…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsComputer Graphics and Visualization Techniques · Advanced Vision and Imaging · Human Motion and Animation
