TL;DR
SMART-Editor is a multi-agent framework that enables human-like, globally coherent design editing across structured and unstructured visual domains, using reward-guided refinement and preference optimization.
Contribution
It introduces SMART-Editor, a novel multi-agent framework with reward-guided strategies for global coherence in design editing, along with a new benchmark for evaluation.
Findings
Outperforms baselines like InstructPix2Pix and HIVE.
RewardDPO achieves up to 15% improvement in structured editing.
Reward-Refine enhances natural image editing quality.
Abstract
We present SMART-Editor, a framework for compositional layout and content editing across structured (posters, websites) and unstructured (natural images) domains. Unlike prior models that perform local edits, SMART-Editor preserves global coherence through two strategies: Reward-Refine, an inference-time rewardguided refinement method, and RewardDPO, a training-time preference optimization approach using reward-aligned layout pairs. To evaluate model performance, we introduce SMARTEdit-Bench, a benchmark covering multi-domain, cascading edit scenarios. SMART-Editor outperforms strong baselines like InstructPix2Pix and HIVE, with RewardDPO achieving up to 15% gains in structured settings and Reward-Refine showing advantages on natural images. Automatic and human evaluations confirm the value of reward-guided planning in producing semantically consistent and visually aligned edits.
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
