FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing

Jiahao Chen; Zhiyong Ma; Wenbiao Du; Qingyuan Chuai

arXiv:2508.16230·cs.CV·August 25, 2025

FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing

Jiahao Chen, Zhiyong Ma, Wenbiao Du, Qingyuan Chuai

PDF

TL;DR

FlexMUSE is a novel multimodal framework for creative writing that unifies textual and visual semantics, enabling flexible, interactive, and semantically consistent illustrated article generation.

Contribution

It introduces a flexible multimodal unification framework with semantic alignment, attention-based fusion, and a new dataset for creative writing tasks.

Findings

01

Demonstrates semantic consistency and coherence in generated articles.

02

Enhances creative writing with flexible multimodal interactions.

03

Achieves promising results in multimodal creative article generation.

Abstract

Multi-modal creative writing (MMCW) aims to produce illustrated articles. Unlike common multi-modal generative (MMG) tasks such as storytelling or caption generation, MMCW is an entirely new and more abstract challenge where textual and visual contexts are not strictly related to each other. Existing methods for related tasks can be forcibly migrated to this track, but they require specific modality inputs or costly training, and often suffer from semantic inconsistencies between modalities. Therefore, the main challenge lies in economically performing MMCW with flexible interactive patterns, where the semantics between the modalities of the output are more aligned. In this work, we propose FlexMUSE with a T2I module to enable optional visual input. FlexMUSE promotes creativity and emphasizes the unification between modalities by proposing the modality semantic alignment gating…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.