DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing

Ke Li; Maoliang Li; Jialiang Chen; Jiayu Chen; Zihao Zheng; Shaoqi Wang; Xiang Chen

arXiv:2604.04875·cs.CV·April 7, 2026

DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing

Ke Li, Maoliang Li, Jialiang Chen, Jiayu Chen, Zihao Zheng, Shaoqi Wang, Xiang Chen

PDF

1 Repo

TL;DR

The paper introduces DIRECT, a hierarchical multi-agent framework for automated video mashup creation that improves visual and auditory coherence, outperforming existing methods in quality and alignment.

Contribution

It formulates video mashup as a Multimodal Coherency Satisfaction Problem and proposes a novel hierarchical multi-agent approach with a new benchmark dataset.

Findings

01

DIRECT outperforms state-of-the-art baselines in objective metrics.

02

Extensive experiments show improved visual continuity and auditory alignment.

03

Human evaluations favor DIRECT's mashup quality.

Abstract

Video mashup creation represents a complex video editing paradigm that recomposes existing footage to craft engaging audio-visual experiences, demanding intricate orchestration across semantic, visual, and auditory dimensions and multiple levels. However, existing automated editing frameworks often overlook the cross-level multimodal orchestration to achieve professional-grade fluidity, resulting in disjointed sequences with abrupt visual transitions and musical misalignment. To address this, we formulate video mashup creation as a Multimodal Coherency Satisfaction Problem (MMCSP) and propose the DIRECT framework. Simulating a professional production pipeline, our hierarchical multi-agent framework decomposes the challenge into three cascade levels: the Screenwriter for source-aware global structural anchoring, the Director for instantiating adaptive editing intent and guidance, and the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

AK-DREAM/DIRECT
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.