Camera Artist: A Multi-Agent Framework for Cinematic Language Storytelling Video Generation
Haobo Hu, Qi Mao, Yuanhang Li, and Libiao Jin

TL;DR
Camera Artist is a multi-agent framework designed to generate narrative videos with explicit cinematic language, improving storytelling coherence and filmic quality.
Contribution
It introduces a Cinematography Shot Agent with recursive storyboarding and cinematic language injection, enhancing narrative continuity and expressiveness in video generation.
Findings
Outperforms baselines in narrative consistency
Enhances dynamic expressiveness of generated videos
Improves perceived film quality through cinematic language
Abstract
We propose Camera Artist, a multi-agent framework that models a real-world filmmaking workflow to generate narrative videos with explicit cinematic language. While recent multi-agent systems have made substantial progress in automating filmmaking workflows from scripts to videos, they often lack explicit mechanisms to structure narrative progression across adjacent shots and deliberate use of cinematic language, resulting in fragmented storytelling and limited filmic quality. To address this, Camera Artist builds upon established agentic pipelines and introduces a dedicated Cinematography Shot Agent, which integrates recursive storyboard generation to strengthen shot-to-shot narrative continuity and cinematic language injection to produce more expressive, film-oriented shot designs. Extensive quantitative and qualitative results demonstrate that our approach consistently outperforms…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
