From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

Bohan Li; Shuojue Yang; Baorui Peng; Xianda Guo; Erli Zhang; Youqi Tao; Junfeng Duan; Daguang Xu; Qi Dou; Xin Jin; Wenjun Zeng; Hao Zhao; Yueming Jin

arXiv:2605.08712·cs.CV·May 12, 2026

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo, Erli Zhang, Youqi Tao, Junfeng Duan, Daguang Xu, Qi Dou, Xin Jin, Wenjun Zeng, Hao Zhao, Yueming Jin

PDF

TL;DR

This paper introduces a hierarchical routing framework for action-conditioned surgical video generation, converting articulated kinematics into control modalities and improving efficiency, stability, and realism.

Contribution

It proposes a novel kinematic-to-visual lifting paradigm and a hierarchically routed control framework with routing loss functions and a budgeted scheme, advancing surgical video synthesis.

Findings

01

Improves action faithfulness and visual fidelity in generated videos.

02

Achieves significant latency reduction with maintained control accuracy.

03

Provides a new benchmark with articulated annotations for surgical videos.

Abstract

Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we propose a kinematic-to-visual lifting paradigm that converts articulated kinematics into a unified set of five image-aligned control modalities. Building on this representation, we introduce a hierarchically routed visual control framework that selectively activates the most relevant control modalities and motion scales. Instead of uniformly applying all control signals, our model performs hierarchical routing to dynamically allocate conditioning capacity. We further design kinematic-prior-guided routing loss functions to ensure physically meaningful, temporally stable, and efficient expert utilization. To improve efficiency, we propose a budgeted…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.