SurgLQA: Scalable Long-Horizon Surgical Video Question Answering

Diandian Guo; Xikai Yang; Ruiyang Li; Jialun Pei; Pheng-Ann Heng

arXiv:2605.17915·cs.CV·May 19, 2026

SurgLQA: Scalable Long-Horizon Surgical Video Question Answering

Diandian Guo, Xikai Yang, Ruiyang Li, Jialun Pei, Pheng-Ann Heng

PDF

1 Repo

TL;DR

SurgLQA introduces a scalable framework for long-horizon surgical video question answering, enabling better modeling of extended procedural dynamics and causal dependencies in intraoperative settings.

Contribution

It proposes novel methods for long-range temporal reasoning and adaptive inference, and restructures a benchmark dataset for systematic evaluation.

Findings

01

Achieves consistent performance improvements in long-range reasoning tasks.

02

Demonstrates effectiveness on Colon-LQA and REAL-Colon-VQA datasets.

Abstract

Surgical Video Question Answering (VideoQA) provides a promising paradigm for dynamic intraoperative interpretation, enabling real-time decision support and context-aware retrieval in clinical environments. Nevertheless, existing approaches are predominantly restricted to images or short clips, limiting their ability to model long-range procedural dynamics and causal dependencies across extended surgical workflows. To address this challenge, we propose SurgLQA, a unified long-horizon VideoQA framework for scalable surgical reasoning. This framework incorporates Faithful Temporal Consolidation (FTC), which leverages intrinsic temporal cues to construct compact long-range representations while preserving fine-grained temporal fidelity. Further, we develop Temporally-Grounded Multi-Policy Scaling (TMS), an adaptive test-time inference paradigm that strategically adjusts policy-level…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

RascalGdd/SurgLQA
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.