KV-Tracker: Real-Time Pose Tracking with Transformers

Marwan Taher; Ignacio Alzugaray; Kirill Mazur; Xin Kong; Andrew J. Davison

arXiv:2512.22581·cs.CV·December 30, 2025

KV-Tracker: Real-Time Pose Tracking with Transformers

Marwan Taher, Ignacio Alzugaray, Kirill Mazur, Xin Kong, Andrew J. Davison

PDF

Open Access

TL;DR

KV-Tracker enables real-time 6-DoF pose tracking and scene reconstruction from monocular RGB videos by caching global self-attention key-value pairs, achieving significant speedups without sacrificing accuracy.

Contribution

It introduces a novel caching strategy for multi-view networks that allows real-time pose tracking and reconstruction without retraining or depth data.

Findings

01

Achieves up to 27 FPS in experiments.

02

Maintains accuracy without drift or forgetting.

03

Applicable to various multi-view networks.

Abstract

Multi-view 3D geometry networks offer a powerful prior but are prohibitively slow for real-time applications. We propose a novel way to adapt them for online use, enabling real-time 6-DoF pose tracking and online reconstruction of objects and scenes from monocular RGB videos. Our method rapidly selects and manages a set of images as keyframes to map a scene or object via $π^{3}$ with full bidirectional attention. We then cache the global self-attention block's key-value (KV) pairs and use them as the sole scene representation for online tracking. This allows for up to $15 \times$ speedup during inference without the fear of drift or catastrophic forgetting. Our caching strategy is model-agnostic and can be applied to other off-the-shelf multi-view networks without retraining. We demonstrate KV-Tracker on both scene-level tracking and the more challenging task of on-the-fly object…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsVideo Surveillance and Tracking Methods · Human Pose and Action Recognition · 3D Shape Modeling and Analysis