VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking

Boyue Xu; Ruichao Hou; Tongwei Ren; Gangshan Wu

arXiv:2605.04574·cs.CV·May 7, 2026

VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking

Boyue Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu

PDF

1 Repo

TL;DR

VL-UniTrack introduces a unified visual-language framework for UAV-ground tracking, enhancing cross-view feature interaction and reliability through prompts and shared encoding.

Contribution

It proposes a novel unified framework with visual-language prompts and a shared encoder to improve UAV-ground visual tracking performance.

Findings

01

Achieves state-of-the-art results on benchmark datasets.

02

Effectively fuses language and visual features for better correspondence.

03

Regularizes training with a mutual distillation loss.

Abstract

UAV-ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two-stream methods suffer from isolated feature extraction and rely heavily on implicit appearance matching, which struggles to establish reliable correspondence under drastic view differences, leading to tracking unreliability. To address these limitations, we propose VL-UniTrack, a fully unified framework enhanced by visual-language prompts. By encoding features from both views within a single shared encoder, our method breaks the barrier of feature isolation to facilitate sufficient cross-view interaction. To overcome the ambiguity caused by relying solely on appearance matching, we design visual-language geometric prompting module, which fuses language descriptions with visual features to generate learnable prompts. These prompts are then fed into…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

xuboyue1999/VL-UniTrack.git
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.