Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens

Yohann Perron; Vladyslav Sydorov; Christophe Pottier; Loic Landrieu

arXiv:2601.05927·cs.CV·January 12, 2026

Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens

Yohann Perron, Vladyslav Sydorov, Christophe Pottier, Loic Landrieu

PDF

Open Access

TL;DR

This paper introduces a multi-scale vision transformer approach with relay tokens for ultra-high resolution semantic segmentation, effectively combining local detail preservation with global context understanding.

Contribution

It presents a simple method integrating relay tokens into standard transformers to enhance multi-scale reasoning without significant parameter increase.

Findings

01

Achieves up to 15% relative mIoU improvement on benchmarks.

02

Works with standard transformer backbones like ViT and Swin.

03

Requires fewer than 2% additional parameters.

Abstract

Current approaches for segmenting ultra high resolution images either slide a window, thereby discarding global context, or downsample and lose fine detail. We propose a simple yet effective method that brings explicit multi scale reasoning to vision transformers, simultaneously preserving local details and global awareness. Concretely, we process each image in parallel at a local scale (high resolution, small crops) and a global scale (low resolution, large crops), and aggregate and propagate features between the two branches with a small set of learnable relay tokens. The design plugs directly into standard transformer backbones (eg ViT and Swin) and adds fewer than 2 % parameters. Extensive experiments on three ultra high resolution segmentation benchmarks, Archaeoscape, URUR, and Gleason, and on the conventional Cityscapes dataset show consistent gains, with up to 15 % relative mIoU…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Neural Network Applications · Advanced Image and Video Retrieval Techniques · Generative Adversarial Networks and Image Synthesis