CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration

Xuecong Liu; Mengzhu Ding; Zixuan Sun; Zhang Li; Xichao Teng

arXiv:2604.05689·cs.CV·April 8, 2026

CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration

Xuecong Liu, Mengzhu Ding, Zixuan Sun, Zhang Li, Xichao Teng

PDF

1 Repo

TL;DR

CRFT is a transformer-based framework that achieves robust cross-modal image registration by learning a modality-independent feature flow through a coarse-to-fine, iterative refinement process.

Contribution

It introduces a unified coarse-to-fine transformer architecture with an iterative discrepancy-guided attention mechanism for improved cross-modal registration.

Findings

01

CRFT outperforms state-of-the-art methods in accuracy and robustness.

02

It effectively handles large affine and scale variations.

03

The approach is applicable to remote sensing, navigation, and medical imaging.

Abstract

We present Consistent-Recurrent Feature Flow Transformer (CRFT), a unified coarse-to-fine framework based on feature flow learning for robust cross-modal image registration. CRFT learns a modality-independent feature flow representation within a transformer-based architecture that jointly performs feature alignment and flow estimation. The coarse stage establishes global correspondences through multi-scale feature correlation, while the fine stage refines local details via hierarchical feature fusion and adaptive spatial reasoning. To enhance geometric adaptability, an iterative discrepancy-guided attention mechanism with a Spatial Geometric Transform (SGT) recurrently refines the flow field, progressively capturing subtle spatial inconsistencies and enforcing feature-level consistency. This design enables accurate alignment under large affine and scale variations while maintaining…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

NEU-Liuxuecong/CRFT
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.