X2Localizer Logo X²Localizer:
Cross-grained Alignment for
Progressive Cross-view Video Geo-localization

BMVC 2026 ✨Oral
1. University College London   2. Karlsruhe Institute of Technology   3. Hunan University
4. University of Alberta   5. Shenzhen University
† Corresponding authors

Abstract

Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on fixed-length inputs and post-hoc refinement, hindering online-oriented localization under partial or dynamic observations. In this work, we formulate Progressive Cross-view Video Geo-localization (PCVG) as a deployment-oriented extension and evaluation protocol of CVG, enabling localization under varying temporal budgets, prefix-based inference, random-start evaluation, and long-range localization with interruptions. To explore PCVG, we introduce X²Localizer, a cross-grained alignment framework that jointly supervises global prefix-to-aerial retrieval and token-aggregated frame--aerial-tile matching with a budget-dependent asymmetric objective. Furthermore, we introduce a Sliding-Window Re-Localization (SWRL) strategy that dynamically refreshes candidate regions for failure recovery and long-range deployment without full-sequence reprocessing. Extensive experiments show that X²Localizer preserves conventional full-video performance, with marginal gains of +0.1 Recall@1 and +0.3 Recall@10, while substantially improving early localization. In the challenging single-frame setting, X²Localizer improves coarse retrieval by +4.7 Recall@1 and +11.5 Recall@10 over the previous state-of-the-art method. With SWRL, our approach further enables robust progressive localization under random-start and long-distance scenarios, narrowing the gap between benchmark evaluation and real-world deployment.

🥰 Acknowledgement

X²Localizer is built upon prior works including GAReT, GAMa, TransGeo, and X-CLIP. We sincerely thank the authors for making their work publicly available.

📃 BibTeX

@article{zeng2026x,
                  title={X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization}, 
                  author={Zeng, Zichao and Fan, Weijia and Chen, Yufan and Goo, June Moh and Zheng, Junwei and Liu, Ruiping and Peng, Kunyu and Zhang, Jiaming and Stiefelhagen, Rainer and Boehm, Jan},
                  journal={arXiv preprint arXiv:2608.16658},
                  year={2026}
                }
}