GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Hosam Elgendy; Ahmed Sharshar; Ahmed Aboeitta; Yasser Ashraf; Mohsen Guizani

arXiv:2410.19552·cs.CV·May 23, 2025

GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Hosam Elgendy, Ahmed Sharshar, Ahmed Aboeitta, Yasser Ashraf, Mohsen Guizani

PDF

Open Access 1 Repo

TL;DR

GeoLLaVA introduces an efficient fine-tuning approach for vision-language models to accurately detect and describe temporal changes in remote sensing imagery, aiding environmental and urban planning.

Contribution

The paper presents a novel dataset and applies advanced fine-tuning techniques to improve vision-language models' ability to analyze temporal geographical changes.

Findings

01

Achieved a BERT score of 0.864 in change detection.

02

Improved ROUGE-1 score to 0.576 for land-use descriptions.

03

Enhanced model performance on remote sensing temporal change tasks.

Abstract

Detecting temporal changes in geographical landscapes is critical for applications like environmental monitoring and urban planning. While remote sensing data is abundant, existing vision-language models (VLMs) often fail to capture temporal dynamics effectively. This paper addresses these limitations by introducing an annotated dataset of video frame pairs to track evolving geographical patterns over time. Using fine-tuning techniques like Low-Rank Adaptation (LoRA), quantized LoRA (QLoRA), and model pruning on models such as Video-LLaVA and LLaVA-NeXT-Video, we significantly enhance VLM performance in processing remote sensing temporal changes. Results show significant improvements, with the best performance achieving a BERT score of 0.864 and ROUGE-1 score of 0.576, demonstrating superior accuracy in describing land-use transformations.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

HosamGen/GeoLLaVA
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsGeographic Information Systems Studies

MethodsRefunds@Expedia|||How do I get a full refund from Expedia? · Attention Is All You Need · Linear Layer · Dropout · Dense Connections · Weight Decay · Layer Normalization · Pruning · Residual Connection · Linear Warmup With Linear Decay