Describing Land Cover Changes via Multi-Temporal Remote Sensing Image Captioning Using LLM, ViT, and LoRA

dc.contributor.authorLamar-Leon, Javier
dc.contributor.authorNogueira, Vitor
dc.contributor.authorSalgueiro, Pedro
dc.contributor.authorQuaresma, Paulo
dc.contributor.editorPan, Jiayi
dc.contributor.editorLi, Xinghua
dc.date.accessioned2026-02-23T15:53:07Z
dc.date.available2026-02-23T15:53:07Z
dc.date.issued2026-01-04
dc.description.abstractDescribing land cover changes from multi-temporal remote sensing imagery requires capturing both visual transformations and their semantic meaning in natural language. Existing methods often struggle to balance visual accuracy with descriptive coherence. We propose MVLT-LoRA-CC (Multi-modal Vision Language Transformer with Low-Rank Adaptation for Change Captioning), a framework that integrates a Vision Transformer (ViT), a Large Language Model (LLM), and Low-Rank Adaptation (LoRA) for efficient multi-modal learning. The model processes paired temporal images through patch embeddings and transformer blocks, aligning visual and textual representations via a multi-modal adapter. To improve efficiency and avoid unnecessary parameter growth, LoRA modules are selectively inserted only into the attention projection layers and cross-modal adapter blocks rather than being uniformly applied to all linear layers. This targeted design preserves general linguistic knowledge while enabling effective adaptation to remote sensing change description. To assess performance, we introduce the Complementary Consistency Score (CCS) framework, which evaluates both descriptive fidelity for change instances and classification accuracy for no change cases. Experiments on the LEVIR-CC test set demonstrate that MVLT-LoRA-CC generates semantically accurate captions, surpassing prior methods in both descriptive richness and temporal change recognition. The approach establishes a scalable solution for multi-modal land cover change description in remote sensing applications.por
dc.identifier.authoremailjlamarleon@uevora.pt
dc.identifier.authoremailvbn@uevora.pt
dc.identifier.authoremailpds@uevora.pt
dc.identifier.authoremailpq@uevora.pt
dc.identifier.citationLeón, Javier Lamar, Vitor Nogueira, Pedro Salgueiro, and Paulo Quaresma. 2026. "Describing Land Cover Changes via Multi-Temporal Remote Sensing Image Captioning Using LLM, ViT, and LoRA" Remote Sensing 18, no. 1: 166. https://doi.org/10.3390/rs18010166por
dc.identifier.doihttps://doi.org/10.3390/rs18010166por
dc.identifier.scientificarea283por
dc.identifier.urihttp://hdl.handle.net/10174/41416
dc.language.isoporpor
dc.peerreviewedyespor
dc.publisherRemote Sensing MDPIpor
dc.rightsopenAccesspor
dc.subjectImage Captioningpor
dc.subjectRemote Sensingpor
dc.subjectLLMpor
dc.subjectLoRApor
dc.titleDescribing Land Cover Changes via Multi-Temporal Remote Sensing Image Captioning Using LLM, ViT, and LoRApor
dc.typearticlepor
degois.publication.issue1por
degois.publication.locationRemote Sensing of Coastal Waters, Land Use/Cover, Lakes, Rivers and Watersheds III)por
degois.publication.titleRemote Sensingpor
degois.publication.volume18por

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Describing Land Cover Changes via Multi-Temporal Remote Sensing Image Captioning Using LLM, ViT, and LoRA - remotesensing-18-00166.pdf
Size:
8.2 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
3.89 KB
Format:
Item-specific license agreed upon to submission
Description: