Pith. sign in

REVIEW 3 cited by

GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19552 v2 pith:RIL5RORU submitted 2024-10-25 cs.CV

classification cs.CV
keywords temporalmodelsremotesensingchangesgeographicallikelora
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Detecting temporal changes in geographical landscapes is critical for applications like environmental monitoring and urban planning. While remote sensing data is abundant, existing vision-language models (VLMs) often fail to capture temporal dynamics effectively. This paper addresses these limitations by introducing an annotated dataset of video frame pairs to track evolving geographical patterns over time. Using fine-tuning techniques like Low-Rank Adaptation (LoRA), quantized LoRA (QLoRA), and model pruning on models such as Video-LLaVA and LLaVA-NeXT-Video, we significantly enhance VLM performance in processing remote sensing temporal changes. Results show significant improvements, with the best performance achieving a BERT score of 0.864 and ROUGE-1 score of 0.576, demonstrating superior accuracy in describing land-use transformations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    VectorLLM, a multimodal LLM that regresses building contour vertices token by token, reports gains of 5.6 to 13.6 AP over prior polygon extraction methods on WHU, WHU-Mix, and CrowdAI.

  2. Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A vision-language model fine-tuned on a new synthetic slide-animation dataset outperforms GPT-4.1 and Gemini-2.5-Pro at describing slide animations, especially on synthetic evaluation data.

  3. Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A remote sensing LVLM that augments visual features with retrieved captions and routes them through level-specific experts improves performance on several RS vision-language benchmarks.

Pith tools