Pith. sign in

REVIEW 4 cited by

Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.09868 v1 pith:D5U2CIR5 submitted 2022-04-21 cs.CV cs.MM

classification cs.CVcs.MM
keywords retrievalimagemulti-scalecross-modalfeaturesremotesensingsimilarity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Remote sensing (RS) cross-modal text-image retrieval has attracted extensive attention for its advantages of flexible input and efficient query. However, traditional methods ignore the characteristics of multi-scale and redundant targets in RS image, leading to the degradation of retrieval accuracy. To cope with the problem of multi-scale scarcity and target redundancy in RS multimodal retrieval task, we come up with a novel asymmetric multimodal feature matching network (AMFMN). Our model adapts to multi-scale feature inputs, favors multi-source retrieval methods, and can dynamically filter redundant features. AMFMN employs the multi-scale visual self-attention (MVSA) module to extract the salient features of RS image and utilizes visual features to guide the text representation. Furthermore, to alleviate the positive samples ambiguity caused by the strong intraclass similarity in RS image, we propose a triplet loss function with dynamic variable margin based on prior similarity of sample pairs. Finally, unlike the traditional RS image-text dataset with coarse text and higher intraclass similarity, we construct a fine-grained and more challenging Remote sensing Image-Text Match dataset (RSITMD), which supports RS image retrieval through keywords and sentence separately and jointly. Experiments on four RS text-image datasets demonstrate that the proposed model can achieve state-of-the-art performance in cross-modal RS text-image retrieval task.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage method (MpGI) produces a 210K-image, 1.26M-caption remote sensing dataset and state-of-the-art CLIP and CoCa models.

  2. TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval

    cs.CV 2025-10 conditional novelty 5.0 of 10

    A new drone video-text retrieval benchmark DVTMD and a multi-granularity CLIP-based model TCMA reach 45.5% R@1 text-to-video on DVTMD, but CapERA results do not consistently beat prior methods.

  3. Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A remote sensing LVLM that augments visual features with retrieved captions and routes them through level-specific experts improves performance on several RS vision-language benchmarks.

  4. A Synthetic-to-Real Dehazing Method based on Domain Unification

    eess.IV 2025-09 conditional novelty 4.0 of 10

    The paper derives a composite atmospheric scattering model for non-ideal clean data and uses a four-loss committee to train a dehazing network, claiming state-of-the-art results on real-world haze benchmarks.

Pith tools