SAT-Cap, a single-stage transformer with spatial-channel attention and cosine-similarity fusion, achieves state-of-the-art CIDEr scores of 140.23% on LEVIR-CC and 97.74% on DUBAI-CCD for remote sensing change captioning.
Pixel-Level Change Detection Pseudo-Label Learning for Remote Sensing Change Captioning
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The existing methods for Remote Sensing Image Change Captioning (RSICC) perform well in simple scenes but exhibit poorer performance in complex scenes. This limitation is primarily attributed to the model's constrained visual ability to distinguish and locate changes. Acknowledging the inherent correlation between change detection (CD) and RSICC tasks, we believe pixel-level CD is significant for describing the differences between images through language. Regrettably, the current RSICC dataset lacks readily available pixel-level CD labels. To address this deficiency, we leverage a model trained on existing CD datasets to derive CD pseudo-labels. We propose an innovative network with an auxiliary CD branch, supervised by pseudo-labels. Furthermore, a semantic fusion augment (SFA) module is proposed to fuse the feature information extracted by the CD branch, thereby facilitating the nuanced description of changes. Experiments demonstrate that our method achieves state-of-the-art performance and validate that learning pixel-level CD pseudo-labels significantly contributes to change captioning. Our code will be available at: https://github.com/Chen-Yang-Liu/Pix4Cap
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Change Captioning in Remote Sensing: Evolution to SAT-Cap -- A Single-Stage Transformer Approach
SAT-Cap, a single-stage transformer with spatial-channel attention and cosine-similarity fusion, achieves state-of-the-art CIDEr scores of 140.23% on LEVIR-CC and 97.74% on DUBAI-CCD for remote sensing change captioning.