REVIEW 3 major objections 5 minor 299 references
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This survey organizes remote-sensing vision-language modeling into three families — contrastive learning, visual instruction tuning, and text-conditioned image generation — and makes datasets a first-class part of its structured review of…
desk verdict A genuinely useful survey of remote sensing VLMs whose comparative tables should not be trusted until the protocol mixing and a few internal inconsistencies are cleaned up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is the three-way taxonomy of vision-language modeling, held together by the two-stage paradigm of pre-training followed by fine-tuning. Within each family the paper identifies a load-bearing mechanism: for contrastive models, the symmetric InfoNCE objective that pulls matched image-text pairs together and pushes unmatched pairs apart in a shared embedding space; for instruction-tuned models, the next-token prediction objective over question-answer pairs, with the connector trained during alignment and the connector plus a low-rank adapter on the language model during fine-tuning; for generative models, the latent-diffusion denoising objective, conditioned on text, metadata such as coordinates and ground sampling distance, and image-form conditions injected through ControlNet-style modules. On the data side, the machinery is a typology of pre-training, instruction-following, and benchmark datasets, analyzed through two construction choices: the image source (open-source image datasets versus public geographic databases) and the caption generation method (manual annotation, rule-based assembly, or prompting a large model).
What would settle it
Re-run the models compared in Tables II, VI, VII, and VIII under one shared protocol — identical prompts, splits, and evaluation code — and check whether the reported gaps, such as S-CLIP and SetCLIP falling more than 20% behind in retrieval or SkySenseGPT's 97.02% zero-shot accuracy on WHU-RS19, survive the re-run; a large reshuffling would overturn the survey's comparative conclusions.
Extended reading notes
Core claim
Stated as the authors would state it: vision-language modeling in remote sensing has, since roughly 2023, matured into a field whose progress is fully captured by the pre-train-then-fine-tune paradigm, and whose models fall cleanly into three families. Contrastive models such as GeoRSCLIP, RemoteCLIP, and SkyCLIP align image and text embeddings in a shared space with InfoNCE-style losses; instruction-tuned models such as GeoChat, LHRS-Bot, and SkySenseGPT connect a frozen vision encoder to a large language model through a connector and learn next-token prediction over instruction-following data; generative models such as DiffusionSat and CRS-Diff learn an implicit image-text joint distribution through latent diffusion, conditioning on text plus metadata and image-form controls. The survey's distinctive move is to treat datasets as first-class objects: it shows that the field's largest pre-training corpus, Git-10M with over ten million pairs, is still two orders of magnitude below the 400-million-pair scale of the original CLIP, and it draws concrete lessons such as long, rich captions mattering more than caption accuracy. It concludes that existing models are far from expert-level and names five open directions: broader cross-modal alignment, comprehension of vague requirements, explanation-driven reliability, continual adaptation, and larger multimodal datasets with more challenging benchmarks.
Load-bearing premise
The comparison tables treat accuracy and retrieval numbers copied from different original papers as directly comparable, and although zero-shot results are marked with daggers, the tables still mix supervised and zero-shot numbers while training data, prompt templates, and evaluation protocols differ across the sources.
Editorial extensions
If this is right
- Any new remote-sensing vision-language model can be classified immediately by which bridge it builds — contrastive alignment, instruction tuning, or generative conditioning — and by which stage of the two-stage paradigm it modifies.
- Dataset construction choices are decisive: the survey reports that a model pre-trained on the noisy, long-caption VersaD outperformed one trained on the cleaner SkyScript, indicating that rich descriptive captions buffer against annotation noise.
- Data-efficient pre-training objectives — pseudo-labels in S-CLIP, distribution matching in Set-CLIP, ground-to-satellite alignment in GRAFT — bring models trained on limited pairs close to large-data models, so more data is not the only route forward.
- Instruction-tuned models now cover region- and point-level understanding, time-series change analysis, quantitative counting, and honest refusal of unanswerable questions, expanding what counts as a useful remote-sensing assistant.
- Generative foundation models have moved beyond image synthesis into captioning, pansharpening, cloud removal, zero-shot SAR target recognition, and urban prediction, making the image-text joint distribution a general-purpose tool.
Reading between the lines
- A standardization problem follows from the survey's own evidence, though the paper does not fully solve it: rebuilding the compared models under one shared harness with identical prompts, splits, and protocols would likely reshuffle several reported rankings, including the large retrieval gap attributed to limited-pair pre-training.
- The documented trend toward multi-sensor (optical, SAR, infrared) and cross-view understanding points to a missing piece the paper leaves implicit: a unified benchmark that evaluates the same model across sensors, resolutions, and viewing angles rather than within single-sensor datasets.
- The vague-requirement direction implies an agentic testbed: instruction data constructed by chaining subtasks — decomposing 'assess the water quality in this area' into water detection followed by quantitative retrieval — would let models learn decomposition rather than simple classification.
- The honesty dataset HnstD suggests a cheap, generalizable evaluation axis: refusal accuracy on unanswerable queries could be reported alongside captioning and retrieval metrics for every instruction-tuned remote-sensing model.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a survey of vision-language modeling (VLM) in remote sensing, organized around the two-stage pre-training and fine-tuning paradigm. It proposes a three-category taxonomy—contrastive learning, visual instruction tuning, and text-conditioned image generation—and for each category reviews architectures, training objectives, and representative works. It also provides a broad catalog of pre-training, instruction-following, and benchmark datasets, describing their construction methods, and concludes with a discussion of future research directions such as cross-modal alignment, explanation-driven reliability, continual adaptation, and larger multimodal datasets.
Significance. If its comparative claims are reliable, this survey would be a valuable structured reference: the taxonomy is clear, the coverage of 2023-2025 works is current, and the dataset catalog (Tables XII-XIV) is more detailed than in earlier surveys. The paper gives due attention to dataset construction methodology, including caption generation pipelines and geographic coverage, which is a genuine strength. However, the quantitative comparisons that support several headline conclusions are currently compromised by heterogeneous evaluation protocols and by internal numerical inconsistencies. The machine-checkable portions of the survey are the tables themselves, and several cells disagree with the text or with their own arithmetic, so the central claim of providing a reliable reference is not yet fully met.
major comments (3)
- [Section III-A, Table II; Section IV-C, Table VII]
- [Section IV-C, Table VIII and Table VII]
- [Section IV-C, Table VIII]
minor comments (5)
- [Section III-B, Prompt-CC paragraph]
- [Section V-B, Wang et al. [155] description]
- [Table XIV, GEOBench-VLM row]
- [Table II footnote]
- [Table XIII, RSVP-3M row]
Circularity Check
Survey is organizational; no derivation chain reduces to its own inputs, and self-citations are standard independent benchmarks.
full rationale
This paper is a survey, not a derivation. Its central claim is that it provides a timely and comprehensive review of vision-language modeling in remote sensing under the pre-training/fine-tuning paradigm. That claim is supported by the paper's structure itself: a taxonomy (contrastive learning, visual instruction tuning, text-conditioned image generation), summaries of prior work, and dataset catalogs. No equation in the paper is used to derive a conclusion that was assumed in its construction. Equations (1)-(13) are quoted from the surveyed literature as descriptions of pre-training objectives, not used to prove novel results. The comparative tables (II, VI, VII, VIII) compile numbers reported by other papers; even if those numbers are protocol-mixed, that is a correctness/reliability concern, not circularity, because the survey does not fit any parameter to those tables and then call the fit a prediction. Some datasets authored by the survey's own co-author (AID, Million-AID, DOTA) appear as benchmark datasets or dataset sources, but they are used as external, publicly available evaluation resources and are not evidence for a novel claim made by this survey. The self-citations are therefore not load-bearing in any circular sense. No instance of the seven enumerated circularity patterns can be exhibited with a quote showing Eq. X reducing to Eq. Y by construction or a fitted parameter renamed as a prediction. Accordingly, the appropriate finding is no significant circularity, score 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The two-stage pre-training then fine-tuning paradigm is the organizing frame for remote sensing VLM.
- ad hoc to paper The three-category taxonomy (contrastive, instruction tuning, generation) is exhaustive for the surveyed literature.
- domain assumption Accuracy numbers copied from original papers are comparable across tables.
Cite this review
Pith. "Pith review of Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives." pith.science (2026). https://pith.science/paper/AZPTPE52
@misc{pith2026250514361,
author = {Pith},
title = {Pith review of: Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/AZPTPE52}},
note = {Machine review of arXiv:2505.14361}
}
read the original abstract
Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, VLM in the remote sensing domain has made significant progress. The resulting models benefit from the absorption of extensive general knowledge and demonstrate strong performance across a variety of remote sensing data analysis tasks. Moreover, they are capable of interacting with users in a conversational manner. In this paper, we aim to provide the remote sensing community with a timely and comprehensive review of the developments in VLM using the two-stage paradigm. Specifically, we first cover a taxonomy of VLM in remote sensing: contrastive learning, visual instruction tuning, and text-conditioned image generation. For each category, we detail the commonly used network architecture and pre-training objectives. Second, we conduct a thorough review of existing works, examining foundation models and task-specific adaptation methods in contrastive-based VLM, architectural upgrades, training strategies and model capabilities in instruction-based VLM, as well as generative foundation models with their representative downstream applications. Third, we summarize datasets used for VLM pre-training, fine-tuning, and evaluation, with an analysis of their construction methodologies (including image sources and caption generation) and key properties, such as scale and task adaptability. Finally, we conclude this survey with insights and discussions on future research directions: cross-modal representation alignment, vague requirement comprehension, explanation-driven model reliability, continually scalable model capabilities, and large-scale datasets featuring richer modalities and greater challenges.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Rsgpt: A remote sensing vision language model and benchmark,
Y . Hu, J. Yuan, C. Wen, X. Lu, Y . Liu, and X. Li, “Rsgpt: A remote sensing vision language model and benchmark,” ISPRS J. of Photogrammetry Remote Sens. , vol. 224, pp. 272–286, 2025
2025
-
[2]
Remoteclip: A vision language foundation model for remote sensing,
F. Liu, D. Chen, Z. Guan, X. Zhou, J. Zhu, Q. Ye, L. Fu, and J. Zhou, “Remoteclip: A vision language foundation model for remote sensing,” IEEE Trans. Geosci. Remote Sens. , vol. 62, pp. 1–16, 2024
2024
-
[3]
Geochat: Grounded large vision-language model for remote sensing,
K. Kuckreja, M. S. Danish, M. Naseer, A. Das, S. Khan, and F. S. Khan, “Geochat: Grounded large vision-language model for remote sensing,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 27831–27840, 2024
2024
-
[4]
Remote sensing vision-language foundation models without annotations via ground remote alignment,
U. Mall, C. P. Phoo, M. K. Liu, C. V ondrick, B. Hariharan, and K. Bala, “Remote sensing vision-language foundation models without annotations via ground remote alignment,” in Proc. Int. Conf. Learn. Representations (ICLR), pp. 1–13, 2024
2024
-
[5]
Skyeyegpt: Unifying remote sens- ing vision-language tasks via instruction tuning with large language model,
Y . Zhan, Z. Xiong, and Y . Yuan, “Skyeyegpt: Unifying remote sens- ing vision-language tasks via instruction tuning with large language model,” ISPRS J. of Photogrammetry Remote Sens. , vol. 221, pp. 64– 77, 2025
2025
-
[6]
Earthgpt: A universal multimodal large language model for multisensor image com- prehension in remote sensing domain,
W. Zhang, M. Cai, T. Zhang, Y . Zhuang, and X. Mao, “Earthgpt: A universal multimodal large language model for multisensor image com- prehension in remote sensing domain,” IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–20, 2024
2024
-
[7]
Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing,
Z. Zhang, T. Zhao, Y . Guo, and J. Yin, “Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing,” IEEE Trans. Geosci. Remote Sens. , vol. 62, pp. 1–23, 2024
2024
-
[8]
Skyscript: A large and semantically diverse vision-language dataset for remote sensing,
Z. Wang, R. Prabha, T. Huang, J. Wu, and R. Rajagopal, “Skyscript: A large and semantically diverse vision-language dataset for remote sensing,” in Proc. AAAI Conf. Artif. Intell. , vol. 38, pp. 5805–5813, 2024
2024
Show all 299 references
-
[9]
Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,
D. Muhtar, Z. Li, F. Gu, X. Zhang, and P. Xiao, “Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,” in Proc. Eur. Conf. Comput. Vis. (ECCV) , pp. 440–457, 2024
2024
-
[10]
Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,
Y . Bazi, L. Bashmal, M. M. Al Rahhal, R. Ricci, and F. Melgani, “Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,” Remote Sens. , vol. 16, no. 9, p. 1477, 2024
2024
-
[11]
Skysensegpt: A fine-grained instruction tuning dataset and model for remote sensing vision-language understanding,
J. Luo, Z. Pang, Y . Zhang, T. Wang, L. Wang, B. Dang, J. Lao, J. Wang, J. Chen, Y . Tan,et al., “Skysensegpt: A fine-grained instruction tuning dataset and model for remote sensing vision-language understanding,” arXiv:2406.10100, 2024
2024 arXiv
-
[12]
Vrsbench: A versatile vision- language benchmark dataset for remote sensing image understanding,
X. Li, J. Ding, and M. Elhoseiny, “Vrsbench: A versatile vision- language benchmark dataset for remote sensing image understanding,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) Datasets Benchmarks Track, 2024
2024
-
[13]
Changeclip: Remote sens- ing change detection with multimodal vision-language representation learning,
S. Dong, L. Wang, B. Du, and X. Meng, “Changeclip: Remote sens- ing change detection with multimodal vision-language representation learning,” ISPRS J. Photogrammetry Remote Sens., vol. 208, pp. 53–69, 2024
2024
-
[14]
Rs-clip: Zero shot remote sensing scene classification via contrastive vision-language supervision,
X. Li, C. Wen, Y . Hu, and N. Zhou, “Rs-clip: Zero shot remote sensing scene classification via contrastive vision-language supervision,” Int. J. Appl. Earth Observ. Geoinf. , vol. 124, p. 103497, 2023
2023
-
[15]
Mind the modality gap: Towards a remote sensing vision-language model via cross-modal alignment,
A. Zavras, D. Michail, B. Demir, and I. Papoutsis, “Mind the modality gap: Towards a remote sensing vision-language model via cross-modal alignment,” arXiv:2402.09816, 2024
2024 arXiv
-
[16]
Towards vision-language geo-foundation model: A survey,
Y . Zhou, L. Feng, Y . Ke, X. Jiang, J. Yan, X. Yang, and W. Zhang, “Towards vision-language geo-foundation model: A survey,” arXiv:2406.09385, 2024
2024
-
[17]
Vision-language models in remote sensing: Current progress and future trends,
X. Li, C. Wen, Y . Hu, Z. Yuan, and X. X. Zhu, “Vision-language models in remote sensing: Current progress and future trends,” IEEE Geosci. Remote Sens. Mag. , vol. 12, no. 2, pp. 32–66, 2024
2024
-
[18]
S-clip: Semi-supervised vision- language learning using few specialist captions,
S. Mo, M. Kim, K. Lee, and J. Shin, “S-clip: Semi-supervised vision- language learning using few specialist captions,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 36, pp. 61187–61212, 2023
2023
-
[19]
Multi-view feature fusion and visual prompt for remote sensing image captioning,
S. Wang, Q. Lin, X. Ye, Y . Liao, D. Quan, Z. Jin, B. Hou, and L. Jiao, “Multi-view feature fusion and visual prompt for remote sensing image captioning,” IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–16, 2024
2024
-
[20]
Vision- language models for zero-shot classification of remote sensing images,
M. M. Al Rahhal, Y . Bazi, H. Elgibreen, and M. Zuair, “Vision- language models for zero-shot classification of remote sensing images,” Appl. Sci., vol. 13, no. 22, p. 12462, 2023
2023
-
[21]
Detecting cloud presence in satellite images using the rgb-based clip vision-language model,
M. Czerkawski, R. Atkinson, and C. Tachtatzis, “Detecting cloud presence in satellite images using the rgb-based clip vision-language model,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. (IGARSS) , pp. 5170–5173, 2023. 29
2023
-
[22]
Chatearthnet: a global- scale image-text dataset empowering vision-language geo-foundation models,
Z. Yuan, Z. Xiong, L. Mou, and X. X. Zhu, “Chatearthnet: a global- scale image-text dataset empowering vision-language geo-foundation models,” Earth Syst. Sci. Data Discuss. , vol. 17, no. 3, pp. 1245–1263, 2025
2025
-
[23]
Bi-modal transformer-based approach for visual question answering in remote sensing imagery,
Y . Bazi, M. M. Al Rahhal, M. L. Mekhalfi, M. A. Al Zuair, and F. Melgani, “Bi-modal transformer-based approach for visual question answering in remote sensing imagery,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–11, 2022
2022
-
[24]
Popeye: A unified visual-language model for multi-source ship detection from remote sensing imagery,
W. Zhang, M. Cai, T. Zhang, G. Lei, Y . Zhuang, and X. Mao, “Popeye: A unified visual-language model for multi-source ship detection from remote sensing imagery,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 17, pp. 20050–20063, 2024
2024
-
[25]
Deep semantic-visual alignment for zero-shot remote sensing image scene classification,
W. Xu, J. Wang, Z. Wei, M. Peng, and Y . Wu, “Deep semantic-visual alignment for zero-shot remote sensing image scene classification,” ISPRS J. Photogrammetry Remote Sens. , vol. 198, pp. 140–152, 2023
2023
-
[26]
Set-clip: Exploring aligned semantic from low-alignment multimodal data through a distribution view,
Z. Song, Z. Zang, Y . Wang, G. Yang, K. Yu, W. Chen, M. Wang, and S. Z. Li, “Set-clip: Exploring aligned semantic from low-alignment multimodal data through a distribution view,” arXiv:2406.05766, 2024
2024 arXiv
-
[27]
Segment change model (scm) for unsupervised change detection in vhr remote sensing images: a case study of buildings,
X. Tan, G. Chen, T. Wang, J. Wang, and X. Zhang, “Segment change model (scm) for unsupervised change detection in vhr remote sensing images: a case study of buildings,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. (IGARSS) , pp. 8577–8580, 2024
2024
-
[28]
A new learning paradigm for foundation model-based remote-sensing change detection,
K. Li, X. Cao, and D. Meng, “A new learning paradigm for foundation model-based remote-sensing change detection,” IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–12, 2024
2024
-
[29]
Vlca: vision-language aligning model with cross-modal attention for bilingual remote sensing image captioning,
T. Wei, W. Yuan, J. Luo, W. Zhang, and L. Lu, “Vlca: vision-language aligning model with cross-modal attention for bilingual remote sensing image captioning,” J. Syst. Eng. Electron. , vol. 34, no. 1, pp. 9–18, 2023
2023
-
[30]
Rs-gpt4v: A unified multimodal instruction-following dataset for remote sensing image understanding,
L. Xu, L. Zhao, W. Guo, Q. Li, K. Long, K. Zou, Y . Wang, and H. Li, “Rs-gpt4v: A unified multimodal instruction-following dataset for remote sensing image understanding,” arXiv:2406.12479, 2024
2024 arXiv
-
[31]
Luojiahog: A hierarchy oriented geo-aware image caption dataset for remote sensing image–text retrieval,
Y . Zhao, M. Zhang, B. Yang, Z. Zhang, J. Kang, and J. Gong, “Luojiahog: A hierarchy oriented geo-aware image caption dataset for remote sensing image–text retrieval,”ISPRS J. Photogrammetry Remote Sens., vol. 222, pp. 130–151, 2025
2025
-
[32]
Satin: A multi-task meta- dataset for classifying satellite imagery using vision-language models,
J. Roberts, K. Han, and S. Albanie, “Satin: A multi-task meta- dataset for classifying satellite imagery using vision-language models,” arXiv:2304.11619, 2023
2023 arXiv
-
[33]
Rsvg: Exploring data and models for visual grounding on remote sensing data,
Y . Zhan, Z. Xiong, and Y . Yuan, “Rsvg: Exploring data and models for visual grounding on remote sensing data,” IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–13, 2023
2023
-
[34]
Good at captioning bad at counting: Bench- marking gpt-4v on earth observation data,
C. Zhang and S. Wang, “Good at captioning bad at counting: Bench- marking gpt-4v on earth observation data,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) Workshops , pp. 7839–7849, 2024
2024
-
[35]
Satclip: Global, general-purpose location embeddings with satellite imagery,
K. Klemmer, E. Rolf, C. Robinson, L. Mackey, and M. Rußwurm, “Satclip: Global, general-purpose location embeddings with satellite imagery,” in Proc. AAAI Conf. Artif. Intell. , vol. 39, pp. 4347–4355, 2025
2025
-
[36]
GeoCLIP: Clip-inspired alignment between locations and images for effective worldwide geo- localization,
V . V . Cepeda, G. K. Nayak, and M. Shah, “GeoCLIP: Clip-inspired alignment between locations and images for effective worldwide geo- localization,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, pp. 8690–8701, 2023
2023
-
[37]
Csp: Self-supervised contrastive spatial pre-training for geospatial-visual representations,
G. Mai, N. Lao, Y . He, J. Song, and S. Ermon, “Csp: Self-supervised contrastive spatial pre-training for geospatial-visual representations,” in Proc. Int. Conf. Mach. Learn. (ICML) , pp. 23498–23515, 2023
2023
-
[38]
Diffusionsat: A generative foundation model for satellite imagery,
S. Khanna, P. Liu, L. Zhou, C. Meng, R. Rombach, M. Burke, D. B. Lobell, and S. Ermon, “Diffusionsat: A generative foundation model for satellite imagery,” in Proc. Int. Conf. Learn. Representations (ICLR) , 2023
2023
-
[39]
Crs-diff: Controllable remote sensing image generation with diffusion model,
D. Tang, X. Cao, X. Hou, Z. Jiang, J. Liu, and D. Meng, “Crs-diff: Controllable remote sensing image generation with diffusion model,” IEEE Trans. Geosci. Remote Sens. , vol. 62, pp. 1–14, 2024
2024
-
[40]
Practical techniques for vision- language segmentation model in remote sensing,
Y . Lin, K. Suzuki, and S. Sogo, “Practical techniques for vision- language segmentation model in remote sensing,” Int. Arch. Pho- togramm. Remote Sens. Spatial Inf. Sci. , vol. 48, pp. 203–210, 2024
2024
-
[41]
Text2seg: Zero-shot remote sensing image semantic segmentation via text-guided visual foundation models,
J. Zhang, Z. Zhou, G. Mai, M. Hu, Z. Guan, S. Li, and L. Mu, “Text2seg: Zero-shot remote sensing image semantic segmentation via text-guided visual foundation models,” in Proc. ACM SIGSPATIAL Int. Workshop AI Geographic Knowl. Discov. (AI GeoKD), p. 63–66, 2024
2024
-
[42]
A decoupling paradigm with prompt learning for remote sensing image change captioning,
C. Liu, R. Zhao, J. Chen, Z. Qi, Z. Zou, and Z. Shi, “A decoupling paradigm with prompt learning for remote sensing image change captioning,” IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–18, 2023
2023
-
[43]
Bootstrapping interactive image-text alignment for remote sensing image captioning,
C. Yang, Z. Li, and L. Zhang, “Bootstrapping interactive image-text alignment for remote sensing image captioning,” IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–12, 2024
2024
-
[44]
Text- guided diverse image synthesis for long-tailed remote sensing object classification,
H. Tang, W. Zhao, G. Hu, Y . Xiao, Y . Li, and H. Wang, “Text- guided diverse image synthesis for long-tailed remote sensing object classification,” IEEE Trans. Geosci. Remote Sens. , vol. 62, pp. 1–13, 2024
2024
-
[45]
Object detection in aerial images: A large-scale benchmark and challenges,
J. Ding, N. Xue, G.-S. Xia, X. Bai, W. Yang, M. Y . Yang, S. Belongie, J. Luo, M. Datcu, M. Pelillo, et al., “Object detection in aerial images: A large-scale benchmark and challenges,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 11, pp. 7778–7796, 2021
2021
-
[46]
InstructBLIP: Towards general-purpose vision-language models with instruction tuning,
W. Dai, J. Li, D. Li, A. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “InstructBLIP: Towards general-purpose vision-language models with instruction tuning,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2023
2023
-
[47]
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
W.-L. Chiang, Z. Li, Z. Lin, Y . Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y . Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023
2023
-
[48]
Learning to rank question answer pairs with holographic dual lstm architecture,
Y . Tay, M. C. Phan, L. A. Tuan, and S. C. Hui, “Learning to rank question answer pairs with holographic dual lstm architecture,” in Proc. Int. ACM SIGIR Conf. Res. Dev. Inf. Retr. (SIGIR), pp. 695–704, 2017
2017
-
[49]
Improved baselines with visual instruction tuning,
H. Liu, C. Li, Y . Li, and Y . J. Lee, “Improved baselines with visual instruction tuning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 26296–26306, 2024
2024
-
[50]
Lora: Low-rank adaptation of large language models,
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in Proc. Int. Conf. Learn. Representations (ICLR) , 2022
2022
-
[51]
Object detection in optical remote sensing images: A survey and a new benchmark,
K. Li, G. Wan, G. Cheng, L. Meng, and J. Han, “Object detection in optical remote sensing images: A survey and a new benchmark,” ISPRS J. Photogrammetry Remote Sens. , vol. 159, pp. 296–307, 2020
2020
-
[52]
Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,
X. Sun, P. Wang, Z. Yan, F. Xu, R. Wang, W. Diao, J. Chen, J. Li, Y . Feng, T. Xu, et al., “Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,” ISPRS J. Photogrammetry Remote Sens. , vol. 184, pp. 116–130, 2022
2022
-
[53]
Remote sensing image scene classifi- cation: Benchmark and state of the art,
G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classifi- cation: Benchmark and state of the art,” Proc. IEEE, vol. 105, no. 10, pp. 1865–1883, 2017
2017
-
[54]
Rsvqa: Visual question answering for remote sensing data,
S. Lobry, D. Marcos, J. Murray, and D. Tuia, “Rsvqa: Visual question answering for remote sensing data,” IEEE Trans. Geosci. Remote Sens., vol. 58, no. 12, pp. 8555–8566, 2020
2020
-
[55]
Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,
M. Rahnemoonfar, T. Chowdhury, A. Sarkar, D. Varshney, M. Yari, and R. R. Murphy, “Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,” IEEE Access, vol. 9, pp. 89644– 89654, 2021
2021
-
[56]
Eva: Exploring the limits of masked visual representation learning at scale,
Y . Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y . Cao, “Eva: Exploring the limits of masked visual representation learning at scale,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , pp. 19358–19369, 2023
2023
-
[57]
Llama 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale,et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv:2307.09288, 2023
2023 arXiv
-
[58]
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, et al., “Llama: Open and efficient foundation language models,” arXiv:2302.13971, 2023
2023 arXiv
-
[59]
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning,
J. Chen, D. Zhu, X. Shen, X. Li, Z. Liu, P. Zhang, R. Krishnamoorthi, V . Chandra, Y . Xiong, and M. Elhoseiny, “Minigpt-v2: large language model as a unified interface for vision-language multi-task learning,” arXiv:2310.09478, 2023
-
[60]
Exploring models and data for remote sensing image caption generation,
X. Lu, B. Wang, X. Zheng, and X. Li, “Exploring models and data for remote sensing image caption generation,” IEEE Trans. Geosci. Remote Sens., vol. 56, no. 4, pp. 2183–2195, 2017
2017
-
[61]
Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,
Z. Yuan, W. Zhang, K. Fu, X. Li, C. Deng, H. Wang, and X. Sun, “Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–19, 2022
2022
-
[62]
Deep semantic understanding of high resolution remote sensing image,
B. Qu, X. Li, D. Tao, and X. Lu, “Deep semantic understanding of high resolution remote sensing image,” in Proc. Int. Conf. Comput. Inf. Telecom. Syst. (CITS) , pp. 1–5, 2016
2016
-
[63]
Nwpu- captions dataset and mlca-net for remote sensing image captioning,
Q. Cheng, H. Huang, Y . Xu, Y . Zhou, H. Li, and Z. Wang, “Nwpu- captions dataset and mlca-net for remote sensing image captioning,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–19, 2022
2022
-
[64]
Capera: Captioning events in aerial videos,
L. Bashmal, Y . Bazi, M. M. Al Rahhal, M. Zuair, and F. Melgani, “Capera: Captioning events in aerial videos,” Remote Sens. , vol. 15, no. 8, p. 2139, 2023
2023
-
[65]
Mutual attention inception network for remote sensing visual question answering,
X. Zheng, B. Wang, X. Du, and X. Lu, “Mutual attention inception network for remote sensing visual question answering,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–14, 2021. 30
2021
-
[66]
Visual grounding in remote sensing images,
Y . Sun, S. Feng, X. Li, Y . Ye, J. Kang, and X. Huang, “Visual grounding in remote sensing images,” in Proc. ACM Int. Conf. Multimedia , pp. 404–412, 2022
2022
-
[67]
Dinov2: Learning robust visual features without supervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khali- dov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al., “Dinov2: Learning robust visual features without supervision,” Trans. Mach. Learn. Res., pp. 1–31, 2024
2024
-
[68]
Multistep question-driven visual question answering for remote sensing,
M. Zhang, F. Chen, and B. Li, “Multistep question-driven visual question answering for remote sensing,” IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–12, 2023
2023
-
[69]
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,
P. Helber, B. Bischke, A. Dengel, and D. Borth, “Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. , vol. 12, no. 7, pp. 2217–2226, 2019
2019
-
[70]
Bag-of-visual-words and spatial extensions for land-use classification,
Y . Yang and S. Newsam, “Bag-of-visual-words and spatial extensions for land-use classification,” in Proc. ACM SIGSPATIAL Int. Conf. Adv. Geogr. Inf. Syst. (SIGSPATIAL), pp. 270–279, 2010
2010
-
[71]
Satellite image classification via two-layer sparse coding with biased image representation,
D. Dai and W. Yang, “Satellite image classification via two-layer sparse coding with biased image representation,” IEEE Geosci. Remote Sens. Lett., vol. 8, no. 1, pp. 173–176, 2010
2010
-
[72]
Deep learning based feature selection for remote sensing scene classification,
Q. Zou, L. Ni, T. Zhang, and Q. Wang, “Deep learning based feature selection for remote sensing scene classification,” IEEE Geosci. Remote Sens. Lett., vol. 12, no. 11, pp. 2321–2325, 2015
2015
-
[73]
A public dataset for ship classification in remote sensing images,
Y . Di, Z. Jiang, H. Zhang, and G. Meng, “A public dataset for ship classification in remote sensing images,” in Proc. SPIE 11155, Image Signal Process. Remote Sens. XXV , vol. 11155, pp. 515–521, 2019
2019
-
[74]
A public dataset for fine-grained ship classification in optical remote sensing images,
Y . Di, Z. Jiang, and H. Zhang, “A public dataset for fine-grained ship classification in optical remote sensing images,” Remote Sens., vol. 13, no. 4, p. 747, 2021
2021
-
[75]
Rotation-insensitive and context- augmented object detection in remote sensing images,
K. Li, G. Cheng, S. Bu, and X. You, “Rotation-insensitive and context- augmented object detection in remote sensing images,” IEEE Trans. Geosci. Remote Sens. , vol. 56, no. 4, pp. 2337–2348, 2017
2017
-
[76]
Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection,
Y . Zhang, Y . Yuan, Y . Feng, and X. Lu, “Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection,” IEEE Trans. Geosci. Remote Sens. , vol. 57, no. 8, pp. 5535–5548, 2019
2019
-
[77]
Orientation robust object detection in aerial images using deep convolutional neural network,
H. Zhu, X. Chen, W. Dai, K. Fu, Q. Ye, and J. Jiao, “Orientation robust object detection in aerial images using deep convolutional neural network,” in Proc. IEEE Int. Conf. Image Process. (ICIP) , pp. 3735– 3739, 2015
2015
-
[78]
Detection and tracking meet drones challenge,
P. Zhu, L. Wen, D. Du, X. Bian, H. Fan, Q. Hu, and H. Ling, “Detection and tracking meet drones challenge,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 11, pp. 7380–7399, 2021
2021
-
[79]
Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,
T. Zhang, X. Zhang, J. Li, X. Xu, B. Wang, X. Zhan, Y . Xu, X. Ke, T. Zeng, H. Su, et al. , “Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,”Remote Sens., vol. 13, no. 18, p. 3690, 2021
2021
-
[80]
Hit-uav: A high-altitude infrared thermal dataset for unmanned aerial vehicle- based object detection,
J. Suo, T. Wang, X. Zhang, H. Chen, W. Zhou, and W. Shi, “Hit-uav: A high-altitude infrared thermal dataset for unmanned aerial vehicle- based object detection,” Sci. Data, vol. 10, no. 1, p. 227, 2023
2023
-
[81]
Sea-shipping,
InfiRay, “Sea-shipping,” 2021
2021
-
[82]
Infrared-security,
InfiRay, “Infrared-security,” 2021
2021
-
[83]
Aerial-mancar,
InfiRay, “Aerial-mancar,” 2021
2021
-
[84]
Double-light-vehicle,
InfiRay, “Double-light-vehicle,” 2021
2021
-
[85]
Oceanic-ship,
C. for Optics Research and E. of Shandong University, “Oceanic-ship,” 2020
2020
-
[86]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., “Learning transferable visual models from natural language supervision,” in Proc. Int. Conf. Mach. Learn. (ICML) , pp. 8748–8763, 2021
2021
-
[87]
A novel svm-based decoder for remote sensing image captioning,
G. Hoxha and F. Melgani, “A novel svm-based decoder for remote sensing image captioning,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–14, 2021
2021
-
[88]
Star: A first- ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery,
Y . Li, L. Wang, T. Wang, X. Yang, J. Luo, Q. Wang, Y . Deng, W. Wang, X. Sun, H. Li, B. Dang, Y . Zhang, Y . Yu, and Y . Junchi, “Star: A first- ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery,” IEEE Trans. Pattern Anal. Mac...
2024
-
[89]
Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,
Y . Han, X. Yang, T. Pu, and Z. Peng, “Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–18, 2021
2021
-
[90]
Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,
S. Wei, X. Zeng, Q. Qu, M. Wang, H. Su, and J. Shi, “Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,” IEEE Access, vol. 8, pp. 120234–120254, 2020
2020
-
[91]
Deepsat: a learning framework for satellite imagery,
S. Basu, S. Ganguly, S. Mukhopadhyay, R. DiBiano, M. Karki, and R. Nemani, “Deepsat: a learning framework for satellite imagery,” in Proc. ACM SIGSPATIAL Int. Conf. Adv. Geogr. Inf. Syst. (SIGSPA- TIAL), pp. 1–10, 2015
2015
-
[92]
Nasc-tg2: Natural scene classification with tiangong-2 remotely sensed imagery,
Z. Zhou, S. Li, W. Wu, W. Guo, X. Li, G. Xia, and Z. Zhao, “Nasc-tg2: Natural scene classification with tiangong-2 remotely sensed imagery,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. , vol. 14, pp. 3228–3242, 2021
2021
-
[93]
Feature significance-based multibag- of-visual-words model for remote sensing image scene classification,
L. Zhao, P. Tang, and L. Huo, “Feature significance-based multibag- of-visual-words model for remote sensing image scene classification,” J. Appl. Remote Sens. , vol. 10, no. 3, p. 035004, 2016
2016
-
[94]
Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,
B. Zhao, Y . Zhong, G.-S. Xia, and L. Zhang, “Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,” IEEE Trans. Geosci. Remote Sens. , vol. 54, no. 4, pp. 2108–2123, 2015
2015
-
[95]
Patternnet: A benchmark dataset for performance evaluation of remote sensing image retrieval,
W. Zhou, S. Newsam, C. Li, and Z. Shao, “Patternnet: A benchmark dataset for performance evaluation of remote sensing image retrieval,” ISPRS J. Photogrammetry Remote Sens. , vol. 145, pp. 197–209, 2018
2018
-
[96]
Accurate object localization in remote sensing images based on convolutional neural networks,
Y . Long, Y . Gong, Z. Xiao, and Q. Liu, “Accurate object localization in remote sensing images based on convolutional neural networks,” IEEE Trans. Geosci. Remote Sens. , vol. 55, no. 5, pp. 2486–2498, 2017
2017
-
[97]
Land-cover classification with high-resolution remote sensing images using transferable deep models,
X.-Y . Tong, G.-S. Xia, Q. Lu, H. Shen, S. Li, S. You, and L. Zhang, “Land-cover classification with high-resolution remote sensing images using transferable deep models,” Remote Sens. Environ. , vol. 237, p. 111322, 2020
2020
-
[98]
Scene classification with recurrent attention of vhr remote sensing images,
Q. Wang, S. Liu, J. Chanussot, and X. Li, “Scene classification with recurrent attention of vhr remote sensing images,” IEEE Trans. Geosci. Remote Sens., vol. 57, no. 2, pp. 1155–1167, 2018
2018
-
[99]
Clrs: Continual learning benchmark for remote sensing image scene classification,
H. Li, H. Jiang, X. Gu, J. Peng, W. Li, L. Hong, and C. Tao, “Clrs: Continual learning benchmark for remote sensing image scene classification,” Sensors, vol. 20, no. 4, p. 1226, 2020
2020
-
[100]
On creating benchmark dataset for aerial image interpreta- tion: Reviews, guidances, and million-aid,
Y . Long, G.-S. Xia, S. Li, W. Yang, M. Y . Yang, X. X. Zhu, L. Zhang, and D. Li, “On creating benchmark dataset for aerial image interpreta- tion: Reviews, guidances, and million-aid,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. , vol. 14, pp. 4205–4230, 2021
2021
-
[101]
Rsi-cb: A large-scale remote sensing image classification benchmark using crowdsourced data,
H. Li, X. Dou, C. Tao, Z. Wu, J. Chen, J. Peng, M. Deng, and L. Zhao, “Rsi-cb: A large-scale remote sensing image classification benchmark using crowdsourced data,” Sensors, vol. 20, no. 6, p. 1594, 2020
2020
-
[102]
Aid: A benchmark data set for performance evaluation of aerial scene classification,
G.-S. Xia, J. Hu, F. Hu, B. Shi, X. Bai, Y . Zhong, L. Zhang, and X. Lu, “Aid: A benchmark data set for performance evaluation of aerial scene classification,” IEEE Trans. Geosci. Remote Sens., vol. 55, no. 7, pp. 3965–3981, 2017
2017
-
[103]
Mlrsnet: A multi-label high spatial resolution remote sensing dataset for semantic scene understanding,
X. Qi, P. Zhu, Y . Wang, L. Zhang, J. Peng, M. Wu, J. Chen, X. Zhao, N. Zang, and P. T. Mathiopoulos, “Mlrsnet: A multi-label high spatial resolution remote sensing dataset for semantic scene understanding,” ISPRS J. Photogrammetry Remote Sens. , vol. 169, pp. 337–350, 2020
2020
-
[104]
Multiscene: A large-scale dataset and benchmark for multiscene recognition in single aerial images,
Y . Hua, L. Mou, P. Jin, and X. X. Zhu, “Multiscene: A large-scale dataset and benchmark for multiscene recognition in single aerial images,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–13, 2021
2021
-
[105]
Airbus wind turbine patches,
A. D. G. S.A., “Airbus wind turbine patches,” 2021
2021
-
[106]
Deep learning based damage detection on post-hurricane satellite imagery,
Q. D. Cao and Y . Choe, “Deep learning based damage detection on post-hurricane satellite imagery,” arXiv:1807.01688, 2018
2018 arXiv
-
[107]
Ships in satellite imagery,
R. Hammell, “Ships in satellite imagery,” 2018
2018
-
[108]
Towards the creation of a canadian land-use dataset for agricultural land classification,
A. A. B. Jacques, A. B. Diallo, and E. Lord, “Towards the creation of a canadian land-use dataset for agricultural land classification,” in Proc. Can. Symp. Remote Sens.: Understanding Our World: Remote Sens. Sustain. Future , vol. 4, p. 6, 2021
2021
-
[109]
Smokenet: Satellite smoke scene detection using convolutional neural network with spatial and channel-wise attention,
R. Ba, C. Chen, J. Yuan, W. Song, and S. Lo, “Smokenet: Satellite smoke scene detection using convolutional neural network with spatial and channel-wise attention,” Remote Sens. , vol. 11, no. 14, p. 1702, 2019
2019
-
[110]
Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?,
O. A. Penatti, K. Nogueira, and J. A. Dos Santos, “Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) Workshops, pp. 44–51, 2015
2015
-
[111]
Towards vegetation species discrimination by using data-driven descriptors,
K. Nogueira, J. A. Dos Santos, T. Fornazari, T. S. F. Silva, L. P. Morel- lato, and R. d. S. Torres, “Towards vegetation species discrimination by using data-driven descriptors,” in Proc. IAPR Workshop Pattern Recognit. Remote Sens. (PRRS) , pp. 1–6, 2016
2016
-
[112]
Wilds: A benchmark of in-the-wild distribution shifts,
P. W. Koh, S. Sagawa, H. Marklund, S. M. Xie, M. Zhang, A. Balsub- ramani, W. Hu, M. Yasunaga, R. L. Phillips, I. Gao, et al., “Wilds: A benchmark of in-the-wild distribution shifts,” in Proc. Int. Conf. Mach. Learn. (ICML), pp. 5637–5664, 2021
2021
-
[113]
Bigearthnet: A large-scale benchmark archive for remote sensing image under- standing,
G. Sumbul, M. Charfuelan, B. Demir, and V . Markl, “Bigearthnet: A large-scale benchmark archive for remote sensing image under- standing,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. (IGARSS) , pp. 5901–5904, 2019. 31
2019
-
[114]
A benchmark dataset for canopy crown detection and delineation in co-registered airborne rgb, lidar and hyperspectral imagery from the national ecological observation network,
B. G. Weinstein, S. J. Graves, S. Marconi, A. Singh, A. Zare, D. Stew- art, S. A. Bohlman, and E. P. White, “A benchmark dataset for canopy crown detection and delineation in co-registered airborne rgb, lidar and hyperspectral imagery from the national ecological observation n...
2021
-
[115]
A large contextual dataset for classification, detection and counting of cars with deep learning,
T. N. Mundhenk, G. Konjevod, W. A. Sakla, and K. Boakye, “A large contextual dataset for classification, detection and counting of cars with deep learning,” in Proc. Eur. Conf. Comput. Vis. (ECCV), pp. 785–800, 2016
2016
-
[116]
Improving the precision and accuracy of animal population estimates with aerial image object detection,
J. A. Eikelboom, J. Wind, E. van de Ven, L. M. Kenana, B. Schroder, H. J. de Knegt, F. van Langevelde, and H. H. Prins, “Improving the precision and accuracy of animal population estimates with aerial image object detection,” Methods Ecol. Evol., vol. 10, no. 11, pp. 1875– 1887, 2019
2019
-
[117]
Functional map of the world,
G. Christie, N. Fendley, J. Wilson, and R. Mukherjee, “Functional map of the world,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 6172–6180, 2018
2018
-
[118]
Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” in Proc. Int. Conf. Mach. Learn. (ICML) , pp. 19730–19742, 2023
2023
-
[119]
Satlaspretrain: A large-scale dataset for remote sensing image un- derstanding,
F. Bastani, P. Wolters, R. Gupta, J. Ferdinando, and A. Kembhavi, “Satlaspretrain: A large-scale dataset for remote sensing image un- derstanding,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , pp. 16772–16782, 2023
2023
-
[120]
Integration of the 3d environment for uav onboard visual object tracking,
S. Vujasinovi ´c, S. Becker, T. Breuer, S. Bullinger, N. Scherer- Negenborn, and M. Arens, “Integration of the 3d environment for uav onboard visual object tracking,” Appl. Sci. , vol. 10, no. 21, p. 7622, 2020
2020
-
[121]
Drone-based object counting by spatially regularized regional proposal network,
M.-R. Hsieh, Y .-L. Lin, and W. H. Hsu, “Drone-based object counting by spatially regularized regional proposal network,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , pp. 4145–4153, 2017
2017
-
[122]
A high resolution optical satellite image dataset for ship recognition and some new baselines,
Z. Liu, L. Yuan, L. Weng, and Y . Yang, “A high resolution optical satellite image dataset for ship recognition and some new baselines,” in Proc. Int. Conf. Pattern Recognit. Appl. Methods (ICPRAM) , vol. 1, pp. 324–331, 2017
2017
-
[123]
A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,
H. Chen and Z. Shi, “A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,” Remote Sens., vol. 12, no. 10, p. 1662, 2020
2020
-
[124]
Learning social etiquette: Human trajectory understanding in crowded scenes,
A. Robicquet, A. Sadeghian, A. Alahi, and S. Savarese, “Learning social etiquette: Human trajectory understanding in crowded scenes,” in Proc. Eur. Conf. Comput. Vis. (ECCV) , pp. 549–565, 2016
2016
-
[125]
isaid: A large-scale dataset for instance segmentation in aerial images,
S. Waqas Zamir, A. Arora, A. Gupta, S. Khan, G. Sun, F. Shah- baz Khan, F. Zhu, L. Shao, G.-S. Xia, and X. Bai, “isaid: A large-scale dataset for instance segmentation in aerial images,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) Workshops , pp. 28–37, 2019
2019
-
[126]
Loveda: A remote sensing land-cover dataset for domain adaptive semantic segmentation,
J. Wang, Z. Zheng, A. Ma, X. Lu, and Y . Zhong, “Loveda: A remote sensing land-cover dataset for domain adaptive semantic segmentation,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) Datasets Benchmarks Track, 2021
2021
-
[127]
2d semantic labeling contest - potsdam,
I. S. for Photogrammetry and R. Sensing, “2d semantic labeling contest - potsdam,” 2012
2012
-
[128]
2d semantic labeling - vaihingen data,
I. S. for Photogrammetry and R. Sensing, “2d semantic labeling - vaihingen data,” 2012
2012
-
[129]
Minigpt-4: En- hancing vision-language understanding with advanced large language models,
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: En- hancing vision-language understanding with advanced large language models,” in Proc. Int. Conf. Learn. Representations (ICLR) , 2024
2024
-
[130]
Spacenet: A remote sensing dataset and challenge series,
A. Van Etten, D. Lindenbaum, and T. M. Bacastow, “Spacenet: A remote sensing dataset and challenge series,” arXiv:1807.01232, 2018
2018 arXiv
-
[131]
Knowledge-aware text-image retrieval for remote sensing images,
L. Mi, X. Dai, J. Castillo-Navarro, and D. Tuia, “Knowledge-aware text-image retrieval for remote sensing images,” IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–13, 2024
2024
-
[132]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in Proc. Int. Conf. Mach. Learn. (ICML) , pp. 12888–12900, 2022
2022
-
[133]
Pir: Remote sensing image-text retrieval with prior instruction representation learning,
J. Pan, M. Ma, Q. Ma, C. Bai, and S. Chen, “Pir: Remote sensing image-text retrieval with prior instruction representation learning,” arXiv:2405.10160, 2024
2024 arXiv
-
[134]
Composed image retrieval for remote sensing,
B. Psomas, I. Kakogeorgiou, N. Efthymiadis, G. Tolias, O. Chum, Y . Avrithis, and K. Karantzalos, “Composed image retrieval for remote sensing,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. (IGARSS) , pp. 8526–8534, 2024
2024
-
[135]
Mgimm: Multi-granularity instruction multimodal model for attribute-guided remote sensing image detailed description,
C. Yang, Z. Li, and L. Zhang, “Mgimm: Multi-granularity instruction multimodal model for attribute-guided remote sensing image detailed description,” IEEE Trans. Geosci. Remote Sens. , vol. 62, pp. 1–13, 2024
2024
-
[136]
Earth- marker: A visual prompting multimodal large language model for remote sensing,
W. Zhang, M. Cai, T. Zhang, Y . Zhuang, J. Li, and X. Mao, “Earth- marker: A visual prompting multimodal large language model for remote sensing,” IEEE Trans. Geosci. Remote Sens. , vol. 63, pp. 1– 19, 2024
2024
-
[137]
Mar20: A benchmark for military aircraft recognition in remote sensing images,
Y . Wenqi, C. Gong, W. Meijun, Y . Yanqing, X. Xingxing, Y . Xiwen, and H. Junwei, “Mar20: A benchmark for military aircraft recognition in remote sensing images,” Natl. Remote Sens. Bull. , vol. 27, no. 12, pp. 2688–2696, 2024
2024
-
[138]
Hi-ucd: A large-scale dataset for urban semantic change detection in remote sensing imagery,
S. Tian, A. Ma, Z. Zheng, and Y . Zhong, “Hi-ucd: A large-scale dataset for urban semantic change detection in remote sensing imagery,” arXiv:2011.03247, 2020
2011 arXiv
-
[139]
Uavid: A semantic segmentation dataset for uav imagery,
Y . Lyu, G. V osselman, G.-S. Xia, A. Yilmaz, and M. Y . Yang, “Uavid: A semantic segmentation dataset for uav imagery,” ISPRS J. Photogrammetry Remote Sens. , vol. 165, pp. 108–119, 2020
2020
-
[140]
Samrs: Scaling-up remote sensing segmentation dataset with segment anything model,
D. Wang, J. Zhang, B. Du, M. Xu, L. Liu, D. Tao, and L. Zhang, “Samrs: Scaling-up remote sensing segmentation dataset with segment anything model,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) Datasets Benchmarks Track, vol. 36, pp. 8815–8827, 2023
2023
-
[141]
Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set,
S. Ji, S. Wei, and M. Lu, “Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set,” IEEE Trans. Geosci. Remote Sens. , vol. 57, no. 1, pp. 574–586, 2018
2018
-
[142]
Ifship: A large vision- language model for interpretable fine-grained ship classification via domain knowledge-enhanced instruction tuning,
M. Guo, M. Wu, Y . Shen, H. Li, and C. Tao, “Ifship: A large vision- language model for interpretable fine-grained ship classification via domain knowledge-enhanced instruction tuning,” arXiv:2408.06631, 2024
2024 arXiv
-
[143]
Rsteller: Scaling up visual lan- guage modeling in remote sensing with rich linguistic semantics from openly available data and large language models,
J. Ge, Y . Zheng, K. Guo, and J. Liang, “Rsteller: Scaling up visual lan- guage modeling in remote sensing with rich linguistic semantics from openly available data and large language models,” arXiv:2408.14744, 2024
2024 arXiv
-
[144]
Mixtral of experts,
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bam- ford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al. , “Mixtral of experts,” arXiv:2401.04088, 2024
2024 arXiv
-
[145]
Rsdiff: Remote sensing image generation from text using diffusion model,
A. Sebaq and M. ElHelw, “Rsdiff: Remote sensing image generation from text using diffusion model,” Neural Comput. Appl. , vol. 36, pp. 23103–23111, 2024
2024
-
[146]
Improved denoising diffusion proba- bilistic models,
A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion proba- bilistic models,” in Proc. Int. Conf. Mach. Learn. (ICML) , pp. 8162– 8171, 2021
2021
-
[147]
Photorealistic text-to-image diffusion models with deep language understanding,
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Den- ton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Sali- mans, et al., “Photorealistic text-to-image diffusion models with deep language understanding,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol....
2022
-
[148]
Empower generalizability for pansharpening through text-modulated diffusion model,
Y . Xing, L. Qu, S. Zhang, J. Feng, X. Zhang, and Y . Zhang, “Empower generalizability for pansharpening through text-modulated diffusion model,” IEEE Trans. Geosci. Remote Sens. , vol. 62, pp. 1–12, 2024
2024
-
[149]
Up-diff: Latent diffusion model for remote sensing urban prediction,
Z. Wang, Z. Hao, Y . Zhang, Y . Feng, and Y . Guo, “Up-diff: Latent diffusion model for remote sensing urban prediction,” IEEE Geosci. Remote Sens. Lett. , vol. 22, pp. 1–5, 2024
2024
-
[150]
Vcc-diffnet: Visual conditional con- trol diffusion network for remote sensing image captioning,
Q. Cheng, Y . Xu, and Z. Huang, “Vcc-diffnet: Visual conditional con- trol diffusion network for remote sensing image captioning,” Remote Sens., vol. 16, no. 16, p. 2961, 2024
2024
-
[151]
Structured denoising diffusion models in discrete state-spaces,
J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg, “Structured denoising diffusion models in discrete state-spaces,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 34, pp. 17981– 17993, 2021
2021
-
[152]
Enhancing remote sensing vision-language models for zero-shot scene classification,
K. El Khoury, M. Zanella, B. G ´erin, T. Godelaine, B. Macq, S. Mah- moudi, C. De Vleeschouwer, and I. B. Ayed, “Enhancing remote sensing vision-language models for zero-shot scene classification,” in Proc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP), pp. 1– 5, 2025
2025
-
[153]
Efficient prompt tuning of large vision-language model for fine-grained ship classifica- tion,
L. Lan, F. Wang, X. Zheng, Z. Wang, and X. Liu, “Efficient prompt tuning of large vision-language model for fine-grained ship classifica- tion,” IEEE Trans. Geosci. Remote Sens. , vol. 63, pp. 1–10, 2024
2024
-
[154]
Diffusion-rscc: Diffusion probabilistic model for change captioning in remote sensing images,
X. Yu, Y . Li, and J. Ma, “Diffusion-rscc: Diffusion probabilistic model for change captioning in remote sensing images,” arXiv:2405.12875, 2024
2024 arXiv
-
[155]
Leveraging visual language model and generative diffusion model for zero-shot sar target recognition,
J. Wang, H. Sun, T. Tang, Y . Sun, Q. He, L. Lin, and K. Ji, “Leveraging visual language model and generative diffusion model for zero-shot sar target recognition,” Remote Sens., vol. 16, no. 16, p. 2927, 2024
2024
-
[156]
Urbench: A comprehensive benchmark for evaluat- ing large multimodal models in multi-view urban scenarios,
B. Zhou, H. Yang, D. Chen, J. Ye, T. Bai, J. Yu, S. Zhang, D. Lin, C. He, and W. Li, “Urbench: A comprehensive benchmark for evaluat- ing large multimodal models in multi-view urban scenarios,” in Proc. AAAI Conf. Artif. Intell. , vol. 39, pp. 10707–10715, 2025
2025
-
[157]
The cityscapes dataset 32 for semantic urban scene understanding,
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Be- nenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset 32 for semantic urban scene understanding,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , pp. 3213–3223, 2016
2016
-
[158]
The mapillary traffic sign dataset for detection and classification on a global scale,
C. Ertler, J. Mislej, T. Ollmann, L. Porzi, G. Neuhold, and Y . Kuang, “The mapillary traffic sign dataset for detection and classification on a global scale,” in Proc. Eur. Conf. Comput. Vis. (ECCV) , pp. 68–84, 2020
2020
-
[159]
Vigor: Cross-view image geo- localization beyond one-to-one retrieval,
S. Zhu, T. Yang, and C. Chen, “Vigor: Cross-view image geo- localization beyond one-to-one retrieval,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , pp. 3640–3649, 2021
2021
-
[160]
Im2gps: estimating geographic information from a single image,
J. Hays and A. A. Efros, “Im2gps: estimating geographic information from a single image,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 1–8, 2008
2008
-
[161]
University-1652: A multi-view multi- source benchmark for drone-based geo-localization,
Z. Zheng, Y . Wei, and Y . Yang, “University-1652: A multi-view multi- source benchmark for drone-based geo-localization,” in Proc. ACM Int. Conf. Multimedia, pp. 1395–1403, 2020
2020
-
[162]
Towards natural language-guided drones: Geotext-1652 benchmark with spatial relation matching,
M. Chu, Z. Zheng, W. Ji, T. Wang, and T.-S. Chua, “Towards natural language-guided drones: Geotext-1652 benchmark with spatial relation matching,” in Proc. Eur. Conf. Comput. Vis. (ECCV) , pp. 213–231, 2025
2025
-
[163]
Multi-grained vision language pre- training: Aligning texts with visual concepts,
Y . Zeng, X. Zhang, and H. Li, “Multi-grained vision language pre- training: Aligning texts with visual concepts,” in Proc. Int. Conf. Mach. Learn. (ICML), pp. 25994–26009, 2022
2022
-
[164]
Language integration in remote sensing: Tasks, datasets, and future directions,
L. Bashmal, Y . Bazi, F. Melgani, M. M. Al Rahhal, and M. A. Al Zuair, “Language integration in remote sensing: Tasks, datasets, and future directions,” IEEE Geosci. Remote Sens. Mag., vol. 11, pp. 63–93, 2023
2023
-
[165]
Retro-remote sensing: Generating images from ancient texts,
M. B. Bejiga, F. Melgani, and A. Vascotto, “Retro-remote sensing: Generating images from ancient texts,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. , vol. 12, no. 3, pp. 950–960, 2019
2019
-
[166]
Textrs: Deep bidirectional triplet network for matching text to remote sensing images,
T. Abdullah, Y . Bazi, M. M. Al Rahhal, M. L. Mekhalfi, L. Rangarajan, and M. Zuair, “Textrs: Deep bidirectional triplet network for matching text to remote sensing images,” Remote Sens. , vol. 12, no. 3, p. 405, 2020
2020
-
[167]
Transforming remote sensing images to textual descriptions,
U. Zia, M. M. Riaz, and A. Ghafoor, “Transforming remote sensing images to textual descriptions,” Int. J. Appl. Earth Observ. Geoinf. , vol. 108, p. 102741, 2022
2022
-
[168]
Vaa: Visual aligning attention model for remote sensing image captioning,
Z. Zhang, W. Zhang, W. Diao, M. Yan, X. Gao, and X. Sun, “Vaa: Visual aligning attention model for remote sensing image captioning,” IEEE Access, vol. 7, pp. 137355–137364, 2019
2019
-
[169]
A multi-level attention model for remote sensing image captions,
Y . Li, S. Fang, L. Jiao, R. Liu, and R. Shang, “A multi-level attention model for remote sensing image captions,” Remote Sens., vol. 12, no. 6, p. 939, 2020
2020
-
[170]
High-resolution remote sensing image captioning based on structured attention,
R. Zhao, Z. Shi, and Z. Zou, “High-resolution remote sensing image captioning based on structured attention,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–14, 2021
2021
-
[171]
Very deep convolutional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv:1409.1556, 2014
2014 arXiv
-
[172]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 770–778, 2016
2016
-
[173]
Long short-term memory,
J. Schmidhuber, S. Hochreiter, et al. , “Long short-term memory,” Neural Comput., vol. 9, no. 8, pp. 1735–1780, 1997
1997
-
[174]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 30, 2017
2017
-
[175]
Dimensionality reduction by learning an invariant mapping,
R. Hadsell, S. Chopra, and Y . LeCun, “Dimensionality reduction by learning an invariant mapping,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), vol. 2, pp. 1735–1742, 2006
2006
-
[176]
Representation learning with contrastive predictive coding,
A. v. d. Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv:1807.03748, 2018
2018 arXiv
-
[177]
Comprehending and ordering semantics for image captioning,
Y . Li, Y . Pan, T. Yao, and T. Mei, “Comprehending and ordering semantics for image captioning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , pp. 17990–17999, 2022
2022
-
[178]
Clipn for zero-shot ood detection: Teaching clip to say no,
H. Wang, Y . Li, H. Yao, and X. Li, “Clipn for zero-shot ood detection: Teaching clip to say no,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 1802–1812, 2023
2023
-
[179]
Zegclip: Towards adapting clip for zero-shot semantic segmentation,
Z. Zhou, Y . Lei, B. Zhang, L. Liu, and Y . Liu, “Zegclip: Towards adapting clip for zero-shot semantic segmentation,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 11175–11185, 2023
2023
-
[180]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. Int. Conf. Learn. Representati...
2021
-
[181]
Multi- scale representation learning for spatial feature distributions using grid cells,
G. Mai, K. Janowicz, B. Yan, R. Zhu, L. Cai, and N. Lao, “Multi- scale representation learning for spatial feature distributions using grid cells,” in Proc. Int. Conf. Learn. Representations (ICLR) , 2020
2020
-
[182]
The benchmarking initiative for multimedia evaluation: Mediaeval 2016,
M. Larson, M. Soleymani, G. Gravier, B. Ionescu, and G. J. Jones, “The benchmarking initiative for multimedia evaluation: Mediaeval 2016,” IEEE MultiMedia, vol. 24, no. 1, pp. 93–96, 2017
2016
-
[183]
Ge- ographic location encoding with spherical harmonics and sinusoidal representation networks,
M. Rußwurm, K. Klemmer, E. Rolf, R. Zbinden, and D. Tuia, “Ge- ographic location encoding with spherical harmonics and sinusoidal representation networks,” in Proc. Int. Conf. Learn. Representations (ICLR), 2024
2024
-
[184]
A simple frame- work for contrastive learning of visual representations,
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple frame- work for contrastive learning of visual representations,” in Proc. Int. Conf. Mach. Learn. (ICML) , pp. 1597–1607, 2020
2020
-
[185]
Fourier features let networks learn high frequency functions in low dimensional domains,
M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Ragha- van, U. Singhal, R. Ramamoorthi, J. Barron, and R. Ng, “Fourier features let networks learn high frequency functions in low dimensional domains,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 33,...
2020
-
[186]
Structural high-resolution satellite image indexing,
G.-S. Xia, W. Yang, J. Delon, Y . Gousseau, H. Sun, and H. Ma ˆıtre, “Structural high-resolution satellite image indexing,” in ISPRS TC VII Symposium-100 Years ISPRS, vol. 38, pp. 298–303, 2010
2010
-
[187]
Segearth-ov: Towards-free open-vocabulary segmentation for remote sensing im- ages,
K. Li, R. Liu, X. Cao, D. Meng, and Z. Wang, “Segearth-ov: Towards-free open-vocabulary segmentation for remote sensing im- ages,” arXiv:2410.01768, 2024
2024 arXiv
-
[188]
Clip-adapter: Better vision-language models with feature adapters,
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y . Zhang, H. Li, and Y . Qiao, “Clip-adapter: Better vision-language models with feature adapters,” Int. J. Comput. Vis. , vol. 132, no. 2, pp. 581–595, 2024
2024
-
[189]
Conceptnet 5.5: An open multilingual graph of general knowledge,
R. Speer, J. Chin, and C. Havasi, “Conceptnet 5.5: An open multilingual graph of general knowledge,” in Proc. AAAI Conf. Artif. Intell., vol. 31, pp. 4444–4451, 2017
2017
-
[190]
Robust deep alignment network with remote sensing knowledge graph for zero-shot and generalized zero-shot remote sensing image scene classification,
Y . Li, D. Kong, Y . Zhang, Y . Tan, and L. Chen, “Robust deep alignment network with remote sensing knowledge graph for zero-shot and generalized zero-shot remote sensing image scene classification,” ISPRS J. Photogrammetry Remote Sens. , vol. 179, pp. 145–158, 2021
2021
-
[191]
Segment any- thing,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, et al. , “Segment any- thing,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , pp. 4015– 4026, 2023
2023
-
[192]
Image segmentation using text and image prompts,
T. L ¨uddecke and A. Ecker, “Image segmentation using text and image prompts,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 7086–7096, 2022
2022
-
[193]
Remote sensing image change detection with transformers,
H. Chen, Z. Qi, and Z. Shi, “Remote sensing image change detection with transformers,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1– 14, 2021
2021
-
[194]
A transformer-based siamese network for change detection,
W. G. C. Bandara and V . M. Patel, “A transformer-based siamese network for change detection,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. (IGARSS), pp. 207–210, 2022
2022
-
[195]
Learning to prompt for vision-language models,
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,” Int. J. Comput. Vis., vol. 130, no. 9, pp. 2337– 2348, 2022
2022
-
[196]
Addressclip: Empowering vision-language models for city-wide image address lo- calization,
S. Xu, C. Zhang, L. Fan, G. Meng, S. Xiang, and J. Ye, “Addressclip: Empowering vision-language models for city-wide image address lo- calization,” in Proc. Eur. Conf. Comput. Vis. (ECCV), pp. 76–92, 2024
2024
-
[197]
Opt: Open pre-trained transformer language models,
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V . Lin, et al. , “Opt: Open pre-trained transformer language models,” arXiv:2205.01068, 2022
2022 arXiv
-
[198]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proc. Conf. North Am. Chapter Assoc. Comput. Linguist.: Hum. Lang. Technol. (NAACL-HLT), vol. 1, pp. 4171–4186, 2019
2019
-
[199]
Beyond self-attention: External attention using two linear layers for visual tasks,
M.-H. Guo, Z.-N. Liu, T.-J. Mu, and S.-M. Hu, “Beyond self-attention: External attention using two linear layers for visual tasks,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 5, pp. 5436–5447, 2022
2022
-
[200]
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019
2019
-
[201]
Visual instruction tuning,
H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” inProc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 36, pp. 34892–34916, 2024
2024
-
[202]
Vhm: Versatile and honest vision language model for remote sensing image analysis,
C. Pang, X. Weng, J. Wu, J. Li, Y . Liu, J. Sun, W. Li, S. Wang, L. Feng, G.-S. Xia, et al. , “Vhm: Versatile and honest vision language model for remote sensing image analysis,” in Proc. AAAI Conf. Artif. Intell. , vol. 39, pp. 6381–6388, 2025
2025
-
[203]
High- resolution image synthesis with latent diffusion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High- resolution image synthesis with latent diffusion models,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , pp. 10684– 10695, 2022
2022
-
[204]
Auto-encoding variational bayes,
D. P. Kingma, “Auto-encoding variational bayes,” arXiv:1312.6114, 2013. 33
2013 arXiv
-
[205]
Diffusion models meet remote sensing: Principles, methods, and perspectives,
Y . Liu, J. Yue, S. Xia, P. Ghamisi, W. Xie, and L. Fang, “Diffusion models meet remote sensing: Principles, methods, and perspectives,” IEEE Trans. Geosci. Remote Sens. , vol. 62, pp. 1–22, 2024
2024
-
[206]
Cross- view meets diffusion: Aerial image synthesis with geometry and text guidance,
A. Arrabi, X. Zhang, W. Sultan, C. Chen, and S. Wshah, “Cross- view meets diffusion: Aerial image synthesis with geometry and text guidance,” in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), pp. 5356–5366, 2025
2025
-
[207]
Hsigene: A foundation model for hyperspectral image generation,
L. Pang, D. Tang, S. Xu, D. Meng, and X. Cao, “Hsigene: A foundation model for hyperspectral image generation,” arXiv:2409.12470, 2024
2024 arXiv
-
[208]
Metaearth: A generative foundation model for global-scale remote sensing image generation,
Z. Yu, C. Liu, L. Liu, Z. Shi, and Z. Zou, “Metaearth: A generative foundation model for global-scale remote sensing image generation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 47, no. 3, 2025
2025
-
[209]
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs,
C. Schuhmann, R. Kaczmarczyk, A. Komatsuzaki, A. Katta, R. Vencu, R. Beaumont, J. Jitsev, T. Coombes, and C. Mullis, “Laion-400m: Open dataset of clip-filtered 400 million image-text pairs,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) Workshop Datacentric AI , no. FZJ-2...
2022
-
[210]
Microsoft coco captions: Data collection and evaluation server,
X. Chen, H. Fang, T.-Y . Lin, R. Vedantam, S. Gupta, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco captions: Data collection and evaluation server,” arXiv:1504.00325, 2015
2015 arXiv
-
[211]
Referitgame: Referring to objects in photographs of natural scenes,
S. Kazemzadeh, V . Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Proc. Conf. Empirical Methods Nat. Lang. Process. (EMNLP) , pp. 787–798, 2014
2014
-
[212]
Modeling context in referring expressions,
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in Proc. Eur. Conf. Comput. Vis. (ECCV), pp. 69–85, 2016
2016
-
[213]
Teochat: A large vision-language assistant for temporal earth observation data,
J. A. Irvin, E. R. Liu, J. C. Chen, I. Dormoy, J. Kim, S. Khanna, Z. Zheng, and S. Ermon, “Teochat: A large vision-language assistant for temporal earth observation data,” in Proc. Int. Conf. Learn. Repre- sentations (ICLR), 2025
2025
-
[214]
Video-llava: Learning united visual representation by alignment before projection,
B. Lin, Y . Ye, B. Zhu, C. Jiaxi, M. Ning, P. Jin, and L. Yuan, “Video-llava: Learning united visual representation by alignment before projection,” in Proc. Conf. Empirical Methods Nat. Lang. Process. (EMNLP), pp. 5971–5984, 2024
2024
-
[215]
Creating xbd: A dataset for assessing building damage from satellite imagery,
R. Gupta, B. Goodman, N. Patel, R. Hosfelt, S. Sajeev, E. Heim, J. Doshi, K. Lucas, H. Choset, and M. Gaston, “Creating xbd: A dataset for assessing building damage from satellite imagery,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) Workshops , pp. 10–17, 2019
2019
-
[216]
S2looking: A satellite side-looking dataset for building change detection,
L. Shen, Y . Lu, H. Chen, H. Wei, D. Xie, J. Yue, R. Chen, S. Lv, and B. Jiang, “S2looking: A satellite side-looking dataset for building change detection,” Remote Sens., vol. 13, no. 24, p. 5094, 2021
2021
-
[217]
Qfabric: Multi-task change detection dataset,
S. Verma, A. Panigrahi, and S. Gupta, “Qfabric: Multi-task change detection dataset,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) Workshops, pp. 1052–1061, 2021
2021
-
[218]
Airborne hyperspectral data over chikusei,
N. Yokoya and A. Iwasaki, “Airborne hyperspectral data over chikusei,” Space Application Laboratory, University of Tokyo , vol. 5, no. 5, p. 5, 2016
2016
-
[219]
Aerial hyperspectral remote sensing classification dataset of xiongan new area (matiwan village),
C. Yi, L. Zhang, X. Zhang, W. Yueming, Q. Wenchao, T. Senlin, and P. Zhang, “Aerial hyperspectral remote sensing classification dataset of xiongan new area (matiwan village),” Natl. Remote Sens. Bull., vol. 24, no. 11, pp. 1299–1306, 2020
2020
-
[220]
A multiscale dataset for understanding complex eco- hydrological processes in a heterogeneous oasis system,
X. Li, S. Liu, Q. Xiao, M. Ma, R. Jin, T. Che, W. Wang, X. Hu, Z. Xu, J. Wen, et al. , “A multiscale dataset for understanding complex eco- hydrological processes in a heterogeneous oasis system,” Sci. Data , vol. 4, no. 1, pp. 1–11, 2017
2017
-
[221]
Houston hyperspectral dataset,
I. G. D. F. Contest, “Houston hyperspectral dataset,” 2013
2013
-
[222]
Houston hyperspectral dataset,
I. G. D. F. Contest, “Houston hyperspectral dataset,” 2018
2018
-
[223]
Geosynth: Contextually-aware high-resolution satellite image synthesis,
S. Sastry, S. Khanal, A. Dhakal, and N. Jacobs, “Geosynth: Contextually-aware high-resolution satellite image synthesis,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) Workshops , pp. 460–470, 2024
2024
-
[224]
Meter-ml: A multi-sensor earth observation benchmark for automated methane source mapping,
B. Zhu, N. Lui, J. Irvin, J. Le, S. Tadwalkar, C. Wang, Z. Ouyang, F. Y . Liu, A. Y . Ng, and R. B. Jackson, “Meter-ml: A multi-sensor earth observation benchmark for automated methane source mapping,” arXiv:2207.11166, 2022
2022 arXiv
-
[225]
Bleu: a method for automatic evaluation of machine translation,
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proc. Annu. Meet. Assoc. Comput. Linguist. (ACL) , pp. 311–318, 2002
2002
-
[226]
Meteor: An automatic metric for mt evalu- ation with improved correlation with human judgments,
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evalu- ation with improved correlation with human judgments,” in Proc. ACL Workshop Intrinsic Extrinsic Eval. Measures Mach. Transl. Summar. , pp. 65–72, 2005
2005
-
[227]
Rouge: A package for automatic evaluation of summaries,
C.-Y . Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out , pp. 74–81, 2004
2004
-
[228]
Cider: Consensus- based image description evaluation,
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus- based image description evaluation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , pp. 4566–4575, 2015
2015
-
[229]
Time-series satellite remote sensing reveals gradually increasing war damage in the gaza strip,
S. Holail, T. Saleh, X. Xiao, J. Xiao, G.-S. Xia, Z. Shao, M. Wang, J. Gong, and D. Li, “Time-series satellite remote sensing reveals gradually increasing war damage in the gaza strip,” Natl. Sci. Rev. , vol. 11, no. 9, p. nwae304, 2024
2024
-
[230]
An automatic change detection method for monitoring newly constructed building areas using time-series multi-view high-resolution optical satellite images,
X. Huang, Y . Cao, and J. Li, “An automatic change detection method for monitoring newly constructed building areas using time-series multi-view high-resolution optical satellite images,” Remote Sens. En- viron., vol. 244, p. 111802, 2020
2020
-
[231]
Sagn: Semantic-aware graph network for remote sensing scene classification,
Y . Yang, X. Tang, Y .-M. Cheung, X. Zhang, and L. Jiao, “Sagn: Semantic-aware graph network for remote sensing scene classification,” IEEE Trans. Image Process. , vol. 32, pp. 1011–1025, 2023
2023
-
[232]
Srsg and s2sg: a model and a dataset for scene graph generation of remote sensing images from segmentation results,
Z. Lin, F. Zhu, Y . Kong, Q. Wang, and J. Wang, “Srsg and s2sg: a model and a dataset for scene graph generation of remote sensing images from segmentation results,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–11, 2022
2022
-
[233]
Llama-adapter v2: Parameter-efficient visual instruction model,
P. Gao, J. Han, R. Zhang, Z. Lin, S. Geng, A. Zhou, W. Zhang, P. Lu, C. He, X. Yue, et al. , “Llama-adapter v2: Parameter-efficient visual instruction model,” arXiv:2304.15010, 2023
2023 arXiv
-
[234]
Adding conditional control to text-to-image diffusion models,
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 3836–3847, 2023
2023
-
[235]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 33, pp. 6840–6851, 2020
2020
-
[236]
Esrgan: Enhanced super-resolution generative adver- sarial networks,
X. Wang, K. Yu, S. Wu, J. Gu, Y . Liu, C. Dong, Y . Qiao, and C. Change Loy, “Esrgan: Enhanced super-resolution generative adver- sarial networks,” in Proc. Eur. Conf. Comput. Vis. (ECCV) Workshops, 2018
2018
-
[237]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, no. 140, pp. 1–67, 2020
2020
-
[238]
Holistically-nested edge detection,
S. Xie and Z. Tu, “Holistically-nested edge detection,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , pp. 1395–1403, 2015
2015
-
[239]
Line segment detection using transformers without edges,
Y . Xu, W. Xu, D. Cheung, and Z. Tu, “Line segment detection using transformers without edges,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 4257–4266, 2021
2021
-
[240]
Learning to simplify: fully convolutional networks for rough sketch cleanup,
E. Simo-Serra, S. Iizuka, K. Sasaki, and H. Ishikawa, “Learning to simplify: fully convolutional networks for rough sketch cleanup,” ACM Trans. Graph., vol. 35, no. 4, pp. 1–11, 2016
2016
-
[241]
Uni-controlnet: All-in-one control to text-to-image diffusion models,
S. Zhao, D. Chen, Y .-C. Chen, J. Bao, S. Hao, L. Yuan, and K.-Y . K. Wong, “Uni-controlnet: All-in-one control to text-to-image diffusion models,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 36, pp. 11127–11150, 2024
2024
-
[242]
Attentional feature fusion,
Y . Dai, F. Gieseke, S. Oehmcke, Y . Wu, and K. Barnard, “Attentional feature fusion,” in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), pp. 3560–3569, 2021
2021
-
[243]
Gans trained by a two time-scale update rule converge to a local nash equilibrium,
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, pp. 6629–6640, 2017
2017
-
[244]
Improved techniques for training gans,
T. Salimans, I. Goodfellow, W. Zaremba, V . Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 29, pp. 2234–2242, 2016
2016
-
[245]
Ediffsr: An efficient diffusion probabilistic model for remote sensing image super- resolution,
Y . Xiao, Q. Yuan, K. Jiang, J. He, X. Jin, and L. Zhang, “Ediffsr: An efficient diffusion probabilistic model for remote sensing image super- resolution,” IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–14, 2023
2023
-
[246]
A conditional diffusion model with fast sampling strategy for remote sensing image super-resolution,
F. Meng, Y . Chen, H. Jing, L. Zhang, Y . Yan, Y . Ren, S. Wu, T. Feng, R. Liu, and Z. Du, “A conditional diffusion model with fast sampling strategy for remote sensing image super-resolution,” IEEE Trans. Geosci. Remote Sens. , vol. 62, pp. 1–16, 2024
2024
-
[247]
Siamese meets diffusion network: Smdnet for enhanced change detection in high-resolution rs imagery,
J. Jia, G. Lee, Z. Wang, L. Zhi, and Y . He, “Siamese meets diffusion network: Smdnet for enhanced change detection in high-resolution rs imagery,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. , vol. 17, pp. 8189–8202, 2024
2024
-
[248]
Transc-gd-cd: Transformer- based conditional generative diffusion change detection model,
Y . Wen, Z. Zhang, Q. Cao, and G. Niu, “Transc-gd-cd: Transformer- based conditional generative diffusion change detection model,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. , vol. 17, pp. 7144– 7158, 2024
2024
-
[249]
Spectraldiff: A generative frame- work for hyperspectral image classification with diffusion models,
N. Chen, J. Yue, L. Fang, and S. Xia, “Spectraldiff: A generative frame- work for hyperspectral image classification with diffusion models,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–16, 2023. 34
2023
-
[250]
Diverse hyperspectral remote sensing image synthesis with diffusion models,
L. Liu, B. Chen, H. Chen, Z. Zou, and Z. Shi, “Diverse hyperspectral remote sensing image synthesis with diffusion models,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–16, 2023
2023
-
[251]
Label freedom: Stable diffusion for remote sensing image semantic segmen- tation data generation,
C. Zhao, Y . Ogawa, S. Chen, Z. Yang, and Y . Sekimoto, “Label freedom: Stable diffusion for remote sensing image semantic segmen- tation data generation,” in Proc. IEEE Int. Conf. Big Data (BigData) , pp. 1022–1030, 2023
2023
-
[252]
Efficient and controllable remote sensing fake sample generation based on diffusion model,
Z. Yuan, C. Hao, R. Zhou, J. Chen, M. Yu, W. Zhang, H. Wang, and X. Sun, “Efficient and controllable remote sensing fake sample generation based on diffusion model,” IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–12, 2023
2023
-
[253]
Cloud removal in remote sensing using sequential- based diffusion models,
X. Zhao and K. Jia, “Cloud removal in remote sensing using sequential- based diffusion models,” Remote Sens., vol. 15, no. 11, p. 2861, 2023
2023
-
[254]
Denoising diffusion probabilistic feature-based network for cloud removal in sentinel-2 imagery,
R. Jing, F. Duan, F. Lu, M. Zhang, and W. Zhao, “Denoising diffusion probabilistic feature-based network for cloud removal in sentinel-2 imagery,” Remote Sens., vol. 15, no. 9, p. 2217, 2023
2023
-
[255]
Remote sensing image change captioning using multi-attentive network with diffusion model,
Y . Yang, T. Liu, Y . Pu, L. Liu, Q. Zhao, and Q. Wan, “Remote sensing image change captioning using multi-attentive network with diffusion model,” Remote Sens., vol. 16, no. 21, p. 4083, 2024
2024
-
[256]
Image super-resolution via iterative refinement,
C. Saharia, J. Ho, W. Chan, T. Salimans, D. J. Fleet, and M. Norouzi, “Image super-resolution via iterative refinement,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 4, pp. 4713–4726, 2022
2022
-
[257]
Pandiff: A novel pansharpening method based on denoising diffusion probabilistic model,
Q. Meng, W. Shi, S. Li, and L. Zhang, “Pandiff: A novel pansharpening method based on denoising diffusion probabilistic model,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–17, 2023
2023
-
[258]
Triposr: Fast 3d object reconstruction from a single image,
D. Tochilkin, D. Pankratz, Z. Liu, Z. Huang, A. Letts, Y . Li, D. Liang, C. Laforte, V . Jampani, and Y .-P. Cao, “Triposr: Fast 3d object reconstruction from a single image,” arXiv:2403.02151, 2024
2024 arXiv
-
[259]
Exploring the capability of text-to- image diffusion models with structural edge guidance for multi-spectral satellite image inpainting,
M. Czerkawski and C. Tachtatzis, “Exploring the capability of text-to- image diffusion models with structural edge guidance for multi-spectral satellite image inpainting,” IEEE Geosci. Remote Sens. Lett. , vol. 21, pp. 1–5, 2024
2024
-
[260]
A convnet for the 2020s,
Z. Liu, H. Mao, C.-Y . Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 11976–11986, 2022
2022
-
[261]
Ddpm-cd: Remote sensing change detection using denoising diffusion probabilistic mod- els,
W. G. C. Bandara, N. G. Nair, and V . M. Patel, “Ddpm-cd: Remote sensing change detection using denoising diffusion probabilistic mod- els,” arXiv:2206.11892, 2022
2022 arXiv
-
[262]
Crowdai,
C. M. Challenge, “Crowdai,” 2018
2018
-
[263]
Lending orientation to neural networks for cross- view geo-localization,
L. Liu and H. Li, “Lending orientation to neural networks for cross- view geo-localization,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 5624–5633, 2019
2019
-
[264]
Wide-area image geolocal- ization with aerial reference imagery,
S. Workman, R. Souvenir, and N. Jacobs, “Wide-area image geolocal- ization with aerial reference imagery,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , pp. 3961–3969, 2015
2015
-
[265]
Gemini: a family of highly capable multimodal models,
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. , “Gemini: a family of highly capable multimodal models,” arXiv:2312.11805, 2023
2023 arXiv
-
[266]
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proc. Annu. Meet. Assoc. Comput. Linguist. (ACL) , pp. 2556–2565, 2018
2018
-
[267]
Chatgpt,
OpenAI, “Chatgpt,” 2024
2024
-
[268]
Detecting building changes with off-nadir aerial images,
C. Pang, J. Wu, J. Ding, C. Song, and G.-S. Xia, “Detecting building changes with off-nadir aerial images,” Sci. China Inf. Sci. , vol. 66, no. 4, p. 140306, 2023
2023
-
[269]
A scene change detection framework for multi-temporal very high resolution remote sensing images,
C. Wu, L. Zhang, and L. Zhang, “A scene change detection framework for multi-temporal very high resolution remote sensing images,” Signal Process., vol. 124, pp. 184–197, 2016
2016
-
[270]
Crtranssar: A visual transformer based on contextual joint representation learning for sar ship detection,
R. Xia, J. Chen, Z. Huang, H. Wan, B. Wu, L. Sun, B. Yao, H. Xiang, and M. Xing, “Crtranssar: A visual transformer based on contextual joint representation learning for sar ship detection,” Remote Sens. , vol. 14, no. 6, p. 1488, 2022
2022
-
[271]
Deepglobe 2018: A challenge to parse the earth through satellite images,
I. Demir, K. Koperski, D. Lindenbaum, G. Pang, J. Huang, S. Basu, F. Hughes, D. Tuia, and R. Raskar, “Deepglobe 2018: A challenge to parse the earth through satellite images,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) Workshops, pp. 172–181, 2018
2018
-
[272]
Enabling country-scale land cover mapping with meter-resolution satellite imagery,
X.-Y . Tong, G.-S. Xia, and X. X. Zhu, “Enabling country-scale land cover mapping with meter-resolution satellite imagery,” ISPRS J. Photogrammetry Remote Sens. , vol. 196, pp. 178–196, 2023
2023
-
[273]
Coreval: A comprehensive and objective benchmark for evaluating the remote sensing capabilities of large vision-language models,
X. An, J. Sun, Z. Gui, and W. He, “Coreval: A comprehensive and objective benchmark for evaluating the remote sensing capabilities of large vision-language models,” arXiv:2411.18145, 2024
2024
-
[274]
Geobench-vlm: Benchmarking vision-language models for geospatial tasks,
M. S. Danish, M. A. Munir, S. R. A. Shah, K. Kuckreja, F. S. Khan, P. Fraccaro, A. Lacoste, and S. Khan, “Geobench-vlm: Benchmarking vision-language models for geospatial tasks,” arXiv:2411.19325, 2024
2024 arXiv
-
[275]
Airound and cv-brct: Novel multiview datasets for scene classification,
G. Machado, E. Ferreira, K. Nogueira, H. Oliveira, M. Brito, P. H. T. Gama, and J. A. dos Santos, “Airound and cv-brct: Novel multiview datasets for scene classification,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. , vol. 14, pp. 488–503, 2020
2020
-
[276]
Similarity learning for land use scene-level change detection,
J. Liu, W. Zhou, H. Guan, and W. Zhao, “Similarity learning for land use scene-level change detection,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. , vol. 17, pp. 6501–6513, 2024
2024
-
[277]
Firerisk: A remote sensing dataset for fire risk assessment with benchmarks using supervised and self-supervised learning,
S. Shen, S. Seneviratne, X. Wanyan, and M. Kirley, “Firerisk: A remote sensing dataset for fire risk assessment with benchmarks using supervised and self-supervised learning,” in Proc. Int. Conf. Digit. Image Comput.: Tech. Appl. (DICTA) , pp. 189–196, 2023
2023
-
[278]
Forest damages-larch casebearer,
S. F. Agency, “Forest damages-larch casebearer,” 2021
2021
-
[279]
Deforestation-satellite-imagery dataset,
CSE499DeforestationSatellite, “Deforestation-satellite-imagery dataset,” 2024
2024
-
[280]
Marine debris dataset for object detection in planetscope imagery,
A. Shah, L. Thomas, and M. Maskey, “Marine debris dataset for object detection in planetscope imagery,” 2021
2021
-
[281]
Rareplanes: Synthetic data takes flight,
J. Shermeyer, T. Hossler, A. Van Etten, D. Hogan, R. Lewis, and D. Kim, “Rareplanes: Synthetic data takes flight,” 2020
2020
-
[282]
Panoptic segmentation of satellite image time series with convolutional temporal attention networks,
V . S. F. Garnot and L. Landrieu, “Panoptic segmentation of satellite image time series with convolutional temporal attention networks,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 4872–4881, 2021
2021
-
[283]
Fpcd: An open aerial vhr dataset for farm pond change detection,
C. Tundia, R. Kumar, O. Damani, and G. Sivakumar, “Fpcd: An open aerial vhr dataset for farm pond change detection,” arXiv:2302.14554, 2023
2023 arXiv
-
[284]
Cross-domain landslide mapping from large-scale remote sensing images using prototype- guided domain-aware progressive representation learning,
X. Zhang, W. Yu, M.-O. Pun, and W. Shi, “Cross-domain landslide mapping from large-scale remote sensing images using prototype- guided domain-aware progressive representation learning,” ISPRS J. Photogrammetry Remote Sens. , vol. 197, pp. 1–17, 2023
2023
-
[285]
Synthesizing optical and sar imagery from land cover maps and auxiliary raster data,
G. Baier, A. Deschemps, M. Schmitt, and N. Yokoya, “Synthesizing optical and sar imagery from land cover maps and auxiliary raster data,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–12, 2021
2021
-
[286]
So2sat lcz42: A benchmark data set for the classification of global local climate zones [software and data sets],
X. X. Zhu, J. Hu, C. Qiu, Y . Shi, J. Kang, L. Mou, H. Bagheri, M. Haberle, Y . Hua, R. Huang, et al. , “So2sat lcz42: A benchmark data set for the classification of global local climate zones [software and data sets],” IEEE Geosci. Remote Sens. Mag., vol. 8, no. 3, pp. 76– 89, 2020
2020
-
[287]
Quakeset: A dataset and low-resource models to monitor earthquakes through sentinel-1,
D. R. Cambrin and P. Garza, “Quakeset: A dataset and low-resource models to monitor earthquakes through sentinel-1,” in Proc. Int. IS- CRAM Conf., 2024
2024
-
[288]
Earthnets: Empowering artificial intelligence for earth observation,
Z. Xiong, F. Zhang, Y . Wang, Y . Shi, and X. X. Zhu, “Earthnets: Empowering artificial intelligence for earth observation,” IEEE Geosci. Remote Sens. Mag. , vol. Early Access, pp. 2–36, 2024
2024
-
[289]
Rapid flood inun- dation mapping using social media, remote sensing and topographic data,
J. F. Rosser, D. G. Leibovici, and M. J. Jackson, “Rapid flood inun- dation mapping using social media, remote sensing and topographic data,” Nat. Hazards, vol. 87, pp. 103–120, 2017
2017
-
[290]
So- cial media: New perspectives to improve remote sensing for emergency response,
J. Li, Z. He, J. Plaza, S. Li, J. Chen, H. Wu, Y . Wang, and Y . Liu, “So- cial media: New perspectives to improve remote sensing for emergency response,” Proc. IEEE, vol. 105, no. 10, pp. 1900–1912, 2017
1900
-
[291]
Investigating the catastrophic forgetting in multimodal large language model fine-tuning,
Y . Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y . J. Lee, and Y . Ma, “Investigating the catastrophic forgetting in multimodal large language model fine-tuning,” in Conf. Parsimony Learn., vol. 234, pp. 202–227, 2024
2024
-
[292]
Model tailor: Mitigating catastrophic forgetting in multi-modal large language models,
D. Zhu, Z. Sun, Z. Li, T. Shen, K. Yan, S. Ding, C. Wu, and K. Kuang, “Model tailor: Mitigating catastrophic forgetting in multi-modal large language models,” in Proc. Int. Conf. Mach. Learn. (ICML), pp. 62581– 62598, 2024
2024
-
[293]
Text2earth: Unlocking text-driven remote sensing image generation with a global-scale dataset and a foundation model,
C. Liu, K. Chen, R. Zhao, Z. Zou, and Z. Shi, “Text2earth: Unlocking text-driven remote sensing image generation with a global-scale dataset and a foundation model,” arXiv:2501.00895, 2025
2025 arXiv
-
[294]
Ssl4eo-s12: A large-scale multimodal, multitemporal dataset for self-supervised learning in earth observation [software and data sets],
Y . Wang, N. A. A. Braham, Z. Xiong, C. Liu, C. M. Albrecht, and X. X. Zhu, “Ssl4eo-s12: A large-scale multimodal, multitemporal dataset for self-supervised learning in earth observation [software and data sets],” IEEE Geosci. Remote Sens. Mag. , vol. 11, no. 3, pp. 98–106, 2023
2023
-
[295]
Towards geospatial foundation models via continual pretraining,
M. Mendieta, B. Han, X. Shi, Y . Zhu, and C. Chen, “Towards geospatial foundation models via continual pretraining,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , pp. 16806–16816, 2023
2023
-
[296]
Rsi-cb: A large scale remote sensing image classification benchmark via crowdsource data,
H. Li, X. Dou, C. Tao, Z. Hou, J. Chen, J. Peng, M. Deng, and L. Zhao, “Rsi-cb: A large scale remote sensing image classification benchmark via crowdsource data,” arXiv:1705.10450, 2017
2017 arXiv
-
[297]
Machine-to-machine visual dia- loguing with chatgpt for enriched textual image description,
R. Ricci, Y . Bazi, and F. Melgani, “Machine-to-machine visual dia- loguing with chatgpt for enriched textual image description,” Remote Sens., vol. 16, no. 3, p. 441, 2024
2024
-
[298]
Language models are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. , “Language models are few-shot learners,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, pp. 1877–1901, 2020. 35
1901
-
[299]
Gqa: Training generalized multi-query transformer models from multi-head checkpoints,
J. Ainslie, J. Lee-Thorp, M. De Jong, Y . Zemlyanskiy, F. Lebr ´on, and S. Sanghai, “Gqa: Training generalized multi-query transformer models from multi-head checkpoints,” in Proc. Conf. Empirical Methods Nat. Lang. Process. (EMNLP) , pp. 4895–4901, 2023
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.