REVIEW 4 major objections 5 minor 74 references
Bridging Scales in Map Generation: A scale-aware cascaded generative mapping framework for seamless and consistent multi-scale cartographic representation
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read SCGM claims that multi-scale tile maps can be generated from remote sensing imagery by one self-cascading latent diffusion model that conditions each finer scale on the previously generated coarser map tile plus textual scale information…
desk verdict A genuinely novel cascaded latent diffusion idea for multi-scale map generation, but the evaluation never specifies whether test-time cascade references are ground truth or self-generated, and one ablation claim directly contradicts its own table. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Three modules carry the argument: ScaleEncoder, a CLIP-based text-to-embedding module that turns scale descriptors into a condition vector; MFEncoder, a dual-branch encoder that fuses remote sensing features with upsampled features from the smaller-scale reference tile using SPADE normalization; and SFAdapter, which projects the fused condition into multi-resolution features that are merged with corresponding U-Net decoder layers. The load-bearing mechanism is the cascade itself: the output tile at scale $k$ is reused as the spatial prior to generate the tile at scale $k+1$, so each finer tile inherits the structure of the coarser map, which is what suppresses edge artifacts and preserves cross-scale consistency.
What would settle it
Run SCGM recursively over several levels, for example from level 14 to level 18, and compare the resulting maps against the same model when ground-truth coarser tiles are used as cascade references instead of generated ones. If quality, measured by FID, MFP, and measured edge continuity, degrades steadily with recursion depth, or if substituting generated references for true references leaves performance unchanged, the cascade claim would be shown to be either harmful or vacuous. A second check is to stitch the generated tiles and measure cross-tile feature continuity against independent tile generation on identical data.
Extended reading notes
Core claim
The discovery is that scale is a conditioning modality rather than a fixed level label. The authors show that by embedding scale information such as map level, spatial resolution, scale ratio, and geospatial descriptions through a pretrained CLIP encoder, and by feeding the previously generated smaller-scale map tile as a cascade reference into the denoising U-Net, the same model can produce maps at successive scales that remain geographically aligned and visually continuous across tile boundaries. This is the SCGM framework. Their experiments claim that this formulation outperforms LACG, SMAPGAN, Pix2Pix, and other compared methods on MLMG-CN, MLMG-US, and their own CSCMG dataset, both in pixel-level quality metrics and in a semantic map-feature perception metric.
Load-bearing premise
The cascade works only if the smaller-scale map tile used as the reference is geographically correct and pixel-aligned with the remote sensing tile for the same area; if that prior strays, later scales inherit its errors, and the paper does not quantify how these errors accumulate across recursive stages.
Editorial extensions
If this is right
- Emergency responders could obtain a fresh multi-scale tile set from a single satellite pass without waiting for manual vector generalization.
- Tile stitching artifacts disappear by construction because each finer tile is generated in the context of the coarser tile that spatially surrounds it.
- The same trained model can in principle be iterated past the trained levels, producing arbitrarily large and seamless map extents.
- The ablation results indicate that a wider cascade span, 4X instead of 2X, improves generation quality, so cascade span is a tunable control over how much cartographic context the model uses.
- The scale encoder actively represents map generalization, so outputs at 1:2,000 contain building-level detail while outputs at 1:8,000 show generalized structure rather than a single texture copied across zoom levels.
Reading between the lines
- If the cascade prior is trustworthy, the same architecture should transfer to other tile hierarchies, such as OpenStreetMap-style or topographic tiles, as long as paired multi-scale imagery and map tiles can be assembled; the paper's CSCMG dataset is a template for that resource.
- The paper compares 2X and 4X cascade spans but does not analyze adaptive span selection; a testable extension is choosing the reference span region-by-region to avoid propagating artifacts in natural landscapes or small scales.
- Because scale descriptors are encoded via CLIP text embeddings, a natural extension is text-driven map style control, where a user specifies a custom scale string and the generator conditions the output accordingly.
- The success of the MFP metric for evaluation raises the possibility of using it as an auxiliary loss during training; the paper's own results suggest pixel-level metrics alone miss cartographic validity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SCGM, a scale-aware cascaded latent diffusion framework for generating multi-scale tile maps from remote sensing imagery. The framework introduces three components: a CLIP-based scale encoder that embeds scale information, an MFEncoder that fuses remote sensing and cascade reference features via SPADE blocks, and an SFAdapter that injects conditional features into the denoising U-Net. The method is evaluated on the MLMG dataset (Table 2) and a new CSCMG dataset (Table 3), reporting improved FID/PSNR/SSIM and a new MFP metric. The paper claims state-of-the-art performance and seamless cross-scale consistency, supported by qualitative visual comparisons and ablations in Table 4.
Significance. The framework addresses a real need in automated cartography, and the idea of conditioning each scale on a coarser prior is principled. The MLMG results are based on standard metrics and an external benchmark, which is a strength. However, the paper's main evidence for geographic fidelity rests on the self-defined MFP metric, whose direction is internally inconsistent, and the test-time conditioning source is ambiguous. These issues currently prevent the strong claims from being fully substantiated.
major comments (4)
- [4.3.2 / Table 4] The text in Section 4.3.2 states that incorporating the ScaleEncoder results in improvements across all metrics compared to the baseline, but Table 4 shows that with the 2X cascade, adding the ScaleEncoder worsens FID from 45.512 to 50.052; only the 4X row supports the claim. The narrative must be corrected to address this interaction between cascade span and scale encoder.
- [3.7 / Table 3] The direction of the MFP metric is contradictory: Section 3.7 says 'Lower MFP values indicate superior semantic fidelity and spatial coherence,' while Table 3 labels the column 'MFP↑' and the text celebrates SCGM's highest MFP of 0.7351. Because MFP is defined as a weighted sum of a similarity term (higher is better) and a distance term (lower is better), the interpretation is ambiguous. This metric is central to the paper's geographic-fidelity claims, so the contradiction must be resolved.
- [3.3 / 4.2] The test-time source of the cascade reference x0^(k) is never specified. The method description in Section 3.1 says the model conditions on the 'previously generated smaller-scale tile,' but the dataset construction (Figure 6) and training protocol use ground-truth coarser tiles. If evaluation uses ground-truth references, SCGM receives oracle information that the baselines do not, and the FID/PSNR gains in Tables 2 and 3 no longer support the claim of autonomous multi-scale generation; if self-generated references are used, error accumulation over recursive stages is unanalyzed. The authors must state the test-time conditioning source and, ideally, report both settings.
- [3.7 / Eq. (15)] The MFP metric is defined in the authors' companion paper (Sun and Bai, 2025) with weights lambda1=10 and lambda2=1 described as 'empirically determined,' and no external validation is provided. Using a self-defined metric as the primary evidence for the central claim of geographic fidelity is circular; the paper should validate MFP against human judgments or established semantic metrics, or de-emphasize it in the conclusions.
minor comments (5)
- [4.2.1] The text refers to 'LCAG' and 'LCAG' while the baseline is named LACG elsewhere; please use a consistent name throughout.
- [Table 3 / 4.2.3] The baseline 'SMAGAN' in Table 3 and Section 4.2.3 should be 'SMAPGAN'.
- [3.2 / Eq. (2), (11)] The formulations y = arg min_y -log p(y|x) and y = arg min_y -log p_theta(x|c) do not correspond to the diffusion training or sampling procedure; they appear to confuse optimization over the output with maximum-likelihood estimation. Please rephrase using the standard noise-prediction objective (Eq. 12) and reverse-process sampling.
- [4.3.2] The phrase 'an 8.3 increase in the FID score' is ambiguous because FID is lower-is-better; the authors likely mean a decrease or improvement and should state this clearly.
- [Figure 9 / Section 5.1] The smoothness and continuity claim is supported only by qualitative visual evidence; consider adding a quantitative measure of tile-boundary continuity to substantiate the 'seamless' claim.
Circularity Check
Partially self-referential evaluation: the cartographic-fidelity metric MFP is defined in the authors' own companion paper with empirically chosen weights, but FID/PSNR on an external benchmark provide independent support.
-
self citation load bearing
[Section 3.7 (Eqs. 13-15), Section 4.2.3, and Section 6]
"To address the limitations of pixel or distribution-based metrics in assessing semantic-level cartographic features, we employ the map feature perception metric MFP (Sun and Bai, 2025). ... MFP = λ1 G + λ2 S with λ1 = 10 and λ2 = 1 empirically determined. Lower MFP values indicate superior semantic fidelity and spatial coherence. A detailed explanation of the interpretability and spatial consistency of the MFP metric is provided in (Sun and Bai, 2025)."
The distinctive claim that SCGM preserves cartographic semantic fidelity and spatial coherence is measured by MFP, a metric defined in the authors' own companion paper (Sun and Bai, 2025), with combination weights that are 'empirically determined' rather than derived or externally validated. The evaluation section then uses this metric as the key evidence that SCGM's scale-aware cartographic generalization is real, and the conclusion repeats 'as measured by the map feature perception metric.' Thus the distinctive part of the quality claim rests on an author-defined, empirically weighted measure rather than on an independently established instrument. This is a load-bearing self-citation.
full rationale
The core derivation chain in Sections 3.1-3.6 is not circular: Eq. (3) defines a conditional generation objective, Eq. (4) specifies how the smaller-scale cascade tile is encoded and fused, and Eq. (12) is a standard diffusion denoising loss. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and the cascaded architecture is a normal conditional-design choice rather than a definitional restatement of the claimed result. The evaluation does contain one self-referential component: the cartographic-fidelity metric MFP is introduced in the authors' companion paper with empirically determined weights, and the paper's conclusion leans on MFP to support the geographic-fidelity claim. That is a genuine but partial self-citation issue. I also note that the manuscript never explicitly states whether the cascade tile used at test time on MLMG is the model's own previous output or a ground-truth lower-scale tile; if it were the latter, the seamless-consistency result would be an oracle-conditioned comparison. However, the paper describes the cascade as 'previously generated smaller-scale tiles,' and no direct quote establishes the ground-truth-at-test reduction needed for a stronger circularity charge under the stated hard rules. Overall, the central performance claim has independent support from FID/PSNR comparisons on the external MLMG dataset, so the score is 4 rather than higher.
Assumptions & free parameters
free parameters (2)
- MFP weights lambda1 and lambda2 =
lambda1 = 10, lambda2 = 1
- Cascade span =
4X (versus 2X)
assumptions (5)
- domain assumption Stable Diffusion 2.1 base latent space can be adapted to cartographic map tiles by fine-tuning a linear layer and adapters.
- domain assumption Lower-scale map tiles are valid spatial priors for the next scale, and cascaded conditioning prevents rather than propagates errors.
- domain assumption CLIP text embeddings of strings such as 'Level:16, resolution: 2.389, scale: 1:8000' capture meaningful cartographic scale semantics.
- domain assumption Google Maps tiles are an acceptable ground truth for cartographic generalization, and the RS-map pairs in CSCMG are correctly aligned.
- ad hoc to paper The MFP metric measures map feature perception and spatial consistency as claimed.
invented entities (1)
-
MFP (Map Feature Perception metric)
Cite this review
Pith. "Pith review of Bridging Scales in Map Generation: A scale-aware cascaded generative mapping framework for seamless and consistent multi-scale cartographic representation." pith.science (2026). https://pith.science/paper/FJDYNF56
@misc{pith2026250204991,
author = {Pith},
title = {Pith review of: Bridging Scales in Map Generation: A scale-aware cascaded generative mapping framework for seamless and consistent multi-scale cartographic representation},
year = {2026},
howpublished = {\url{https://pith.science/paper/FJDYNF56}},
note = {Machine review of arXiv:2502.04991}
}
read the original abstract
Multi-scale tile maps are essential for geographic information services, serving as fundamental outcomes of surveying and cartographic workflows. While existing image generation networks can produce map-like outputs from remote sensing imagery, their emphasis on replicating texture rather than preserving geospatial features limits cartographic validity. Current approaches face two fundamental challenges: inadequate integration of cartographic generalization principles with dynamic multi-scale generation and spatial discontinuities arising from tile-wise generation. To address these limitations, we propose a scale-aware cartographic generation framework (SCGM) that leverages conditional guided diffusion and a multi-scale cascade architecture. The framework introduces three key innovations: a scale modality encoding mechanism to formalize map generalization relationships, a scale-driven conditional encoder for robust feature fusion, and a cascade reference mechanism ensuring cross-scale visual consistency. By hierarchically constraining large-scale map synthesis with small-scale structural priors, SCGM effectively mitigates edge artifacts while maintaining geographic fidelity. Comprehensive evaluations on cartographic benchmarks confirm the framework's ability to generate seamless multi-scale tile maps with enhanced spatial coherence and generalization-aware representation, demonstrating significant potential for emergency mapping and automated cartography applications.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
I. Goodfellow, J. Pouget-Abadie , M. Mirza, B. Xu, D. Warde-Farley , S. Ozair, A. Courville, and Y. Bengio, ``Generative Adversarial Nets ,'' in Advances in Neural Information Processing Systems 27 , Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, Eds. 1em plus 0.5em minus 0.4em Curran Associates, Inc., 2014, pp. 2672--2680
work page 2014
-
[2]
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, ``Image-to- Image Translation with Conditional Adversarial Networks ,'' in 2017 IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) . 1em plus 0.5em minus 0.4em Honolulu, HI: IEEE, Jul. 2017, pp. 5967--5976
work page 2017
-
[3]
T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, A. Tao, J. Kautz, and B. Catanzaro, ``High- Resolution Image Synthesis and Semantic Manipulation with Conditional GANs ,'' arXiv:1711.11585 [cs], Aug. 2018
arXiv 2018
-
[4]
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, ``Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks ,'' arXiv:1703.10593 [cs], Nov. 2018
arXiv 2018
-
[5]
S. Ganguli, P. Garzon, and N. Glaser, `` GeoGAN : A Conditional GAN with Reconstruction and Style Loss to Generate Standard Layer of Maps from Satellite Images ,'' Apr. 2019
work page 2019
-
[6]
X. Chen, S. Chen, T. Xu, B. Yin, J. Peng, X. Mei, and H. Li, `` SMAPGAN : Generative Adversarial Network-Based Semisupervised Styled Map Tile Generation Method ,'' IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 5, pp. 4388--4406, May 2021
work page 2021
-
[7]
Y. Fu, S. Liang, D. Chen, and Z. Chen, ``Translation of Aerial Image Into Digital Map via Discriminative Segmentation and Creative Generation ,'' IEEE Transactions on Geoscience and Remote Sensing, pp. 1--15, 2021
work page 2021
-
[8]
Y. Liu, W. Wang, F. Fang, L. Zhou, C. Sun, Y. Zheng, and Z. Chen, `` CscGAN : Conditional Scale-Consistent Generation Network for Multi-Level Remote Sensing Image to Map Translation ,'' Remote Sensing, vol. 13, no. 10, p. 1936, May 2021
work page 1936
Show all 74 references
-
[9]
Y. Fu, Z. Fang, L. Chen, T. Song, and D. Lin, ``Level- Aware Consistent Multilevel Map Translation From Satellite Imagery ,'' IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1--14, 2023
2023
-
[10]
J. Ho, A. Jain, and P. Abbeel, ``Denoising Diffusion Probabilistic Models ,'' Dec. 2020
2020
-
[11]
J. Song, C. Meng, and S. Ermon, ``Denoising Diffusion Implicit Models ,'' Oct. 2022
2022
-
[12]
Dhariwal and A
P. Dhariwal and A. Nichol, ``Diffusion Models Beat GANs on Image Synthesis ,'' Jun. 2021
2021
-
[13]
T. Wang, T. Zhang, B. Zhang, H. Ouyang, D. Chen, Q. Chen, and F. Wen, ``Pretraining is All You Need for Image-to-Image Translation ,'' May 2022
2022
-
[14]
Saharia, W
C. Saharia, W. Chan, H. Chang, C. Lee, J. Ho, T. Salimans, D. Fleet, and M. Norouzi, ``Palette: Image-to-Image Diffusion Models ,'' in Special Interest Group on Computer Graphics and Interactive Techniques Conference Proceedings . 1em plus 0.5em minus 0.4em Vancouver BC Canada...
2022
-
[15]
W. Wang, J. Bao, W. Zhou, D. Chen, D. Chen, L. Yuan, and H. Li, ``Semantic Image Synthesis via Diffusion Models ,'' Nov. 2022
2022
-
[16]
Zhong, R
L. Zhong, R. Onishi, L. Wang, L. Ruan, and S. J. Tan, ``A Scalable Blockchain-based High-Definition Map Update Management System ,'' in 2021 IEEE International Smart Cities Conference ( ISC2 ) , Sep. 2021, pp. 1--4
2021
-
[17]
R. Wang, H. Jiang, and Y. Li, `` UPerNet with ConvNeXt for Semantic Segmentation ,'' in 2023 IEEE 3rd International Conference on Electronic Technology , Communication and Information ( ICETCI ) , May 2023, pp. 764--769
2023
-
[18]
X. Ma, X. Zhang, M.-O. Pun, and M. Liu, ``A Multilevel Multimodal Fusion Transformer for Remote Sensing Semantic Segmentation ,'' IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1--15, 2024
2024
-
[19]
Zhang, X
J. Zhang, X. Yang, R. Jiang, W. Shao, and L. Zhang, `` RSAM-Seg : A SAM-based Approach with Prior Knowledge Integration for Remote Sensing Image Semantic Segmentation ,'' Feb. 2024
2024
-
[20]
Toker, M
A. Toker, M. Eisenberger, D. Cremers, and L. Leal-Taix \'e , `` SatSynth : Augmenting image-mask pairs through diffusion models for aerial semantic segmentation,'' Mar. 2024
2024
-
[21]
Z. Zhao, H. Bai, J. Zhang, Y. Zhang, S. Xu, Z. Lin, R. Timofte, and L. Van Gool, `` CDDFuse : Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2023,...
2023
-
[22]
D. Peng, P. Hu, Q. Ke, and J. Liu, ``Diffusion-based Image Translation with Label Guidance for Domain Adaptive Semantic Segmentation ,'' in Proceedings of the IEEE / CVF International Conference on Computer Vision , 2023, pp. 808--820
2023
-
[23]
J. Lu, G. He, H. Dou, Q. Gao, L. Fang, and Y. Deng, `` ScoreSeg : Leveraging Score-based Generative Model for Self-Supervised Semantic Segmentation of Remote Sensing ,'' IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, pp. 1--16, 2023
2023
-
[24]
Ayala, R
C. Ayala, R. Sesma, C. Aranda, and M. Galar, ``Diffusion models for remote sensing imagery semantic segmentation,'' in IGARSS 2023 - 2023 IEEE International Geoscience and Remote Sensing Symposium , Jul. 2023, pp. 5654--5657
2023
-
[25]
Z. Chen, D. Li, W. Fan, H. Guan, C. Wang, and J. Li, ``Self- Attention in Reconstruction Bias U-Net for Semantic Segmentation of Building Rooftops in Optical Remote Sensing Images ,'' Remote Sensing, vol. 13, no. 13, p. 2524, Jun. 2021
2021
-
[26]
Cheng, I
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, ``Masked- Attention Mask Transformer for Universal Image Segmentation ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 1290--1299
2022
-
[27]
T. Shen, Y. Zhang, L. Qi, J. Kuen, X. Xie, J. Wu, Z. Lin, and J. Jia, ``High Quality Segmentation for Ultra High-Resolution Images ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 1310--1319
2022
-
[28]
Jiang, C
J. Jiang, C. Lyu, S. Liu, Y. He, and X. Hao, `` RWSNet : A semantic segmentation network based on SegNet combined with random walk for remote sensing,'' International Journal of Remote Sensing, vol. 41, no. 2, pp. 487--505, Jan. 2020
2020
-
[29]
J. Yao, B. Zhang, C. Li, D. Hong, and J. Chanussot, ``Extended Vision Transformer ( ExViT ) for Land Use and Land Cover Classification : A Multimodal Deep Learning Framework ,'' IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1--15, 2023
2023
-
[30]
F. Wang, J. Ji, and Y. Wang, `` DSViT : Dynamically Scalable Vision Transformer for Remote Sensing Image Segmentation and Classification ,'' IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 16, pp. 5441--5452, 2023
2023
-
[31]
X. Chen, Z. Liu, H. Tang, L. Yi, H. Zhao, and S. Han, `` SparseViT : Revisiting Activation Sparsity for Efficient High-Resolution Vision Transformer ,'' Mar. 2023
2023
-
[32]
Y. Li, J. Luo, Y. Zhang, Y. Tan, J.-G. Yu, and S. Bai, ``Learning to Holistically Detect Bridges From Large-Size VHR Remote Sensing Imagery ,'' IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1--18, 2024
2024
-
[33]
S. Zhao, H. Chen, X. Zhang, P. Xiao, L. Bai, and W. Ouyang, `` RS-mamba for large remote sensing image dense prediction,'' Mar. 2024
2024
-
[34]
S. Guo, L. Liu, Z. Gan, Y. Wang, W. Zhang, C. Wang, G. Jiang, W. Zhang, R. Yi, L. Ma, and K. Xu, `` ISDNet : Integrating Shallow and Deep Networks for Efficient Ultra-high Resolution Segmentation ,'' in 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition ( CV...
2022
-
[35]
J. Xi, O. K. Ersoy, J. Fang, M. Cong, T. Wu, C. Zhao, and Z. Li, ``Wide Sliding Window and Subsampling Network for Hyperspectral Image Classification ,'' Remote Sensing, vol. 13, no. 7, p. 1290, Mar. 2021
2021
-
[36]
P. Luc, C. Couprie, S. Chintala, and J. Verbeek, ``Semantic Segmentation using Adversarial Networks ,'' Nov. 2016
2016
-
[37]
Z. Chen, C. Wang, J. Li, N. Xie, Y. Han, and J. Du, ``Reconstruction Bias U-Net for Road Extraction From Optical Remote Sensing Images ,'' IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 2284--2294, 2021
2021
-
[38]
X. Yang, J. Yang, J. Yan, Y. Zhang, T. Zhang, Z. Guo, X. Sun, and K. Fu, `` SCRDet : Towards More Robust Detection for Small , Cluttered and Rotated Objects ,'' in Proceedings of the IEEE / CVF International Conference on Computer Vision , 2019, pp. 8232--8241
2019
-
[39]
J. Ding, N. Xue, Y. Long, G.-S. Xia, and Q. Lu, ``Learning RoI Transformer for Oriented Object Detection in Aerial Images ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 2849--2858
2019
-
[40]
S. Yin, H. Li, and L. Teng, ``Airport Detection Based on Improved Faster RCNN in Large Scale Remote Sensing Images ,'' Sensing and Imaging, vol. 21, no. 1, p. 49, Dec. 2020
2020
-
[41]
Li and C
P. Li and C. Che, `` SeMo-YOLO : A Multiscale Object Detection Network in Satellite Remote Sensing Images ,'' in 2021 International Joint Conference on Neural Networks ( IJCNN ) . 1em plus 0.5em minus 0.4em Shenzhen, China: IEEE, Jul. 2021, pp. 1--8
2021
-
[42]
J. Han, J. Ding, N. Xue, and G.-S. Xia, `` ReDet : A Rotation-Equivariant Detector for Aerial Object Detection ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 2786--2795
2021
-
[43]
C. Xia, X. Wang, F. Lv, X. Hao, and Y. Shi, `` ViT-CoMer : Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 5493--5502
2024
-
[44]
Muhtar, Z
D. Muhtar, Z. Li, F. Gu, X. Zhang, and P. Xiao, `` LHRS-Bot : Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model ,'' Feb. 2024
2024
-
[45]
Zhang, M
W. Zhang, M. Cai, T. Zhang, Y. Zhuang, and X. Mao, `` EarthGPT : A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing Domain ,'' IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1--20, 2024
2024
-
[46]
Kuckreja, M
K. Kuckreja, M. S. Danish, M. Naseer, A. Das, S. Khan, and F. S. Khan, `` GeoChat : Grounded Large Vision-Language Model for Remote Sensing ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 27\,831--27\,840
2024
-
[47]
X. Guo, J. Lao, B. Dang, Y. Zhang, L. Yu, L. Ru, L. Zhong, Z. Huang, K. Wu, D. Hu, H. He, J. Wang, J. Chen, M. Yang, Y. Zhang, and Y. Li, `` SkySense : A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery ,'' in Proceedin...
2024
-
[48]
U. Mall, C. P. Phoo, M. K. Liu, C. Vondrick, B. Hariharan, and K. Bala, ``Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment ,'' Dec. 2023
2023
-
[49]
Zhang, Y
Y. Zhang, Y. Yin, R. Zimmermann, G. Wang, J. Varadarajan, and S.-K. Ng, ``An Enhanced GAN Model for Automatic Satellite-to-Map Image Conversion ,'' IEEE Access, vol. 8, pp. 176\,704--176\,716, 2020
2020
-
[50]
J. Li, Z. Chen, X. Zhao, and L. Shao, `` MapGAN : An Intelligent Generation Model for Network Tile Maps ,'' Sensors, vol. 20, no. 11, p. 3119, May 2020
2020
-
[51]
Tasar, S
O. Tasar, S. L. Happy, Y. Tarabalka, and P. Alliez, `` ColorMapGAN : Unsupervised Domain Adaptation for Semantic Segmentation Using Color Mapping Generative Adversarial Networks ,'' IEEE Transactions on Geoscience and Remote Sensing, vol. 58, no. 10, pp. 7178--7193, Oct. 2020
2020
-
[52]
X. Chen, B. Yin, S. Chen, H. Li, and T. Xu, ``Generating Multiscale Maps From Satellite Images via Series Generative Adversarial Networks ,'' IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1--5, 2022
2022
-
[53]
Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo, `` StarGAN : Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation ,'' in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8789--8797
2018
-
[54]
Sohl-Dickstein , E
J. Sohl-Dickstein , E. Weiss, N. Maheswaranathan, and S. Ganguli, ``Deep Unsupervised Learning using Nonequilibrium Thermodynamics ,'' in Proceedings of the 32nd International Conference on Machine Learning . 1em plus 0.5em minus 0.4em PMLR, Jun. 2015, pp. 2256--2265
2015
-
[55]
Nichol and P
A. Nichol and P. Dhariwal, ``Improved Denoising Diffusion Probabilistic Models ,'' Feb. 2021
2021
-
[56]
J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans, ``Cascaded Diffusion Models for High Fidelity Image Generation ,'' Journal of Machine Learning Research, vol. 23, no. 47, pp. 1--33, 2022
2022
-
[57]
Saharia, J
C. Saharia, J. Ho, W. Chan, T. Salimans, D. J. Fleet, and M. Norouzi, ``Image Super-Resolution Via Iterative Refinement ,'' IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1--14, 2022
2022
-
[58]
Zhang, A
L. Zhang, A. Rao, and M. Agrawala, ``Adding Conditional Control to Text-to-Image Diffusion Models ,'' in 2023 IEEE / CVF International Conference on Computer Vision ( ICCV ) . 1em plus 0.5em minus 0.4em Paris, France: IEEE, Oct. 2023, pp. 3813--3824
2023
-
[59]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, ``High- Resolution Image Synthesis With Latent Diffusion Models ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition . 1em plus 0.5em minus 0.4em arXiv, 2022, pp. 10\,684--10\,695
2022
-
[60]
Saharia, W
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi, ``Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding ,'' Advances in Neural Informatio...
2022
-
[61]
X. Wu, D. Zhang, R. Gan, J. Lu, Z. Wu, R. Sun, J. Zhang, P. Zhang, and Y. Song, ``Taiyi-diffusion- XL : Advancing bilingual text-to-image generation with large vision-language model support,'' Jun. 2024
2024
-
[62]
Avrahami, D
O. Avrahami, D. Lischinski, and O. Fried, ``Blended Diffusion for Text-Driven Editing of Natural Images ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18\,208--18\,218
2022
-
[63]
S. Nie, H. A. Guo, C. Lu, Y. Zhou, C. Zheng, and C. Li, ``The Blessing of Randomness : SDE Beats ODE in General Diffusion-based Image Editing ,'' Nov. 2023
2023
-
[64]
Y. Shi, C. Xue, J. H. Liew, J. Pan, H. Yan, W. Zhang, V. Y. F. Tan, and S. Bai, `` DragDiffusion : Harnessing Diffusion Models for Interactive Point-based Image Editing ,'' Jun. 2023
2023
-
[65]
Hertz, R
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or , ``Prompt-to- Prompt Image Editing with Cross Attention Control ,'' Aug. 2022
2022
-
[66]
H. Li, Y. Yang, M. Chang, S. Chen, H. Feng, Z. Xu, Q. Li, and Y. Chen, `` SRDiff : Single image super-resolution with diffusion probabilistic models,'' Neurocomputing, vol. 479, pp. 47--59, Mar. 2022
2022
-
[67]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, ``Learning Transferable Visual Models From Natural Language Supervision ,'' Feb. 2021
2021
-
[68]
Cherti, R
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev, ``Reproducible scaling laws for contrastive language-image learning,'' in 2023 IEEE / CVF Conference on Computer Vision and Pattern Recognition ( CVPR ) , Jun. 202...
2023
-
[69]
Vaswani, ``Attention is all you need,'' Advances in Neural Information Processing Systems, 2017
A. Vaswani, ``Attention is all you need,'' Advances in Neural Information Processing Systems, 2017
2017
-
[70]
Z. Wang, A. Bovik, H. Sheikh, and E. Simoncelli, ``Image quality assessment: From error visibility to structural similarity,'' IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600--612, Apr. 2004
2004
-
[71]
Heusel, H
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, `` GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium ,'' in Advances in Neural Information Processing Systems , vol. 30. 1em plus 0.5em minus 0.4em Curran Associates, Inc., 2017
2017
-
[72]
Park, M.-Y
T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu, ``Semantic Image Synthesis With Spatially-Adaptive Normalization ,'' in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition . 1em plus 0.5em minus 0.4em Proceedings of the IEEE/CVF Conference on Com...
2019
-
[73]
Ledig, L
C. Ledig, L. Theis, F. Huszar, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, and W. Shi, ``Photo- Realistic Single Image Super-Resolution Using a Generative Adversarial Network ,'' in Proceedings of the IEEE Conference on Computer Vision and P...
2017
-
[74]
Y. Pei, Y. Huang, Q. Zou, Y. Lu, and S. Wang, ``Does Haze Removal Help CNN-based Image Classification ?'' in Proceedings of the European Conference on Computer Vision ( ECCV ) , 2018, pp. 682--697
2018
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.