Pith. sign in

REVIEW 1 cited by

A Complex-valued SAR Foundation Model Based on Physically Inspired Representation Learning

T0 review · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read A complex-valued SAR foundation model, pre-trained with polarimetric decomposition losses, improves segmentation, detection, and classification on six radar benchmarks.

arxiv 2504.11999 v1 pith:QN3UEMQP submitted 2025-04-16 cs.CV

classification cs.CV
keywords foundationscatteringmodelcoefficientspowercomplex-valueddecompositiondownstream
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Radar images from satellites are complex: each pixel has an amplitude and a phase, and full polarization gives four channels (HH, HV, VH, VV). Most AI models for radar throw away everything except the amplitude. This paper keeps the full complex signal and builds a large pretrained network, a foundation model, that learns to describe each pixel as a mix of basic scattering behaviors: surface bounce, double bounce, volume scattering, helix scattering, and other finer mechanisms. This is the same idea as the Yamaguchi decomposition used by radar scientists.

The network is trained without labels on 400,000 radar images. It uses two losses: one that forces its predicted scattering categories to match the classic Yamaguchi decomposition (converted into simple yes/no labels), and one that forces the predicted scattering powers to add up to the total radar power of the pixel. The authors design special 'scattering queries', vectors that are supposed to represent each physical scattering mechanism, and the network learns to match image regions to these queries.

After pretraining, the encoder part of the network is used for downstream tasks: semantic segmentation, few-shot segmentation, unsupervised classification, ship detection, aircraft detection, and segmentation of ordinary amplitude-only SAR images. The paper reports consistent improvements over existing radar and remote sensing foundation models.

The main caveats are that the comparisons are not always apples-to-apples, no error bars are given, and the code and data are not released. The physical interpretability claim rests on an unusual step where the queries are initialized by feeding random matrix equations through a language model, which is not rigorously justified.

Extended reading notes

Core claim

The paper's load-bearing assertion is that simulating the physical process of polarimetric decomposition during self-supervised pretraining on complex-valued SAR data yields a foundation model that achieves state-of-the-art performance across six downstream tasks and generalizes even in data-scarce conditions (Abstract and Section V). If true, it shows that physically inspired pretext tasks on full complex-valued SAR data provide better transferable representations than amplitude-based masked image modeling.

Load-bearing premise

The scattering queries, which the paper claims represent independent physical scattering bases, are initialized by generating random sample pairs satisfying X=TY, encoding them with BERT, and averaging the resulting vectors (Section IV-C1, Fig. 6). The paper assumes this language-model encoding preserves the physical semantics of the nine scattering matrices. If it does not, the 'physically meaningful' queries are essentially arbitrary initializations, the interpretability claim loses its foundation, and the method reduces to a self-supervised pretraining with a Yamaguchi-binarization loss.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

The central claim rests on two ad hoc modeling choices: the BERT-based query initialization and the median-based binarization thresholds. The Yamaguchi decomposition and SPAN identities are standard physics borrowed from prior literature. The adaptive scattering basis is an invented component with no independent corroboration.

free parameters (3)
  • alpha (power loss weight) = 0.1
    Constant balancing power self-supervision loss in Eq. 13; set manually with no sensitivity analysis.
  • theta_i (Yamaguchi binarization thresholds) = Value at cumulative probability 0.5 per component, not reported numerically
    Thresholds are computed from the pre-training data distribution to binarize the four Yamaguchi coefficients into BCE targets (Eq. 8, Section IV-D1). They are data-derived and per-component.
  • Number of scattering bases N=10 = 10 (9 physical + 1 adaptive)
    Chosen based on seven- and eight-component decomposition literature; the adaptive basis is added ad hoc to absorb residual power (Section IV-C1).
assumptions (4)
  • domain assumption Yamaguchi decomposition provides ground-truth scattering coefficients (Eq. 4).
    The paper treats the deterministic Yamaguchi algorithm as the physical ground truth for the training targets in the polarimetric decomposition loss (Section IV-D1).
  • standard math SPAN equals the sum of ten scattering powers (Eq. 10) and equals the sum of squared channel magnitudes (Eq. 11).
    Standard polarimetric identity from Cloude and Pottier [18]; used to define the power self-supervision loss.
  • ad hoc to paper BERT encoding of random samples X=TY produces feature vectors that preserve the semantics of the nine physical scattering bases (Section IV-C1).
    No theoretical or empirical validation beyond a t-SNE plot; the mapping from matrix equations to language-model embeddings is unjustified.
  • ad hoc to paper Scattering values follow a Rayleigh distribution, so cumulative probability 0.5 is an appropriate binarization threshold (Section IV-D1).
    The paper fits Rayleigh distributions to the data and uses the median as threshold; this is a modeling choice not derived from scattering physics.
invented entities (1)
  • Adaptive scattering basis [Ta] (Pa)
    purpose: Accounts for residual power beyond the nine physical components, allowing the power sum to match SPAN.
    Introduced as a tenth basis to absorb unexplained power; no physical mechanism or external prediction is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Complex-valued SAR Foundation Model Based on Physically Inspired Representation Learning." pith.science (2026). https://pith.science/paper/QN3UEMQP

@misc{pith2026250411999,
  author       = {Pith},
  title        = {Pith review of: A Complex-valued SAR Foundation Model Based on Physically Inspired Representation Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QN3UEMQP}},
  note         = {Machine review of arXiv:2504.11999}
}
read the original abstract

Vision foundation models in remote sensing have been extensively studied due to their superior generalization on various downstream tasks. Synthetic Aperture Radar (SAR) offers all-day, all-weather imaging capabilities, providing significant advantages for Earth observation. However, establishing a foundation model for SAR image interpretation inevitably encounters the challenges of insufficient information utilization and poor interpretability. In this paper, we propose a remote sensing foundation model based on complex-valued SAR data, which simulates the polarimetric decomposition process for pre-training, i.e., characterizing pixel scattering intensity as a weighted combination of scattering bases and scattering coefficients, thereby endowing the foundation model with physical interpretability. Specifically, we construct a series of scattering queries, each representing an independent and meaningful scattering basis, which interact with SAR features in the scattering query decoder and output the corresponding scattering coefficient. To guide the pre-training process, polarimetric decomposition loss and power self-supervision loss are constructed. The former aligns the predicted coefficients with Yamaguchi coefficients, while the latter reconstructs power from the predicted coefficients and compares it to the input image's power. The performance of our foundation model is validated on six typical downstream tasks, achieving state-of-the-art results. Notably, the foundation model can extract stable feature representations and exhibits strong generalization, even in data-scarce conditions.

Figures

Figures reproduced from arXiv: 2504.11999 by the authors.

Figure 1
Figure 1. The paradigm of SAR foundation model + downstream tasks. The foundation model explores unlabeled SAR data to extract the semantics of interested areas and classes, and the pre-trained foundation model can be broadly applied across various downstream tasks and applications. modeling [4] and contrastive learning [5]. As illustrated in [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Amplitude and phase information in complex-valued SAR imagery. The radar sensor transmits electromagnetic (EM) waves and re￾ceives complex-valued radar signals. The amplitude image highlights strong scattering points and clear contours, while the phase image contains more detailed information about the structure of objects, such as aircraft. is essential for improving model performance. As illustrated in Fig.2, most… view at source ↗
Figure 4
Figure 4. Schematic diagram of the Yamaguchi decomposition components. The top line exhibits fully PolSAR images, arranged from left to right with polarization modes HH, HV, VH, and VV respectively. The subsequent line depicts the results of four components from T, illustrating Dbl scattering, Hlx scattering, S.F. scattering, and Vol scattering components respectively. two-dimensional polarization scattering matrix S. The rec… view at source ↗
Figures from the paper (8 more)
Figure 5
Figure 5. Figure 5: Overview of our foundation model based on complex-valued SAR images. Deep network simulates the physical process of polarimetric decomposition, constrained by the polarimetric decomposition loss and power self-supervised loss during the pre-training phase. The embeddin…
Figure 6
Figure 6. Figure 6: The process of scattering query initialization. Nine physical characteristics are represented by corresponding scattering mathematical matrices. Samples are generated based on X = T Y and converted into feature vectors by BERT. Scattering queries are then obtained thro…
Figure 7
Figure 7. Figure 7: Feature distribution for the scattering queries. The features are dimensionality reduced through t-SNE and are independent of each other. supplement the remaining power. The results indicate that the scattering query vectors can effectively represent the semantics of e…
Figure 8
Figure 8. Figure 8: Statistical distribution of four Yamaguchi scattering components. Each subgraph provides cumulative density function located at 1/2, 1/4, and 3/4 and scale parameter µ for fitting the Rayleigh distribution. distribution is not symmetric. In this case, a probability val…
Figure 9
Figure 9. Figure 9: The loss curve during the pre-training phase. The total loss, polari￾metric decomposition loss and power self-supervised loss steadily decrease to convergence [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Quantitative results of polarimetric decomposition coefficients during the pre-training phase. OA, mIoU, mAcc gradually increase until convergence. and a Poly learning strategy with a parameter of 0.9 is adopted. During the pre-training phase, the convergence process …
Figure 13
Figure 13. Figure 13: Unsupervised classification visualization results for the SAR￾ClsL1 dataset. Taking San Francisco scenario as an example, our foundation model has significant advantages without fine-tuning. Downstream tasks including object detection and semantic segmentation are use…
Figure 14
Figure 14. Figure 14: Visualization of ship detection on the HRSID dataset. The results of different pre-trained foundation models with DDQ DETR are shown. The last column demonstrates that the detection results align more closely with the ground truth when using our foundation model. imen…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SARATR-X-v2: Scale-Aware Structural Pre-Training for SAR Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A learnable-weighted fusion of six fixed, speckle-robust structural operators as the masked pre-training target transfers better than pixel targets on 10 of 12 SAR benchmarks.

Reference graph

Works this paper leans on

69 extracted references · 57 canonical work pages · cited by 1 Pith paper

  1. [1]

    An empirical study of remote sensing pretraining,

    D. Wang, J. Zhang, B. Du, G.-S. Xia, and D. Tao, “An empirical study of remote sensing pretraining,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–20, 2023

  2. [2]

    Advancing plain vision transformer toward remote sensing foundation model,

    D. Wang, Q. Zhang, Y . Xu, J. Zhang, B. Du, D. Tao, and L. Zhang, “Advancing plain vision transformer toward remote sensing foundation model,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–15, 2023

  3. [3]

    Mtp: Advancing remote sensing foundation model via multi-task pretraining,

    D. Wang, J. Zhang, M. Xu, L. Liu, D. Wang, E. Gao, C. Han, H. Guo, B. Du, D. Tao, and L. Zhang, “Mtp: Advancing remote sensing foundation model via multi-task pretraining,” IEEE Jour . Select. Topi. Appli. Earth Obser . Remote Sens. , pp. 1–24, 2024

  4. [4]

    Simmim: A simple framework for masked image modeling,

    Z. Xie, Z. Zhang, Y . Cao, Y . Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu, “Simmim: A simple framework for masked image modeling,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 9653–9663

  5. [5]

    Dense contrastive learning for self-supervised visual pre-training,

    X. Wang, R. Zhang, C. Shen, T. Kong, and L. Li, “Dense contrastive learning for self-supervised visual pre-training,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. , 2021, pp. 3024–3033

  6. [6]

    Ringmo: A remote sensing foundation model with masked image modeling,

    X. Sun, P. Wang, W. Lu, Z. Zhu, X. Lu, Q. He, J. Li, X. Rong, Z. Yang, H. Chang et al. , “Ringmo: A remote sensing foundation model with masked image modeling,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–22, 2022

  7. [7]

    Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning,

    C. J. Reed, R. Gupta, S. Li, S. Brockman, C. Funk, B. Clipp, K. Keutzer, S. Candido, M. Uyttendaele, and T. Darrell, “Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning,” in Proc. IEEE Int. Conf. Comput. Vis. , 2023, pp. 4088–4099

  8. [8]

    Spectralgpt: Spectral foundation model,

    D. Hong, B. Zhang, X. Li, Y . Li, C. Li, J. Yao, N. Yokoya, H. Li, X. Jia, A. Plaza et al., “Spectralgpt: Spectral foundation model,” arXiv preprint arXiv:2311.07113, 2023

Show all 69 references
  1. [10]

    Scattering prompt tuning: A fine-tuned foundation model for sar object recognition,

    W. Guo, S. Li, and J. Yang, “Scattering prompt tuning: A fine-tuned foundation model for sar object recognition,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 3056–3065

  2. [11]

    Croma: Remote sensing represen- tations with contrastive radar-optical masked autoencoders,

    A. Fuller, K. Millard, and J. Green, “Croma: Remote sensing represen- tations with contrastive radar-optical masked autoencoders,” Adv. Neural Inform. Process. Syst. , vol. 36, 2024

  3. [12]

    Deep learning meets sar: Concepts, models, pitfalls, and perspectives,

    X. X. Zhu, S. Montazeri, M. Ali, Y . Hua, Y . Wang, L. Mou, Y . Shi, F. Xu, and R. Bamler, “Deep learning meets sar: Concepts, models, pitfalls, and perspectives,” IEEE Geosci. Remote Sens. Magazine , vol. 9, no. 4, pp. 143–172, 2021

  4. [13]

    Mcanet: A joint semantic segmentation framework of optical and sar images for land use classification,

    X. Li, G. Zhang, H. Cui, S. Hou, S. Wang, X. Li, Y . Chen, Z. Li, and L. Zhang, “Mcanet: A joint semantic segmentation framework of optical and sar images for land use classification,” Inter . Jour . Appli. Earth Obser . Geoinf., vol. 106, p. 102638, 2022

  5. [14]

    Denet: Double- encoder network with feature refinement and region adaption for terrain segmentation in polsar images,

    X. Zeng, Z. Wang, X. Sun, Z. Chang, and X. Gao, “Denet: Double- encoder network with feature refinement and region adaption for terrain segmentation in polsar images,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–19, 2021

  6. [15]

    Sar automatic target recognition method based on multi-stream complex-valued networks,

    Z. Zeng, J. Sun, Z. Han, and W. Hong, “Sar automatic target recognition method based on multi-stream complex-valued networks,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–18, 2022

  7. [16]

    Interpretable deep learning: Interpretation, interpretability, trustworthi- ness, and beyond,

    X. Li, H. Xiong, X. Li, X. Wu, X. Zhang, J. Liu, J. Bian, and D. Dou, “Interpretable deep learning: Interpretation, interpretability, trustworthi- ness, and beyond,” Knowledge and Information Systems , vol. 64, no. 12, pp. 3197–3234, 2022

  8. [17]

    Four- component scattering model for polarimetric sar image decomposition,

    Y . Yamaguchi, T. Moriyama, M. Ishido, and H. Yamada, “Four- component scattering model for polarimetric sar image decomposition,” IEEE Trans. Geosci. Remote Sens. , vol. 43, no. 8, pp. 1699–1706, 2005

  9. [18]

    A review of target decomposition theorems in radar polarimetry,

    S. R. Cloude and E. Pottier, “A review of target decomposition theorems in radar polarimetry,” IEEE Trans. Geosci. Remote Sens. , vol. 34, no. 2, pp. 498–518, 1996

  10. [19]

    Llama: Open and efficient foundation language models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar et al. , “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023

  11. [20]

    Masked autoencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 16 000–16 009

  12. [21]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Loet al., “Segment anything,” in Proc. IEEE Int. Conf. Comput. Vis. , 2023, pp. 4015–4026

  13. [22]

    Satmae: Pre-training transformers for temporal and multi-spectral satellite imagery,

    Y . Cong, S. Khanna, C. Meng, P. Liu, E. Rozi, Y . He, M. Burke, D. Lo- bell, and S. Ermon, “Satmae: Pre-training transformers for temporal and multi-spectral satellite imagery,” Adv. Neural Inform. Process. Syst. , vol. 35, pp. 197–211, 2022. JOURNAL OF LATEX CLASS FILES, VOL...

  14. [23]

    Feature guided masked autoencoder for self-supervised learning in remote sens- ing,

    Y . Wang, H. H. Hern ´andez, C. M. Albrecht, and X. X. Zhu, “Feature guided masked autoencoder for self-supervised learning in remote sens- ing,” arXiv preprint arXiv:2310.18653 , 2023

  15. [24]

    Self-supervised vision transformers for joint sar-optical representation learning,

    Y . Wang, C. M. Albrecht, and X. X. Zhu, “Self-supervised vision transformers for joint sar-optical representation learning,” in IEEE Int. Geosci. Remote Sens. Sympo. IEEE, 2022, pp. 139–142

  16. [25]

    Skysense: A multi-modal remote sensing foundation model towards universal interpretation for earth observation imagery,

    X. Guo, J. Lao, B. Dang, Y . Zhang, L. Yu, L. Ru, L. Zhong, Z. Huang, K. Wu, D. Hu et al. , “Skysense: A multi-modal remote sensing foundation model towards universal interpretation for earth observation imagery,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. , 2024,...

  17. [26]

    Polarimetric convo- lutional network for polsar image classification,

    X. Liu, L. Jiao, X. Tang, Q. Sun, and D. Zhang, “Polarimetric convo- lutional network for polsar image classification,” IEEE Trans. Geosci. Remote Sens. , vol. 57, no. 5, pp. 3040–3054, 2018

  18. [27]

    Radar polaritnetry for geoscience applica- tions,

    F. T. Ulaby and C. Elachi, “Radar polaritnetry for geoscience applica- tions,” 1990

  19. [28]

    New decomposition of the radar target scattering matrix,

    E. Krogager, “New decomposition of the radar target scattering matrix,” Electronics Letters, vol. 18, no. 26, pp. 1525–1527, 1990

  20. [29]

    Simulated polari- metric signatures of primitive geometrical shapes,

    W. L. Cameron, N. N. Youssef, and L. K. Leung, “Simulated polari- metric signatures of primitive geometrical shapes,” IEEE Trans. Geosci. Remote Sens. , vol. 34, no. 3, pp. 793–803, 1996

  21. [30]

    A review of polarimetry in the context of synthetic aperture radar: Concepts and information extraction,

    R. Touzi, W. Boerner, J. Lee, and E. Lueneburg, “A review of polarimetry in the context of synthetic aperture radar: Concepts and information extraction,” Canadian Journal of Remote Sensing , vol. 30, no. 3, pp. 380–407, 2004

  22. [31]

    Eigen-decomposition-based four-component decomposition for polsar data,

    B. Zou, D. Lu, L. Zhang, and W. M. Moon, “Eigen-decomposition-based four-component decomposition for polsar data,” IEEE Jour . Select. Topi. Appli. Earth Obser . Remote Sens. , vol. 9, no. 3, pp. 1286–1296, 2016

  23. [32]

    Advanced polarimetric target decomposition,

    S. Chen, X. Wang, S. Xiao, and M. Sato, “Advanced polarimetric target decomposition,” Target Scatt. Mechan. Polari. Synth. Aper . Radar: Interpr . Appli., pp. 43–106, 2018

  24. [33]

    Phenomenological theory of radar targets,

    J. R. Huynen, “Phenomenological theory of radar targets,” 1970

  25. [34]

    Three-component scattering model to describe polarimetric sar data,

    A. Freeman and S. L. Durden, “Three-component scattering model to describe polarimetric sar data,” in Radar Polarimetry, vol. 1748. SPIE, 1993, pp. 213–224

  26. [35]

    Seven-component scattering power decomposition of polsar coherency matrix,

    G. Singh, R. Malik, S. Mohanty, V . S. Rathore, K. Yamada, M. Umemura, and Y . Yamaguchi, “Seven-component scattering power decomposition of polsar coherency matrix,” IEEE Trans. Geosci. Remote Sens., vol. 57, no. 11, pp. 8371–8382, 2019

  27. [36]

    Exploring fine polarimetric decomposition technique for built-up area monitoring,

    S. Quan, T. Zhang, W. Wang, G. Kuang, X. Wang, and B. Zeng, “Exploring fine polarimetric decomposition technique for built-up area monitoring,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–19, 2023

  28. [37]

    Polarimetric decomposition-based unified manmade target scattering characterization with mathematical programming strategies,

    S. Quan, Y . Qin, D. Xiang, W. Wang, and X. Wang, “Polarimetric decomposition-based unified manmade target scattering characterization with mathematical programming strategies,” IEEE Trans. Geosci. Re- mote Sens. , vol. 60, pp. 1–18, 2022

  29. [38]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of naacL-HLT , vol. 1. Minneapolis, Minnesota, 2019, p. 2

  30. [39]

    Pyramid scene parsing network,

    H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. , 2017, pp. 2881–2890

  31. [40]

    Unified perceptual parsing for scene understanding,

    T. Xiao, Y . Liu, B. Zhou, Y . Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in Proc. Eur . Conf. Comput. Vis., 2018, pp. 418–434

  32. [41]

    Encoder- decoder with atrous separable convolution for semantic image segmen- tation,

    L.-C. Chen, Y . Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder- decoder with atrous separable convolution for semantic image segmen- tation,” arXiv preprint arXiv:1802.02611 , 2018

  33. [42]

    Asymmetric non-local neural networks for semantic segmentation,

    Z. Zhu, M. Xu, S. Bai, T. Huang, and X. Bai, “Asymmetric non-local neural networks for semantic segmentation,” in Proc. IEEE Int. Conf. Comput. Vis., 2019, pp. 593–602

  34. [43]

    Ccnet: Criss-cross attention for semantic segmentation,

    Z. Huang, X. Wang, L. Huang, C. Huang, Y . Wei, and W. Liu, “Ccnet: Criss-cross attention for semantic segmentation,” in Proc. IEEE Int. Conf. Comput. Vis. , 2019, pp. 603–612

  35. [44]

    K-net: Towards unified image segmentation,

    W. Zhang, J. Pang, K. Chen, and C. C. Loy, “K-net: Towards unified image segmentation,” Adv. Neural Inform. Process. Syst. , vol. 34, pp. 10 326–10 338, 2021

  36. [45]

    Segnext: Rethinking convolutional attention design for semantic segmentation. arxiv 2022,

    M. Guo, C. Lu, Q. Hou, Z. Liu, M. Cheng, and S. Hu, “Segnext: Rethinking convolutional attention design for semantic segmentation. arxiv 2022,” arXiv preprint arXiv:2209.08575 , 2022

  37. [46]

    Masked-attention mask transformer for universal image segmentation,

    B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit., 2022, pp. 1290– 1299

  38. [47]

    Deep dual-resolution networks for real-time and accurate semantic segmentation of traffic scenes,

    H. Pan, Y . Hong, W. Sun, and Y . Jia, “Deep dual-resolution networks for real-time and accurate semantic segmentation of traffic scenes,” IEEE Trans. Intell. Transpor . Syst., vol. 24, no. 3, pp. 3448–3460, 2023

  39. [48]

    Agmtr: Agent mining transformer for few-shot segmentation in remote sensing,

    H. Bi, Y . Feng, Y . Mao, J. Pei, W. Diao, H. Wang, and X. Sun, “Agmtr: Agent mining transformer for few-shot segmentation in remote sensing,” Int. J. Comput. Vis. , pp. 1–28, 2024

  40. [49]

    Not just learning from others but relying on yourself: A new perspective on few-shot segmentation in remote sensing,

    H. Bi, Y . Feng, Z. Yan, Y . Mao, W. Diao, H. Wang, and X. Sun, “Not just learning from others but relying on yourself: A new perspective on few-shot segmentation in remote sensing,” IEEE Trans. Geosci. Remote Sens., 2023

  41. [50]

    Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,

    S. Wei, X. Zeng, Q. Qu, M. Wang, H. Su, and J. Shi, “Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,” IEEE Access , vol. 8, pp. 120 234–120 254, 2020

  42. [52]

    Air-polsar-seg: A large- scale data set for terrain segmentation in complex-scene polsar images,

    Z. Wang, X. Zeng, Z. Yan, J. Kang, and X. Sun, “Air-polsar-seg: A large- scale data set for terrain segmentation in complex-scene polsar images,” IEEE Jour . Select. Topi. Appli. Earth Obser . Remote Sens. , vol. 15, pp. 3830–3841, 2022

  43. [53]

    Mask r-cnn,

    K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proc. IEEE Int. Conf. Comput. Vis. , 2017, pp. 2961–2969

  44. [54]

    Cascade r-cnn: Delving into high quality object detection,

    Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high quality object detection,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recog- nit., 2018, pp. 6154–6162

  45. [55]

    Tood: Task- aligned one-stage object detection,

    C. Feng, Y . Zhong, Y . Gao, M. R. Scott, and W. Huang, “Tood: Task- aligned one-stage object detection,” in Proc. IEEE Int. Conf. Comput. Vis. IEEE Computer Society, 2021, pp. 3490–3499

  46. [56]

    Swin-paff: A sar ship detection network with contextual cross-information fusion

    Y . Zhang, D. Han et al. , “Swin-paff: A sar ship detection network with contextual cross-information fusion.” Computers, Materials & Continua , vol. 77, no. 2, 2023

  47. [57]

    A novel anchor- free detector using global context-guide feature balance pyramid and united attention for sar ship detection,

    L. Bai, C. Yao, Z. Ye, D. Xue, X. Lin, and M. Hui, “A novel anchor- free detector using global context-guide feature balance pyramid and united attention for sar ship detection,” IEEE Geosci. Remote Sens. Lett. , vol. 20, pp. 1–5, 2023

  48. [58]

    Dense distinct query for end-to-end object detection,

    S. Zhang, X. Wang, J. Wang, J. Pang, C. Lyu, W. Zhang, P. Luo, and K. Chen, “Dense distinct query for end-to-end object detection,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. , 2023, pp. 7329–7338

  49. [59]

    Yolox: Exceeding yolo series in 2021,

    Z. Ge, S. Liu, F. Wang, Z. Li, and J. Sun, “Yolox: Exceeding yolo series in 2021,” arXiv preprint arXiv:2107.08430 , 2021

  50. [60]

    Sar-aircraft-1.0: High-resolution sar aircraft detection and recognition dataset,

    W. Zhirui, K. Yuzhuo, Z. Xuan, W. Yuelei, Z. Ting, and S. Xian, “Sar-aircraft-1.0: High-resolution sar aircraft detection and recognition dataset,” Journal of Radars , vol. 12, no. 4, pp. 906–922, 2023

  51. [61]

    Scattering-keypoint-guided network for oriented ship detection in high-resolution and large-scale sar images,

    K. Fu, J. Fu, Z. Wang, and X. Sun, “Scattering-keypoint-guided network for oriented ship detection in high-resolution and large-scale sar images,” IEEE Jour . Select. Topi. Appli. Earth Obser . Remote Sens. , vol. 14, pp. 11 162–11 178, 2021

  52. [62]

    Yolov5 by ultralytics,

    “Yolov5 by ultralytics,” https://github.com/ultralytics/yolov5

  53. [63]

    Mlsdnet: Multi-class lightweight sar detection network based on adaptive scale distribution attention,

    H. Chang, X. Fu, J. Dong, J. Liu, and Z. Zhou, “Mlsdnet: Multi-class lightweight sar detection network based on adaptive scale distribution attention,” IEEE Geosci. Remote Sens. Lett. , 2023

  54. [64]

    Diffusiondet: Diffusion model for object detection,

    S. Chen, P. Sun, Y . Song, and P. Luo, “Diffusiondet: Diffusion model for object detection,” in Proc. IEEE Int. Conf. Comput. Vis. , 2023, pp. 19 830–19 843

  55. [65]

    Diffdet4sar: Diffusion-based aircraft target detection network for sar images,

    J. Zhou, C. Xiao, B. Peng, Z. Liu, L. Liu, Y . Liu, and X. Li, “Diffdet4sar: Diffusion-based aircraft target detection network for sar images,” IEEE Geosci. Remote Sens. Lett. , 2024

  56. [66]

    Saratr-x: A foundation model for synthetic aperture radar images target recognition,

    W. Yang, Y . Hou, L. Liu, Y . Liu, X. Li et al. , “Saratr-x: A foundation model for synthetic aperture radar images target recognition,” arXiv preprint arXiv:2405.09365, 2024

  57. [67]

    Unleashing channel potential: Space-frequency selection convolution for sar object detection,

    K. Li, D. Wang, Z. Hu, W. Zhu, S. Li, and Q. Wang, “Unleashing channel potential: Space-frequency selection convolution for sar object detection,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 17 323–17 332

  58. [68]

    Non-local neural net- works,

    X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural net- works,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. , 2018, pp. 7794–7803

  59. [69]

    Expectation- maximization attention networks for semantic segmentation,

    X. Li, Z. Zhong, J. Wu, Y . Yang, Z. Lin, and H. Liu, “Expectation- maximization attention networks for semantic segmentation,” in Proc. IEEE Int. Conf. Comput. Vis. , 2019, pp. 9167–9176

  60. [70]

    Object-level semantic segmentation on the high-resolution gaofen-3 fusar-map dataset,

    X. Shi, S. Fu, J. Chen, F. Wang, and F. Xu, “Object-level semantic segmentation on the high-resolution gaofen-3 fusar-map dataset,” IEEE Jour . Select. Topi. Appli. Earth Obser . Remote Sens., vol. 14, pp. 3107– 3119, 2021

  61. [71]

    Safe: a sar feature extractor based on self-supervised learning and masked siamese vits,

    M. Muzeau, J. Frontera-Pons, C. Ren, and J.-P. Ovarlez, “Safe: a sar feature extractor based on self-supervised learning and masked siamese vits,” arXiv preprint arXiv:2407.00851 , 2024

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.