Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Beyond Motion Cues and Structural Sparsity: Revisiting Small Moving Target Detection

T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read TenRPCANet treats dim moving targets as outliers against a low-rank background model.

desk verdict Solid empirical paper on two hard detection tasks; the claimed low-rank-theoretic mechanism is an unproven analogy, and the SOTA phrasing overreaches on the synthetic set. read the letter →

arxiv 2509.07654 v1 pith:LC7NMIN2 submitted 2025-09-09 cs.CV

classification cs.CV
keywords infraredsmalltargetdetectionspaceobjectlow-rankandsparsedecompositiontensorRPCAself-attentionvisiontransformermulti-framevideobackgroundmodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

TenRPCANet claims that the reliable signal in small moving target detection is not the target itself but the background: even cluttered infrared scenes and star fields have strong low-rank structure, so treating detection as tensor low-rank and sparse decomposition should transfer across tasks. The paper proposes a deep network that encodes this prior without iterative optimization, using a tokenization strategy that makes self-attention behave like diffusion on a smooth embedding manifold, and a Progressive Feature Refinement module inspired by the sparse-update step of tensor RPCA. On the NUDT-MIRSDT, NUDT-MIRSDT-HiNo, and Real Optical Set benchmarks, the method is reported to reach the best detection probability, AUC, and IoU among the compared methods. If correct, the paper establishes that background discrimination, rather than target-specific features or motion cues, is the transferable principle for dim and anomalous object detection.

What carries the argument

The load-bearing identity is the manifold-diffusion description of self-attention: if the embedding manifold is smooth, local inner products approximate geodesic distances, $q_i^\top k_j \approx -\frac{1}{2} d_M^2(z_i,z_j)$, attention weights become a heat kernel $\alpha_{ij}\approx \exp(-d_M^2(z_i,z_j)/(2\sigma^2))$, and self-attention acts as diffusion $Af \approx e^{-\sigma^2\Delta_M}f$. The Locally Subspace Embedding (LSE) module implements the tokenization: a 3D patch branch approximates the fourth-order spatiotemporal patch tensor, while a 2D multi-scale branch enforces the third-order spatial patch tensor, with no activations so the low-rank inductive bias is preserved. The encoder-decoder uses Video Swin Transformer blocks to estimate the low-rank background, and the Progressive Feature Refinement (PFR) module performs the sparse update of Tensor RPCA, $S^*=\operatorname{Soft}_{\lambda/\mu}(M(X-L_t)+\Lambda_{t-1}/\mu)$, realized by linear residual blocks plus sigmoid and hard thresholding. Together the modules are designed to solve $\min_{L,S,N} \operatorname{rank}(L)+\lambda J_S(S)+\eta J_N(N)$ subject to $X=L+S+N$ implicitly, without iterative optimization.

What would settle it

Directly test the manifold premise by measuring, on the trained model, the correlation between attention logits $q_i^\top k_j$ and the squared embedding distance $-\frac{1}{2}d_M^2(z_i,z_j)$ over background patches in a validation sequence; a near-zero correlation would falsify the claimed mechanism, since the diffusion identity in Eqs. (13)-(15) is what connects LSE to low-rank background modeling.

Watch

Extended reading notes

Core claim

The paper's central discovery is that small moving target detection and background modeling are two sides of the same coin: the background of a video, whether infrared clutter or a dense star field, is highly redundant and hence approximately low-rank in a suitable patch-tensor representation, while a moving target is a compact, locally smooth spatiotemporal outlier. TenRPCANet therefore implements an implicit tensor RPCA: the Locally Subspace Embedding module tokenizes the input through 3D and multi-scale 2D patch branches without nonlinear activations, shaping the self-attention operator into a diffusion process on a smooth embedding manifold; the Video Swin encoder-decoder estimates the low-rank background; and the Progressive Feature Refinement module updates the sparse target component by a soft-threshold-like operation that combines the residual between input and estimated background with confidence gating. The paper reports the best detection probability, false-alarm rate, AUC, and IoU among compared methods on multi-frame infrared small target detection and space object detection, and argues that because the background prior requires no assumptions about target appearance, motion, or structure, the method generalizes across two very different observation modalities.

Load-bearing premise

The entire low-rank mechanism rests on the premise that the learned embedding space is smooth enough that attention weights truly behave like heat diffusion; the paper offers no direct measurement of that smoothness, so if the learned embeddings are not that well-behaved, the claimed implicit low-rank enforcement is unsupported even though the architecture might still work for other reasons.

Editorial extensions

If this is right

  • Detection performance should transfer across imaging modalities because the low-rank background prior is not tied to infrared or optical target appearance; the paper demonstrates this on infrared sequences and spaceborne optical sequences.
  • Motion-cue-based and target-sparsity-based detectors can be outperformed by a background-discrimination approach in low-SNR and non-Gaussian-noise conditions, as reported on NUDT-MIRSDT-HiNo and the Real Optical Set.
  • No initialization or frame-discarding stage is needed: all frames in an 8-frame temporal window are processed at once, so even targets visible in only two frames can be detected.
  • Keeping the embedding and refinement modules linear rather than inserting activation functions is part of the design; the ablation study shows that adding ReLU reduces detection probability and increases false alarms.
  • Longer temporal windows provide richer context for estimating the background and improve detection performance, motivating the use of 8 input frames in the experiments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: The reported gains may owe as much to the transformer's larger effective receptive field and multi-scale tokenization as to the low-rank/diffusion mechanism; an architecture-matched comparison with a non-geometric stem would separate the two.
  • Editorial inference: If the background low-rank prior is truly modality-agnostic, TenRPCANet should also apply to other video-anomaly tasks where the background is highly redundant, such as tiny drone detection, medical video artifact detection, or surveillance of stationary scenes.
  • Editorial inference: The manifold-diffusion premise in Eqs. (13)-(15) is testable directly: computing the correlation between $q_i^\top k_j$ and $-\frac{1}{2}d_M^2(z_i,z_j)$ over background tokens would convert the analogy into a measured property of the trained model.
  • Editorial inference: The failure case where a target overlaps a star suggests the model relies on the topological consistency of the background; a natural extension is a temporal gating mechanism that down-weights background consistency when a target is being occulted.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes TenRPCANet, a deep network for small moving target detection that frames the task as tensor-based low-rank and sparse decomposition. It introduces a Locally Subspace Embedding (LSE) tokenization stem intended to implicitly enforce multi-order tensor low-rank priors through self-attention, and a Progressive Feature Refinement (PFR) module inspired by the sparse update step of tensor RPCA. The network uses Video Swin Transformer blocks in an encoder-decoder with skip connections and is evaluated on multi-frame infrared small target detection (NUDT-MIRSDT, NUDT-MIRSDT-HiNo) and space object detection (Synthesis Set, Real Optical Set). The authors report state-of-the-art results on several metrics and provide ablations, hyperparameter studies, and failure-case analysis.

Significance. If the empirical results hold, the paper is a practically strong contribution: the model is compact (1.78M parameters), fast (176 FPS), and shows promising cross-task generalization across two quite different detection tasks. The experimental comparison is extensive, and the inclusion of ablations, hyperparameter studies, and honest failure cases is a strength. The main weakness is that the central theoretical mechanism — that LSE makes self-attention act as diffusion and thereby implicitly enforces tensor low-rank priors — is asserted rather than derived or empirically validated, and some state-of-the-art claims need qualification. The paper would be suitable for publication after the theoretical claims are either supported by evidence or appropriately de-emphasized.

major comments (4)
  1. [III-E, Eqs. (13)-(15)] The load-bearing theoretical claim that LSE implicitly enforces low-rank priors is not established. Eq. (13) identifies the learned query-key bilinear form q_i^T k_j with -1/2 d_M^2(z_i,z_j), but q_i^T k_j = z_i^T W_q^T W_k z_j contains no constraint such as W_q^T W_k = -I and no centering/norm condition that would make a bilinear form equal a squared distance. Eq. (14) ignores the softmax partition function in Eq. (12), so attention weights cannot equal the heat kernel exp(-d_M^2/(2 sigma^2)) even if Eq. (13) held. Eq. (15) asserts a diffusion-semigroup interpretation, but no argument connects the spectrum of the attention operator to the tensor rank of the background in Eq. (1). Consequently, the statement that LSE 'implicitly regularizes self-attention' and enforces multi-order tensor low-rank priors is currently an analogy, not a derivation. I would ask the authors either to provide empirical checks (e.g., spectral analysis of learned features, verification that attention maps approximate heat kernels) or to reframe the paper as presenting a heuristic architecture validated solely by experiments.
  2. [Table II / Abstract] The abstract's claim of state-of-the-art performance on both tasks is too strong as written. On the Synthesis Set in Table II, TenRPCANet is not the best on Recall or F1: it reports Rt=87.81% and F1=89.73%, while DeepPro reports Rt=93.94% and F1=93.31% and LMAFormer reports Rt=92.85% and F1=93.53%. Only IoU (74.79%) is best. Please qualify the claim to the Real Optical Set, or to specific metrics on the Synthesis Set, and state explicitly which metric is being optimized.
  3. [Tables I and II, Section IV-B] No error bars, multiple runs, or statistical significance tests are reported for any table. The margins over the strongest baselines on key metrics are small (e.g., Table I NUDT-MIRSDT AUC 0.9978 vs DeepPro 0.9973; Table II Real Optical IoU 80.78% vs DnT-Net 79.04%), so the state-of-the-art claims cannot currently be distinguished from training variance. Please report mean and standard deviation over at least three seeds, or provide significance testing, for the main comparison tables.
  4. [III-E, Eqs. (16)-(17)] There is a gap between the announced mechanism and the actual module. The LSE implementation is a Conv3D branch with stride (1,4,4) plus a multi-scale Conv2D concatenation followed by a 1x1 convolution (Eqs. (16)-(17)); nothing in these operations explicitly minimizes tensor rank or computes a subspace embedding, and the module named 'Locally Linear Embedding' is not the classical LLE procedure referenced in [86] and [108]. If the low-rank interpretation is retained, the paper should explain precisely how these linear convolutions implement subspace embedding and provide evidence (e.g., singular-value spectra of features with and without LSE) that low-rankness is actually induced.
minor comments (6)
  1. [Table III] The table header says 'LSE and RFR' but the module is called PFR; also footnote (†) says the metrics are recorded on the multi-frame infrared task while the same table includes Real Optical Set columns. Please correct these inconsistencies.
  2. [Eq. (26)] The loss function uses i as the summation index for both height and width; the second sum should be over j (or a separate index) to avoid ambiguity.
  3. [III-F, Eq. (20)] The minimization is written as 'arg min_T' but the variable being optimized is S; this appears to be a typo.
  4. [Tables IV-VI] The dataset name is written 'NUDT-MRISTD' in several places; it should be 'NUDT-MIRSDT'.
  5. [References] References [3] and [8] appear to be duplicates of the same paper; one should be removed.
  6. [Reproducibility] No code or data availability statement is provided; please add one if possible, as the reported compact model and strong results would benefit from reproducibility checks.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the architecture is motivated by tensor RPCA but not derived from it, and the reported results are external benchmark comparisons.

full rationale

Walking the paper's derivation chain, the central claim is that TenRPCANet implicitly performs low-rank and sparse decomposition. This claim is supported by motivation rather than by a derivation that reduces to its own inputs. The tensor RPCA objective in Eq. (1) is posed as background; Section III-E introduces LSE with the manifold/heat-kernel analogy in Eqs. (12)-(15), explicitly labeled as an assumption ('Assuming M is a smooth manifold ...'). That analogy is unverified and is a correctness risk, but it is not a fitted parameter and no prediction is computed from it. The PFR module derives the soft-threshold update in Eqs. (18)-(21), but then implements a heuristic approximation in Eq. (25) (sigmoid plus confidence threshold) and explicitly says it 'mimics' soft-thresholding rather than reproducing it. Thus no equation-level reduction is present. The empirical claims rely on comparisons against external methods on NUDT-MIRSDT, NUDT-MIRSDT-HiNo, and the Real Optical Set; the last benchmark originates in the authors' prior work [15], but this is a normal benchmark comparison and does not make the derivation circular. Because no step can be exhibited where a prediction reduces by construction to a fitted input or to a self-citation chain, the correct circularity finding is no significant circularity.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The network's success is claimed to rest on several structural premises: low-rank video backgrounds, stable false-alarm contexts, PSF-induced target compactness, unstructured noise, and a smooth embedding manifold for self-attention. The free parameters are standard deep learning hyperparameters, but the threshold tau and temporal window T directly control output sparsity and are not given fixed justified values. No new physical entities are introduced.

free parameters (3)
  • Temporal window size T = 8 frames
    Set to 8 frames in Section IV-A.3; Table V shows that larger windows such as 8x7x7 improve NUDT Pd and AUC, so this choice is tuned to practical constraints and validation results.
  • Confidence threshold tau in PFR = not reported
    In Eq. (25), tau gates the output to enforce sparsity; its value is not stated and must be tuned, directly affecting false alarm rate and detection confidence.
  • Channel width of network = 24
    Table VI shows channel size 24 gives the best Pd and AUC balance on NUDT among 12, 24, and 48; this is a model-capacity hyperparameter selected on validation data.
assumptions (5)
  • domain assumption Video backgrounds of small moving target detection are low-rank in tensor patch representations, rank(X) much less than min(hwt, P).
    Section III-A, Eq. (5): the central premise for modeling background as low-rank; asserted as Phenomenon 1 and not empirically quantified on the test datasets.
  • domain assumption False alarm sources have stable contextual patterns over time while true targets have unstable contexts due to motion.
    Section III-A, Phenomenon 2 and Fig. 3: this motivates the 2D patch branch and the treatment of targets as contextually inconsistent outliers; it can fail when targets overlap stars or when star intensity fluctuates.
  • ad hoc to paper The token embedding manifold M is smooth enough that self-attention weights approximate a heat kernel and the attention operator acts as Laplace-Beltrami diffusion.
    Section III-E, Eqs. (13)-(15): the central theoretical justification for LSE is an analogy from manifold learning, not a proven property of learned transformer features.
  • domain assumption Targets exhibit local intensity smoothness and spatial compactness due to the point spread function and motion blur.
    Section III-B, Eqs. (7)-(8): used to justify the PFR module's 3x3 convolution and structural sparsity; standard optics but may not hold for very fast or unresolved targets.
  • domain assumption Noise is general and unstructured, with no semantic consistency or learnable patterns.
    Section III-C: used to justify the deep prior Y=f_theta(z); later the paper notes real optical data has structured, non-Gaussian sensor noise, so this assumption is idealized.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Motion Cues and Structural Sparsity: Revisiting Small Moving Target Detection." pith.science (2026). https://pith.science/paper/LC7NMIN2

@misc{pith2026250907654,
  author       = {Pith},
  title        = {Pith review of: Beyond Motion Cues and Structural Sparsity: Revisiting Small Moving Target Detection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LC7NMIN2}},
  note         = {Machine review of arXiv:2509.07654}
}
read the original abstract

Small moving target detection is crucial for many defense applications but remains highly challenging due to low signal-to-noise ratios, ambiguous visual cues, and cluttered backgrounds. In this work, we propose a novel deep learning framework that differs fundamentally from existing approaches, which often rely on target-specific features or motion cues and tend to lack robustness in complex environments. Our key insight is that small target detection and background discrimination are inherently coupled, even cluttered video backgrounds often exhibit strong low-rank structures that can serve as stable priors for detection. We reformulate the task as a tensor-based low-rank and sparse decomposition problem and conduct a theoretical analysis of the background, target, and noise components to guide model design. Building on these insights, we introduce TenRPCANet, a deep neural network that requires minimal assumptions about target characteristics. Specifically, we propose a tokenization strategy that implicitly enforces multi-order tensor low-rank priors through a self-attention mechanism. This mechanism captures both local and non-local self-similarity to model the low-rank background without relying on explicit iterative optimization. In addition, inspired by the sparse component update in tensor RPCA, we design a feature refinement module to enhance target saliency. The proposed method achieves state-of-the-art performance on two highly distinct and challenging tasks: multi-frame infrared small target detection and space object detection. These results demonstrate both the effectiveness and the generalizability of our approach.

Figures

Figures reproduced from arXiv: 2509.07654 by the authors.

Figure 1
Figure 1. Illustration of the problem formulation. The target tensor contains [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. With the increasing frequency of space activities, the number of [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. An important phenomenon is that false alarm sources may exhibit [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The overall architecture of the proposed TenRPCANet is designed based on our theoretical framework and consists of three core components. First, the [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: ROC curve on the NUDT-MIRSDT dataset (SNR ≤ 3). to converge. Although the DTUM method focuses on the spatiotemporal motion cues of the target, accurately extracting such cues from weak targets under heavy noise is highly challenging. The DeepPro method enhances target …
Figure 6
Figure 6. Figure 6: Visual comparison on the NUDT-MIRSDT dataset ( [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Visual comparisons on five sequences from real optical dataset are provided, where the raw images and detection results are overlaid for better [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Typical failure cases. (a) The detection fails when the target overlaps [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. You Only Look Omni Gradient Backpropagation for Moving Infrared Small Target Detection

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A feature pyramid with a gradient-detached low-level shortcut and parameter-shared fusion (BP-FPN) lifts moving infrared small-target detection on IRDST from 71.6 to 82.9 mAP50 at +30K parameters.

Reference graph

Works this paper leans on

109 extracted references · 68 canonical work pages · cited by 1 Pith paper

  1. [86]

    Multilinear supervised neighbor- hood embedding of a local descriptor tensor for scene/object recognition,

    X.-H. Han, Y .-W. Chen, and X. Ruan, “Multilinear supervised neighbor- hood embedding of a local descriptor tensor for scene/object recognition,” IEEE Transactions on Image Processing, vol. 21, no. 3, pp. 1314–1326, 2012

  2. [108]

    Multi- linear discriminant analysis for face recognition,

    S. Yan, D. Xu, Q. Yang, L. Zhang, X. Tang, and H.-J. Zhang, “Multi- linear discriminant analysis for face recognition,”IEEE Transactions on Image Processing, vol. 16, no. 1, pp. 212–220, 2007

  3. [1]

    Visible- thermal tiny object detection: A benchmark dataset and baselines,

    X. Ying, C. Xiao, W. An, R. Li, X. He, B. Li, X. Cao, Z. Li, Y . Wang, M. Hu, Q. Xu, Z. Lin, M. Li, S. Zhou, L. Liu, and W. Sheng, “Visible- thermal tiny object detection: A benchmark dataset and baselines,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 7, pp. 6088–6096, 2025

  4. [2]

    Infrared patch-image model for small target detection in a single image,

    C. Gao, D. Meng, Y . Yang, Y . Wang, X. Zhou, and A. G. Hauptmann, “Infrared patch-image model for small target detection in a single image,” IEEE Transactions on Image Processing, vol. 22, no. 12, pp. 4996–5009, 2013

  5. [4]

    Uiu-net: U-net in u-net for infrared small object detection,

    X. Wu, D. Hong, and J. Chanussot, “Uiu-net: U-net in u-net for infrared small object detection,”IEEE Transactions on Image Processing, vol. 32, pp. 364–376, 2023

  6. [5]

    Dense nested attention network for infrared small target detection,

    B. Li, C. Xiao, L. Wang, Y . Wang, Z. Lin, M. Li, W. An, and Y . Guo, “Dense nested attention network for infrared small target detection,” IEEE Transactions on Image Processing, vol. 32, pp. 1745–1758, 2023

  7. [6]

    A new pattern for detection of streak-like space target from single optical images,

    B. Lin, L. Zhong, S. Zhuge, X. Yang, Y . Yang, K. Wang, and X. Zhang, “A new pattern for detection of streak-like space target from single optical images,”IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–13, 2021

  8. [7]

    Tnlrs: Target-aware non-local low-rank modeling with saliency filtering regularization for infrared small target detection,

    H. Zhu, H. Ni, S. Liu, G. Xu, and L. Deng, “Tnlrs: Target-aware non-local low-rank modeling with saliency filtering regularization for infrared small target detection,”IEEE Transactions on Image Processing, vol. 29, pp. 9546–9558, 2020

Show all 109 references
  1. [8]

    Infrared small and dim target detection with transformer under complex backgrounds,

    F. Liu, C. Gao, F. Chen, D. Meng, W. Zuo, and X. Gao, “Infrared small and dim target detection with transformer under complex backgrounds,” IEEE Transactions on Image Processing, vol. 32, pp. 5921–5932, 2023

  2. [9]

    Learning contrast-enhanced shape-biased representations for infrared small target detection,

    F. Lin, K. Bao, Y . Li, D. Zeng, and S. Ge, “Learning contrast-enhanced shape-biased representations for infrared small target detection,”IEEE Transactions on Image Processing, vol. 33, pp. 3047–3058, 2024

  3. [10]

    Mtmlnet: Multi-task mutual learning network for infrared small target detection and segmentation,

    B. Yang, F. Li, S. Zhao, W. Wang, J. Luo, H. Pu, M. Zhou, and Y . Pi, “Mtmlnet: Multi-task mutual learning network for infrared small target detection and segmentation,”IEEE Transactions on Image Processing, vol. 34, pp. 4414–4425, 2025

  4. [11]

    Osformer: One-step trans- former for infrared video small object detection,

    H. Qin, T. Xu, Y . Tang, F. Xu, and J. Li, “Osformer: One-step trans- former for infrared video small object detection,”IEEE Transactions on Image Processing, pp. 1–1, 2025

  5. [12]

    Infrared small target detection with scale and location sensitivity,

    Q. Liu, R. Liu, B. Zheng, H. Wang, and Y . Fu, “Infrared small target detection with scale and location sensitivity,” inProceedings of the IEEE/CVF Computer Vision and Pattern Recognition, 2024

  6. [13]

    Random access memories: A new paradigm for target detection in high resolution aerial remote sensing images,

    Z. Zou and Z. Shi, “Random access memories: A new paradigm for target detection in high resolution aerial remote sensing images,”IEEE Transactions on Image Processing, vol. 27, no. 3, pp. 1100–1111, 2018

  7. [14]

    Direction- coded temporal u-shape module for multiframe infrared small target detection,

    R. Li, W. An, C. Xiao, B. Li, Y . Wang, M. Li, and Y . Guo, “Direction- coded temporal u-shape module for multiframe infrared small target detection,”IEEE Transactions on Neural Networks and Learning Systems, 2023

  8. [15]

    Convolutional long-short term memory network for space debris detection and tracking,

    S. Chen, H. Wang, Z. Shen, K. Wang, and X. Zhang, “Convolutional long-short term memory network for space debris detection and tracking,” Knowledge-Based Systems, vol. 304, p. 112535, 2024

  9. [16]

    Total variation regularized tensor rpca for background subtraction from compressive measurements,

    W. Cao, Y . Wang, J. Sun, D. Meng, C. Yang, A. Cichocki, and Z. Xu, “Total variation regularized tensor rpca for background subtraction from compressive measurements,”IEEE Transactions on Image Processing, vol. 25, no. 9, pp. 4075–4090, 2016

  10. [17]

    Complex background subtraction by pursuing dynamic spatio-temporal models,

    L. Lin, Y . Xu, X. Liang, and J. Lai, “Complex background subtraction by pursuing dynamic spatio-temporal models,”IEEE Transactions on Image Processing, vol. 23, no. 7, pp. 3191–3202, 2014

  11. [18]

    Universal background subtraction based on arithmetic distribution neural network,

    C. Zhao, K. Hu, and A. Basu, “Universal background subtraction based on arithmetic distribution neural network,”IEEE Transactions on Image Processing, vol. 31, pp. 2934–2949, 2022

  12. [19]

    Foreground segmentation with tree- structured sparse rpca,

    S. E. Ebadi and E. Izquierdo, “Foreground segmentation with tree- structured sparse rpca,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 9, pp. 2273–2280, 2018

  13. [20]

    Moving object detection in video via hierarchical modeling and alternating optimization,

    L. Li, Q. Hu, and X. Li, “Moving object detection in video via hierarchical modeling and alternating optimization,”IEEE Transactions on Image Processing, vol. 28, no. 4, pp. 2021–2036, 2019

  14. [21]

    Block-sparse rpca for salient motion detection,

    Z. Gao, L.-F. Cheong, and Y .-X. Wang, “Block-sparse rpca for salient motion detection,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 36, no. 10, pp. 1975–1987, 2014

  15. [22]

    Moving object detection by detecting contiguous outliers in the low-rank representation,

    X. Zhou, C. Yang, and W. Yu, “Moving object detection by detecting contiguous outliers in the low-rank representation,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 3, pp. 597– 610, 2013. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 12

  16. [23]

    Receptive-field and direction induced attention network for infrared dim small target detection with a large-scale dataset irdst,

    H. Sun, J. Bai, F. Yang, and X. Bai, “Receptive-field and direction induced attention network for infrared dim small target detection with a large-scale dataset irdst,”IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–13, 2023

  17. [24]

    Gated mechanism attention transformer based on wavelet enhanced optical flow field estimation for foreground detection,

    Z. Ge, H. Chen, X. Liang, and L. Ma, “Gated mechanism attention transformer based on wavelet enhanced optical flow field estimation for foreground detection,”IEEE Transactions on Circuits and Systems for Video Technology, 2024

  18. [25]

    Nonconvex robust high-order tensor completion using randomized low- rank approximation,

    W. Qin, H. Wang, F. Zhang, W. Ma, J. Wang, and T. Huang, “Nonconvex robust high-order tensor completion using randomized low- rank approximation,”IEEE Transactions on Image Processing, vol. 33, pp. 2835–2850, 2024

  19. [26]

    Irregular tensor low-rank represen- tation for hyperspectral image representation,

    B. Han, Y . Jia, H. Liu, and J. Hou, “Irregular tensor low-rank represen- tation for hyperspectral image representation,”IEEE Transactions on Image Processing, vol. 34, pp. 3239–3252, 2025

  20. [27]

    Low-rank tensor subspace learning for rgb-d action recognition,

    C. Jia and Y . Fu, “Low-rank tensor subspace learning for rgb-d action recognition,”IEEE Transactions on Image Processing, vol. 25, no. 10, pp. 4641–4652, 2016

  21. [28]

    Multi-branch tensor network structure for tensor-train discriminant analysis,

    S. E. Sofuoglu and S. Aviyente, “Multi-branch tensor network structure for tensor-train discriminant analysis,”IEEE Transactions on Image Processing, vol. 30, pp. 8926–8938, 2021

  22. [29]

    Multiplex transformed tensor decomposition for multidimensional image recovery,

    L. Feng, C. Zhu, Z. Long, J. Liu, and Y . Liu, “Multiplex transformed tensor decomposition for multidimensional image recovery,”IEEE Transactions on Image Processing, vol. 32, pp. 3397–3412, 2023

  23. [30]

    Tensor factorization for low-rank tensor completion,

    P. Zhou, C. Lu, Z. Lin, and C. Zhang, “Tensor factorization for low-rank tensor completion,”IEEE Transactions on Image Processing, vol. 27, no. 3, pp. 1152–1163, 2018

  24. [31]

    Learnable scaled gradient descent for guaranteed robust tensor pca,

    L. Feng, C. Zhu, Y . Liu, S. Ravishankar, and L. Huang, “Learnable scaled gradient descent for guaranteed robust tensor pca,”arXiv preprint arXiv:2501.04565, 2025

  25. [32]

    Guaranteed tensor recovery fused low-rankness and smoothness,

    H. Wang, J. Peng, W. Qin, J. Wang, and D. Meng, “Guaranteed tensor recovery fused low-rankness and smoothness,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 9, pp. 10 990– 11 007, 2023

  26. [33]

    Low-rank tensor function representation for multi-dimensional data recovery,

    Y . Luo, X. Zhao, Z. Li, M. K. Ng, and D. Meng, “Low-rank tensor function representation for multi-dimensional data recovery,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 5, pp. 3351–3369, 2024

  27. [34]

    Revisiting nonlocal self-similarity from continuous representation,

    Y . Luo, X. Zhao, and D. Meng, “Revisiting nonlocal self-similarity from continuous representation,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 1, pp. 450–468, 2025

  28. [35]

    Tensor robust principal component analysis with a new tensor nuclear norm,

    C. Lu, J. Feng, Y . Chen, W. Liu, Z. Lin, and S. Yan, “Tensor robust principal component analysis with a new tensor nuclear norm,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 4, pp. 925–938, 2020

  29. [36]

    Attention-guided low-rank tensor completion,

    T. T. N. Mai, E. Y . Lam, and C. Lee, “Attention-guided low-rank tensor completion,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 9818–9833, 2024

  30. [37]

    Interpretable neural networks for video separation: Deep unfolding rpca with foreground masking,

    B. Joukovsky, Y . C. Eldar, and N. Deligiannis, “Interpretable neural networks for video separation: Deep unfolding rpca with foreground masking,”IEEE Transactions on Image Processing, vol. 33, pp. 108–122, 2023

  31. [38]

    Image denoising using green channel prior,

    Z. Kong, F. Deng, and X. Yang, “Image denoising using green channel prior,”IEEE Transactions on Image Processing, vol. 34, pp. 4869–4884, 2025

  32. [39]

    Non-local robust quaternion matrix completion for large-scale color image and video inpainting,

    Z. Jia, Q. Jin, M. K. Ng, and X.-L. Zhao, “Non-local robust quaternion matrix completion for large-scale color image and video inpainting,” IEEE Transactions on Image Processing, vol. 31, pp. 3868–3883, 2022

  33. [40]

    A robust space target extraction algorithm based on standardized correlation space construction,

    H. Wang, S. Chen, Z. Shen, K. Wang, M. Duan, W. Yang, B. Lin, and X. Zhang, “A robust space target extraction algorithm based on standardized correlation space construction,”IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 10 188–1...

  34. [41]

    Space debris and satellite laser ranging combined using a megahertz system,

    M. A. Steindorfer, P. Wang, F. Koidl, and G. Kirchner, “Space debris and satellite laser ranging combined using a megahertz system,”Nature Communications, vol. 16, no. 1, p. 575, 2025

  35. [42]

    Learning dynamic local context representations for infrared small target detection,

    G. Zhang, G. Xu, S. Chen, H. Wang, and X. Zhang, “Learning dynamic local context representations for infrared small target detection,”IEEE Transactions on Geoscience and Remote Sensing, pp. 1–1, 2025

  36. [43]

    Sstnet: Sliced spatio-temporal network with cross-slice convlstm for moving infrared dim-small target detection,

    S. Chen, L. Ji, J. Zhu, M. Ye, and X. Yao, “Sstnet: Sliced spatio-temporal network with cross-slice convlstm for moving infrared dim-small target detection,”IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–12, 2024

  37. [44]

    Seqcsist: Sequential closely-spaced infrared small target unmixing,

    X. Zhai, B. Xu, Y . Chen, H. Wang, K. Guo, and Y . Dai, “Seqcsist: Sequential closely-spaced infrared small target unmixing,”IEEE Trans- actions on Geoscience and Remote Sensing, vol. 63, pp. 1–15, 2025

  38. [45]

    Infrared small target detection via nonconvex tensor tucker decomposition with factor prior,

    T. Liu, J. Yang, B. Li, Y . Wang, and W. An, “Infrared small target detection via nonconvex tensor tucker decomposition with factor prior,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–17, 2023

  39. [46]

    Infrared small target detection in satellite videos: A new dataset and a novel recurrent feature refinement framework,

    X. Ying, L. Liu, Z. Lin, Y . Shi, Y . Wang, R. Li, X. Cao, B. Li, S. Zhou, and W. An, “Infrared small target detection in satellite videos: A new dataset and a novel recurrent feature refinement framework,”IEEE Transactions on Geoscience and Remote Sensing, 2025

  40. [47]

    Probing deep into temporal profile makes the infrared small target detector much better,

    R. Li, W. An, X. Ying, Y . Wang, Y . Dai, L. Wang, M. Li, Y . Guo, and L. Liu, “Probing deep into temporal profile makes the infrared small target detector much better,”arXiv preprint arXiv:2506.12766, 2025

  41. [48]

    Learn to represent and suppress smears with imaging mechanism guided neural network,

    H. Wang, X. Zhang, S. Chen, T. Zhang, G. Zhang, B. Lin, and X. Xu, “Learn to represent and suppress smears with imaging mechanism guided neural network,”IEEE Transactions on Industrial Informatics, pp. 1–10, 2025

  42. [49]

    Anomalous individuals searching framework for space debris detection in single optical astronomical image,

    H. Wang, G. Zhang, L. Wang, S. Chen, Z. Shen, X. Yang, X. Xu, and X. Zhang, “Anomalous individuals searching framework for space debris detection in single optical astronomical image,”Advances in Space Research, vol. 75, no. 4, pp. 3820–3837, 2025

  43. [50]

    Topological sweep for multi-target detection of geostationary space objects,

    D. Liu, B. Chen, T.-J. Chin, and M. G. Rutten, “Topological sweep for multi-target detection of geostationary space objects,”IEEE Transactions on Signal Processing, vol. 68, pp. 5166–5177, 2020

  44. [51]

    Space target detection in complicated situations for wide-field surveillance,

    M. Li, C. Yan, C. Hu, C. Liu, and L. Xu, “Space target detection in complicated situations for wide-field surveillance,”IEEE Access, vol. 7, pp. 123 658–123 670, 2019

  45. [52]

    Space target extraction and detection for wide-field surveillance,

    D. Liu, X. Wang, Z. Xu, Y . Li, and W. Liu, “Space target extraction and detection for wide-field surveillance,”Astronomy and Computing, vol. 32, p. 100408, 2020

  46. [53]

    Space debris detection in optical image sequences,

    J. Xi, D. Wen, O. K. Ersoy, H. Yi, D. Yao, Z. Song, and S. Xi, “Space debris detection in optical image sequences,”Applied optics, vol. 55, no. 28, pp. 7929–7940, 2016

  47. [54]

    Using moffat profiles to register astronomical images,

    M. Schuckman, R. Prouty, D. Chapman, and D. Engel, “Using moffat profiles to register astronomical images,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 80–95

  48. [55]

    Improving the space surveillance telescope’s performance using multi-hypothesis testing,

    J. C. Zingarelli, E. Pearce, R. Lambour, T. Blake, C. J. Peterson, and S. Cain, “Improving the space surveillance telescope’s performance using multi-hypothesis testing,”The Astronomical Journal, vol. 147, no. 5, p. 111, 2014

  49. [56]

    Sdebrisnet: A spatial–temporal saliency network for space debris detection,

    J. Tao, Y . Cao, and M. Ding, “Sdebrisnet: A spatial–temporal saliency network for space debris detection,”Applied Sciences, vol. 13, no. 8, p. 4955, 2023

  50. [57]

    Early convolutions help transformers see better,

    T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollar, and R. Girshick, “Early convolutions help transformers see better,” inAdvances in Neural Information Processing Systems, vol. 34, 2021, pp. 30 392–30 400

  51. [58]

    Hiri-vit: Scaling vision transformer with high resolution inputs,

    T. Yao, Y . Li, Y . Pan, and T. Mei, “Hiri-vit: Scaling vision transformer with high resolution inputs,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 9, pp. 6431–6442, 2024

  52. [59]

    Mb-taylorformer v2: Improved multi-branch linear transformer expanded by taylor formula for image restoration,

    Z. Jin, Y . Qiu, K. Zhang, H. Li, and W. Luo, “Mb-taylorformer v2: Improved multi-branch linear transformer expanded by taylor formula for image restoration,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 7, pp. 5990–6005, 2025

  53. [60]

    Csformer: Bridging convolution and transformer for compressive sensing,

    D. Ye, Z. Ni, H. Wang, J. Zhang, S. Wang, and S. Kwong, “Csformer: Bridging convolution and transformer for compressive sensing,”IEEE Transactions on Image Processing, vol. 32, pp. 2827–2842, 2023

  54. [61]

    Revivediff: A universal diffusion model for restoring images in adverse weather conditions,

    W. Huang, G. Xu, W. Jia, S. Perry, and G. Gao, “Revivediff: A universal diffusion model for restoring images in adverse weather conditions,” IEEE Transactions on Image Processing, vol. 34, pp. 4706–4720, 2025

  55. [62]

    A survey of the vision transformers and their cnn-transformer based variants,

    A. Khan, Z. Rauf, A. Sohail, A. R. Khan, H. Asif, A. Asif, and U. Farooq, “A survey of the vision transformers and their cnn-transformer based variants,”Artificial Intelligence Review, vol. 56, no. Suppl 3, pp. 2917– 2970, 2023

  56. [63]

    Scaled relu matters for training vision transformers,

    P. Wang, X. Wang, H. Luo, J. Zhou, Z. Zhou, F. Wang, H. Li, and R. Jin, “Scaled relu matters for training vision transformers,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 3, pp. 2495–2503, Jun. 2022

  57. [64]

    Framelet representation of tensor nuclear norm for third-order tensor completion,

    T.-X. Jiang, M. K. Ng, X.-L. Zhao, and T.-Z. Huang, “Framelet representation of tensor nuclear norm for third-order tensor completion,” IEEE Transactions on Image Processing, vol. 29, pp. 7233–7244, 2020

  58. [65]

    Self-supervised nonlinear transform-based tensor nuclear norm for multi- dimensional image recovery,

    Y .-S. Luo, X.-L. Zhao, T.-X. Jiang, Y . Chang, M. K. Ng, and C. Li, “Self-supervised nonlinear transform-based tensor nuclear norm for multi- dimensional image recovery,”IEEE Transactions on Image Processing, vol. 31, pp. 3793–3808, 2022

  59. [66]

    Nonlocal patch tensor sparse representation for hyperspectral image super-resolution,

    Y . Xu, Z. Wu, J. Chanussot, and Z. Wei, “Nonlocal patch tensor sparse representation for hyperspectral image super-resolution,”IEEE Transactions on Image Processing, vol. 28, no. 6, pp. 3034–3047, 2019. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13

  60. [67]

    Multi-dimensional visual data restoration: Uncovering the global discrepancy in transformed high- order tensor singular values,

    C. He, Y . Xu, Z. Wu, S. Zheng, and Z. Wei, “Multi-dimensional visual data restoration: Uncovering the global discrepancy in transformed high- order tensor singular values,”IEEE Transactions on Image Processing, vol. 33, pp. 6409–6424, 2024

  61. [68]

    Low-rank high-order tensor completion with applications in visual data,

    W. Qin, H. Wang, F. Zhang, J. Wang, X. Luo, and T. Huang, “Low-rank high-order tensor completion with applications in visual data,”IEEE Transactions on Image Processing, vol. 31, pp. 2433–2448, 2022

  62. [69]

    Kronecker-basis-representation based tensor sparsity and its applications to tensor recovery,

    Q. Xie, Q. Zhao, D. Meng, and Z. Xu, “Kronecker-basis-representation based tensor sparsity and its applications to tensor recovery,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 8, pp. 1888–1902, 2018

  63. [70]

    A registration algorithm for astronomical images based on geometric constraints and homography,

    B. Lin, X. Xu, Z. Shen, X. Yang, L. Zhong, and X. Zhang, “A registration algorithm for astronomical images based on geometric constraints and homography,”Remote Sensing, vol. 15, no. 7, p. 1921, 2023

  64. [71]

    Dtnet: A specialized dual-tuning network for infrared vehicle detection in aerial images,

    N. Zhang, Y . Liu, H. Liu, T. Tian, J. Ma, and J. Tian, “Dtnet: A specialized dual-tuning network for infrared vehicle detection in aerial images,”IEEE Transactions on Geoscience and Remote Sensing, 2024

  65. [72]

    Fast and accurate tensor completion with total variation regularized tensor trains,

    C.-Y . Ko, K. Batselier, L. Daniel, W. Yu, and N. Wong, “Fast and accurate tensor completion with total variation regularized tensor trains,” IEEE Transactions on Image Processing, vol. 29, pp. 6918–6931, 2020

  66. [73]

    Exploring video denoising in thermal infrared imaging: Physics-inspired noise generator, dataset, and model,

    L. Cai, X. Dong, K. Zhou, and X. Cao, “Exploring video denoising in thermal infrared imaging: Physics-inspired noise generator, dataset, and model,”IEEE Transactions on Image Processing, vol. 33, pp. 3839–3854, 2024

  67. [74]

    Deep variational network toward blind image restoration,

    Z. Yue, H. Yong, Q. Zhao, L. Zhang, D. Meng, and K.-Y . K. Wong, “Deep variational network toward blind image restoration,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 11, pp. 7011–7026, 2024

  68. [75]

    Is denoising dead?

    P. Chatterjee and P. Milanfar, “Is denoising dead?”IEEE Transactions on Image Processing, vol. 19, no. 4, pp. 895–911, 2010

  69. [76]

    Multi-scale patch-based image restoration,

    V . Papyan and M. Elad, “Multi-scale patch-based image restoration,” IEEE Transactions on Image Processing, vol. 25, no. 1, pp. 249–261, 2016

  70. [77]

    Image denoising via sparse and redundant representations over learned dictionaries,

    M. Elad and M. Aharon, “Image denoising via sparse and redundant representations over learned dictionaries,”IEEE Transactions on Image Processing, vol. 15, no. 12, pp. 3736–3745, 2006

  71. [78]

    Deep image prior,

    D. Ulyanov, A. Vedaldi, and V . Lempitsky, “Deep image prior,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 9446–9454

  72. [79]

    Deep image prior,

    U. Dmitry, A. Vedaldi, and L. Victor, “Deep image prior,”International Journal of Computer Vision, vol. 128, no. 7, pp. 1867–1888, 2020

  73. [80]

    Weighted tensor rank-1 decomposition for nonlocal image denoising,

    Y . Wu, L. Fang, and S. Li, “Weighted tensor rank-1 decomposition for nonlocal image denoising,”IEEE Transactions on Image Processing, vol. 28, no. 6, pp. 2719–2730, 2019

  74. [81]

    Video swin transformer,

    Z. Liu, J. Ning, Y . Cao, Y . Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 3202–3211

  75. [82]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,”ICLR, 2021

  76. [83]

    Manifold learning: What, how, and why,

    M. Meil ˘a and H. Zhang, “Manifold learning: What, how, and why,” Annual Review of Statistics and Its Application, vol. 11, no. 1, pp. 393–417, 2024

  77. [84]

    Primal-attention: Self- attention through asymmetric kernel svd in primal representation,

    Y . Chen, Q. Tao, F. Tonin, and J. Suykens, “Primal-attention: Self- attention through asymmetric kernel svd in primal representation,” Advances in Neural Information Processing Systems, vol. 36, pp. 65 088– 65 101, 2023

  78. [85]

    Diffusion kernels on statistical manifolds

    J. Lafferty, G. Lebanon, and T. Jaakkola, “Diffusion kernels on statistical manifolds.”Journal of Machine Learning Research, vol. 6, no. 1, 2005

  79. [87]

    Denoising hyper- spectral image with non-i.i.d. noise structure,

    Y . Chen, X. Cao, Q. Zhao, D. Meng, and Z. Xu, “Denoising hyper- spectral image with non-i.i.d. noise structure,”IEEE Transactions on Cybernetics, vol. 48, no. 3, pp. 1054–1066, 2018

  80. [88]

    Robust low-rank matrix factorization under general mixture noise distributions,

    X. Cao, Q. Zhao, D. Meng, Y . Chen, and Z. Xu, “Robust low-rank matrix factorization under general mixture noise distributions,”IEEE Transactions on Image Processing, vol. 25, no. 10, pp. 4677–4690, 2016

  81. [89]

    Robust multiview subspace learning with nonindependently and nonidentically distributed complex noise,

    Z. Yue, H. Yong, D. Meng, Q. Zhao, Y . Leung, and L. Zhang, “Robust multiview subspace learning with nonindependently and nonidentically distributed complex noise,”IEEE Transactions on Neural Networks and Learning Systems, vol. 31, no. 4, pp. 1070–1083, 2020

  82. [90]

    Multiscale patch-based contrast measure for small infrared target detection,

    Y . Wei, X. You, and H. Li, “Multiscale patch-based contrast measure for small infrared target detection,”Pattern Recognition, vol. 58, pp. 216–226, 2016

  83. [91]

    Infrared dim and small target detection via multiple subspace learning and spatial-temporal patch-tensor model,

    Y . Sun, J. Yang, and W. An, “Infrared dim and small target detection via multiple subspace learning and spatial-temporal patch-tensor model,” IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 5, pp. 3737–3752, 2020

  84. [92]

    IMNN-LWEC: A novel infrared small target detection based on spatial-temporal tensor model,

    Y . Luo, X. Li, S. Chen, C. Xia, and L. Zhao, “IMNN-LWEC: A novel infrared small target detection based on spatial-temporal tensor model,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–22, 2022

  85. [93]

    Sparse regularization-based spatial-temporal twist tensor model for infrared small target detection,

    J. Li, P. Zhang, L. Zhang, and Z. Zhang, “Sparse regularization-based spatial-temporal twist tensor model for infrared small target detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–17, 2023

  86. [94]

    4DST-BTMD: An infrared small target detection method based on 4-D data-sphered space,

    Y . Luo, X. Li, S. Chen, and C. Xia, “4DST-BTMD: An infrared small target detection method based on 4-D data-sphered space,”IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–20, 2023

  87. [95]

    Spatial-temporal tensor represen- tation learning with priors for infrared small target detection,

    Y . Luo, X. Li, Y . Yan, and C. Xia, “Spatial-temporal tensor represen- tation learning with priors for infrared small target detection,”IEEE Transactions on Aerospace and Electronic Systems, vol. 59, no. 6, pp. 9598–9620, 2023

  88. [96]

    Infrared small target detection using spatiotemporal 4-D tensor train and ring unfolding,

    F. Wu, H. Yu, A. Liu, J. Luo, and Z. Peng, “Infrared small target detection using spatiotemporal 4-D tensor train and ring unfolding,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–22, 2023

  89. [97]

    Asymmetric contextual modulation for infrared small target detection,

    Y . Dai, Y . Wu, F. Zhou, and K. Barnard, “Asymmetric contextual modulation for infrared small target detection,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), January 2021, pp. 950–959

  90. [98]

    Attentional local contrast networks for infrared small target detection,

    ——, “Attentional local contrast networks for infrared small target detection,”IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 11, pp. 9813–9824, 2021

  91. [99]

    Weighted Res-UNet for high- quality retina vessel segmentation,

    X. Xiao, S. Lian, Z. Luo, and S. Li, “Weighted Res-UNet for high- quality retina vessel segmentation,” inInternational Conference on Information Technology in Medicine and Education, 2018, pp. 327–331

  92. [100]

    Isnet: Shape matters for infrared small target detection,

    M. Zhang, R. Zhang, Y . Yang, H. Bai, J. Zhang, and J. Guo, “Isnet: Shape matters for infrared small target detection,” in2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 867–876

  93. [101]

    Attention-guided pyramid context networks for detecting infrared small target under complex background,

    T. Zhang, L. Li, S. Cao, T. Pu, and Z. Peng, “Attention-guided pyramid context networks for detecting infrared small target under complex background,”IEEE Transactions on Aerospace and Electronic Systems, vol. 59, no. 4, pp. 4250–4261, 2023

  94. [102]

    Sctransnet: Spatial- channel cross transformer network for infrared small target detection,

    S. Yuan, H. Qin, X. Yan, N. Akhtar, and A. Mian, “Sctransnet: Spatial- channel cross transformer network for infrared small target detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–15, 2024

  95. [103]

    Mim-istd: Mamba-in-mamba for efficient infrared small-target detection,

    T. Chen, Z. Ye, Z. Tan, T. Gong, Y . Wu, Q. Chu, B. Liu, N. Yu, and J. Ye, “Mim-istd: Mamba-in-mamba for efficient infrared small-target detection,”IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–13, 2024

  96. [104]

    Rpcanet: Deep unfolding rpca based infrared small target detection,

    F. Wu, T. Zhang, L. Li, Y . Huang, and Z. Peng, “Rpcanet: Deep unfolding rpca based infrared small target detection,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), January 2024, pp. 4809–4818

  97. [105]

    STDMANet: Spatio-temporal differential multiscale attention network for small moving infrared target detection,

    P. Yan, R. Hou, X. Duan, C. Yue, X. Wang, and X. Cao, “STDMANet: Spatio-temporal differential multiscale attention network for small moving infrared target detection,”IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–16, 2023

  98. [106]

    Lmaformer: Local motion aware transformer for small moving infrared target detection,

    Y . Huang, X. Zhi, J. Hu, L. Yu, Q. Han, W. Chen, and W. Zhang, “Lmaformer: Local motion aware transformer for small moving infrared target detection,”IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–17, 2024

  99. [107]

    Single-frame infrared small-target detection: A survey,

    M. Zhao, W. Li, L. Li, J. Hu, P. Ma, and R. Tao, “Single-frame infrared small-target detection: A survey,”IEEE Geoscience and Remote Sensing Magazine, vol. 10, no. 2, pp. 87–119, 2022

  100. [109]

    A linear support higher- order tensor machine for classification,

    Z. Hao, L. He, B. Chen, and X. Yang, “A linear support higher- order tensor machine for classification,”IEEE Transactions on Image Processing, vol. 22, no. 7, pp. 2911–2920, 2013

  101. [110]

    Lsknet: A foundation lightweight backbone for remote sensing,

    Y . Li, X. Li, Y . Dai, Q. Hou, L. Liu, Y . Liu, M.-M. Cheng, and J. Yang, “Lsknet: A foundation lightweight backbone for remote sensing,” International Journal of Computer Vision, pp. 1–22, 2024

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.