Pith. sign in

REVIEW 3 major objections 64 references

LCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target Detection

T0 review · 3 major / 0 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read Infrared small-target detection works better when low-rank/sparse unfolding is done in latent space with consistent proximal updates and shared stage memory.

desk verdict Solid RPCANet-line engineering: latent proximal updates + shared memory cut false alarms hard; SOTA claim is real but single-run and conservative on Pd. read the letter →

arxiv 2607.04603 v1 pith:MIQ4NXPU submitted 2026-07-06 cs.CV cs.AI

classification cs.CVcs.AI
keywords infraredsmalltargetdetectiondeepunfoldinglow-rankandsparsedecompositionlatent-spaceoptimizationproximalsolversharedmemoryremotesensing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Infrared small target detection must find dim, few-pixel objects against bright, structured clutter. Pure feed-forward networks learn image-to-mask maps but ignore the physical idea that an infrared frame is background plus sparse target plus noise. Earlier deep-unfolding methods put that idea into a network, yet they still work mostly in image space and update variables in ways that break continuity with the optimization. This paper claims that the same low-rank prior still holds after a learned lift into latent features, so the whole iterative separation can run there without repeatedly crushing intermediate states back to a single channel. It then replaces residual-style reconstruction with a Latent Consistent Proximal solver that evolves each variable from its own previous state, and replaces branch-local memory with one Shared Optimization Memory that guides background, target, and noise together. On four public benchmarks the resulting LCPNet raises overlap accuracy while driving false alarms down and keeping runtime competitive.

What carries the argument

Latent Consistent Proximal (LCP) unfolding: after verifying the low-rank prior in latent features, ADMM-style stages run in that space; each variable is updated from its previous state via a learnable proximal surrogate (with group and spectral normalization), while Shared Optimization Memory supplies a single gated historical state to all branches.

What would settle it

If, on held-out infrared scenes, the latent tensors after the paper's encoders do not show rapidly decaying singular values (or Tucker rank), or if forcing the latent decomposition constraint measurably raises false-alarm rate versus an identical architecture without that constraint, the central physical-prior claim fails.

Watch

Extended reading notes

Core claim

The authors establish that low-rank/sparse decomposition remains valid in a learned latent space, and that a deep-unfolding network built on a Latent Consistent Proximal solver plus Shared Optimization Memory yields more accurate, lower-false-alarm infrared small-target detection than prior HVS, optimization, deep, and deep-unfolding methods on four public benchmarks.

Load-bearing premise

The claim rests on the premise that after a learned image-to-latent map, the background is still meaningfully low-rank and the target still sparse in that latent space, so the imposed decomposition constraint still encodes real scene structure rather than an encoder artifact.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper proposes LCPNet, a deep-unfolding IRSTD method that lifts low-rank/sparse decomposition from the image domain into a multi-channel latent space (Eqs. 2–3), derives a Latent Consistent Proximal (LCP) solver that updates each variable from its previous state via a Lipschitz majorization of an unknown regularizer (Eqs. 9–17, Appendices A–B), and introduces Shared Optimization Memory (SOM) as a gated recurrent state shared by all decomposition variables (Eqs. 18–20, Appendix C). After K latent ADMM-style stages, the target latent is decoded to a detection mask. On NUDT-SIRST, IRSTD-1K, SIRST, and SIRST-Aug, LCPNet-4/6 report best or near-best IoU/F1, the lowest Fa, competitive AUC/ROC, and moderate params/FLOPs/latency versus HVS, optimization, deep, and prior unfolding baselines (Tables I–II, Figs. 6–11), with ablations on domain, solver style, updater regularization, memory type, and stage count (Tables III–V).

Significance. If the empirical gains hold under stronger evaluation, the work is a solid incremental contribution to interpretable IRSTD: it couples a latent-domain physical constraint with a proximal-style update and system-level memory, and it ships code plus detailed ADMM/majorization/SOM derivations. The combination of latent unfolding, GN+SN updater design, and shared memory is practically useful for false-alarm-sensitive remote sensing, and the multi-benchmark comparison against RPCANet-family and strong non-unfolding nets is valuable even if the absolute novelty over prior deep RPCA is moderate.

major comments (3)
  1. Abstract / §IV-B / Table I: the unqualified claim that LCPNet “outperforms state-of-the-art methods” is not fully supported by the reported metrics. On IRSTD-1K, LCPNet-6 IoU/F1 are best, but Pd (87.63%) is below DRPCANet (92.09%), DNANet (92.44%), and MSHNet (92.78%); on SIRST, Pd is 96.33% vs RPCANet++ 100% and DRPCANet 99.08%, while Fa is driven near zero. The operating point is therefore more conservative than uniformly superior. Please restate claims in terms of the IoU–Fa tradeoff (and/or fixed-Pd Fa), report multi-seed means±std or at least repeated runs for the headline margins, and avoid “outperforms SOTA” language where Pd is materially lower.
  2. §III-A / Eqs. 2–3 / Fig. 2 / Appendix D: the load-bearing premise that the low-rank prior “remains valid” after the learned lift is only partially evidenced. Rapid singular-value / Tucker-rank decay shows compressibility of the latent tensor, not that the learned encoders preserve a physically meaningful background–target–noise split under X=B+T+N. Without controls that freeze or ablate the encoders, or that measure reconstruction fidelity of B/T/N under the latent constraint, the model-driven interpretation of the unfolded stages remains partly circular. A short encoder-control experiment or quantitative latent-rank vs. image-rank comparison under fixed encoders would substantially strengthen this claim.
  3. Table V / §IV-C5: stage-depth behavior is non-monotonic (NUDT IoU peaks at K=5 then drops at K=6; Fa fluctuates 0.011–0.092), yet LCPNet-4 and LCPNet-6 are presented as primary models without a selection protocol, validation criterion, or uncertainty. Because K is a free hyperparameter that changes both accuracy and cost, the paper should either fix K a priori, select it on a held-out split with a stated rule, or report the full K-curve with variance so the SOTA numbers are not depth-tuned post hoc.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: architecture is optimization-motivated with learnable surrogates; SOTA claims are external-benchmark evaluations, not results forced by definition or self-citation.

full rationale

LCPNet’s load-bearing claims are empirical (Table I IoU/F1/Pd/Fa and ROC/AUC on four public IRSTD benchmarks against independent and prior methods) and architectural (latent ADMM-style unfolding, LCP proximal surrogate, SOM). The LCP “derivation” (Eqs. 6–17, Appendices A–B) starts from a standard ADMM Lagrangian, assumes an unknown Lipschitz regularizer, majorizes it, and obtains a proximal-style update form; the ideal direction is then replaced by a learnable network Ψ, so the solver is not a closed prediction that equals its inputs by construction. Low-rank validity in latent space is an empirical design premise (Fig. 2, Appendix D), not a tautology that forces the reported metrics. Self-citations to RPCANet/RPCANet++/DRPCANet supply motivation and baselines, not uniqueness theorems or fitted quantities renamed as predictions. Training is ordinary supervised SoftIoU learning on held-out splits; nothing in the chain reduces Eq. X to Eq. Y by definition or fits a parameter then “predicts” a statistically forced sibling quantity. Score 0 is appropriate.

Assumptions & free parameters 5 free parameters · 4 assumptions · 3 invented entities

The central empirical claim rests on standard IRSTD modeling (low-rank background + sparse target), ADMM-style alternating proximal steps, and several architectural choices (latent lift, stage depth, GN/SN, SoftIoU training) that are free design parameters rather than derived constants. Invented modules (LCP solver, SOM) are engineering constructs justified by ablations, not new physical entities.

free parameters (5)
  • Unfolding stage count K = 4 or 6 (main models)
    Chosen and ablated (K=2..6); reported models use K=4 and K=6. Depth is not derived from theory.
  • Latent channel width C and encoder/decoder architecture
    Defines the latent manifold where decomposition is enforced; selected as network design, not measured from physics.
  • Group count g in GroupNorm and spectral-normalized conv gains
    Task-adaptive normalization/gain control hyperparameters inside the LCP updater.
  • Training hyperparameters (lr=1e-4, poly power 0.9, batch=8, epochs 400/800, SoftIoU) = lr 1e-4; SoftIoU; 800/400 epochs
    Fully determine the learned proximal surrogates and final metrics; standard but free.
  • ADMM penalty / step-size related coefficients (μ, η via L+μ)
    Appear in the derived update; in the network they are absorbed into learned modules rather than fixed from first principles.
assumptions (4)
  • domain assumption Infrared observations admit a low-rank background + sparse target + noise decomposition (image and, after encoding, latent).
    Inherited from IPI/RPCA literature and re-asserted for latents via Fig. 2 and Appendix D; load-bearing for the physical-constraint claim.
  • standard math Unknown latent regularizers have Lipschitz-continuous gradients, enabling the quadratic majorization used to derive the LCP closed form.
    Descent-lemma style assumption (Eqs. 9–13, Appendix B); standard optimization tool, not verified for the learned latent regularizer.
  • domain assumption ADMM alternating updates in latent space remain a valid algorithmic skeleton once analytical proximal maps are replaced by neural surrogates.
    Standard deep-unfolding premise; convergence of the learned finite-stage network is not proved.
  • domain assumption Public single-frame IRSTD splits (NUDT-SIRST 1:1, IRSTD-1K, SIRST, SIRST-Aug) are adequate proxies for real long-range detection performance.
    All quantitative claims rest on these benchmarks under SoftIoU training.
invented entities (3)
  • Latent Consistent Proximal (LCP) solver
    purpose: Replace residual-style reconstruction with a previous-state-anchored proximal surrogate update for B,T,N in latent space.
    New algorithmic module derived then parameterized by conv+GN+SN; evidence is ablation and SOTA tables, not external theory.
  • Shared Optimization Memory (SOM)
    purpose: Provide one gated recurrent historical state shared by all decomposition variables across stages.
    System-level memory distinct from branch-only ConvLSTM in RPCANet++; justified by Table IV, not by independent physical measurement.
  • Latent-space IRSTD unfolding formulation (X=B+T+N in R^{H×W×C})
    purpose: Keep physical additivity while avoiding repeated image-domain compression between stages.
    Core modeling move of the paper; low-rank checks are internal visualizations/rank plots, not external validation of the latent physics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target Detection." pith.science (2026). https://pith.science/paper/MIQ4NXPU

@misc{pith2026260704603,
  author       = {Pith},
  title        = {Pith review of: LCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target Detection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MIQ4NXPU}},
  note         = {Machine review of arXiv:2607.04603}
}
read the original abstract

Infrared small target detection (IRSTD) aims to identify long distance small targets from complex infrared backgrounds, and is a fundamental task in remote sensing. Deep learning methods have improved IRSTD by learning discriminative image-to-mask mappings, but such feed-forward designs often underuse physical decomposition structure between targets and backgrounds. Deep unfolding methods partially address this issue by embedding model-driven iterations into neural networks, yet existing designs still operate mainly in image domain and use updates and memory mechanisms that are not fully coupled with underlying optimization process. To address these limitations, we propose Latent Consistent Proximal unfolding network (LCPNet). First, we verify that low-rank prior remains valid in latent representations and perform unfolding in this space, preserving physical constraint while avoiding repeated compression of intermediate states. Second, we derive a Latent Consistent Proximal (LCP) solver that evolves each latent variable from its previous state rather than reconstructing through an indirect residual, and stabilizes small target updates through task-adaptive normalization and gain control. Third, we introduce Shared Optimization Memory (SOM), a common historical state shared by all decomposition variables to provide coordinated guidance across unfolding stages. Extensive experiments on four public benchmarks demonstrate that LCPNet outperforms state-of-the-art methods while achieving accurate and robust detection with low false alarms and competitive efficiency. Model and code are available at https://github.com/Tianfang-Zhang/LCPNet.

Figures

Figures reproduced from arXiv: 2607.04603 by the authors.

Figure 1
Figure 1. Comparison of Fa(10−5 )-IoU scatter plots for infrared small target detection algorithms on IRSTD-1k [7]. Circle size indicates the parameter number. Points closer to top-left indicate better performance. cal value and is one of the most prominent and challenging tasks in low-level vision and remote sensing. The difficulty of IRSTD stems from the inherent imaging characteristics of infrared small targets. In long-ra… view at source ↗
Figure 2
Figure 2. Visualization of the low-rank property in latent domain. From top [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 4
Figure 4. Overall architecture of LCPNet. An infrared image is first lifted into the latent domain, where latent-space decomposition is performed through [PITH_FULL_IMAGE:figures/full_fig_p003_4.png] view at source ↗
Figures from the paper (12 more)
Figure 5
Figure 5. Figure 5: Details of our custom-designed modules. (a) Illustration of the [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Visualization of the last layer feature maps produced by different methods. For each method, the feature map is projected by PCA, and the most [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: ROC curves of different state-of-the-art methods across four infrared small target detection datasets. Our method is represented by the red line. [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Qualitative comparison on image XDU685. Green boxes denote correct detections, while red boxes denote false alarms or missed targets. Best view in color. NUDT-SIRST: 000660.png [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Qualitative comparison on image 000660. Green boxes denote correct detections, while red boxes denote false alarms or missed targets. Best view in color. inference. Deep learning-based methods improve inference efficiency by replacing iterative solvers with feed-forwar…
Figure 10
Figure 10. Figure 10: Qualitative comparison on image 000971. Green boxes denote correct detections, while red boxes denote false alarms or missed targets. Best view in color. NUDT-SIRST: 000523.png [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]
Figure 11
Figure 11. Figure 11: Qualitative comparison on image 000523. Green boxes denote correct detections, while red boxes denote false alarms or missed targets. Best view in color. result indicates that the expanded latent representation can better separate target and background components when…
Figure 12
Figure 12. Figure 12: Supplementary on the visualization of last layer feature maps produced by different methods. For each method, the feature map is projected by [PITH_FULL_IMAGE:figures/full_fig_p016_12.png]
Figure 13
Figure 13. Figure 13: Visualization of the low-rank property in latent domain. From left to right are the original infrared image, a visualized top three channels of the [PITH_FULL_IMAGE:figures/full_fig_p017_13.png]
Figure 14
Figure 14. Figure 14: Qualitative comparison on image XDU709. Green boxes denote correct detections, while red boxes denote false alarms or missed targets. Best view in color. IRSTD-1K: XDU343.png [PITH_FULL_IMAGE:figures/full_fig_p017_14.png]
Figure 15
Figure 15. Figure 15: Qualitative comparison on image XDU343. Green boxes denote correct detections, while red boxes denote false alarms or missed targets. Best view in color. Therefore, hk is not determined only by hk−1 or a fixed short window of previous states. Instead, it contains a ga…
Figure 16
Figure 16. Figure 16: Qualitative comparison on image 001176. Green boxes denote correct detections, while red boxes denote false alarms or missed targets. Best view in color. APPENDIX D MORE EXPERIMENTAL RESULTS This section provides additional qualitative results to com￾plement the visua…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 2 linked inside Pith

  1. [1]

    Single-frame infrared small-target detection: A survey,

    M. Zhao, W. Li, L. Li, J. Hu, P. Ma, and R. Tao, “Single-frame infrared small-target detection: A survey,”IEEE Geoscience and Remote Sensing Magazine, vol. 10, no. 2, pp. 87–119, 2022

  2. [2]

    Infrared small target segmentation networks: A survey,

    R. Kou, C. Wang, Z. Peng, Z. Zhao, Y . Chen, J. Han, F. Huang, Y . Yu, and Q. Fu, “Infrared small target segmentation networks: A survey,” Pattern Recognition, vol. 143, p. 109788, 2023

  3. [3]

    Miss detection vs. false alarm: Adversarial learning for small object segmentation in infrared images,

    H. Wang, L. Zhou, and L. Wang, “Miss detection vs. false alarm: Adversarial learning for small object segmentation in infrared images,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 8509–8518

  4. [4]

    A local contrast method for small infrared target detection,

    C. P. Chen, H. Li, Y . Wei, T. Xia, and Y . Y . Tang, “A local contrast method for small infrared target detection,”IEEE Transactions on Geoscience and Remote Sensing, vol. 52, no. 1, pp. 574–581, 2013

  5. [5]

    Infrared patch-image model for small target detection in a single image,

    C. Gao, D. Meng, Y . Yang, Y . Wang, X. Zhou, and A. G. Hauptmann, “Infrared patch-image model for small target detection in a single image,”IEEE Transactions on Image Processing, vol. 22, no. 12, pp. 4996–5009, 2013

  6. [6]

    Attentional local contrast networks for infrared small target detection,

    Y . Dai, Y . Wu, F. Zhou, and K. Barnard, “Attentional local contrast networks for infrared small target detection,”IEEE transactions on geoscience and remote sensing, vol. 59, no. 11, pp. 9813–9824, 2021

  7. [7]

    Isnet: Shape matters for infrared small target detection,

    M. Zhang, R. Zhang, Y . Yang, H. Bai, J. Zhang, and J. Guo, “Isnet: Shape matters for infrared small target detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 877–886

  8. [8]

    Asymmetric contextual modulation for infrared small target detection,

    Y . Dai, Y . Wu, F. Zhou, and K. Barnard, “Asymmetric contextual modulation for infrared small target detection,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 950–959

Show all 64 references
  1. [9]

    Dense nested attention network for infrared small target detection,

    B. Li, C. Xiao, L. Wang, Y . Wang, Z. Lin, M. Li, W. An, and Y . Guo, “Dense nested attention network for infrared small target detection,” IEEE Transactions on Image Processing, vol. 32, pp. 1745–1758, 2022

  2. [10]

    Uiu-net: U-net in u-net for infrared small object detection,

    X. Wu, D. Hong, and J. Chanussot, “Uiu-net: U-net in u-net for infrared small object detection,”IEEE Transactions on Image Processing, vol. 32, pp. 364–376, 2022

  3. [11]

    Attention-guided pyramid context networks for detecting infrared small target under complex background,

    T. Zhang, L. Li, S. Cao, T. Pu, and Z. Peng, “Attention-guided pyramid context networks for detecting infrared small target under complex background,”IEEE Transactions on Aerospace and Electronic Systems, vol. 59, no. 4, pp. 4250–4261, 2023

  4. [12]

    Infrared small target detection with scale and location sensitivity,

    Q. Liu, R. Liu, B. Zheng, H. Wang, and Y . Fu, “Infrared small target detection with scale and location sensitivity,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 17 490–17 499

  5. [13]

    Robust principal component analysis?

    E. J. Cand `es, X. Li, Y . Ma, and J. Wright, “Robust principal component analysis?”Journal of the ACM, vol. 58, no. 3, pp. 1–37, 2011

  6. [14]

    Infrared small target detection via non-convex rank approximation minimization joint l 2, 1 norm,

    L. Zhang, L. Peng, T. Zhang, S. Cao, and Z. Peng, “Infrared small target detection via non-convex rank approximation minimization joint l 2, 1 norm,”Remote Sensing, vol. 10, no. 11, p. 1821, 2018

  7. [15]

    Reweighted infrared patch-tensor model with both nonlocal and local priors for single-frame small target detection,

    Y . Dai and Y . Wu, “Reweighted infrared patch-tensor model with both nonlocal and local priors for single-frame small target detection,”IEEE journal of selected topics in applied earth observations and remote sensing, vol. 10, no. 8, pp. 3752–3767, 2017

  8. [16]

    Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing,

    V . Monga, Y . Li, and Y . C. Eldar, “Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing,”IEEE Signal Processing Magazine, vol. 38, no. 2, pp. 18–44, 2021

  9. [17]

    Rpcanet: Deep unfolding rpca based infrared small target detection,

    F. Wu, T. Zhang, L. Li, Y . Huang, and Z. Peng, “Rpcanet: Deep unfolding rpca based infrared small target detection,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 4809–4818

  10. [18]

    Rpcanet++: Deep interpretable robust pca for sparse object segmenta- tion,

    F. Wu, Y . Dai, T. Zhang, Y . Ding, J. Yang, M.-M. Cheng, and Z. Peng, “Rpcanet++: Deep interpretable robust pca for sparse object segmenta- tion,”arXiv preprint arXiv:2508.04190, 2025

  11. [19]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift,

    S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” inInternational conference on machine learning. pmlr, 2015, pp. 448–456

  12. [20]

    Group normalization,

    Y . Wu and K. He, “Group normalization,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 3–19

  13. [21]

    Parseval networks: Improving robustness to adversarial examples,

    M. Cisse, P. Bojanowski, E. Grave, Y . Dauphin, and N. Usunier, “Parseval networks: Improving robustness to adversarial examples,” in International conference on machine learning. PMLR, 2017, pp. 854– 863

  14. [22]

    Spectral normalization for generative adversarial networks,

    T. Miyato, T. Kataoka, M. Koyama, and Y . Yoshida, “Spectral normalization for generative adversarial networks,”arXiv preprint arXiv:1802.05957, 2018

  15. [23]

    Convolutional lstm network: A machine learning approach for precipitation nowcasting,

    X. Shi, Z. Chen, H. Wang, D.-Y . Yeung, W.-K. Wong, and W.-C. Woo, “Convolutional lstm network: A machine learning approach for precipitation nowcasting,” inAdvances in Neural Information Processing Systems, vol. 28, 2015, pp. 802–810

  16. [24]

    Infrared small-target detection based on multi-directional multi-scale high-boost response,

    L. Peng, T. Zhang, S. Huang, T. Pu, Y . Liu, Y . Lv, Y . Zheng, and Z. Peng, “Infrared small-target detection based on multi-directional multi-scale high-boost response,”Optical Review, vol. 26, no. 6, pp. 568–582, 2019

  17. [25]

    Morphology-based algorithm for point target detection in infrared backgrounds,

    V . T. Tom, T. Peli, M. Leung, and J. E. Bondaryk, “Morphology-based algorithm for point target detection in infrared backgrounds,” inSignal and Data Processing of Small Targets 1993, vol. 1954. International Society for Optics and Photonics, 1993, pp. 2–11

  18. [26]

    Max- mean and max-median filters for detection of small targets,

    S. D. Deshpande, M. H. Er, R. Venkateswarlu, and P. Chan, “Max- mean and max-median filters for detection of small targets,” inSignal and Data Processing of Small Targets 1999, vol. 3809. International Society for Optics and Photonics, 1999, pp. 74–83

  19. [27]

    Multiscale patch-based contrast measure for small infrared target detection,

    Y . Wei, X. You, and H. Li, “Multiscale patch-based contrast measure for small infrared target detection,”Pattern Recognition, vol. 58, pp. 216–226, 2016

  20. [28]

    Irsam: Advancing segment anything model for infrared small target detection,

    M. Zhang, Y . Wang, J. Guo, Y . Li, X. Gao, and J. Zhang, “Irsam: Advancing segment anything model for infrared small target detection,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 233– 249

  21. [29]

    Saliency at the helm: Steering infrared small target detection with learnable kernels,

    F. Wu, A. Liu, T. Zhang, L. Zhang, J. Luo, and Z. Peng, “Saliency at the helm: Steering infrared small target detection with learnable kernels,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1– 14, 2024

  22. [30]

    Pinwheel- shaped convolution and scale-based dynamic loss for infrared small JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 target detection,

    J. Yang, S. Liu, J. Wu, X. Su, N. Hai, and X. Huang, “Pinwheel- shaped convolution and scale-based dynamic loss for infrared small JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 target detection,” inProceedings of the AAAI Conference on Artificial Intelligence, v...

  23. [31]

    Ilnet: Low-level matters for salient infrared small target detection,

    H. Li, J. Yang, R. Wang, and Y . Xu, “Ilnet: Low-level matters for salient infrared small target detection,”IEEE Transactions on Aerospace and Electronic Systems, 2025

  24. [32]

    Learning fast approximations of sparse cod- ing,

    K. Gregor and Y . LeCun, “Learning fast approximations of sparse cod- ing,” inProceedings of the 27th International Conference on Machine Learning. Omnipress, 2010, pp. 399–406

  25. [33]

    Optimization-inspired cumulative trans- mission network for image compressive sensing,

    T. Zhang, L. Li, and Z. Peng, “Optimization-inspired cumulative trans- mission network for image compressive sensing,”Knowledge-Based Systems, vol. 279, p. 110963, 2023

  26. [34]

    Ista-net: Interpretable optimization-inspired deep network for image compressive sensing,

    J. Zhang and B. Ghanem, “Ista-net: Interpretable optimization-inspired deep network for image compressive sensing,” in2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 2018, pp. 1828–1837

  27. [35]

    Learned D-AMP: Principled neural network based compressive image recovery,

    C. A. Metzler, A. Mousavi, and R. G. Baraniuk, “Learned D-AMP: Principled neural network based compressive image recovery,” 2017

  28. [36]

    AMP-Net: Denoising based deep unfolding for compressive image sensing,

    Z. Zhang, Y . Liu, J. Liu, F. Wen, and C. Zhu, “AMP-Net: Denoising based deep unfolding for compressive image sensing,”IEEE Transac- tions on Image Processing, vol. 30, pp. 1487–1500, 2021

  29. [37]

    Deep admm-net for compressive sensing mri,

    Y . Yang, J. Sun, H. Li, and Z. Xu, “Deep admm-net for compressive sensing mri,” inAdvances in Neural Information Processing Systems, vol. 29. Curran Associates, Inc., 2016

  30. [38]

    Modl: Model-based deep learning architecture for inverse problems,

    H. K. Aggarwal, M. P. Mani, and M. Jacob, “Modl: Model-based deep learning architecture for inverse problems,”IEEE Transactions on Medical Imaging, vol. 38, no. 2, pp. 394–405, 2018

  31. [39]

    Learning a variational network for reconstruction of accelerated MRI data,

    K. Hammernik, T. Klatzer, E. Kobler, M. P. Recht, D. K. Sodickson, T. Pock, and F. Knoll, “Learning a variational network for reconstruction of accelerated MRI data,”Magnetic Resonance in Medicine, vol. 79, no. 6, pp. 3055–3071, 2017

  32. [40]

    Learned primal-dual reconstruction,

    J. Adler and O. ¨Oktem, “Learned primal-dual reconstruction,”IEEE Transactions on Medical Imaging, vol. 37, no. 6, pp. 1322–1332, 2018

  33. [41]

    End-to-end variational networks for accelerated MRI reconstruction,

    A. Sriram, J. Zbontar, T. Murrell, A. Defazio, C. L. Zitnick, N. Yakubova, F. Knoll, and P. Johnson, “End-to-end variational networks for accelerated MRI reconstruction,” inMedical Image Computing and Computer Assisted Intervention – MICCAI 2020. Springer International Publish...

  34. [42]

    Learned robust pca: A scalable deep unfolding approach for high-dimensional outlier detection,

    H. Cai, J. Liu, and W. Yin, “Learned robust pca: A scalable deep unfolding approach for high-dimensional outlier detection,” inAdvances in Neural Information Processing Systems, vol. 34. Curran Associates, Inc., 2021, pp. 16 977–16 989

  35. [43]

    Deep unsupervised learning using nonequilibrium thermodynamics,

    J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” 2015

  36. [44]

    Generative modeling by estimating gradients of the data distribution,

    Y . Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” 2020

  37. [45]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” inAdvances in Neural Information Processing Systems, vol. 33. Curran Associates, Inc., 2020, pp. 6840–6851

  38. [46]

    Improved denoising diffusion probabilistic models,

    A. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” 2021

  39. [47]

    Denoising diffusion implicit models,

    J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” 2022

  40. [48]

    Score-based generative modeling through stochastic differ- ential equations,

    Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differ- ential equations,” inInternational Conference on Learning Representa- tions, 2021

  41. [49]

    High- resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High- resolution image synthesis with latent diffusion models,” 2022

  42. [50]

    Nesterov,Introductory Lectures on Convex Optimization: A Basic Course, ser

    Y . Nesterov,Introductory Lectures on Convex Optimization: A Basic Course, ser. Applied Optimization. Springer, 2004, vol. 87

  43. [51]

    Long short-term memory,

    S. Hochreiter and J. Schmidhuber, “Long short-term memory,”Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997

  44. [52]

    Learning phrase representations using rnn encoder–decoder for statistical machine translation,

    K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y . Bengio, “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” inProceedings of the 2014 Conference on Empirical Methods in Natural Language Process...

  45. [53]

    Recurrent inference machines for solving inverse problems,

    P. Putzky and M. Welling, “Recurrent inference machines for solving inverse problems,” 2017

  46. [54]

    Dense recurrent neural networks for accelerated MRI: History- cognizant unrolling of optimization algorithms,

    S. A. H. Hosseini, B. Yaman, S. Moeller, M. Hong, and M. Ak- cakaya, “Dense recurrent neural networks for accelerated MRI: History- cognizant unrolling of optimization algorithms,”IEEE Journal of Se- lected Topics in Signal Processing, vol. 14, no. 6, pp. 1280–1291, 2020

  47. [55]

    Recurrent variational network: A deep learning inverse problem solver applied to the task of accelerated MRI reconstruction,

    G. Yiasemis, J.-J. Sonke, C. Sanchez, and J. Teuwen, “Recurrent variational network: A deep learning inverse problem solver applied to the task of accelerated MRI reconstruction,” in2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 2022, pp. 722–731

  48. [56]

    Memory-augmented deep unfolding network for compressive sensing,

    J. Song, B. Chen, and J. Zhang, “Memory-augmented deep unfolding network for compressive sensing,” inProceedings of the 29th ACM International Conference on Multimedia. ACM, 2021, pp. 4249–4258

  49. [57]

    Efficiently modeling long sequences with structured state spaces,

    A. Gu, K. Goel, and C. R ´e, “Efficiently modeling long sequences with structured state spaces,” 2022

  50. [58]

    Mamba: Linear-time sequence modeling with selective state spaces,

    A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” 2024

  51. [59]

    Vision mamba: Efficient visual representation learning with bidirectional state space model,

    L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, “Vision mamba: Efficient visual representation learning with bidirectional state space model,” 2024

  52. [60]

    Vmamba: Visual state space model,

    Y . Liu, Y . Tian, Y . Zhao, H. Yu, L. Xie, Y . Wang, Q. Ye, J. Jiao, and Y . Liu, “Vmamba: Visual state space model,” 2024

  53. [61]

    Distributed optimization and statistical learning via the alternating direction method of multipliers,

    S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, “Distributed optimization and statistical learning via the alternating direction method of multipliers,”Foundations and Trends in Machine Learning, vol. 3, no. 1, pp. 1–122, 2011

  54. [62]

    Infrared small target detection based on non-convex optimization with lp-norm constraint,

    T. Zhang, H. Wu, Y . Liu, L. Peng, C. Yang, and Z. Peng, “Infrared small target detection based on non-convex optimization with lp-norm constraint,”Remote Sensing, vol. 11, no. 5, p. 559, 2019

  55. [63]

    Sctransnet: Spatial- channel cross transformer network for infrared small target detection,

    S. Yuan, H. Qin, X. Yan, N. Akhtar, and A. Mian, “Sctransnet: Spatial- channel cross transformer network for infrared small target detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1– 15, 2024

  56. [64]

    Drpca-net: Make robust pca great again for infrared small target detection,

    Z. Xiong, F. Zhou, F. Wu, S. Yuan, M. Fu, Z. Peng, J. Yang, and Y . Dai, “Drpca-net: Make robust pca great again for infrared small target detection,”IEEE Transactions on Geoscience and Remote Sensing, 2025. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 15 APPENDIX...

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.