Pith. sign in

REVIEW 1 major objections 6 minor 116 references

Rapid Salient Object Detection with Difference Convolutional Neural Networks

T0 review · 1 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Image and video saliency can run in real time on embedded hardware if the network trains with contrast-encoding difference convolutions and then folds them losslessly into a single standard convolution at inference, yielding…

desk verdict A clean, honest engineering paper whose central reparameterization trick is exactly correct; the main risks are reproducibility and measurement detail, not correctness. read the letter →

arxiv 2507.01182 v1 pith:XD2IYJAL submitted 2025-07-01 cs.CV

classification cs.CV
keywords salientobjectdetectionreal-timemodelspixeldifferenceconvolutionreparameterizationspatiotemporallightweightnetworksembeddeddeployment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Salient object detection (SOD)—finding the most visually distinctive regions in an image or video—is a pre-attentive task that classical methods solved with hand-crafted contrast cues, but modern deep detectors pay for accuracy with heavy networks. This paper tries to get both: a sub-1M-parameter CNN that encodes the classic center-surround contrast idea through difference convolutions, and a reparameterization step that folds those contrast operators into ordinary convolutions so inference costs nothing extra. The authors claim the resulting SDNet and STDNet run at 46 FPS on an embedded GPU for images and 150 FPS for videos, two to three times faster than the best lightweight competitors with equal or better accuracy. If right, this makes real-time SOD practical as a front-end for downstream tasks on resource-limited devices, and it suggests contrast-aware training can be a free lunch for small vision networks.

What carries the argument

Pixel Difference Convolution (PDC) computes an inner product between learnable weights and pixel differences of selected pairs (central, angular, and radial patterns) rather than raw intensities, acting as a learned high-pass filter that encodes center-surround contrast. Difference Convolution Reparameterization (DCR) is the identity that makes PDC free at inference: each PDC branch $i$ with kernel $\theta_i$ is first converted to an equivalent standard convolution with kernel $\theta'_i$ (Equation 3), then all branches' outputs weighted by learned coefficients $\alpha_i$ are summed into one standard convolution with kernel $\theta' = \sum_i \alpha_i \theta'_i$ (Equation 4). SpatioTemporal Difference Convolution (STDC) extends PDC to video by slicing the 3D volume into H-T and W-T planes and applying central and angular difference patterns there; because those are still linear convolutions, DCR folds them away at inference, leaving a plain 3D standard convolution.

What would settle it

Compare SDNet against its own Baseline-Rep variant—an identical architecture whose every training branch is a standard convolution—across all six test datasets and several training seeds; if the PDC-trained folds do not consistently beat the standard-trained folds, the claim that difference operators, rather than reparameterization itself, cause the accuracy gain is falsified.

Watch

Extended reading notes

Core claim

The central claim is that contrast captures can be added to a CNN during training and then removed from the architecture at inference without losing their benefit. The paper's load-bearing identity, Equation (4), states that a weighted sum of standard convolution and multiple pixel difference convolutions (PDCs) equals a single standard convolution whose kernel is the coefficient-weighted sum of the individual kernels, $\theta' = \sum_i \alpha_i \theta'_i$, after each PDC is rewritten in standard form. Consequently the deployed SDNet and STDNet backbones are plain convolutional networks; the difference operators exist only during training, where they force the model to learn high-frequency contrast cues that survive in the folded kernel. For video, the same logic is extended to spatiotemporal difference convolutions (STDCs) on orthogonal W-T and H-T planes, so motion and appearance contrasts are learned without adding inference cost.

Load-bearing premise

The entire efficiency story rests on the premise that the contrast information the network learns during its multi-branch training phase stays inside the single folded kernel at inference; the paper checks this only on salient-object-detection datasets with one fixed training recipe.

Editorial extensions

If this is right

  • Because DCR collapses any number of PDC or STDC branches into one convolution, a system designer can add contrast-aware training branches to arbitrary layers of a small CNN with zero added inference parameters or FLOPs.
  • On the embedded GPU tested, the image model's 46 FPS and the video model's 150 FPS make salient object detection fast enough to run as a front-end for real-time downstream applications rather than a system bottleneck.
  • The ablation shows that reparameterizing standard-convolution branches alone yields no accuracy gain, so the reported benefit is specifically attributable to difference operators encoding contrast, not to over-parameterization per se.
  • For video, STDC-based temporal modules improve temporal consistency over spatial-only processing and outperform a lightweight temporal attention baseline on the evaluated benchmarks, suggesting contrast cues are a strong alternative to attention for motion-aware saliency.
  • Scaling up the STDNet backbone to a larger EfficientNet-B5 shifts the trade-off toward accuracy on DAVSOD, VOS, and DAVIS while keeping the model far faster than prior real-time methods, so the design leaves room for accuracy-first deployments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The DCR identity in Equation (4) is purely algebraic, so the same train-then-fold recipe could be applied to any dense prediction task where local contrast is informative—edge detection, tracking, remote sensing, or image enhancement—without altering the deployed architecture; this is an extension the paper only gestures at in its conclusion.
  • The paper's finding that video accuracy saturates at 8 input frames suggests the temporal receptive field of the STDM is the current bottleneck; a longer-range temporal contrast mechanism might push VSOD accuracy further, though it would have to keep the folded-inference trick to preserve the reported speed.
  • Because the folded kernel is mathematically identical to the trained multi-branch layer in exact arithmetic, any residual difference between the two in practice comes from floating-point rounding; comparing the two at inference would tell whether the apparent 'free lunch' is fully realized in finite precision.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 6 minor

Summary. The manuscript proposes SDNet and STDNet, lightweight CNNs for image and video salient object detection, built on Pixel Difference Convolutions (PDCs) and their video extension SpatioTemporal Difference Convolution (STDC). The central technical device is Difference Convolution Reparameterization (DCR): during training a layer contains several difference-convolution branches plus a standard convolution branch, and at inference the weighted branch kernels are algebraically collapsed into a single standard convolution kernel via Eq. (4), so the deployed model is exactly the same function as the trained multi-branch model with no extra parameters or FLOPs. Experiments on six image SOD benchmarks and three video SOD benchmarks, including measurements on an RTX 2080 Ti, Jetson AGX Orin, and Xavier NX, report sub-million-parameter models running at 46 FPS (image) and 150 FPS (video) on Orin with accuracy competitive with much larger models.

Significance. The core contribution is a clean, exact reparameterization identity: Eqs. (3)-(4) are a genuine algebraic equivalence, not an approximation, and the manuscript supports the empirical benefit with a well-designed ablation (Baseline vs. Baseline-Rep vs. SDNet in Table 6), where the failure of the standard-convolution-only Baseline-Rep to improve over Baseline actually strengthens the attribution of the gain to the PDC branches. The hardware efficiency claims are concrete and falsifiable: the authors report batch-size-1 FPS on two real devices and use a unified evaluation code for all accuracy metrics. The main limitation is scope: the DCR training benefit is demonstrated only for SOD, and the authors themselves note in Sec. 4.3 that simply reparameterizing standard convolutions gives no gain in this setting; generalizing the conclusion to other tasks requires additional evidence. I agree with the stress-test assessment that Eq. (4) is exact as written for the linear operators, and I do not see a circularity problem in the derivation.

major comments (1)
  1. [Section 3.2, Eqs. (3)-(4)] The DCR identity is stated for branches that are pure linear convolutions, but the backbone is described as PiDiNet-based, and standard PiDiNet blocks typically contain batch normalization. If the training-time branches include BN, the paper must show how the BN affine parameters are folded into theta'_i before the weighted sum; otherwise the claim that the inference-time model is literally the same function as the training-time model is under-specified. Please state explicitly whether BN is used in the DCR branches and, if so, give the folding equations and state where BN is placed after fusion.
minor comments (6)
  1. [General] There are several typos that should be corrected: "contribuions" in the Introduction, "unprecented" in the Introduction, and "Tempral attentions" in Table 7.
  2. [Section 4.1] The DAVIS validation set is described as "randomly choose 8 videos from its training set," but no seed or video list is given; providing this detail would make the reported DAVIS numbers reproducible.
  3. [Section 4.1] The sentence "we left pad the last remaining frames" should be "we left-pad the last remaining frames" or "we pad the last remaining frames on the left."
  4. [Section 3.2, Fig. 11] The manuscript should state explicitly whether the learned coefficients alpha_i are included in the reported parameter counts and whether they are trained with a softmax constraint in all layers; Fig. 11 suggests softmax is used, but the text does not specify this in Sec. 3.2.
  5. [Fig. 1] The caption says the circle size represents parameter count "illustrated separately for the two figures," meaning the circle scales are not directly comparable between the ISOD and VSOD panels; this should be acknowledged in the caption so readers are not misled.
  6. [Abstract and Introduction] The speed-up claims are stated as "more than 2x and 3x" in the abstract and "4~6 and 2~4 times faster" in the Introduction; these formulations should be harmonized so that the numbers are immediately consistent.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Eq. (4) is an exact algebraic equivalence, and the PDC training benefit is directly ablated in Table 6.

full rationale

The paper's central technical claim is an exact linear-algebraic equivalence, not a statistical prediction. Eq. (3) rewrites a PDC branch f(Δx, θ_i) = Σ_j w_{i,j}(x_j - x'_j) by redistributing each difference pair's weights onto the two involved pixels, giving an equivalent standard convolution with kernel θ'_i; Eq. (4) then uses linearity to absorb the learned branch coefficients α_i into a single kernel θ' = Σ_i α_i θ'_i. The inference-time single-branch network computes literally the same function as the training-time multi-branch network, so no fitted parameter is renamed as a prediction and no result is assumed in its own derivation. The empirical accuracy gain of using PDC branches is directly tested in Table 6, where Baseline-Rep (standard-convolution branches reparameterized the same way) fails to improve over Baseline while SDNet improves, showing the gain is attributed to PDC content rather than to the reparameterization procedure itself. The paper does cite the authors' prior PiDiNet work for PDC definitions, the 3×3 RPDC-to-5×5 conversion, and the CDCM/CSAM modules, but the load-bearing DCR derivation is reproduced in full in Eqs. (3)-(4) and does not depend on those citations for its validity. Remaining concerns (single-run timing, code availability, transferability to other tasks) are verification and generality issues, not circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 1 invented entities

The paper introduces one new operator (STDC) and a reparameterization trick, both grounded in prior work. The free parameters are all learned or chosen via ablations; there are no unexplained physical constants. The central claim rests on the empirical validity of the SOD benchmarks and the assumption that the contrast prior helps.

free parameters (4)
  • branch coefficients alpha_i = learned per layer
    The weights given to each PDC branch and the standard convolution branch in each layer are learned during training. They are part of the model parameters, not free hyper-parameters, but they are not derived from theory; they are fitted to the SOD data.
  • number of frames (8) in video input = 8
    The video input clip length is set to 8 frames, and the ablation (Tab. 7) shows that 8 is better than 2, 4, or 16. This is a hyper-parameter chosen by experiment.
  • input resolution (320x320 for ISOD, 256x256 for VSOD) = 320x320/256x256
    The input size is a design choice, and the paper uses a single resolution for comparison, but it is not a theoretical constant; it affects FLOPs and accuracy.
  • loss weighting (L = Lbce + LIoU + Lssim) = equal weights of 1
    The loss is a sum of three losses with equal weight, which is a design choice not justified by derivation.
assumptions (4)
  • standard math The reparameterization identity in Eq. (3) holds for any pixel pair selection strategy.
    This is a simple algebraic manipulation: a linear combination of pixel differences can be rewritten as a linear combination of pixel intensities with adjusted weights. The paper states this without formal proof, but it is easily verified.
  • domain assumption PDC acts as a high-pass filter, and standard convolution preserves low frequencies.
    This is a standard observation in the frequency domain, cited from the authors' prior work [43]. The paper relies on it to justify combining PDC and standard convolution.
  • domain assumption Salient objects are characterized by high-frequency spatial contrast and center-surround interactions.
    This is the classical prior for SOD, backed by biological findings in the introduction, but it is an assumption about the data, not a mathematical theorem.
  • ad hoc to paper Slicing the 3D volume into H-T and W-T planes preserves enough spatiotemporal information for VSOD.
    This is a design choice in the STDC module (Section 3.3), similar to LBP-TOP [44]; the paper does not derive optimality, and the ablation only compares it to a few alternatives, not to full 3D operators.
invented entities (1)
  • SpatioTemporal Difference Convolution (STDC) independent evidence
    purpose: To capture contrast in the temporal dimension by applying PDC-like operators on H-T and W-T planes.
    The paper provides empirical evidence via ablations (Tab. 7) that STDC improves VSOD performance compared to standard convolution. The concept is not a physical entity, but a new computational operator; it has a falsifiable handle in the form of accuracy improvements on standard benchmarks.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Rapid Salient Object Detection with Difference Convolutional Neural Networks." pith.science (2026). https://pith.science/paper/XD2IYJAL

@misc{pith2026250701182,
  author       = {Pith},
  title        = {Pith review of: Rapid Salient Object Detection with Difference Convolutional Neural Networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XD2IYJAL}},
  note         = {Machine review of arXiv:2507.01182}
}
abstract

This paper addresses the challenge of deploying salient object detection (SOD) on resource-constrained devices with real-time performance. While recent advances in deep neural networks have improved SOD, existing top-leading models are computationally expensive. We propose an efficient network design that combines traditional wisdom on SOD and the representation power of modern CNNs. Like biologically-inspired classical SOD methods relying on computing contrast cues to determine saliency of image regions, our model leverages Pixel Difference Convolutions (PDCs) to encode the feature contrasts. Differently, PDCs are incorporated in a CNN architecture so that the valuable contrast cues are extracted from rich feature maps. For efficiency, we introduce a difference convolution reparameterization (DCR) strategy that embeds PDCs into standard convolutions, eliminating computation and parameters at inference. Additionally, we introduce SpatioTemporal Difference Convolution (STDC) for video SOD, enhancing the standard 3D convolution with spatiotemporal contrast capture. Our models, SDNet for image SOD and STDNet for video SOD, achieve significant improvements in efficiency-accuracy trade-offs. On a Jetson Orin device, our models with $<$ 1M parameters operate at 46 FPS and 150 FPS on streamed images and videos, surpassing the second-best lightweight models in our experiments by more than $2\times$ and $3\times$ in speed with superior accuracy. Code will be available at https://github.com/hellozhuo/stdnet.git.

Figures

Figures reproduced from arXiv: 2507.01182 by the authors.

Figure 1
Figure 1. Comparing accuracy-runtime trade-offs for different [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. (a) Formulation for selecting pixel pairs in PDC; (b-d): [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Architecture overview. (a) PiDiNet backbone [ [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Our proposed DCR pipeline. In this example, we employ three different PDC operators and a standard convolutional operator. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: We visualize the intermediate feature maps in layer 4 of [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Illustration of the proposed STDC in H-T and W-T planes. [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Illustration of a STDC layer with DCR that can be executed [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: The proposed STDNet architecture. STDC (W-T) and STDC (H-T) are implemented following Fig. [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Qualitative comparison on ISOD. The first two columns are the input images and corresponding ground truth images [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Qualitative comparison on VSOD. The first two columns [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]
Figure 11
Figure 11. Figure 11: For both figures, each column represents a certain layer [PITH_FULL_IMAGE:figures/full_fig_p012_11.png]
Figure 13
Figure 13. Figure 13: Visualization of spatiotemporal feature maps by different [PITH_FULL_IMAGE:figures/full_fig_p013_13.png]
Figure 12
Figure 12. Figure 12: Predicted saliency maps w/ and w/o temporal modules. [PITH_FULL_IMAGE:figures/full_fig_p013_12.png]
Figure 14
Figure 14. Figure 14: Temporal attentions. Features of each frame is regarded [PITH_FULL_IMAGE:figures/full_fig_p014_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

116 extracted references · 77 canonical work pages

  1. [43]

    Lightweight pixel difference networks for efficient visual representation learning,

    Z. Su, J. Zhang, L. Wang, H. Zhang, Z. Liu, M. Pietik ¨ainen, and L. Liu, “Lightweight pixel difference networks for efficient visual representation learning,” IEEE Trans. Pattern Anal. Mach. Intell. , pp. 1–18, 2023

  2. [1]

    A model of saliency-based visual attention for rapid scene analysis,

    L. Itti, C. Koch, and E. Niebur, “A model of saliency-based visual attention for rapid scene analysis,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 20, no. 11, pp. 1254–1259, 1998

  3. [2]

    Learning to detect a salient object,

    T. Liu, Z. Yuan, J. Sun, J. Wang, N. Zheng, X. Tang, and H.-Y. Shum, “Learning to detect a salient object,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 33, no. 2, pp. 353–367, 2010

  4. [3]

    Attentive and pre-attentive aspects of figural processing,

    L. G. Appelbaum and A. M. Norcia, “Attentive and pre-attentive aspects of figural processing,” J. Vis., vol. 9, no. 11, pp. 18–18, 2009

  5. [4]

    Salient object detection: A survey,

    A. Borji, M.-M. Cheng, Q. Hou, H. Jiang, and J. Li, “Salient object detection: A survey,” Computational Visual Media , vol. 5, pp. 117–150, 2019

  6. [5]

    SALISA: Saliency-based input sampling for efficient video object detection,

    B. E. Bejnordi, A. Habibian, F. Porikli, and A. Ghodrati, “SALISA: Saliency-based input sampling for efficient video object detection,” in Eur. Conf. Comput. Vis., 2022

  7. [6]

    Is bottom-up attention useful for object recognition?

    U. Rutishauser, D. Walther, C. Koch, and P . Perona, “Is bottom-up attention useful for object recognition?” in IEEE Conf. Comput. Vis. Pattern Recog., 2004

  8. [7]

    Region-based saliency detection and its application in object recognition,

    Z. Ren, S. Gao, L.-T. Chia, and I. W.-H. Tsang, “Region-based saliency detection and its application in object recognition,” IEEE Trans. Circuits Syst. Video Technol., vol. 24, no. 5, pp. 769–779, 2013

Show all 116 references
  1. [8]

    Highly efficient and unsupervised framework for moving object detection in satellite videos,

    C. Xiao, W. An, Y. Zhang, Z. Su, M. Li, W. Sheng, M. Pietik ¨ainen, and L. Liu, “Highly efficient and unsupervised framework for moving object detection in satellite videos,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 12, pp. 11 532–11 539, 2024

  2. [9]

    Boosting low-data instance segmentation by unsupervised pre-training with saliency prompt,

    H. Li, D. Zhang, N. Liu, L. Cheng, Y. Dai, C. Zhang, X. Wang, and J. Han, “Boosting low-data instance segmentation by unsupervised pre-training with saliency prompt,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023

  3. [10]

    Large- scale unsupervised semantic segmentation,

    S. Gao, Z.-Y. Li, M.-H. Yang, M.-M. Cheng, J. Han, and P . Torr, “Large- scale unsupervised semantic segmentation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 6, pp. 7457–7476, 2022

  4. [11]

    Full-duplex strategy for video object segmentation,

    G.-P . Ji, K. Fu, Z. Wu, D.-P . Fan, J. Shen, and L. Shao, “Full-duplex strategy for video object segmentation,” in Int. Conf. Comput. Vis., 2021

  5. [12]

    Online tracking by learning discriminative saliency map with convolutional neural network,

    S. Hong, T. You, S. Kwak, and B. Han, “Online tracking by learning discriminative saliency map with convolutional neural network,” in ICML, 2015

  6. [13]

    Weighted attentional blocks for probabilistic object tracking,

    H. Wu, G. Li, and X. Luo, “Weighted attentional blocks for probabilistic object tracking,” Vis. Comput., vol. 30, pp. 229–243, 2014

  7. [14]

    Mobile product search with bag of hash bits and boundary reranking,

    J. He, J. Feng, X. Liu, T. Cheng, T.-H. Lin, H. Chung, and S.-F. Chang, “Mobile product search with bag of hash bits and boundary reranking,” in IEEE Conf. Comput. Vis. Pattern Recog., 2012

  8. [15]

    SalientShape: group saliency in image collections,

    M.-M. Cheng, N. J. Mitra, X. Huang, and S.-M. Hu, “SalientShape: group saliency in image collections,” Vis. Comput., vol. 30, no. 4, pp. 443–453, 2014

  9. [16]

    Saliency driven perceptual image compression,

    Y. Patel, S. Appalaraju, and R. Manmatha, “Saliency driven perceptual image compression,” in WACV, 2021

  10. [17]

    Automatic foveation for video compression using a neurobiological model of visual attention,

    L. Itti, “Automatic foveation for video compression using a neurobiological model of visual attention,” IEEE Trans. Image Process., vol. 13, no. 10, pp. 1304–1318, 2004

  11. [19]

    Deep saliency prior for reducing visual distraction,

    K. Aberman, J. He, Y. Gandelsman, I. Mosseri, D. E. Jacobs, K. Kohlhoff, Y. Pritch, and M. Rubinstein, “Deep saliency prior for reducing visual distraction,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022

  12. [20]

    Realistic saliency guided image enhancement,

    S. M. H. Miangoleh, Z. Bylinskii, E. Kee, E. Shechtman, and Y. Aksoy, “Realistic saliency guided image enhancement,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023

  13. [21]

    SaliencyMix: a saliency guided data augmentation strategy for better regularization,

    A. S. Uddin, M. S. Monira, W. Shin, T. Chung, and S.-H. Bae, “SaliencyMix: a saliency guided data augmentation strategy for better regularization,” in Int. Conf. Learn. Represent., 2021

  14. [22]

    Quantitative analysis of human- model agreement in visual saliency modeling: A comparative study,

    A. Borji, D. N. Sihite, and L. Itti, “Quantitative analysis of human- model agreement in visual saliency modeling: A comparative study,” IEEE Trans. Image Process., vol. 22, no. 1, pp. 55–69, 2012

  15. [23]

    Mesh saliency,

    C. H. Lee, A. Varshney, and D. W. Jacobs, “Mesh saliency,” in ACM SIGGRAPH, 2005

  16. [24]

    SMU- Net: saliency-guided morphology-aware u-net for breast lesion segmentation in ultrasound image,

    Z. Ning, S. Zhong, Q. Feng, W. Chen, and Y. Zhang, “SMU- Net: saliency-guided morphology-aware u-net for breast lesion segmentation in ultrasound image,” IEEE Trans. Med. Imaging, vol. 41, no. 2, pp. 476–490, 2021

  17. [25]

    Salient object detection in optical remote sensing images driven by transformer,

    G. Li, Z. Bai, Z. Liu, X. Zhang, and H. Ling, “Salient object detection in optical remote sensing images driven by transformer,” IEEE Trans. Image Process., 2023

  18. [26]

    Lightweight salient object detection in optical remote-sensing images via semantic matching and edge alignment,

    G. Li, Z. Liu, X. Zhang, and W. Lin, “Lightweight salient object detection in optical remote-sensing images via semantic matching and edge alignment,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–11, 2023

  19. [27]

    A shape-based approach for salient object detection using deep learning,

    J. Kim and V . Pavlovic, “A shape-based approach for salient object detection using deep learning,” in Eur. Conf. Comput. Vis., 2016. ACCEPTED IN IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 15

  20. [28]

    Salient object detection via fast r-cnn and low-level cues,

    X. Wang, H. Ma, and X. Chen, “Salient object detection via fast r-cnn and low-level cues,” in IEEE Int. Conf. Image Process., 2016

  21. [29]

    Visual saliency transformer,

    N. Liu, N. Zhang, K. Wan, L. Shao, and J. Han, “Visual saliency transformer,” in Int. Conf. Comput. Vis., 2021

  22. [30]

    Salient object detection via integrity learning,

    M. Zhuge, D.-P . Fan, N. Liu, D. Zhang, D. Xu, and L. Shao, “Salient object detection via integrity learning,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 3, pp. 3738–3752, 2023

  23. [31]

    SAMNet: Stereoscopically attentive multi-scale network for lightweight salient object detection,

    Y. Liu, X.-Y. Zhang, J.-W. Bian, L. Zhang, and M.-M. Cheng, “SAMNet: Stereoscopically attentive multi-scale network for lightweight salient object detection,” IEEE Trans. Image Process. , vol. 30, pp. 3804–3814, 2021

  24. [32]

    Lightweight salient object detection via hierarchical visual perception learning,

    Y. Liu, Y.-C. Gu, X.-Y. Zhang, W. Wang, and M.-M. Cheng, “Lightweight salient object detection via hierarchical visual perception learning,” IEEE Trans. Cybern., vol. 51, no. 9, pp. 4439–4449, 2020

  25. [33]

    PoolNet+: Exploring the potential of pooling for salient object detection,

    J.-J. Liu, Q. Hou, Z.-A. Liu, and M.-M. Cheng, “PoolNet+: Exploring the potential of pooling for salient object detection,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 1, pp. 887–904, 2022

  26. [34]

    EDN: Salient object detection via extremely-downsampled network,

    Y.-H. Wu, Y. Liu, L. Zhang, M.-M. Cheng, and B. Ren, “EDN: Salient object detection via extremely-downsampled network,” IEEE Trans. Image Process., vol. 31, pp. 3125–3136, 2022

  27. [35]

    A highly efficient model to study the semantics of salient object detection,

    M.-M. Cheng, S. Gao, A. Borji, Y.-Q. Tan, Z. Lin, and M. Wang, “A highly efficient model to study the semantics of salient object detection,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 11, pp. 8006–8021, 2021

  28. [36]

    Contrast-based image attention analysis by using fuzzy growing,

    Y.-F. Ma and H.-J. Zhang, “Contrast-based image attention analysis by using fuzzy growing,” in ACM Int. Conf. Multimedia, 2003

  29. [37]

    Saliency filters: Contrast based filtering for salient region detection,

    F. Perazzi, P . Kr ¨ahenb ¨uhl, Y. Pritch, and A. Hornung, “Saliency filters: Contrast based filtering for salient region detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2012

  30. [38]

    Salient region detection by ufo: Uniqueness, focusness and objectness,

    P . Jiang, H. Ling, J. Yu, and J. Peng, “Salient region detection by ufo: Uniqueness, focusness and objectness,” in Int. Conf. Comput. Vis., 2013

  31. [39]

    Efficient salient region detection with soft image abstraction,

    M.-M. Cheng, J. Warrell, W.-Y. Lin, S. Zheng, V . Vineet, and N. Crook, “Efficient salient region detection with soft image abstraction,” in Int. Conf. Comput. Vis., 2013

  32. [40]

    Spatial attention modulates center-surround interactions in macaque visual area v4,

    K. A. Sundberg, J. F. Mitchell, and J. H. Reynolds, “Spatial attention modulates center-surround interactions in macaque visual area v4,” Neuron, vol. 61, no. 6, pp. 952–963, 2009

  33. [41]

    The neural basis of visual function: Vision and visual dysfunction,

    V . Casagrande, T. Norton, and A. Leventhal, “The neural basis of visual function: Vision and visual dysfunction,” 1991

  34. [42]

    Pixel difference networks for efficient edge detection,

    Z. Su, W. Liu, Z. Yu, D. Hu, Q. Liao, Q. Tian, M. Pietik ¨ainen, and L. Liu, “Pixel difference networks for efficient edge detection,” in Int. Conf. Comput. Vis., 2021

  35. [44]

    Dynamic texture recognition using local binary patterns with an application to facial expressions,

    G. Zhao and M. Pietikainen, “Dynamic texture recognition using local binary patterns with an application to facial expressions,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 29, no. 6, pp. 915–928, 2007

  36. [45]

    Exploring rich and efficient spatial temporal interactions for real-time video salient object detection,

    C. Chen, G. Wang, C. Peng, Y. Fang, D. Zhang, and H. Qin, “Exploring rich and efficient spatial temporal interactions for real-time video salient object detection,” IEEE Trans. Image Process., vol. 30, pp. 3995– 4007, 2021

  37. [46]

    Dynamic context-sensitive filtering network for video salient object detection,

    M. Zhang, J. Liu, Y. Wang, Y. Piao, S. Yao, W. Ji, J. Li, H. Lu, and Z. Luo, “Dynamic context-sensitive filtering network for video salient object detection,” in Int. Conf. Comput. Vis., 2021

  38. [47]

    Mobilesal: Extremely efficient RGB-D salient object detection,

    Y.-H. Wu, Y. Liu, J. Xu, J.-W. Bian, Y.-C. Gu, and M.-M. Cheng, “Mobilesal: Extremely efficient RGB-D salient object detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 12, pp. 10 261–10 269, 2022

  39. [48]

    Depth quality- inspired feature manipulation for efficient RGB-D salient object detection,

    W. Zhang, G.-P . Ji, Z. Wang, K. Fu, and Q. Zhao, “Depth quality- inspired feature manipulation for efficient RGB-D salient object detection,” in ACM Int. Conf. Multimedia, 2021

  40. [49]

    A simple pooling- based design for real-time salient object detection,

    J.-J. Liu, Q. Hou, M.-M. Cheng, J. Feng, and J. Jiang, “A simple pooling- based design for real-time salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019

  41. [50]

    EGNet: Edge guidance network for salient object detection,

    J.-X. Zhao, J.-J. Liu, D.-P . Fan, Y. Cao, J. Yang, and M.-M. Cheng, “EGNet: Edge guidance network for salient object detection,” in Int. Conf. Comput. Vis., 2019

  42. [51]

    Amulet: Aggregating multi-level convolutional features for salient object detection,

    P . Zhang, D. Wang, H. Lu, H. Wang, and X. Ruan, “Amulet: Aggregating multi-level convolutional features for salient object detection,” in Int. Conf. Comput. Vis., 2017

  43. [52]

    Detect globally, refine locally: A novel approach to saliency detection,

    T. Wang, L. Zhang, S. Wang, H. Lu, G. Yang, X. Ruan, and A. Borji, “Detect globally, refine locally: A novel approach to saliency detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018

  44. [53]

    Salient object detection with pyramid attention and salient edges,

    W. Wang, S. Zhao, J. Shen, S. C. Hoi, and A. Borji, “Salient object detection with pyramid attention and salient edges,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019

  45. [54]

    MobileNetV2: Inverted residuals and linear bottlenecks,

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018

  46. [55]

    RepVGG: Making vgg-style convnets great again,

    X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, and J. Sun, “RepVGG: Making vgg-style convnets great again,” in IEEE Conf. Comput. Vis. Pattern Recog., 2021

  47. [56]

    Semi- supervised video salient object detection using pseudo-labels,

    P . Yan, G. Li, Y. Xie, Z. Li, C. Wang, T. Chen, and L. Lin, “Semi- supervised video salient object detection using pseudo-labels,” in Int. Conf. Comput. Vis., 2019

  48. [57]

    Shifting more attention to video salient object detection,

    D.-P . Fan, W. Wang, M.-M. Cheng, and J. Shen, “Shifting more attention to video salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019

  49. [58]

    Flow guided recurrent neural encoder for video salient object detection,

    G. Li, Y. Xie, T. Wei, K. Wang, and L. Lin, “Flow guided recurrent neural encoder for video salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018

  50. [59]

    Pyramid dilated deeper convlstm for video salient object detection,

    H. Song, W. Wang, S. Zhao, J. Shen, and K.-M. Lam, “Pyramid dilated deeper convlstm for video salient object detection,” in Eur. Conf. Comput. Vis., 2018

  51. [60]

    Deeply supervised 3d recurrent fcn for salient object detection in videos

    T.-N. Le and A. Sugimoto, “Deeply supervised 3d recurrent fcn for salient object detection in videos.” in Brit. Mach. Vis. Conf., 2017

  52. [61]

    A novel long-term iterative mining scheme for video salient object detection,

    C. Chen, H. Wang, Y. Fang, and C. Peng, “A novel long-term iterative mining scheme for video salient object detection,” IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 11, pp. 7662–7676, 2022

  53. [62]

    Depth-cooperated trimodal network for video salient object detection,

    Y. Lu, D. Min, K. Fu, and Q. Zhao, “Depth-cooperated trimodal network for video salient object detection,” in IEEE Int. Conf. Image Process., 2022

  54. [63]

    Run, don’t walk: chasing higher FLOPS for faster neural networks,

    J. Chen, S.-h. Kao, H. He, W. Zhuo, S. Wen, C.-H. Lee, and S.-H. G. Chan, “Run, don’t walk: chasing higher FLOPS for faster neural networks,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023

  55. [64]

    Lightweight vision transformer with bidirectional interaction,

    Q. Fan, H. Huang, X. Zhou, and R. He, “Lightweight vision transformer with bidirectional interaction,” in Adv. Neural Inform. Process. Syst., 2024

  56. [65]

    Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction,

    H. Cai, J. Li, M. Hu, C. Gan, and S. Han, “Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction,” inInt. Conf. Comput. Vis., 2023

  57. [66]

    EfficientFormer: Vision transformers at mobilenet speed,

    Y. Li, G. Yuan, Y. Wen, E. Hu, G. Evangelidis, S. Tulyakov, Y. Wang, and J. Ren, “EfficientFormer: Vision transformers at mobilenet speed,” in Adv. Neural Inform. Process. Syst., 2022

  58. [67]

    Dynamic group convolution for accelerating convolutional neural networks,

    Z. Su, L. Fang, W. Kang, D. Hu, M. Pietik ¨ainen, and L. Liu, “Dynamic group convolution for accelerating convolutional neural networks,” in Eur. Conf. Comput. Vis., 2020

  59. [68]

    Structured pruning for deep convolutional neural networks: A survey,

    Y. He and L. Xiao, “Structured pruning for deep convolutional neural networks: A survey,” IEEE Trans. Pattern Anal. Mach. Intell., 2023

  60. [69]

    Llm-pruner: On the structural pruning of large language models,

    X. Ma, G. Fang, and X. Wang, “Llm-pruner: On the structural pruning of large language models,” in Adv. Neural Inform. Process. Syst., 2023

  61. [70]

    Boosting convolutional neural networks with middle spectrum grouped convolution,

    Z. Su, J. Zhang, T. Liu, Z. Liu, S. Zhang, M. Pietik ¨ainen, and L. Liu, “Boosting convolutional neural networks with middle spectrum grouped convolution,” IEEE Trans. Neural Netw. Learn. Syst., 2024

  62. [71]

    Knowledge distillation: A survey,

    J. Gou, B. Yu, S. J. Maybank, and D. Tao, “Knowledge distillation: A survey,” Int. J. Comput. Vis., vol. 129, no. 6, pp. 1789–1819, 2021

  63. [72]

    One-for- all: Bridge the gap between heterogeneous architectures in knowledge distillation,

    Z. Hao, J. Guo, K. Han, Y. Tang, H. Hu, Y. Wang, and C. Xu, “One-for- all: Bridge the gap between heterogeneous architectures in knowledge distillation,” in Adv. Neural Inform. Process. Syst., 2024

  64. [73]

    Curriculum temperature for knowledge distillation,

    Z. Li, X. Li, L. Yang, B. Zhao, R. Song, L. Luo, J. Li, and J. Yang, “Curriculum temperature for knowledge distillation,” in AAAI, 2023

  65. [74]

    Dynamic binary neural network by learning channel-wise thresholds,

    J. Zhang, Z. Su, Y. Feng, X. Lu, M. Pietik ¨ainen, and L. Liu, “Dynamic binary neural network by learning channel-wise thresholds,” in ICASSP, 2022

  66. [75]

    SVNet: Where so (3) equivariance meets binarization on point cloud representation,

    Z. Su, M. Welling, L. Liu et al. , “SVNet: Where so (3) equivariance meets binarization on point cloud representation,” in 3DV, 2022

  67. [76]

    Q-dm: An efficient low-bit quantized diffusion model,

    Y. Li, S. Xu, X. Cao, X. Sun, and B. Zhang, “Q-dm: An efficient low-bit quantized diffusion model,” in Adv. Neural Inform. Process. Syst., 2024

  68. [77]

    Oscillation-free quantization for low-bit vision transformers,

    S.-Y. Liu, Z. Liu, and K.-T. Cheng, “Oscillation-free quantization for low-bit vision transformers,” in ICML, 2023

  69. [78]

    Smoothquant: Accurate and efficient post-training quantization for large language models,

    G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in ICML, 2023

  70. [79]

    Median pixel difference convolutional network for robust face recognition,

    J. Zhang, Z. Su, and L. Liu, “Median pixel difference convolutional network for robust face recognition,” in Brit. Mach. Vis. Conf., 2022

  71. [80]

    Multiresolution gray-scale and rotation invariant texture classification with local binary patterns,

    T. Ojala, M. Pietikainen, and T. Maenpaa, “Multiresolution gray-scale and rotation invariant texture classification with local binary patterns,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 24, no. 7, pp. 971–987, 2002

  72. [81]

    Frequency- tuned salient region detection,

    R. Achanta, S. Hemami, F. Estrada, and S. Susstrunk, “Frequency- tuned salient region detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2009

  73. [82]

    Natural scene statistics at the centre of gaze,

    P . Reinagel and A. M. Zador, “Natural scene statistics at the centre of gaze,” Network: Computation in Neural Systems , vol. 10, no. 4, p. 341, 1999

  74. [83]

    Separable self-attention for mobile vision transformers,

    S. Mehta and M. Rastegari, “Separable self-attention for mobile vision transformers,” Trans. Mach. Learn. Res., 2023

  75. [84]

    ExpandNets: Linear over- parameterization to train compact convolutional networks,

    S. Guo, J. M. Alvarez, and M. Salzmann, “ExpandNets: Linear over- parameterization to train compact convolutional networks,” in Adv. Neural Inform. Process. Syst., 2020

  76. [85]

    Deeply supervised salient object detection with short connections,

    Q. Hou, M.-M. Cheng, X. Hu, A. Borji, Z. Tu, and P . Torr, “Deeply supervised salient object detection with short connections,”IEEE Trans. ACCEPTED IN IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 16 Pattern Anal. Mach. Intell., vol. 41, no. 4, pp. 815–828, 2019

  77. [86]

    ESPNetv2: A light-weight, power efficient, and general purpose convolutional neural network,

    S. Mehta, M. Rastegari, L. Shapiro, and H. Hajishirzi, “ESPNetv2: A light-weight, power efficient, and general purpose convolutional neural network,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019

  78. [87]

    BiSeNet v2: Bilateral network with guided aggregation for real-time semantic segmentation,

    C. Yu, C. Gao, J. Wang, G. Yu, C. Shen, and N. Sang, “BiSeNet v2: Bilateral network with guided aggregation for real-time semantic segmentation,” Int. J. Comput. Vis., vol. 129, no. 11, pp. 3051–3068, 2021

  79. [88]

    ENet: A deep neural network architecture for real-time semantic segmentation,

    A. Paszke, A. Chaurasia, S. Kim, and E. Culurciello, “ENet: A deep neural network architecture for real-time semantic segmentation,” arXiv preprint arXiv:1606.02147, 2016

  80. [89]

    DABNet: Depth-wise asymmetric bottleneck for real-time semantic segmentation,

    G. Li and J. Kim, “DABNet: Depth-wise asymmetric bottleneck for real-time semantic segmentation,” in Brit. Mach. Vis. Conf., 2019

  81. [90]

    EdgeNeXt: efficiently amalgamated cnn-transformer architecture for mobile vision applications,

    M. Maaz, A. Shaker, H. Cholakkal, S. Khan, S. W. Zamir, R. M. Anwer, and F. S. Khan, “EdgeNeXt: efficiently amalgamated cnn-transformer architecture for mobile vision applications,” in Eur. Conf. Comput. Vis., 2022

  82. [91]

    MobileViT: Light-weight, general- purpose, and mobile-friendly vision transformer,

    S. Mehta and M. Rastegari, “MobileViT: Light-weight, general- purpose, and mobile-friendly vision transformer,” in Int. Conf. Learn. Represent., 2022

  83. [92]

    Learning uncertain convolutional features for accurate saliency detection,

    P . Zhang, D. Wang, H. Lu, H. Wang, and B. Yin, “Learning uncertain convolutional features for accurate saliency detection,” in Int. Conf. Comput. Vis., 2017

  84. [93]

    A stagewise refinement model for detecting salient objects in images,

    T. Wang, A. Borji, L. Zhang, P . Zhang, and H. Lu, “A stagewise refinement model for detecting salient objects in images,” in Int. Conf. Comput. Vis., 2017

  85. [94]

    Learning to detect salient objects with image-level supervision,

    L. Wang, H. Lu, Y. Wang, M. Feng, D. Wang, B. Yin, and X. Ruan, “Learning to detect salient objects with image-level supervision,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017

  86. [95]

    Hierarchical saliency detection,

    Q. Yan, L. Xu, J. Shi, and J. Jia, “Hierarchical saliency detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2013

  87. [96]

    The secrets of salient object segmentation,

    Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille, “The secrets of salient object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2014

  88. [97]

    Saliency detection via graph-based manifold ranking,

    C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang, “Saliency detection via graph-based manifold ranking,” in IEEE Conf. Comput. Vis. Pattern Recog., 2013

  89. [98]

    Design and perceptual validation of performance measures for salient object segmentation,

    V . Movahedi and J. H. Elder, “Design and perceptual validation of performance measures for salient object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2010

  90. [99]

    Visual saliency based on multiscale deep features,

    G. Li and Y. Yu, “Visual saliency based on multiscale deep features,” in IEEE Conf. Comput. Vis. Pattern Recog., 2015

  91. [100]

    A benchmark dataset and saliency-guided stacked autoencoders for video-based salient object detection,

    J. Li, C. Xia, and X. Chen, “A benchmark dataset and saliency-guided stacked autoencoders for video-based salient object detection,” IEEE Trans. Image Process., vol. 27, no. 1, pp. 349–364, 2018

  92. [101]

    A benchmark dataset and evaluation methodology for video object segmentation,

    F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “A benchmark dataset and evaluation methodology for video object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016

  93. [102]

    Adam: A method for stochastic optimization,

    D. P . Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Int. Conf. Learn. Represent., 2015

  94. [103]

    PyTorch: An imperative style, high- performance deep learning library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: An imperative style, high-...

  95. [104]

    Richer convolutional features for edge detection,

    Y. Liu, M.-M. Cheng, X. Hu, J.-W. Bian, L. Zhang, X. Bai, and J. Tang, “Richer convolutional features for edge detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, no. 8, pp. 1939–1946, 2019

  96. [105]

    ImageNet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in IEEE Conf. Comput. Vis. Pattern Recog., 2009

  97. [106]

    UnitBox: An advanced object detection network,

    J. Yu, Y. Jiang, Z. Wang, Z. Cao, and T. Huang, “UnitBox: An advanced object detection network,” in ACM Int. Conf. Multimedia, 2016

  98. [107]

    Image quality assessment: from error visibility to structural similarity,

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P . Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, 2004

  99. [108]

    Structure-measure: A new way to evaluate foreground maps,

    D.-P . Fan, M.-M. Cheng, Y. Liu, T. Li, and A. Borji, “Structure-measure: A new way to evaluate foreground maps,” in Int. Conf. Comput. Vis. , 2017

  100. [109]

    Salient object detection in the deep learning era: An in-depth survey,

    W. Wang, Q. Lai, H. Fu, J. Shen, H. Ling, and R. Yang, “Salient object detection in the deep learning era: An in-depth survey,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 6, pp. 3239–3259, 2021

  101. [110]

    Real-time salient object detection with a minimum spanning tree,

    W.-C. Tu, S. He, Q. Yang, and S.-Y. Chien, “Real-time salient object detection with a minimum spanning tree,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016

  102. [111]

    SCOM: Spatiotemporal constrained optimization for salient object detection,

    Y. Chen, W. Zou, Y. Tang, X. Li, C. Xu, and N. Komodakis, “SCOM: Spatiotemporal constrained optimization for salient object detection,” IEEE Trans. Image Process., vol. 27, no. 7, pp. 3345–3357, 2018

  103. [112]

    Motion guided attention for video salient object detection,

    H. Li, G. Chen, G. Li, and Y. Yu, “Motion guided attention for video salient object detection,” in Int. Conf. Comput. Vis., 2019

  104. [113]

    Multi-stream attention-aware graph convolution network for video salient object detection,

    M. Xu, P . Fu, B. Liu, and J. Li, “Multi-stream attention-aware graph convolution network for video salient object detection,” IEEE Trans. Image Process., vol. 30, pp. 4183–4197, 2021

  105. [114]

    Progressively real-time video salient object detection via cascaded fully convolutional networks with motion attention,

    Q. Zheng, Y. Li, L. Zheng, and Q. Shen, “Progressively real-time video salient object detection via cascaded fully convolutional networks with motion attention,” Neurocomputing, vol. 467, pp. 465–475, 2022

  106. [115]

    Pyramid constrained self-attention network for fast video salient object detection,

    Y. Gu, L. Wang, Z. Wang, Y. Liu, M.-M. Cheng, and S.-P . Lu, “Pyramid constrained self-attention network for fast video salient object detection,” in AAAI, 2020

  107. [116]

    EfficientNet: Rethinking model scaling for convolutional neural networks,

    M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in ICML, 2019

  108. [117]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inform. Process. Syst., 2017

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.