REVIEW 1 major objections 6 minor 116 references
Rapid Salient Object Detection with Difference Convolutional Neural Networks
T0 review · 1 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Image and video saliency can run in real time on embedded hardware if the network trains with contrast-encoding difference convolutions and then folds them losslessly into a single standard convolution at inference, yielding…
desk verdict A clean, honest engineering paper whose central reparameterization trick is exactly correct; the main risks are reproducibility and measurement detail, not correctness. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Pixel Difference Convolution (PDC) computes an inner product between learnable weights and pixel differences of selected pairs (central, angular, and radial patterns) rather than raw intensities, acting as a learned high-pass filter that encodes center-surround contrast. Difference Convolution Reparameterization (DCR) is the identity that makes PDC free at inference: each PDC branch $i$ with kernel $\theta_i$ is first converted to an equivalent standard convolution with kernel $\theta'_i$ (Equation 3), then all branches' outputs weighted by learned coefficients $\alpha_i$ are summed into one standard convolution with kernel $\theta' = \sum_i \alpha_i \theta'_i$ (Equation 4). SpatioTemporal Difference Convolution (STDC) extends PDC to video by slicing the 3D volume into H-T and W-T planes and applying central and angular difference patterns there; because those are still linear convolutions, DCR folds them away at inference, leaving a plain 3D standard convolution.
What would settle it
Compare SDNet against its own Baseline-Rep variant—an identical architecture whose every training branch is a standard convolution—across all six test datasets and several training seeds; if the PDC-trained folds do not consistently beat the standard-trained folds, the claim that difference operators, rather than reparameterization itself, cause the accuracy gain is falsified.
Extended reading notes
Core claim
The central claim is that contrast captures can be added to a CNN during training and then removed from the architecture at inference without losing their benefit. The paper's load-bearing identity, Equation (4), states that a weighted sum of standard convolution and multiple pixel difference convolutions (PDCs) equals a single standard convolution whose kernel is the coefficient-weighted sum of the individual kernels, $\theta' = \sum_i \alpha_i \theta'_i$, after each PDC is rewritten in standard form. Consequently the deployed SDNet and STDNet backbones are plain convolutional networks; the difference operators exist only during training, where they force the model to learn high-frequency contrast cues that survive in the folded kernel. For video, the same logic is extended to spatiotemporal difference convolutions (STDCs) on orthogonal W-T and H-T planes, so motion and appearance contrasts are learned without adding inference cost.
Load-bearing premise
The entire efficiency story rests on the premise that the contrast information the network learns during its multi-branch training phase stays inside the single folded kernel at inference; the paper checks this only on salient-object-detection datasets with one fixed training recipe.
Editorial extensions
If this is right
- Because DCR collapses any number of PDC or STDC branches into one convolution, a system designer can add contrast-aware training branches to arbitrary layers of a small CNN with zero added inference parameters or FLOPs.
- On the embedded GPU tested, the image model's 46 FPS and the video model's 150 FPS make salient object detection fast enough to run as a front-end for real-time downstream applications rather than a system bottleneck.
- The ablation shows that reparameterizing standard-convolution branches alone yields no accuracy gain, so the reported benefit is specifically attributable to difference operators encoding contrast, not to over-parameterization per se.
- For video, STDC-based temporal modules improve temporal consistency over spatial-only processing and outperform a lightweight temporal attention baseline on the evaluated benchmarks, suggesting contrast cues are a strong alternative to attention for motion-aware saliency.
- Scaling up the STDNet backbone to a larger EfficientNet-B5 shifts the trade-off toward accuracy on DAVSOD, VOS, and DAVIS while keeping the model far faster than prior real-time methods, so the design leaves room for accuracy-first deployments.
Reading between the lines
- The DCR identity in Equation (4) is purely algebraic, so the same train-then-fold recipe could be applied to any dense prediction task where local contrast is informative—edge detection, tracking, remote sensing, or image enhancement—without altering the deployed architecture; this is an extension the paper only gestures at in its conclusion.
- The paper's finding that video accuracy saturates at 8 input frames suggests the temporal receptive field of the STDM is the current bottleneck; a longer-range temporal contrast mechanism might push VSOD accuracy further, though it would have to keep the folded-inference trick to preserve the reported speed.
- Because the folded kernel is mathematically identical to the trained multi-branch layer in exact arithmetic, any residual difference between the two in practice comes from floating-point rounding; comparing the two at inference would tell whether the apparent 'free lunch' is fully realized in finite precision.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes SDNet and STDNet, lightweight CNNs for image and video salient object detection, built on Pixel Difference Convolutions (PDCs) and their video extension SpatioTemporal Difference Convolution (STDC). The central technical device is Difference Convolution Reparameterization (DCR): during training a layer contains several difference-convolution branches plus a standard convolution branch, and at inference the weighted branch kernels are algebraically collapsed into a single standard convolution kernel via Eq. (4), so the deployed model is exactly the same function as the trained multi-branch model with no extra parameters or FLOPs. Experiments on six image SOD benchmarks and three video SOD benchmarks, including measurements on an RTX 2080 Ti, Jetson AGX Orin, and Xavier NX, report sub-million-parameter models running at 46 FPS (image) and 150 FPS (video) on Orin with accuracy competitive with much larger models.
Significance. The core contribution is a clean, exact reparameterization identity: Eqs. (3)-(4) are a genuine algebraic equivalence, not an approximation, and the manuscript supports the empirical benefit with a well-designed ablation (Baseline vs. Baseline-Rep vs. SDNet in Table 6), where the failure of the standard-convolution-only Baseline-Rep to improve over Baseline actually strengthens the attribution of the gain to the PDC branches. The hardware efficiency claims are concrete and falsifiable: the authors report batch-size-1 FPS on two real devices and use a unified evaluation code for all accuracy metrics. The main limitation is scope: the DCR training benefit is demonstrated only for SOD, and the authors themselves note in Sec. 4.3 that simply reparameterizing standard convolutions gives no gain in this setting; generalizing the conclusion to other tasks requires additional evidence. I agree with the stress-test assessment that Eq. (4) is exact as written for the linear operators, and I do not see a circularity problem in the derivation.
major comments (1)
- [Section 3.2, Eqs. (3)-(4)] The DCR identity is stated for branches that are pure linear convolutions, but the backbone is described as PiDiNet-based, and standard PiDiNet blocks typically contain batch normalization. If the training-time branches include BN, the paper must show how the BN affine parameters are folded into theta'_i before the weighted sum; otherwise the claim that the inference-time model is literally the same function as the training-time model is under-specified. Please state explicitly whether BN is used in the DCR branches and, if so, give the folding equations and state where BN is placed after fusion.
minor comments (6)
- [General] There are several typos that should be corrected: "contribuions" in the Introduction, "unprecented" in the Introduction, and "Tempral attentions" in Table 7.
- [Section 4.1] The DAVIS validation set is described as "randomly choose 8 videos from its training set," but no seed or video list is given; providing this detail would make the reported DAVIS numbers reproducible.
- [Section 4.1] The sentence "we left pad the last remaining frames" should be "we left-pad the last remaining frames" or "we pad the last remaining frames on the left."
- [Section 3.2, Fig. 11] The manuscript should state explicitly whether the learned coefficients alpha_i are included in the reported parameter counts and whether they are trained with a softmax constraint in all layers; Fig. 11 suggests softmax is used, but the text does not specify this in Sec. 3.2.
- [Fig. 1] The caption says the circle size represents parameter count "illustrated separately for the two figures," meaning the circle scales are not directly comparable between the ISOD and VSOD panels; this should be acknowledged in the caption so readers are not misled.
- [Abstract and Introduction] The speed-up claims are stated as "more than 2x and 3x" in the abstract and "4~6 and 2~4 times faster" in the Introduction; these formulations should be harmonized so that the numbers are immediately consistent.
Circularity Check
No significant circularity: Eq. (4) is an exact algebraic equivalence, and the PDC training benefit is directly ablated in Table 6.
full rationale
The paper's central technical claim is an exact linear-algebraic equivalence, not a statistical prediction. Eq. (3) rewrites a PDC branch f(Δx, θ_i) = Σ_j w_{i,j}(x_j - x'_j) by redistributing each difference pair's weights onto the two involved pixels, giving an equivalent standard convolution with kernel θ'_i; Eq. (4) then uses linearity to absorb the learned branch coefficients α_i into a single kernel θ' = Σ_i α_i θ'_i. The inference-time single-branch network computes literally the same function as the training-time multi-branch network, so no fitted parameter is renamed as a prediction and no result is assumed in its own derivation. The empirical accuracy gain of using PDC branches is directly tested in Table 6, where Baseline-Rep (standard-convolution branches reparameterized the same way) fails to improve over Baseline while SDNet improves, showing the gain is attributed to PDC content rather than to the reparameterization procedure itself. The paper does cite the authors' prior PiDiNet work for PDC definitions, the 3×3 RPDC-to-5×5 conversion, and the CDCM/CSAM modules, but the load-bearing DCR derivation is reproduced in full in Eqs. (3)-(4) and does not depend on those citations for its validity. Remaining concerns (single-run timing, code availability, transferability to other tasks) are verification and generality issues, not circularity.
Assumptions & free parameters
free parameters (4)
- branch coefficients alpha_i =
learned per layer
- number of frames (8) in video input =
8
- input resolution (320x320 for ISOD, 256x256 for VSOD) =
320x320/256x256
- loss weighting (L = Lbce + LIoU + Lssim) =
equal weights of 1
assumptions (4)
- standard math The reparameterization identity in Eq. (3) holds for any pixel pair selection strategy.
- domain assumption PDC acts as a high-pass filter, and standard convolution preserves low frequencies.
- domain assumption Salient objects are characterized by high-frequency spatial contrast and center-surround interactions.
- ad hoc to paper Slicing the 3D volume into H-T and W-T planes preserves enough spatiotemporal information for VSOD.
invented entities (1)
-
SpatioTemporal Difference Convolution (STDC)
independent evidence
Cite this review
Pith. "Pith review of Rapid Salient Object Detection with Difference Convolutional Neural Networks." pith.science (2026). https://pith.science/paper/XD2IYJAL
@misc{pith2026250701182,
author = {Pith},
title = {Pith review of: Rapid Salient Object Detection with Difference Convolutional Neural Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/XD2IYJAL}},
note = {Machine review of arXiv:2507.01182}
}
abstract
This paper addresses the challenge of deploying salient object detection (SOD) on resource-constrained devices with real-time performance. While recent advances in deep neural networks have improved SOD, existing top-leading models are computationally expensive. We propose an efficient network design that combines traditional wisdom on SOD and the representation power of modern CNNs. Like biologically-inspired classical SOD methods relying on computing contrast cues to determine saliency of image regions, our model leverages Pixel Difference Convolutions (PDCs) to encode the feature contrasts. Differently, PDCs are incorporated in a CNN architecture so that the valuable contrast cues are extracted from rich feature maps. For efficiency, we introduce a difference convolution reparameterization (DCR) strategy that embeds PDCs into standard convolutions, eliminating computation and parameters at inference. Additionally, we introduce SpatioTemporal Difference Convolution (STDC) for video SOD, enhancing the standard 3D convolution with spatiotemporal contrast capture. Our models, SDNet for image SOD and STDNet for video SOD, achieve significant improvements in efficiency-accuracy trade-offs. On a Jetson Orin device, our models with $<$ 1M parameters operate at 46 FPS and 150 FPS on streamed images and videos, surpassing the second-best lightweight models in our experiments by more than $2\times$ and $3\times$ in speed with superior accuracy. Code will be available at https://github.com/hellozhuo/stdnet.git.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[43]
Lightweight pixel difference networks for efficient visual representation learning,
Z. Su, J. Zhang, L. Wang, H. Zhang, Z. Liu, M. Pietik ¨ainen, and L. Liu, “Lightweight pixel difference networks for efficient visual representation learning,” IEEE Trans. Pattern Anal. Mach. Intell. , pp. 1–18, 2023
work page 2023
-
[1]
A model of saliency-based visual attention for rapid scene analysis,
L. Itti, C. Koch, and E. Niebur, “A model of saliency-based visual attention for rapid scene analysis,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 20, no. 11, pp. 1254–1259, 1998
1998
-
[2]
Learning to detect a salient object,
T. Liu, Z. Yuan, J. Sun, J. Wang, N. Zheng, X. Tang, and H.-Y. Shum, “Learning to detect a salient object,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 33, no. 2, pp. 353–367, 2010
2010
-
[3]
Attentive and pre-attentive aspects of figural processing,
L. G. Appelbaum and A. M. Norcia, “Attentive and pre-attentive aspects of figural processing,” J. Vis., vol. 9, no. 11, pp. 18–18, 2009
2009
-
[4]
Salient object detection: A survey,
A. Borji, M.-M. Cheng, Q. Hou, H. Jiang, and J. Li, “Salient object detection: A survey,” Computational Visual Media , vol. 5, pp. 117–150, 2019
2019
-
[5]
SALISA: Saliency-based input sampling for efficient video object detection,
B. E. Bejnordi, A. Habibian, F. Porikli, and A. Ghodrati, “SALISA: Saliency-based input sampling for efficient video object detection,” in Eur. Conf. Comput. Vis., 2022
2022
-
[6]
Is bottom-up attention useful for object recognition?
U. Rutishauser, D. Walther, C. Koch, and P . Perona, “Is bottom-up attention useful for object recognition?” in IEEE Conf. Comput. Vis. Pattern Recog., 2004
2004
-
[7]
Region-based saliency detection and its application in object recognition,
Z. Ren, S. Gao, L.-T. Chia, and I. W.-H. Tsang, “Region-based saliency detection and its application in object recognition,” IEEE Trans. Circuits Syst. Video Technol., vol. 24, no. 5, pp. 769–779, 2013
2013
Show all 116 references
-
[8]
Highly efficient and unsupervised framework for moving object detection in satellite videos,
C. Xiao, W. An, Y. Zhang, Z. Su, M. Li, W. Sheng, M. Pietik ¨ainen, and L. Liu, “Highly efficient and unsupervised framework for moving object detection in satellite videos,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 12, pp. 11 532–11 539, 2024
2024
-
[9]
Boosting low-data instance segmentation by unsupervised pre-training with saliency prompt,
H. Li, D. Zhang, N. Liu, L. Cheng, Y. Dai, C. Zhang, X. Wang, and J. Han, “Boosting low-data instance segmentation by unsupervised pre-training with saliency prompt,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023
2023
-
[10]
Large- scale unsupervised semantic segmentation,
S. Gao, Z.-Y. Li, M.-H. Yang, M.-M. Cheng, J. Han, and P . Torr, “Large- scale unsupervised semantic segmentation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 6, pp. 7457–7476, 2022
2022
-
[11]
Full-duplex strategy for video object segmentation,
G.-P . Ji, K. Fu, Z. Wu, D.-P . Fan, J. Shen, and L. Shao, “Full-duplex strategy for video object segmentation,” in Int. Conf. Comput. Vis., 2021
2021
-
[12]
Online tracking by learning discriminative saliency map with convolutional neural network,
S. Hong, T. You, S. Kwak, and B. Han, “Online tracking by learning discriminative saliency map with convolutional neural network,” in ICML, 2015
2015
-
[13]
Weighted attentional blocks for probabilistic object tracking,
H. Wu, G. Li, and X. Luo, “Weighted attentional blocks for probabilistic object tracking,” Vis. Comput., vol. 30, pp. 229–243, 2014
2014
-
[14]
Mobile product search with bag of hash bits and boundary reranking,
J. He, J. Feng, X. Liu, T. Cheng, T.-H. Lin, H. Chung, and S.-F. Chang, “Mobile product search with bag of hash bits and boundary reranking,” in IEEE Conf. Comput. Vis. Pattern Recog., 2012
2012
-
[15]
SalientShape: group saliency in image collections,
M.-M. Cheng, N. J. Mitra, X. Huang, and S.-M. Hu, “SalientShape: group saliency in image collections,” Vis. Comput., vol. 30, no. 4, pp. 443–453, 2014
2014
-
[16]
Saliency driven perceptual image compression,
Y. Patel, S. Appalaraju, and R. Manmatha, “Saliency driven perceptual image compression,” in WACV, 2021
2021
-
[17]
Automatic foveation for video compression using a neurobiological model of visual attention,
L. Itti, “Automatic foveation for video compression using a neurobiological model of visual attention,” IEEE Trans. Image Process., vol. 13, no. 10, pp. 1304–1318, 2004
2004
-
[19]
Deep saliency prior for reducing visual distraction,
K. Aberman, J. He, Y. Gandelsman, I. Mosseri, D. E. Jacobs, K. Kohlhoff, Y. Pritch, and M. Rubinstein, “Deep saliency prior for reducing visual distraction,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022
2022
-
[20]
Realistic saliency guided image enhancement,
S. M. H. Miangoleh, Z. Bylinskii, E. Kee, E. Shechtman, and Y. Aksoy, “Realistic saliency guided image enhancement,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023
2023
-
[21]
SaliencyMix: a saliency guided data augmentation strategy for better regularization,
A. S. Uddin, M. S. Monira, W. Shin, T. Chung, and S.-H. Bae, “SaliencyMix: a saliency guided data augmentation strategy for better regularization,” in Int. Conf. Learn. Represent., 2021
2021
-
[22]
Quantitative analysis of human- model agreement in visual saliency modeling: A comparative study,
A. Borji, D. N. Sihite, and L. Itti, “Quantitative analysis of human- model agreement in visual saliency modeling: A comparative study,” IEEE Trans. Image Process., vol. 22, no. 1, pp. 55–69, 2012
2012
-
[23]
Mesh saliency,
C. H. Lee, A. Varshney, and D. W. Jacobs, “Mesh saliency,” in ACM SIGGRAPH, 2005
2005
-
[24]
SMU- Net: saliency-guided morphology-aware u-net for breast lesion segmentation in ultrasound image,
Z. Ning, S. Zhong, Q. Feng, W. Chen, and Y. Zhang, “SMU- Net: saliency-guided morphology-aware u-net for breast lesion segmentation in ultrasound image,” IEEE Trans. Med. Imaging, vol. 41, no. 2, pp. 476–490, 2021
2021
-
[25]
Salient object detection in optical remote sensing images driven by transformer,
G. Li, Z. Bai, Z. Liu, X. Zhang, and H. Ling, “Salient object detection in optical remote sensing images driven by transformer,” IEEE Trans. Image Process., 2023
2023
-
[26]
Lightweight salient object detection in optical remote-sensing images via semantic matching and edge alignment,
G. Li, Z. Liu, X. Zhang, and W. Lin, “Lightweight salient object detection in optical remote-sensing images via semantic matching and edge alignment,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–11, 2023
2023
-
[27]
A shape-based approach for salient object detection using deep learning,
J. Kim and V . Pavlovic, “A shape-based approach for salient object detection using deep learning,” in Eur. Conf. Comput. Vis., 2016. ACCEPTED IN IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 15
2016
-
[28]
Salient object detection via fast r-cnn and low-level cues,
X. Wang, H. Ma, and X. Chen, “Salient object detection via fast r-cnn and low-level cues,” in IEEE Int. Conf. Image Process., 2016
2016
-
[29]
Visual saliency transformer,
N. Liu, N. Zhang, K. Wan, L. Shao, and J. Han, “Visual saliency transformer,” in Int. Conf. Comput. Vis., 2021
2021
-
[30]
Salient object detection via integrity learning,
M. Zhuge, D.-P . Fan, N. Liu, D. Zhang, D. Xu, and L. Shao, “Salient object detection via integrity learning,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 3, pp. 3738–3752, 2023
2023
-
[31]
SAMNet: Stereoscopically attentive multi-scale network for lightweight salient object detection,
Y. Liu, X.-Y. Zhang, J.-W. Bian, L. Zhang, and M.-M. Cheng, “SAMNet: Stereoscopically attentive multi-scale network for lightweight salient object detection,” IEEE Trans. Image Process. , vol. 30, pp. 3804–3814, 2021
2021
-
[32]
Lightweight salient object detection via hierarchical visual perception learning,
Y. Liu, Y.-C. Gu, X.-Y. Zhang, W. Wang, and M.-M. Cheng, “Lightweight salient object detection via hierarchical visual perception learning,” IEEE Trans. Cybern., vol. 51, no. 9, pp. 4439–4449, 2020
2020
-
[33]
PoolNet+: Exploring the potential of pooling for salient object detection,
J.-J. Liu, Q. Hou, Z.-A. Liu, and M.-M. Cheng, “PoolNet+: Exploring the potential of pooling for salient object detection,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 1, pp. 887–904, 2022
2022
-
[34]
EDN: Salient object detection via extremely-downsampled network,
Y.-H. Wu, Y. Liu, L. Zhang, M.-M. Cheng, and B. Ren, “EDN: Salient object detection via extremely-downsampled network,” IEEE Trans. Image Process., vol. 31, pp. 3125–3136, 2022
2022
-
[35]
A highly efficient model to study the semantics of salient object detection,
M.-M. Cheng, S. Gao, A. Borji, Y.-Q. Tan, Z. Lin, and M. Wang, “A highly efficient model to study the semantics of salient object detection,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 11, pp. 8006–8021, 2021
2021
-
[36]
Contrast-based image attention analysis by using fuzzy growing,
Y.-F. Ma and H.-J. Zhang, “Contrast-based image attention analysis by using fuzzy growing,” in ACM Int. Conf. Multimedia, 2003
2003
-
[37]
Saliency filters: Contrast based filtering for salient region detection,
F. Perazzi, P . Kr ¨ahenb ¨uhl, Y. Pritch, and A. Hornung, “Saliency filters: Contrast based filtering for salient region detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2012
2012
-
[38]
Salient region detection by ufo: Uniqueness, focusness and objectness,
P . Jiang, H. Ling, J. Yu, and J. Peng, “Salient region detection by ufo: Uniqueness, focusness and objectness,” in Int. Conf. Comput. Vis., 2013
2013
-
[39]
Efficient salient region detection with soft image abstraction,
M.-M. Cheng, J. Warrell, W.-Y. Lin, S. Zheng, V . Vineet, and N. Crook, “Efficient salient region detection with soft image abstraction,” in Int. Conf. Comput. Vis., 2013
2013
-
[40]
Spatial attention modulates center-surround interactions in macaque visual area v4,
K. A. Sundberg, J. F. Mitchell, and J. H. Reynolds, “Spatial attention modulates center-surround interactions in macaque visual area v4,” Neuron, vol. 61, no. 6, pp. 952–963, 2009
2009
-
[41]
The neural basis of visual function: Vision and visual dysfunction,
V . Casagrande, T. Norton, and A. Leventhal, “The neural basis of visual function: Vision and visual dysfunction,” 1991
1991
-
[42]
Pixel difference networks for efficient edge detection,
Z. Su, W. Liu, Z. Yu, D. Hu, Q. Liao, Q. Tian, M. Pietik ¨ainen, and L. Liu, “Pixel difference networks for efficient edge detection,” in Int. Conf. Comput. Vis., 2021
2021
-
[44]
Dynamic texture recognition using local binary patterns with an application to facial expressions,
G. Zhao and M. Pietikainen, “Dynamic texture recognition using local binary patterns with an application to facial expressions,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 29, no. 6, pp. 915–928, 2007
2007
-
[45]
Exploring rich and efficient spatial temporal interactions for real-time video salient object detection,
C. Chen, G. Wang, C. Peng, Y. Fang, D. Zhang, and H. Qin, “Exploring rich and efficient spatial temporal interactions for real-time video salient object detection,” IEEE Trans. Image Process., vol. 30, pp. 3995– 4007, 2021
2021
-
[46]
Dynamic context-sensitive filtering network for video salient object detection,
M. Zhang, J. Liu, Y. Wang, Y. Piao, S. Yao, W. Ji, J. Li, H. Lu, and Z. Luo, “Dynamic context-sensitive filtering network for video salient object detection,” in Int. Conf. Comput. Vis., 2021
2021
-
[47]
Mobilesal: Extremely efficient RGB-D salient object detection,
Y.-H. Wu, Y. Liu, J. Xu, J.-W. Bian, Y.-C. Gu, and M.-M. Cheng, “Mobilesal: Extremely efficient RGB-D salient object detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 12, pp. 10 261–10 269, 2022
2022
-
[48]
Depth quality- inspired feature manipulation for efficient RGB-D salient object detection,
W. Zhang, G.-P . Ji, Z. Wang, K. Fu, and Q. Zhao, “Depth quality- inspired feature manipulation for efficient RGB-D salient object detection,” in ACM Int. Conf. Multimedia, 2021
2021
-
[49]
A simple pooling- based design for real-time salient object detection,
J.-J. Liu, Q. Hou, M.-M. Cheng, J. Feng, and J. Jiang, “A simple pooling- based design for real-time salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019
2019
-
[50]
EGNet: Edge guidance network for salient object detection,
J.-X. Zhao, J.-J. Liu, D.-P . Fan, Y. Cao, J. Yang, and M.-M. Cheng, “EGNet: Edge guidance network for salient object detection,” in Int. Conf. Comput. Vis., 2019
2019
-
[51]
Amulet: Aggregating multi-level convolutional features for salient object detection,
P . Zhang, D. Wang, H. Lu, H. Wang, and X. Ruan, “Amulet: Aggregating multi-level convolutional features for salient object detection,” in Int. Conf. Comput. Vis., 2017
2017
-
[52]
Detect globally, refine locally: A novel approach to saliency detection,
T. Wang, L. Zhang, S. Wang, H. Lu, G. Yang, X. Ruan, and A. Borji, “Detect globally, refine locally: A novel approach to saliency detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018
2018
-
[53]
Salient object detection with pyramid attention and salient edges,
W. Wang, S. Zhao, J. Shen, S. C. Hoi, and A. Borji, “Salient object detection with pyramid attention and salient edges,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019
2019
-
[54]
MobileNetV2: Inverted residuals and linear bottlenecks,
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018
2018
-
[55]
RepVGG: Making vgg-style convnets great again,
X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, and J. Sun, “RepVGG: Making vgg-style convnets great again,” in IEEE Conf. Comput. Vis. Pattern Recog., 2021
2021
-
[56]
Semi- supervised video salient object detection using pseudo-labels,
P . Yan, G. Li, Y. Xie, Z. Li, C. Wang, T. Chen, and L. Lin, “Semi- supervised video salient object detection using pseudo-labels,” in Int. Conf. Comput. Vis., 2019
2019
-
[57]
Shifting more attention to video salient object detection,
D.-P . Fan, W. Wang, M.-M. Cheng, and J. Shen, “Shifting more attention to video salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019
2019
-
[58]
Flow guided recurrent neural encoder for video salient object detection,
G. Li, Y. Xie, T. Wei, K. Wang, and L. Lin, “Flow guided recurrent neural encoder for video salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018
2018
-
[59]
Pyramid dilated deeper convlstm for video salient object detection,
H. Song, W. Wang, S. Zhao, J. Shen, and K.-M. Lam, “Pyramid dilated deeper convlstm for video salient object detection,” in Eur. Conf. Comput. Vis., 2018
2018
-
[60]
Deeply supervised 3d recurrent fcn for salient object detection in videos
T.-N. Le and A. Sugimoto, “Deeply supervised 3d recurrent fcn for salient object detection in videos.” in Brit. Mach. Vis. Conf., 2017
2017
-
[61]
A novel long-term iterative mining scheme for video salient object detection,
C. Chen, H. Wang, Y. Fang, and C. Peng, “A novel long-term iterative mining scheme for video salient object detection,” IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 11, pp. 7662–7676, 2022
2022
-
[62]
Depth-cooperated trimodal network for video salient object detection,
Y. Lu, D. Min, K. Fu, and Q. Zhao, “Depth-cooperated trimodal network for video salient object detection,” in IEEE Int. Conf. Image Process., 2022
2022
-
[63]
Run, don’t walk: chasing higher FLOPS for faster neural networks,
J. Chen, S.-h. Kao, H. He, W. Zhuo, S. Wen, C.-H. Lee, and S.-H. G. Chan, “Run, don’t walk: chasing higher FLOPS for faster neural networks,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023
2023
-
[64]
Lightweight vision transformer with bidirectional interaction,
Q. Fan, H. Huang, X. Zhou, and R. He, “Lightweight vision transformer with bidirectional interaction,” in Adv. Neural Inform. Process. Syst., 2024
2024
-
[65]
Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction,
H. Cai, J. Li, M. Hu, C. Gan, and S. Han, “Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction,” inInt. Conf. Comput. Vis., 2023
2023
-
[66]
EfficientFormer: Vision transformers at mobilenet speed,
Y. Li, G. Yuan, Y. Wen, E. Hu, G. Evangelidis, S. Tulyakov, Y. Wang, and J. Ren, “EfficientFormer: Vision transformers at mobilenet speed,” in Adv. Neural Inform. Process. Syst., 2022
2022
-
[67]
Dynamic group convolution for accelerating convolutional neural networks,
Z. Su, L. Fang, W. Kang, D. Hu, M. Pietik ¨ainen, and L. Liu, “Dynamic group convolution for accelerating convolutional neural networks,” in Eur. Conf. Comput. Vis., 2020
2020
-
[68]
Structured pruning for deep convolutional neural networks: A survey,
Y. He and L. Xiao, “Structured pruning for deep convolutional neural networks: A survey,” IEEE Trans. Pattern Anal. Mach. Intell., 2023
2023
-
[69]
Llm-pruner: On the structural pruning of large language models,
X. Ma, G. Fang, and X. Wang, “Llm-pruner: On the structural pruning of large language models,” in Adv. Neural Inform. Process. Syst., 2023
2023
-
[70]
Boosting convolutional neural networks with middle spectrum grouped convolution,
Z. Su, J. Zhang, T. Liu, Z. Liu, S. Zhang, M. Pietik ¨ainen, and L. Liu, “Boosting convolutional neural networks with middle spectrum grouped convolution,” IEEE Trans. Neural Netw. Learn. Syst., 2024
2024
-
[71]
Knowledge distillation: A survey,
J. Gou, B. Yu, S. J. Maybank, and D. Tao, “Knowledge distillation: A survey,” Int. J. Comput. Vis., vol. 129, no. 6, pp. 1789–1819, 2021
2021
-
[72]
One-for- all: Bridge the gap between heterogeneous architectures in knowledge distillation,
Z. Hao, J. Guo, K. Han, Y. Tang, H. Hu, Y. Wang, and C. Xu, “One-for- all: Bridge the gap between heterogeneous architectures in knowledge distillation,” in Adv. Neural Inform. Process. Syst., 2024
2024
-
[73]
Curriculum temperature for knowledge distillation,
Z. Li, X. Li, L. Yang, B. Zhao, R. Song, L. Luo, J. Li, and J. Yang, “Curriculum temperature for knowledge distillation,” in AAAI, 2023
2023
-
[74]
Dynamic binary neural network by learning channel-wise thresholds,
J. Zhang, Z. Su, Y. Feng, X. Lu, M. Pietik ¨ainen, and L. Liu, “Dynamic binary neural network by learning channel-wise thresholds,” in ICASSP, 2022
2022
-
[75]
SVNet: Where so (3) equivariance meets binarization on point cloud representation,
Z. Su, M. Welling, L. Liu et al. , “SVNet: Where so (3) equivariance meets binarization on point cloud representation,” in 3DV, 2022
2022
-
[76]
Q-dm: An efficient low-bit quantized diffusion model,
Y. Li, S. Xu, X. Cao, X. Sun, and B. Zhang, “Q-dm: An efficient low-bit quantized diffusion model,” in Adv. Neural Inform. Process. Syst., 2024
2024
-
[77]
Oscillation-free quantization for low-bit vision transformers,
S.-Y. Liu, Z. Liu, and K.-T. Cheng, “Oscillation-free quantization for low-bit vision transformers,” in ICML, 2023
2023
-
[78]
Smoothquant: Accurate and efficient post-training quantization for large language models,
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in ICML, 2023
2023
-
[79]
Median pixel difference convolutional network for robust face recognition,
J. Zhang, Z. Su, and L. Liu, “Median pixel difference convolutional network for robust face recognition,” in Brit. Mach. Vis. Conf., 2022
2022
-
[80]
Multiresolution gray-scale and rotation invariant texture classification with local binary patterns,
T. Ojala, M. Pietikainen, and T. Maenpaa, “Multiresolution gray-scale and rotation invariant texture classification with local binary patterns,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 24, no. 7, pp. 971–987, 2002
2002
-
[81]
Frequency- tuned salient region detection,
R. Achanta, S. Hemami, F. Estrada, and S. Susstrunk, “Frequency- tuned salient region detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2009
2009
-
[82]
Natural scene statistics at the centre of gaze,
P . Reinagel and A. M. Zador, “Natural scene statistics at the centre of gaze,” Network: Computation in Neural Systems , vol. 10, no. 4, p. 341, 1999
1999
-
[83]
Separable self-attention for mobile vision transformers,
S. Mehta and M. Rastegari, “Separable self-attention for mobile vision transformers,” Trans. Mach. Learn. Res., 2023
2023
-
[84]
ExpandNets: Linear over- parameterization to train compact convolutional networks,
S. Guo, J. M. Alvarez, and M. Salzmann, “ExpandNets: Linear over- parameterization to train compact convolutional networks,” in Adv. Neural Inform. Process. Syst., 2020
2020
-
[85]
Deeply supervised salient object detection with short connections,
Q. Hou, M.-M. Cheng, X. Hu, A. Borji, Z. Tu, and P . Torr, “Deeply supervised salient object detection with short connections,”IEEE Trans. ACCEPTED IN IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 16 Pattern Anal. Mach. Intell., vol. 41, no. 4, pp. 815–828, 2019
2019
-
[86]
ESPNetv2: A light-weight, power efficient, and general purpose convolutional neural network,
S. Mehta, M. Rastegari, L. Shapiro, and H. Hajishirzi, “ESPNetv2: A light-weight, power efficient, and general purpose convolutional neural network,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019
2019
-
[87]
BiSeNet v2: Bilateral network with guided aggregation for real-time semantic segmentation,
C. Yu, C. Gao, J. Wang, G. Yu, C. Shen, and N. Sang, “BiSeNet v2: Bilateral network with guided aggregation for real-time semantic segmentation,” Int. J. Comput. Vis., vol. 129, no. 11, pp. 3051–3068, 2021
2021
-
[88]
ENet: A deep neural network architecture for real-time semantic segmentation,
A. Paszke, A. Chaurasia, S. Kim, and E. Culurciello, “ENet: A deep neural network architecture for real-time semantic segmentation,” arXiv preprint arXiv:1606.02147, 2016
2016 arXiv
-
[89]
DABNet: Depth-wise asymmetric bottleneck for real-time semantic segmentation,
G. Li and J. Kim, “DABNet: Depth-wise asymmetric bottleneck for real-time semantic segmentation,” in Brit. Mach. Vis. Conf., 2019
2019
-
[90]
EdgeNeXt: efficiently amalgamated cnn-transformer architecture for mobile vision applications,
M. Maaz, A. Shaker, H. Cholakkal, S. Khan, S. W. Zamir, R. M. Anwer, and F. S. Khan, “EdgeNeXt: efficiently amalgamated cnn-transformer architecture for mobile vision applications,” in Eur. Conf. Comput. Vis., 2022
2022
-
[91]
MobileViT: Light-weight, general- purpose, and mobile-friendly vision transformer,
S. Mehta and M. Rastegari, “MobileViT: Light-weight, general- purpose, and mobile-friendly vision transformer,” in Int. Conf. Learn. Represent., 2022
2022
-
[92]
Learning uncertain convolutional features for accurate saliency detection,
P . Zhang, D. Wang, H. Lu, H. Wang, and B. Yin, “Learning uncertain convolutional features for accurate saliency detection,” in Int. Conf. Comput. Vis., 2017
2017
-
[93]
A stagewise refinement model for detecting salient objects in images,
T. Wang, A. Borji, L. Zhang, P . Zhang, and H. Lu, “A stagewise refinement model for detecting salient objects in images,” in Int. Conf. Comput. Vis., 2017
2017
-
[94]
Learning to detect salient objects with image-level supervision,
L. Wang, H. Lu, Y. Wang, M. Feng, D. Wang, B. Yin, and X. Ruan, “Learning to detect salient objects with image-level supervision,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017
2017
-
[95]
Hierarchical saliency detection,
Q. Yan, L. Xu, J. Shi, and J. Jia, “Hierarchical saliency detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2013
2013
-
[96]
The secrets of salient object segmentation,
Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille, “The secrets of salient object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2014
2014
-
[97]
Saliency detection via graph-based manifold ranking,
C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang, “Saliency detection via graph-based manifold ranking,” in IEEE Conf. Comput. Vis. Pattern Recog., 2013
2013
-
[98]
Design and perceptual validation of performance measures for salient object segmentation,
V . Movahedi and J. H. Elder, “Design and perceptual validation of performance measures for salient object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2010
2010
-
[99]
Visual saliency based on multiscale deep features,
G. Li and Y. Yu, “Visual saliency based on multiscale deep features,” in IEEE Conf. Comput. Vis. Pattern Recog., 2015
2015
-
[100]
A benchmark dataset and saliency-guided stacked autoencoders for video-based salient object detection,
J. Li, C. Xia, and X. Chen, “A benchmark dataset and saliency-guided stacked autoencoders for video-based salient object detection,” IEEE Trans. Image Process., vol. 27, no. 1, pp. 349–364, 2018
2018
-
[101]
A benchmark dataset and evaluation methodology for video object segmentation,
F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “A benchmark dataset and evaluation methodology for video object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016
2016
-
[102]
Adam: A method for stochastic optimization,
D. P . Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Int. Conf. Learn. Represent., 2015
2015
-
[103]
PyTorch: An imperative style, high- performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: An imperative style, high-...
2019
-
[104]
Richer convolutional features for edge detection,
Y. Liu, M.-M. Cheng, X. Hu, J.-W. Bian, L. Zhang, X. Bai, and J. Tang, “Richer convolutional features for edge detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, no. 8, pp. 1939–1946, 2019
1939
-
[105]
ImageNet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in IEEE Conf. Comput. Vis. Pattern Recog., 2009
2009
-
[106]
UnitBox: An advanced object detection network,
J. Yu, Y. Jiang, Z. Wang, Z. Cao, and T. Huang, “UnitBox: An advanced object detection network,” in ACM Int. Conf. Multimedia, 2016
2016
-
[107]
Image quality assessment: from error visibility to structural similarity,
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P . Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, 2004
2004
-
[108]
Structure-measure: A new way to evaluate foreground maps,
D.-P . Fan, M.-M. Cheng, Y. Liu, T. Li, and A. Borji, “Structure-measure: A new way to evaluate foreground maps,” in Int. Conf. Comput. Vis. , 2017
2017
-
[109]
Salient object detection in the deep learning era: An in-depth survey,
W. Wang, Q. Lai, H. Fu, J. Shen, H. Ling, and R. Yang, “Salient object detection in the deep learning era: An in-depth survey,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 6, pp. 3239–3259, 2021
2021
-
[110]
Real-time salient object detection with a minimum spanning tree,
W.-C. Tu, S. He, Q. Yang, and S.-Y. Chien, “Real-time salient object detection with a minimum spanning tree,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016
2016
-
[111]
SCOM: Spatiotemporal constrained optimization for salient object detection,
Y. Chen, W. Zou, Y. Tang, X. Li, C. Xu, and N. Komodakis, “SCOM: Spatiotemporal constrained optimization for salient object detection,” IEEE Trans. Image Process., vol. 27, no. 7, pp. 3345–3357, 2018
2018
-
[112]
Motion guided attention for video salient object detection,
H. Li, G. Chen, G. Li, and Y. Yu, “Motion guided attention for video salient object detection,” in Int. Conf. Comput. Vis., 2019
2019
-
[113]
Multi-stream attention-aware graph convolution network for video salient object detection,
M. Xu, P . Fu, B. Liu, and J. Li, “Multi-stream attention-aware graph convolution network for video salient object detection,” IEEE Trans. Image Process., vol. 30, pp. 4183–4197, 2021
2021
-
[114]
Progressively real-time video salient object detection via cascaded fully convolutional networks with motion attention,
Q. Zheng, Y. Li, L. Zheng, and Q. Shen, “Progressively real-time video salient object detection via cascaded fully convolutional networks with motion attention,” Neurocomputing, vol. 467, pp. 465–475, 2022
2022
-
[115]
Pyramid constrained self-attention network for fast video salient object detection,
Y. Gu, L. Wang, Z. Wang, Y. Liu, M.-M. Cheng, and S.-P . Lu, “Pyramid constrained self-attention network for fast video salient object detection,” in AAAI, 2020
2020
-
[116]
EfficientNet: Rethinking model scaling for convolutional neural networks,
M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in ICML, 2019
2019
-
[117]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inform. Process. Syst., 2017
2017
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.