Pith. sign in

REVIEW 3 major objections 8 minor 1 cited by

GAPNet: A Lightweight Framework for Image and Video Salient Object Detection via Granularity-Aware Paradigm

T0 review · 3 major / 8 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read GAPNet claims that a lightweight salient object detector beats prior lightweight models and rivals heavy ones when supervision granularity is matched to feature scale—coarse center maps for high-level features, fine boundary maps for low-le

desk verdict Useful lightweight SOD with a clever granularity supervision split, but video SOTA claim overreaches and ablations tune on the test set. read the letter →

arxiv 2508.07585 v1 pith:QSYIREFZ submitted 2025-08-11 cs.CV

classification cs.CV
keywords salientobjectdetectionlightweightmodelgranularity-awareparadigmmulti-scalefeaturefusionvideocross-scaleattentiondeepsupervision
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

GAPNet is a lightweight network for salient object detection, the task of locating the most attention-grabbing object or region in an image or video frame. Its central idea is that a small network should supervise each decoder stage with a saliency target of matching granularity: coarse object-center maps for high-level features and fine boundary maps for low-level features, rather than giving every stage the full saliency map. The paper reports that this granularity-aware supervision, implemented with granular pyramid convolution and cross-scale attention modules plus a compact global self-attention head, sets a new state of the art among lightweight image and video SOD models—for instance, best F-measure 0.867 versus 0.856 for EDN-Lite on DUTS-TE—while running at hundreds of frames per second. If the claim is right, it means that well-chosen supervision signals can substantially close the accuracy gap between tiny networks and much heavier ones, which matters for deployment on phones and edge devices.

What carries the argument

Granularity-aware paradigm: decompose the ground-truth saliency foreground into boundary (within 5 pixels of background), center (top 20% of distance to background), and others, then supervise high-level decoder outputs with the coarse center map, low-level outputs with the fine boundary-plus-others map, and the final output with the full map. The implementation rests on two fusion blocks: Granular Pyramid Convolution (GPC), an attention-refined multi-scale atrous convolution with channel splits [1/8,1/8,1/4,1/2] and an attention branch that pools to m=7 before computing a compact self-attention; and Cross-Scale Attention (CSA), where the query is computed from both scales but the keys and v

What would settle it

Retrain GAPNet from the released code, but select the pooling size and split ratios on a held-out validation split of the training set rather than on DUTS-TE; if the $F_\beta^{\max}$ margin over EDN-Lite drops below roughly 0.005 or the ranking against other lightweight baselines changes, the quantified advantage of the granularity-aware paradigm is not replicated.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that supervising all decoder side-outputs with the full saliency map wastes the limited feature richness of lightweight backbones. GAPNet decomposes the ground truth into center, boundary, and others regions by pixel distance to the nearest background, then applies granularity-aware deep supervision: the high-level side-output D2 is trained with the coarse center map, the low-level side-output D1 with the boundary-plus-others map, and the final output D3 with the full map. With two compact modules—granular pyramid convolution for low-level fusion and cross-scale attention for high-level fusion—plus a self-attention global extractor, a MobileNetV

Load-bearing premise

The reported results rest on hyperparameters (pooling size $m=7$ and split ratios $[1/8,1/8,1/4,1/2]$) tuned on the DUTS-TE test set, so the claim of a new state of the art depends on those choices not being overfit to that test set.

Editorial extensions

If this is right

  • Lightweight SOD models can be made more accurate by replacing uniform full-map deep supervision with granularity-matched supervision; this design rule is likely to transfer to other small backbones and decoders.
  • The margin over EDN-Lite on DUTS-TE ($F_\beta^{\max}$ 0.867 vs 0.856) indicates that the earlier lightweight state of the art had remaining decoder capacity, not that lightweight encoders are the primary bottleneck.
  • The same decoder, fed with RGB plus optical flow, achieves competitive video results ($S_\alpha$ 0.893 on DAVIS) at roughly 350 FPS, extending the paradigm to spatio-temporal saliency.
  • The reported hyperparameters (m=7, split ratios [1/8,1/8,1/4,1/2]) offer concrete defaults for future lightweight SOD training; ablations show both smaller and larger pooling sizes reduce accuracy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the thesis would be to retrofit other lightweight U-Net style decoders (e.g., SAMNet, HVPNet) with granularity-matched side-output supervision while leaving their architecture unchanged; the paper's ablations suggest part of the gain comes from the supervision scheme, but the paper does not isolate that attribution outside its own modules.
  • Because the hyperparameters were selected using the DUTS-TE test set, the published margins could be optimistic under independent re-evaluation; tuning on a held-out split during ablation would make the state-of-the-art claim more robust.
  • The center/boundary/others decomposition uses fixed thresholds (5 pixels, top 20% of distance). Varying these thresholds or making the decomposition learnable would clarify whether the specific definitions matter or just the coarse-vs-fine distinction.
  • The video pipeline depends on offline optical flow from an external estimator, which is itself a cost and an error source; an online variant that learns motion jointly would show how much of the video improvement is attributable to the granularity paradigm rather than to the optical-flow preprocessing.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. The paper proposes GAPNet, a lightweight encoder-decoder for image and video salient object detection. The decoder uses granularity-aware connections: low-level features are supervised by a boundary/others map, high-level features by a center map, and the final output by the full saliency map. Two fusion modules are introduced: GPC (pyramid atrous convolution with pooled self-attention) for low-level fusion and CSA (cross-scale attention) for high-level fusion, with a lightweight self-attention global feature extractor on top of MobileNet-V2. In video, RGB and FlowNet-2 optical flow features are fused in a two-stream variant. The authors report state-of-the-art results among lightweight models on five image SOD datasets and four video SOD datasets, at high speed (571/349 FPS), and provide code.

Significance. If the reported results are reproducible, GAPNet is a useful contribution to edge-deployable SOD: it offers a simple, modular recipe for matching granularities of supervision to decoder scales, with clear ablations (Tables 3–6) and thorough comparison on standard benchmarks. The image-SOD claim is well supported: GAPNet improves over EDN-Lite on all five datasets across all six metrics, and the gains are not restricted to one benchmark. The paper is also transparent about evaluation: comparisons use official saliency maps where available, and code is promised. The central conceptual claim—that coarse supervision at high levels and fine supervision at low levels improves lightweight decoders—is plausible and falsifiable. However, the video SOTA claim is not supported as stated, and the image margin is attenuated by test-set hyperparameter selection and lack of statistical confidence. I view the core idea as sound and the manuscript as close, but needing revision.

major comments (3)
  1. [Sec. 4.2.1, Table 2; Abstract] The unqualified claim of 'a new state-of-the-art performance among lightweight image and video SOD models' (Abstract, echoed in Conclusion) is contradicted by the paper's own Table 2 on DAVSOD: GAPNet trails JL-DCF-Light on all three reported metrics (S-alpha 0.706 vs 0.728, Fmax-beta 0.597 vs 0.630, MAE 0.089 vs 0.088). Sec. 4.2.1 concedes this explicitly. DAVSOD is a standard video SOD benchmark, so the headline must be qualified to the datasets where GAPNet actually leads, or an aggregate justification (e.g., average rank with significance testing) must be supplied.
  2. [Sec. 4.3.1, Table 3; Sec. 4.3.2, Table 4] The two hyperparameters that define the GPC module—adaptive pooling size m=7 and pyramid split ratios [1/8, 1/8, 1/4, 1/2]—are selected by ablation on DUTS-TE, which is also the main test dataset (Sec. 4.1). The headline margin over EDN-Lite on DUTS-TE Fmax-beta is only 1.1% (0.867 vs 0.856); tuning on the test set can inflate this margin. Since the improvement also appears consistently across the four other datasets, this does not overturn the image claim, but an independent validation split or a sensitivity analysis is needed to establish that the configuration generalizes rather than being optimized to DUTS-TE.
  3. [Tables 1–2, Sec. 4.2.1] All comparisons are single-run, with no error bars or significance tests, and several decisive margins are very small (e.g., DAVIS S-alpha 0.893 vs 0.892 and Fmax-beta 0.864 vs 0.863 against JL-DCF-Light; DUT-OMRON MAE 0.057 vs 0.058). For a state-of-the-art claim, at least a standard deviation over multiple seeds, or a paired test over samples for the closest competitor, should be reported. This is especially important in the video results, where the DAVIS advantage over JL-DCF-Light is within 0.001–0.004 on two metrics.
minor comments (8)
  1. [Section 5] Typo: 'newstate-of-the-art' should be 'new state-of-the-art'.
  2. [Fig. 6] Axis labels read 'DA VIS' and the caption cites 'DAVIS [93]'; should be 'DAVIS' and the correct reference [100].
  3. [Sec. 4.2] The text says 'six lightweight models' in the comparison, but Table 1 lists seven lightweight competitors (HVPNet, CSNet, SAMNet, EDN-Lite, ELWNet, LARNet, ADMNet+).
  4. [References [26]] In Sec. 2, 'lightweight backbones like EfficientNet-B0 [26]' cites ref. [26], which is a 2013 salient-region-detection paper, not EfficientNet-B0. The correct Tan & Le (2019) EfficientNet reference appears to be missing.
  5. [Eq. (11)] The explanation of the dot-product and L1-norm symbols is garbled ('·' and '·' are both rendered as middle dots). Spell out the operations in words or use distinct symbols.
  6. [Sec. 3.2.2, Eq. (10)] Cross-scale attention is described only in words. Since E3 and E4 have different spatial resolutions (1/16 and 1/32), the flatten-and-concatenate step should be stated explicitly, including how the different sequence lengths are handled in the attention computation.
  7. [Sec. 3.2.3] The video feature-fusion mechanism is described verbally but no equations, parameter counts, or ablations are given for the two-stream variant. A schematic or pseudocode would aid reproducibility.
  8. [Sec. 4.2] For ELWNet and LARNet, the paper states that numbers were extracted from published papers because no official code/maps are available. This limits the controlled comparison on those rows, although it does not affect the main conclusions.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the central architecture is evaluated on external benchmarks; self-citations are baselines or standard components, not load-bearing premises.

full rationale

GAPNet's derivation is self-contained: the decoder is defined by explicit equations (Eqs. 1-10), the loss uses ground-truth decompositions derived from the dataset, and all performance claims are tested on standard external benchmarks (DUTS, DUT-OMRON, HKU-IS, ECSSD, PASCAL-S, DAVIS, DAVSOD, etc.) with official evaluation code. The only self-citations ([18] EDN, [86] MobileSal) are used as baselines/comparisons or as the source of the geometric GT decomposition; neither is an unverified premise that assumes GAPNet's results. The decomposition of GT into center/boundary/others is a definition derived from the ground-truth map, not from the model output, so the supervision does not reduce to a fitted constant. The hyperparameters m=7 and split ratios are selected in ablations on DUTS-TE, which is a test-set tuning concern (and a threat to the strength of the DUTS-TE margin), but it is not circular in the sense of a prediction being equivalent to its inputs. Likewise, the DAVSOD result contradicts the unqualified 'state-of-the-art' claim, but that is a correctness/consistency issue, not circularity. No circular step satisfying the strict quote-and-reduce standard is present.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper adds no new physical or mathematical entities. The free parameters are standard network design choices, with the most consequential being m=7 and the split ratios, both tuned on the DUTS-TE test set. The axioms are domain assumptions about the supervision decomposition, benchmark validity, backbone choice, and the treatment of optical flow.

free parameters (2)
  • GPC adaptive pooling size m = 7
    Selected in ablation Table 3 based on DUTS-TE test set performance; the paper reports m=7 gives +0.3% Fmax over no-attention, while m=28 hurts.
  • Pyramid convolution split ratios = [1/8, 1/8, 1/4, 1/2]
    Chosen over identical split [1/4,1/4,1/4,1/4] based on DUTS-TE test set in Table 4, gaining 0.4% Fmax.
assumptions (4)
  • domain assumption The Euclidean-distance decomposition of the foreground into center (top 20% farthest pixels), boundary (< 5 px from background), and others provides supervision signals that improve lightweight SOD training.
    Sec 3.3.1; this decomposition defines all intermediate training targets. The paper's own Table 6 shows it only helps in one configuration, so this assumption is load-bearing for the novelty and is only weakly supported.
  • domain assumption Standard SOD metrics (F_beta, S_alpha, E_xi, MAE) faithfully rank methods on these datasets.
    Sec 4.1; used both for selecting hyperparameters and for the SOTA claim.
  • domain assumption MobileNetV2 is a sufficient encoder; removing the final pool and FC layers preserves dense prediction capability.
    Sec 3.1.1; the architecture is not tested with other backbones.
  • domain assumption Offline FlowNet 2.0 optical flow can be treated as free input for the video model and does not count toward lightweight efficiency.
    Sec 4.1 video setup; the 300 FPS and lightweight video claims exclude flow computation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GAPNet: A Lightweight Framework for Image and Video Salient Object Detection via Granularity-Aware Paradigm." pith.science (2026). https://pith.science/paper/QSYIREFZ

@misc{pith2026250807585,
  author       = {Pith},
  title        = {Pith review of: GAPNet: A Lightweight Framework for Image and Video Salient Object Detection via Granularity-Aware Paradigm},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QSYIREFZ}},
  note         = {Machine review of arXiv:2508.07585}
}
read the original abstract

Recent salient object detection (SOD) models predominantly rely on heavyweight backbones, incurring substantial computational cost and hindering their practical application in various real-world settings, particularly on edge devices. This paper presents GAPNet, a lightweight network built on the granularity-aware paradigm for both image and video SOD. We assign saliency maps of different granularities to supervise the multi-scale decoder side-outputs: coarse object locations for high-level outputs and fine-grained object boundaries for low-level outputs. Specifically, our decoder is built with granularity-aware connections which fuse high-level features of low granularity and low-level features of high granularity, respectively. To support these connections, we design granular pyramid convolution (GPC) and cross-scale attention (CSA) modules for efficient fusion of low-scale and high-scale features, respectively. On top of the encoder, a self-attention module is built to learn global information, enabling accurate object localization with negligible computational cost. Unlike traditional U-Net-based approaches, our proposed method optimizes feature utilization and semantic interpretation while applying appropriate supervision at each processing stage. Extensive experiments show that the proposed method achieves a new state-of-the-art performance among lightweight image and video SOD models. Code is available at https://github.com/yuhuan-wu/GAPNet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MiqroForge: An Intelligent Workflow Platform for Quantum-Enhanced Computational Chemistry

    physics.chem-ph 2025-08 unverdicted novelty 4.0 of 10

    The paper introduces MiqroForge, a workflow platform with AI scheduling and a visual interface for quantum-enhanced computational chemistry.

Reference graph

Works this paper leans on

111 extracted references · 79 canonical work pages · cited by 1 Pith paper

  1. [1]

    Rgb-d salient object detec- tion: A survey,

    T. Zhou, D.-P. Fan, M.-M. Cheng, J. Shen, and L. Shao, “Rgb-d salient object detec- tion: A survey,” Computational Visual Media, vol. 7, pp. 37–69, 2021

  2. [2]

    Advanced deep-learning tech- niques for salient and category-specific object detection: A survey,

    J. Han, D. Zhang, G. Cheng, N. Liu, and D. Xu, “Advanced deep-learning tech- niques for salient and category-specific object detection: A survey,” IEEE Signal Process. Mag. (SPM) , vol. 35, no. 1, pp. 84–100, 2018

  3. [3]

    Salient object detection driven by fixation prediction,

    W. Wang, J. Shen, X. Dong, and A. Borji, “Salient object detection driven by fixation prediction,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 1711–1720

  4. [4]

    Spatial-temporal initialization dilemma: towards realistic visual tracking,

    C. Liu, Y. Yuan, X. Chen, H. Lu, and D. Wang, “Spatial-temporal initialization dilemma: towards realistic visual tracking,” Visual Intelligence, vol. 2, no. 1, p. 35, 2024

  5. [5]

    Pattern-affinitive propagation across depth, surface normal and semantic segmentation,

    Z. Zhang, Z. Cui, C. Xu, Y. Yan, N. Sebe, and J. Yang, “Pattern-affinitive propagation across depth, surface normal and semantic segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 4106–4115

  6. [6]

    Leveraging instance-, image-and dataset-level informa- tion for weakly supervised instance segmen- tation,

    Y. Liu, Y.-H. Wu, P.-S. Wen, Y.-J. Shi, Y. Qiu, and M.-M. Cheng, “Leveraging instance-, image-and dataset-level informa- tion for weakly supervised instance segmen- tation,” IEEE Trans. Pattern Anal. Mach. Intell., 2020

  7. [7]

    Repfinder: find- ing approximately repeated scene elements for image editing,

    M.-M. Cheng, F.-L. Zhang, N. J. Mitra, X. Huang, and S.-M. Hu, “Repfinder: find- ing approximately repeated scene elements for image editing,” ACM Trans. Graphics (TOG), vol. 29, no. 4, pp. 1–8, 2010

  8. [8]

    JCS: An explainable COVID-19 diagnosis Springer Nature 2021 LATEX template GAPNet 15 system by joint classification and segmenta- tion,

    Y.-H. Wu, S.-H. Gao, J. Mei, J. Xu, D.- P. Fan, R.-G. Zhang, and M.-M. Cheng, “JCS: An explainable COVID-19 diagnosis Springer Nature 2021 LATEX template GAPNet 15 system by joint classification and segmenta- tion,” IEEE Trans. Image Process., vol. 30, pp. 3113–3126, 2021

Show all 111 references
  1. [9]

    Environment exploration for object-based visual saliency learning,

    C. Craye, D. Filliat, and J.-F. Goudou, “Environment exploration for object-based visual saliency learning,” in Int. Conf. Robot. Autom. (ICRA) . IEEE, 2016, pp. 2303–2309

  2. [10]

    Salient object detec- tion: A discriminative regional feature inte- gration approach,

    H. Jiang, J. Wang, Z. Yuan, Y. Wu, N. Zheng, and S. Li, “Salient object detec- tion: A discriminative regional feature inte- gration approach,” in IEEE Conf. Comput. Vis. Pattern Recog., 2013, pp. 2083–2090

  3. [11]

    Salient object detection in the deep learning era: An in-depth survey,

    W. Wang, Q. Lai, H. Fu, J. Shen, H. Ling, and R. Yang, “Salient object detection in the deep learning era: An in-depth survey,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 6, pp. 3239–3259, 2021

  4. [12]

    DHSNet: Deep hier- archical saliency network for salient object detection,

    N. Liu and J. Han, “DHSNet: Deep hier- archical saliency network for salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 678–686

  5. [13]

    Amulet: Aggregating multi- level convolutional features for salient object detection,

    P. Zhang, D. Wang, H. Lu, H. Wang, and X. Ruan, “Amulet: Aggregating multi- level convolutional features for salient object detection,” in Int. Conf. Comput. Vis. , 2017, pp. 202–211

  6. [14]

    EGNet: Edge guidance network for salient object detec- tion,

    J.-X. Zhao, J. Liu, D.-P. Fan, Y. Cao, J. Yang, and M.-M. Cheng, “EGNet: Edge guidance network for salient object detec- tion,” in Int. Conf. Comput. Vis. , 2019, pp. 8779–8788

  7. [15]

    A simple pooling-based design for real-time salient object detec- tion,

    J.-J. Liu, Q. Hou, M.-M. Cheng, J. Feng, and J. Jiang, “A simple pooling-based design for real-time salient object detec- tion,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 3917–3926

  8. [16]

    Multi-scale interactive network for salient object detection,

    Y. Pang, X. Zhao, L. Zhang, and H. Lu, “Multi-scale interactive network for salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 9413–9422

  9. [17]

    Interactive two-stream decoder for accurate and fast saliency detection,

    H. Zhou, X. Xie, J.-H. Lai, Z. Chen, and L. Yang, “Interactive two-stream decoder for accurate and fast saliency detection,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 9141–9150

  10. [18]

    Edn: Salient object detec- tion via extremely-downsampled network,

    Y.-H. Wu, Y. Liu, L. Zhang, M.-M. Cheng, and B. Ren, “Edn: Salient object detec- tion via extremely-downsampled network,” IEEE Trans. Image Process. , vol. 31, pp. 3125–3136, 2022

  11. [19]

    Salient object detec- tion via integrity learning,

    M. Zhuge, D.-P. Fan, N. Liu, D. Zhang, D. Xu, and L. Shao, “Salient object detec- tion via integrity learning,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 3, pp. 3738–3752, 2022

  12. [20]

    Very deep convolutional networks for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Int. Conf. Learn. Repre- sent., 2015

  13. [21]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 770–778

  14. [22]

    P2T: Pyramid pooling transformer for scene understanding,

    Y.-H. Wu, Y. Liu, X. Zhan, and M.-M. Cheng, “P2T: Pyramid pooling transformer for scene understanding,” IEEE Trans. Pat- tern Anal. Mach. Intell. , vol. 45, no. 11, pp. 12 760–12 771, 2023

  15. [23]

    Vision transformers with hierarchical attention,

    Y. Liu, Y.-H. Wu, G. Sun, L. Zhang, A. Chhatkuli, and L. Van Gool, “Vision transformers with hierarchical attention,” Machine Intelligence Research , vol. 21, no. 4, pp. 670–683, 2024

  16. [24]

    Low-resolution self-attention for semantic segmentation,

    Y.-H. Wu, S.-C. Zhang, Y. Liu, L. Zhang, X. Zhan, D. Zhou, J. Feng, M.-M. Cheng, and L. Zhen, “Low-resolution self-attention for semantic segmentation,” IEEE Trans. Pattern Anal. Mach. Intell. , 2025

  17. [25]

    Samnet: Stereoscop- ically attentive multi-scale network for lightweight salient object detection,

    Y. Liu, X.-Y. Zhang, J.-W. Bian, L. Zhang, and M.-M. Cheng, “Samnet: Stereoscop- ically attentive multi-scale network for lightweight salient object detection,” IEEE Trans. Image Process. , vol. 30, pp. 3804– 3814, 2021. Springer Nature 2021 LATEX template 16 GAPNet

  18. [26]

    Effi- cient salient region detection with soft image abstraction,

    M.-M. Cheng, J. Warrell, W.-Y. Lin, S. Zheng, V. Vineet, and N. Crook, “Effi- cient salient region detection with soft image abstraction,” in Int. Conf. Comput. Vis. , 2013, pp. 1529–1536

  19. [27]

    MobileNetV2: Inverted residuals and linear bottlenecks,

    M. Sandler, A. Howard, M. Zhu, A. Zhmogi- nov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 4510–4520

  20. [28]

    Deep contrast learning for salient object detection,

    G. Li and Y. Yu, “Deep contrast learning for salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 478– 487

  21. [29]

    Deeply supervised salient object detection with short connec- tions,

    Q. Hou, M.-M. Cheng, X. Hu, A. Borji, Z. Tu, and P. Torr, “Deeply supervised salient object detection with short connec- tions,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, no. 4, pp. 815–828, 2019

  22. [30]

    Revisit- ing computer-aided tuberculosis diagnosis,

    Y. Liu, Y.-H. Wu, S.-C. Zhang, L. Liu, M. Wu, and M.-M. Cheng, “Revisit- ing computer-aided tuberculosis diagnosis,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 46, no. 4, pp. 2316–2332, 2024

  23. [31]

    Video polyp segmentation: A deep learning per- spective,

    G.-P. Ji, G. Xiao, Y.-C. Chou, D.-P. Fan, K. Zhao, G. Chen, and L. Van Gool, “Video polyp segmentation: A deep learning per- spective,” Machine Intelligence Research , vol. 19, no. 6, pp. 531–549, 2022

  24. [32]

    Frontiers in intelligent colonoscopy,

    G.-P. Ji, J. Liu, P. Xu, N. Barnes, F. S. Khan, S. Khan, and D.-P. Fan, “Frontiers in intelligent colonoscopy,” arXiv preprint arXiv:2410.17241, 2024

  25. [33]

    Camouflaged object detection,

    D.-P. Fan, G.-P. Ji, G. Sun, M.-M. Cheng, J. Shen, and L. Shao, “Camouflaged object detection,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2777–2787

  26. [34]

    Deep gradi- ent learning for efficient camouflaged object detection,

    G.-P. Ji, D.-P. Fan, Y.-C. Chou, D. Dai, A. Liniger, and L. Van Gool, “Deep gradi- ent learning for efficient camouflaged object detection,” Machine Intelligence Research , vol. 20, no. 1, pp. 92–108, 2023

  27. [35]

    Superpixel-based spatiotemporal saliency detection,

    Z. Liu, X. Zhang, S. Luo, and O. Le Meur, “Superpixel-based spatiotemporal saliency detection,” IEEE Trans. Circ. Syst. Video Technol. (TCSVT), vol. 24, no. 9, pp. 1522– 1540, 2014

  28. [36]

    Saliency detection via graph- based manifold ranking,

    C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang, “Saliency detection via graph- based manifold ranking,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2013, pp. 3166–3173

  29. [37]

    50 fps object-level saliency detec- tion via maximally stable region,

    X. Huang, Y. Zheng, J. Huang, and Y.-J. Zhang, “50 fps object-level saliency detec- tion via maximally stable region,” IEEE Trans. Image Process. , vol. 29, pp. 1384– 1396, 2019

  30. [38]

    Saliency detection via dense and sparse reconstruction,

    X. Li, H. Lu, L. Zhang, X. Ruan, and M.-H. Yang, “Saliency detection via dense and sparse reconstruction,” in Int. Conf. Comput. Vis., 2013, pp. 2976–2983

  31. [39]

    Global contrast based salient region detection,

    M.-M. Cheng, N. J. Mitra, X. Huang, P. H. Torr, and S.-M. Hu, “Global contrast based salient region detection,” IEEE Trans. Pat- tern Anal. Mach. Intell. , vol. 37, no. 3, pp. 569–582, 2015

  32. [40]

    Salient object detec- tion: A discriminative regional feature inte- gration approach,

    J. Wang, H. Jiang, Z. Yuan, M.-M. Cheng, X. Hu, and N. Zheng, “Salient object detec- tion: A discriminative regional feature inte- gration approach,” Int. J. Comput. Vis., vol. 123, no. 2, pp. 251–268, 2017

  33. [41]

    Decompo- sition and completion network for salient object detection,

    Z. Wu, L. Su, and Q. Huang, “Decompo- sition and completion network for salient object detection,” IEEE Trans. Image Pro- cess., vol. 30, pp. 6226–6239, 2021

  34. [42]

    Salient object detection with pyramid attention and salient edges,

    W. Wang, S. Zhao, J. Shen, S. C. Hoi, and A. Borji, “Salient object detection with pyramid attention and salient edges,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 1448–1457

  35. [43]

    An iterative and cooperative top- down and bottom-up inference network for salient object detection,

    W. Wang, J. Shen, M.-M. Cheng, and L. Shao, “An iterative and cooperative top- down and bottom-up inference network for salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 5968–5977. Springer Nature 2021 LATEX template GAPNet 17

  36. [44]

    Stacked u-shape network with channel-wise atten- tion for salient object detection,

    J. Li, Z. Pan, Q. Liu, and Z. Wang, “Stacked u-shape network with channel-wise atten- tion for salient object detection,” IEEE Trans. Multimedia, vol. 23, pp. 1397–1409, 2020

  37. [45]

    Boundary informa- tion progressive guidance network for salient object detection,

    Z. Yao and L. Wang, “Boundary informa- tion progressive guidance network for salient object detection,” IEEE Trans. Multimedia, vol. 24, pp. 4236–4249, 2021

  38. [46]

    Feature specific progressive improvement for salient object detection,

    X. Wang, Z. Liu, V. Liesaputra, and Z. Huang, “Feature specific progressive improvement for salient object detection,” Pattern Recognition , vol. 147, p. 110085, 2024

  39. [47]

    A simple yet effective network based on vision transformer for camouflaged object and salient object detection,

    C. Hao, Z. Yu, X. Liu, J. Xu, H. Yue, and J. Yang, “A simple yet effective network based on vision transformer for camouflaged object and salient object detection,” IEEE Trans. Image Process., 2025

  40. [48]

    Calibnet: Dual- branch cross-modal calibration for rgb- d salient instance segmentation,

    J. Pei, T. Jiang, H. Tang, N. Liu, Y. Jin, D.-P. Fan, and P.-A. Heng, “Calibnet: Dual- branch cross-modal calibration for rgb- d salient instance segmentation,” IEEE Transactions on Image Processing, 2024

  41. [49]

    Reverse attention for salient object detec- tion,

    S. Chen, X. Tan, B. Wang, and X. Hu, “Reverse attention for salient object detec- tion,” in Eur. Conf. Comput. Vis., 2018, pp. 234–250

  42. [50]

    Com- plementary trilateral decoder for fast and accurate salient object detection,

    Z. Zhao, C. Xia, C. Xie, and J. Li, “Com- plementary trilateral decoder for fast and accurate salient object detection,” in ACM Int. Conf. Multimedia, 2021, pp. 4967–4975

  43. [51]

    Towards a complete and detail-preserved salient object detec- tion,

    Y. K. Yun and W. Lin, “Towards a complete and detail-preserved salient object detec- tion,” IEEE Trans. Multimedia, 2023

  44. [52]

    Rethinking lightweight salient object detection via network depth-width tradeoff,

    J. Li, S. Qiao, Z. Zhao, C. Xie, X. Chen, and C. Xia, “Rethinking lightweight salient object detection via network depth-width tradeoff,” IEEE Trans. Image Process. , 2023

  45. [53]

    Lightweight salient object detection via hierarchical visual per- ception learning,

    Y. Liu, Y.-C. Gu, X.-Y. Zhang, W. Wang, and M.-M. Cheng, “Lightweight salient object detection via hierarchical visual per- ception learning,” IEEE Trans. Cybernetics (TCYB), vol. 51, no. 9, pp. 4439–4449, 2021

  46. [54]

    Densely nested top- down flows for salient object detection,

    C. Fang, H. Tian, D. Zhang, Q. Zhang, J. Han, and J. Han, “Densely nested top- down flows for salient object detection,”Sci- ence China Information Sciences , vol. 65, no. 8, p. 182103, 2022

  47. [55]

    ADM- Net: Attention-guided densely multi-scale network for lightweight salient object detec- tion,

    X. Zhou, K. Shen, and Z. Liu, “ADM- Net: Attention-guided densely multi-scale network for lightweight salient object detec- tion,” IEEE Trans. Multimedia, 2024

  48. [56]

    A highly efficient model to study the semantics of salient object detection,

    M.-M. Cheng, S.-H. Gao, A. Borji, Y.- Q. Tan, Z. Lin, and M. Wang, “A highly efficient model to study the semantics of salient object detection,” IEEE Trans. Pat- tern Anal. Mach. Intell. , vol. 44, no. 11, pp. 8006–8021, 2021

  49. [57]

    Elwnet: An extremely lightweight approach for real- time salient object detection,

    Z. Wang, Y. Zhang, Y. Liu, D. Zhu, S. A. Coleman, and D. Kerr, “Elwnet: An extremely lightweight approach for real- time salient object detection,” IEEE Trans. Circ. Syst. Video Technol. (TCSVT) , 2023

  50. [58]

    Larnet: Towards lightweight, accurate and real-time salient object detection,

    Z. Wang, Y. Zhang, Y. Liu, C. Qin, S. A. Coleman, and D. Kerr, “Larnet: Towards lightweight, accurate and real-time salient object detection,” IEEE Trans. Multimedia, 2023

  51. [59]

    Cnn-based encoder-decoder networks for salient object detection: A comprehensive review and recent advances,

    Y. Ji, H. Zhang, Z. Zhang, and M. Liu, “Cnn-based encoder-decoder networks for salient object detection: A comprehensive review and recent advances,” Information Sciences, vol. 546, pp. 835–857, 2021

  52. [60]

    Regularized densely- connected pyramid network for salient instance segmentation,

    Y.-H. Wu, Y. Liu, L. Zhang, W. Gao, and M.-M. Cheng, “Regularized densely- connected pyramid network for salient instance segmentation,” IEEE Trans. Image Process., vol. 30, pp. 3897–3907, 2021

  53. [61]

    Collaborative compensative transformer network for salient object detection,

    J. Chen, H. Zhang, M. Gong, and Z. Gao, “Collaborative compensative transformer network for salient object detection,” Pat- tern Recognition, vol. 154, p. 110600, 2024. Springer Nature 2021 LATEX template 18 GAPNet

  54. [62]

    Transformer-based efficient salient instance segmentation networks with ori- entative query,

    J. Pei, T. Cheng, H. Tang, and C. Chen, “Transformer-based efficient salient instance segmentation networks with ori- entative query,” IEEE Transactions on Multimedia, vol. 25, pp. 1964–1978, 2022

  55. [63]

    Holistically-nested edge detection,

    S. Xie and Z. Tu, “Holistically-nested edge detection,” in Int. Conf. Comput. Vis. , 2015, pp. 1395–1403

  56. [64]

    Deeply supervised salient object detection with short connec- tions,

    Q. Hou, M.-M. Cheng, X. Hu, A. Borji, Z. Tu, and P. Torr, “Deeply supervised salient object detection with short connec- tions,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 5300–5309

  57. [65]

    Cascaded partial decoder for fast and accurate salient object detection,

    Z. Wu, L. Su, and Q. Huang, “Cascaded partial decoder for fast and accurate salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 3907–3916

  58. [66]

    DNA: Deeply- supervised nonlinear aggregation for salient object detection,

    Y. Liu, M.-M. Cheng, X.-Y. Zhang, G.- Y. Nie, and M. Wang, “DNA: Deeply- supervised nonlinear aggregation for salient object detection,” IEEE Trans. Cybernetics (TCYB), 2021

  59. [67]

    Deeply-supervised nets,

    C.-Y. Lee, S. Xie, P. Gallagher, Z. Zhang, and Z. Tu, “Deeply-supervised nets,” in Artificial intelligence and statistics . Pmlr, 2015, pp. 562–570

  60. [68]

    Label decoupling framework for salient object detection,

    J. Wei, S. Wang, Z. Wu, C. Su, Q. Huang, and Q. Tian, “Label decoupling framework for salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 13 025–13 034

  61. [69]

    Bbs-net: Rgb-d salient object detection with a bifurcated backbone strat- egy network,

    D.-P. Fan, Y. Zhai, A. Borji, J. Yang, and L. Shao, “Bbs-net: Rgb-d salient object detection with a bifurcated backbone strat- egy network,” in European conference on computer vision. Springer, 2020, pp. 275– 292

  62. [70]

    Vscode: General visual salient and camouflaged object detection with 2d prompt learning,

    Z. Luo, N. Liu, W. Zhao, X. Yang, D. Zhang, D.-P. Fan, F. Khan, and J. Han, “Vscode: General visual salient and camouflaged object detection with 2d prompt learning,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 17 169–17 180

  63. [71]

    Exploring salient object detection with adder neural net- works,

    B.-W. Yin and Z. Lin, “Exploring salient object detection with adder neural net- works,” in AAAI Conf. Artif. Intell., vol. 39, no. 9, 2025, pp. 9490–9498

  64. [72]

    Shifting more attention to video salient object detection,

    D.-P. Fan, W. Wang, M.-M. Cheng, and J. Shen, “Shifting more attention to video salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 8554–8564

  65. [73]

    Learning long-term structural dependen- cies for video salient object detection,

    B. Wang, W. Liu, G. Han, and S. He, “Learning long-term structural dependen- cies for video salient object detection,” IEEE Trans. Image Process. , vol. 29, pp. 9017–9031, 2020

  66. [74]

    Exploring rich and efficient spatial temporal interactions for real-time video salient object detection,

    C. Chen, G. Wang, C. Peng, Y. Fang, D. Zhang, and H. Qin, “Exploring rich and efficient spatial temporal interactions for real-time video salient object detection,” IEEE Trans. Image Process. , vol. 30, pp. 3995–4007, 2021

  67. [75]

    Full-duplex strategy for video object segmentation,

    G.-P. Ji, K. Fu, Z. Wu, D.-P. Fan, J. Shen, and L. Shao, “Full-duplex strategy for video object segmentation,” in Int. Conf. Comput. Vis., 2021, pp. 4922–4933

  68. [76]

    Weakly super- vised video salient object detection via point supervision,

    S. Gao, H. Xing, W. Zhang, Y. Wang, Q. Guo, and W. Zhang, “Weakly super- vised video salient object detection via point supervision,” in ACM Int. Conf. Multime- dia, 2022, pp. 3656–3665

  69. [77]

    Psnet: Parallel symmetric network for video salient object detection,

    R. Cong, W. Song, J. Lei, G. Yue, Y. Zhao, and S. Kwong, “Psnet: Parallel symmetric network for video salient object detection,” IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 7, no. 2, pp. 402–414, 2022

  70. [78]

    Con- trollable augmentations for video represen- tation learning,

    R. Qian, W. Lin, J. See, and D. Li, “Con- trollable augmentations for video represen- tation learning,” Visual Intelligence, vol. 2, no. 1, p. 1, 2024

  71. [79]

    Dynamic message propagation network for rgb-d and video Springer Nature 2021 LATEX template GAPNet 19 salient object detection,

    B. Chen, Z. Chen, X. Hu, J. Xu, H. Xie, J. Qin, and M. Wei, “Dynamic message propagation network for rgb-d and video Springer Nature 2021 LATEX template GAPNet 19 salient object detection,” ACM Transac- tions on Multimedia Computing, Communi- cations and Applications, vol. 20,...

  72. [80]

    Tenet: Triple excitation network for video salient object detection,

    S. Ren, C. Han, X. Yang, G. Han, and S. He, “Tenet: Triple excitation network for video salient object detection,” in Eur. Conf. Comput. Vis. Springer, 2020, pp. 212–228

  73. [81]

    Dynamic context-sensitive filtering network for video salient object detection,

    M. Zhang, J. Liu, Y. Wang, Y. Piao, S. Yao, W. Ji, J. Li, H. Lu, and Z. Luo, “Dynamic context-sensitive filtering network for video salient object detection,” in Int. Conf. Com- put. Vis., 2021, pp. 1553–1563

  74. [82]

    Motion-aware mem- ory network for fast video salient object detection,

    X. Zhao, H. Liang, P. Li, G. Sun, D. Zhao, R. Liang, and X. He, “Motion-aware mem- ory network for fast video salient object detection,” IEEE Trans. Image Process. , vol. 33, pp. 709–721, 2024

  75. [83]

    Learning complementary spatial– temporal transformer for video salient object detection,

    N. Liu, K. Nan, W. Zhao, X. Yao, and J. Han, “Learning complementary spatial– temporal transformer for video salient object detection,” IEEE Trans. Neur. Net. Learn. Syst. , vol. 35, no. 8, pp. 10 663– 10 673, 2024

  76. [84]

    A novel divide and conquer solution for long-term video salient object detection,

    Y.-X. Li, C.-L.-Z. Chen, S. Li, A.-M. Hao, and H. Qin, “A novel divide and conquer solution for long-term video salient object detection,” Machine Intelligence Research , vol. 21, no. 4, pp. 684–703, 2024

  77. [85]

    Picanet: Pixel-wise contextual attention learning for accurate saliency detection,

    N. Liu, J. Han, and M.-H. Yang, “Picanet: Pixel-wise contextual attention learning for accurate saliency detection,” IEEE Trans. Image Process. , vol. 29, pp. 6438–6451, 2020

  78. [86]

    Mobile- Sal: Extremely efficient rgb-d salient object detection,

    Y.-H. Wu, Y. Liu, J. Xu, J.-W. Bian, Y.-C. Gu, and M.-M. Cheng, “Mobile- Sal: Extremely efficient rgb-d salient object detection,” IEEE Trans. Pattern Anal. Mach. Intell. , 2021

  79. [87]

    V-Net: Fully convolutional neural networks for volumetric medical image segmenta- tion,

    F. Milletari, N. Navab, and S.-A. Ahmadi, “V-Net: Fully convolutional neural networks for volumetric medical image segmenta- tion,” in International Conference on 3D Vision. IEEE, 2016, pp. 565–571

  80. [88]

    Visual saliency transformer,

    N. Liu, N. Zhang, K. Wan, L. Shao, and J. Han, “Visual saliency transformer,” in Int. Conf. Comput. Vis. , 2021, pp. 4722– 4732

  81. [89]

    PyTorch: An imperative style, high-performance deep learning library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “PyTorch: An imperative style, high-performance deep learning library,” in Adv. Neural Inform. Process. Syst., 2019, pp. 8026–8037

  82. [90]

    Adam: A method for stochastic optimization,

    D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Int. Conf. Learn. Represent., 2015

  83. [91]

    Siamese network for rgb-d salient object detection and beyond,

    K. Fu, D.-P. Fan, G.-P. Ji, Q. Zhao, J. Shen, and C. Zhu, “Siamese network for rgb-d salient object detection and beyond,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 9, pp. 5541–5559, 2022

  84. [92]

    Flownet 2.0: Evolution of optical flow estimation with deep networks,

    E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox, “Flownet 2.0: Evolution of optical flow estimation with deep networks,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 2462–2470

  85. [93]

    Learn- ing to detect salient objects with image-level supervision,

    L. Wang, H. Lu, Y. Wang, M. Feng, D. Wang, B. Yin, and X. Ruan, “Learn- ing to detect salient objects with image-level supervision,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 136–145

  86. [94]

    Visual saliency based on multiscale deep features,

    G. Li and Y. Yu, “Visual saliency based on multiscale deep features,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2015, pp. 5455–5463

  87. [95]

    Hier- archical saliency detection,

    Q. Yan, L. Xu, J. Shi, and J. Jia, “Hier- archical saliency detection,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2013, pp. 1155–1162

  88. [96]

    The secrets of salient object segmentation,

    Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille, “The secrets of salient object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2014, pp. 280–287. Springer Nature 2021 LATEX template 20 GAPNet

  89. [97]

    Detect globally, refine locally: A novel approach to saliency detection,

    T. Wang, L. Zhang, S. Wang, H. Lu, G. Yang, X. Ruan, and A. Borji, “Detect globally, refine locally: A novel approach to saliency detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 3127–3135

  90. [98]

    Learning to promote saliency detectors,

    Y. Zeng, H. Lu, L. Zhang, M. Feng, and A. Borji, “Learning to promote saliency detectors,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 1644–1653

  91. [99]

    Pixels, regions, and objects: Mul- tiple enhancement for salient object detec- tion,

    Y. Wang, R. Wang, X. Fan, T. Wang, and X. He, “Pixels, regions, and objects: Mul- tiple enhancement for salient object detec- tion,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 10 031–10 040

  92. [100]

    A benchmark dataset and eval- uation methodology for video object seg- mentation,

    F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine- Hornung, “A benchmark dataset and eval- uation methodology for video object seg- mentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 724–732

  93. [101]

    Video segmentation by track- ing many figure-ground segments,

    F. Li, T. Kim, A. Humayun, D. Tsai, and J. M. Rehg, “Video segmentation by track- ing many figure-ground segments,” in Int. Conf. Comput. Vis. , 2013, pp. 2192–2199

  94. [102]

    Consis- tent video saliency using local gradient flow optimization and global refinement,

    W. Wang, J. Shen, and L. Shao, “Consis- tent video saliency using local gradient flow optimization and global refinement,” IEEE Trans. Image Process., vol. 24, no. 11, pp. 4185–4196, 2015

  95. [103]

    How to evaluate foreground maps?

    R. Margolin, L. Zelnik-Manor, and A. Tal, “How to evaluate foreground maps?” in IEEE Conf. Comput. Vis. Pattern Recog. , 2014, pp. 248–255

  96. [104]

    Structure-measure: A new way to evaluate foreground maps,

    D.-P. Fan, M.-M. Cheng, Y. Liu, T. Li, and A. Borji, “Structure-measure: A new way to evaluate foreground maps,” in Int. Conf. Comput. Vis., 2017, pp. 4548–4557

  97. [105]

    Enhanced-alignment measure for binary foreground map evalua- tion,

    D.-P. Fan, C. Gong, Y. Cao, B. Ren, M.-M. Cheng, and A. Borji, “Enhanced-alignment measure for binary foreground map evalua- tion,” in IJCAI, 2018, pp. 698–704

  98. [106]

    Is depth really necessary for salient object detection?

    J. Zhao, Y. Zhao, J. Li, and X. Chen, “Is depth really necessary for salient object detection?” in ACM Int. Conf. Multimedia , 2020, pp. 1745–1754

  99. [107]

    Motion guided attention for video salient object detection,

    H. Li, G. Chen, G. Li, and Y. Yu, “Motion guided attention for video salient object detection,” in Int. Conf. Comput. Vis. , 2019, pp. 7274–7283

  100. [108]

    Semi-supervised video salient object detection using pseudo- labels,

    P. Yan, G. Li, Y. Xie, Z. Li, C. Wang, T. Chen, and L. Lin, “Semi-supervised video salient object detection using pseudo- labels,” in Int. Conf. Comput. Vis. , 2019, pp. 7284–7293

  101. [109]

    Weakly supervised video salient object detection,

    W. Zhao, J. Zhang, L. Li, N. Barnes, N. Liu, and J. Han, “Weakly supervised video salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 16 826–16 835. Yu-Huan Wu received his Ph.D. degree from Nankai University in

  102. [2018]

    His research interests include machine learning and optimization

    He is a senior scientist and group manager at the Institute of High Performance Computing (IHPC), A*STAR, Singapore. His research interests include machine learning and optimization. He has led/co-led multiple research initiatives in robust multimodal learning. His research fi...

  103. [2022]

    He has published 10+ papers on top-tier conferences and journals such as IEEE TPAMI/TIP/TNNL- S/CVPR/ICCV

    He is a research scientist at the Institute of High Performance Computing (IHPC), A*STAR, Singa- pore. He has published 10+ papers on top-tier conferences and journals such as IEEE TPAMI/TIP/TNNL- S/CVPR/ICCV. His research interests include computer vision and deep learning. E...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.