REVIEW 3 major objections 5 minor 102 references
Inspiring the Next Generation of Segment Anything Models: Comprehensively Evaluate SAM and SAM 2 with Diverse Prompts Towards Context-Dependent Concepts under Different Scenes
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This evaluation maps SAM and SAM 2's performance boundaries on 11 context-dependent segmentation concepts, finding box prompts generally best and SAM 2's edge confined to video and 3D.
desk verdict Broad, useful SAM/SAM2 evaluation map, but headline performance boundaries are GT-conditioned upper envelopes and two summary claims contradict the paper's own tables. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the unified evaluation framework itself: six prompting settings (everything, mask, box, point, in-context learning, and prompt robustness) applied to SAM and SAM 2 through three prompt-generation strategies. The prediction-based propagated prompt turns the previous frame's predicted mask into the next frame's prompt, letting an image-trained SAM behave as a video segmenter; the bidirectional inference strategy picks the slice with the largest ground-truth foreground mask as an anchor and propagates masks in both directions through a 3D volume; and the in-context learning mode feeds SAM 2 the first 20 training images and masks of a concept as exemplars. The everything-mode outputs are filtered by an overlap filtering strategy that keeps only predicted entities whose overlap with the ground truth exceeds 90%. These mechanisms together convert frozen, off-the-shelf SAMs into measurable systems across heterogeneous data without task-specific fine-tuning.
What would settle it
Rerun the headline settings on a held-out subset of these datasets with no oracle: prompt points and boxes from an off-the-shelf detector or saliency model, no overlap filtering, and no ground-truth-based anchor selection; if SAM 2's everything-mode and point-mode deficits and its 3D advantage shrink or reverse, the central conclusions are artifacts of ground-truth-conditioned prompting. Independently, re-extract the specialist baselines in Table 15a from the original papers, since identical scores across all four modalities would indicate transcription error.
Extended reading notes
Core claim
The paper's central claim is that SAM and SAM 2 have clear, measurable performance boundaries on context-dependent concepts, and that those boundaries depend more on prompt type, data modality, and target material than on model scale. In static images, box prompts dominate on 10 of the 11 concepts and rescue performance on hard cases like camouflage and medical lesions. In everything mode and point-click mode, SAM 2 often performs worse than SAM, which the authors attribute to temporal memory and dynamic prompt attention adding unfavorable bias when no temporal context is available. In video and 3D, the picture flips: SAM 2 with prediction-based propagated prompts, multi-frame mask prompts, and a bidirectional inference strategy for volumetric data surpasses task-specific medical segmentors on brain tumor and multiple sclerosis lesion segmentation. The paper also claims SAM 2 shows real but incomplete in-context learning ability, and that both models are highly sensitive to imperfect prompts, degrading sharply under small perturbations.
Load-bearing premise
The headline numbers assume that ground-truth-derived prompts and ground-truth-filtered outputs are the right way to measure SAMs' context-dependent segmentation ability, and if those oracle conditions are removed, the reported performance boundaries may not describe realistic zero-shot or interactive use.
Editorial extensions
If this is right
- Box prompts should be the default interface when deploying SAMs on context-dependent image tasks; point and everything modes should be treated as unreliable outside well-separated objects.
- SAM 2's temporal memory pays off only when there is temporal structure; for single-frame inference, the original SAM remains competitive or better, so upgrading to SAM 2 is not automatically an improvement.
- In video and 3D medical segmentation, SAM 2 with mask prompts and multi-frame propagation can serve as a strong zero-shot baseline, competitive with or better than task-specific networks.
- Evaluation practices should include prompt-robustness testing, because small perturbations to boxes, points, or masks move scores by percentages that change qualitative conclusions.
- The in-context learning results suggest that a few exemplars can steer SAM 2 toward new concepts, though not yet to the level of existing unified models.
Reading between the lines
- If the oracle-conditioned prompts inflate performance, then the real-world gap between SAMs and specialized models on camouflage, shadows, and medical lesions is larger than the tables suggest; a fully automatic prompt-setting evaluation would test this directly.
- The finding that SAM 2 underperforms SAM on everything mode in static images hints that the memory module or dynamic prompt attention introduces a temporal prior that is harmful without temporal input; inspecting memory-bank activity on single frames could localize the cause.
- The box-prompt dominance across 10 of 11 concepts suggests SAMs act essentially as strong object proposal matchers: given spatial bounds they segment well, but their ability to discover context-dependent regions on their own is limited, so adding a saliency or anomaly proposal module as an automatic prompt generator could be a practical route toward a better next-generation model.
- The 3D success with bidirectional multi-frame mask prompts implies that medical volumes can be treated as videos without architectural changes; testing on CT and MRI modalities beyond the ones reported would establish whether that advantage generalizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents a large-scale empirical evaluation of SAM and SAM 2 on eleven context-dependent (CD) segmentation concepts across 33 datasets, covering 2D images, videos, and 3D medical volumes in natural, medical, and industrial scenes. The authors propose a unified prompting framework that includes everything, box, point, mask, in-context learning, and prompt-robustness modes, and they distill the results into design insights for a future SAM 3. The headline conclusions are that box prompts are generally the most advantageous prompt type, that SAM 2 is not uniformly better than SAM and is worse on everything and point modes, that SAM 2 shows in-context-learning potential, and that bidirectional inference with multi-frame mask prompts enables SAM 2 to surpass specialized 3D medical segmentation models.
Significance. The benchmark is timely and potentially useful: it is the most comprehensive CD-concept evaluation of SAM and SAM 2 that I am aware of, it covers underexplored medical and industrial modalities, and it ships a unified evaluation toolkit and code repository, which supports reproducibility. The robustness analysis in Section 3.8 and the in-context-learning exploration in Section 3.4.2 go beyond earlier SAM evaluations and provide practically relevant information for deployment. If the headline findings were established by the reported data, they would give the community a reference map for prompt selection on CD concepts and a concrete evidence base for SAM 3 design; that is a meaningful contribution. However, as detailed below, several load-bearing conclusions outrun the ground-truth-conditioned protocols that produced the numbers, and one stated conclusion is contradicted by the paper's own tables.
major comments (3)
- [§3.4.1 and §3.9, Items I, II, V] The headline performance conclusions are derived under ground-truth-conditioned protocols rather than zero-shot or realistic interactive conditions. In automatic everything mode, only predicted entities whose overlap with the GT is greater than 90% are retained and merged (Sec. 3.4.1); in point mode, corrective clicks are placed at the largest GT-error region until IoU reaches 0.9 or six clicks are used (Sec. 3.4.1); and in 3D, the anchor slice is selected as the one with the largest GT foreground (Sec. 3.4.4). Consequently, claims such as "box prompts are generally the most advantageous" and "SAM 2 performs worse on everything and point prompts" describe behavior under oracle feedback, not the open-world or zero-shot behavior implied by the abstract and Section 1. The everything-mode comparison may also be confounded by the OFS granularity test: SAM 2's more fragmented automatic masks would be discarded by the >90% per-entity overlap rule even when their union is correct. Section 5's limitations do not disclose this GT-conditioning. Please either re-run automatic and point modes without GT-based filtering/correction (or with a realistic interaction model), re-evaluate the 3D anchor choice without GT, or restrict the conclusions to the oracle-conditioned setting and revise the abstract and Section 3.9 accordingly.
- [§3.5 and Tables 4, 9] The statement that "SAM 2 (point) and SAM 2 (everything) are consistently weaker than their corresponding SAM variants" is contradicted by the paper's own tables. For point mode, Table 4 reports COD10K F_beta of 0.864 for SAM 2 versus 0.823 for SAM, and Table 9 reports SAM 2 with higher Dice on COVID-19 (0.687 vs 0.352), BUSI (0.783 vs 0.694), ISIC-2018 (0.641 vs 0.504), and Polyp-Five (0.862 vs 0.641). For everything mode, the comparison is also uneven across tasks and datasets. The claim in Section 3.9 Item II therefore needs to be revised into a concept- and dataset-specific statement, or the tables must be corrected.
- [§3.7, Table 15(a)] The specialist baselines 3D U-Net and EoFormer are reported with identical Dice values across all four MRI modalities (Flair, T1ce, T1, T2); for example, 3D U-Net shows 0.900/0.807/0.792 repeated verbatim in every column. Identical multi-modality results are implausible and suggest the scores were transcribed from a modality-independent source or copied incorrectly. Since Section 3.7 and Section 3.9 Item V use this comparison to claim that SAM 2 surpasses specialized models such as DRU-Net and 3D U-Net, the baseline values must be verified against the original publications and corrected; otherwise the superiority claim is unsupported.
minor comments (5)
- [§3.4.1] The OFS overlap threshold of 90% is introduced without any sensitivity analysis; a threshold sweep (for example, 50%, 75%, 90%) would clarify how much the everything-mode rankings depend on this arbitrary choice.
- [Tables 3–15] The meaning of "—" differs across tables: in Table 7 it indicates zero prediction accuracy per the PBD protocol, while in Table 9 it means the specialized model does not support that lesion type. A unified caption-level explanation would improve readability.
- [§3.6] The claim that SAM's image-based predictions "fluctuate frame by frame" is supported only by qualitative visualizations; a temporal consistency metric (for example, average inter-frame IoU of predicted masks) would make the comparison quantitative.
- [References [13, 14]] References [13] and [14] appear to be the same paper (Implicit Motion Handling for Video Camouflaged Object Detection) with different page ranges; this looks like a duplicate citation.
- [Prompt icons throughout] The prompt-type icons (for example, "/buromobelexperte", "/d⌢t-circle", "/border-s◎yle") render as odd symbols in the text; please use standard notation or clearly typeset glyphs in the final version.
Circularity Check
No circularity: the paper's claims are empirical measurements under clearly stated oracle-assisted evaluation protocols, not derivations that reduce to their own inputs.
full rationale
This is an empirical evaluation paper, not a derivation. The headline claims (box prompts generally most advantageous; SAM 2 worse on everything and point prompts; bidirectional inference helps 3D) are summaries of measured numbers under protocols that are explicitly described in Sec. 3.4. The everything-mode overlap filtering strategy (OFS) uses the ground-truth mask to retain predicted entities with >90% overlap, and the point-mode simulation uses GT-guided corrective clicks until IoU exceeds 0.9; the 3D anchor slice is chosen by largest GT foreground. These protocols condition on ground truth, which raises a validity question about whether the numbers describe zero-shot or realistic interactive behavior, but they do not make the results circular: the final predictions still depend on SAM/SAM2's raw outputs, and no parameter is fitted from the target quantity and then reported as a prediction. The paper also cites prior work by the same authors (Spider, [98]) for the context-dependent/context-independent taxonomy and for averaging protocols, but that citation is background framing rather than a load-bearing derivation of any experimental result. The identical baseline scores in Table 15a across modalities suggest a transcription concern, not circularity. Because no claimed result is equivalent by construction to its inputs and the evaluation methods are openly stated, the paper is self-contained as an empirical benchmark; concerns about oracle conditioning belong to experimental validity, not circularity.
Assumptions & free parameters
free parameters (5)
- OFS overlap threshold =
0.90 (90%)
- Click stopping IoU and max clicks =
IoU 0.9, max 6 clicks
- Perturbation magnitudes =
box 10% of shorter side, point 10 px, mask 5 erosion/dilation iterations
- ICL exemplar count =
20 training image-mask pairs
- 3D anchor slice selection =
slice with largest foreground mask
assumptions (5)
- domain assumption The CI versus CD concept taxonomy of Barsalou, adopted via the authors' own Spider, is a valid organizing principle for segmentation benchmarks.
- domain assumption Ground-truth-derived prompts and GT-overlap filtering are an appropriate protocol for measuring segmentation capability.
- ad hoc to paper SAM 2's memory mechanism supports genuine in-context learning from 20 exemplar image-mask pairs.
- domain assumption Quoted specialist baselines from original papers are comparable to the SAM results measured here.
- ad hoc to paper The largest-foreground slice is the optimal starting anchor for 3D bidirectional propagation.
Cite this review
Pith. "Pith review of Inspiring the Next Generation of Segment Anything Models: Comprehensively Evaluate SAM and SAM 2 with Diverse Prompts Towards Context-Dependent Concepts under Different Scenes." pith.science (2026). https://pith.science/paper/H72YTHPZ
@misc{pith2026241201240,
author = {Pith},
title = {Pith review of: Inspiring the Next Generation of Segment Anything Models: Comprehensively Evaluate SAM and SAM 2 with Diverse Prompts Towards Context-Dependent Concepts under Different Scenes},
year = {2026},
howpublished = {\url{https://pith.science/paper/H72YTHPZ}},
note = {Machine review of arXiv:2412.01240}
}
read the original abstract
As large-scale foundation models trained on billions of image--mask pairs covering a vast diversity of scenes, objects, and contexts, SAM and its upgraded version, SAM~2, have significantly influenced multiple fields within computer vision. Leveraging such unprecedented data diversity, they exhibit strong open-world segmentation capabilities, with SAM~2 further enhancing these capabilities to support high-quality video segmentation. While SAMs (SAM and SAM~2) have demonstrated excellent performance in segmenting context-independent concepts like people, cars, and roads, they overlook more challenging context-dependent (CD) concepts, such as visual saliency, camouflage, industrial defects, and medical lesions. CD concepts rely heavily on global and local contextual information, making them susceptible to shifts in different contexts, which requires strong discriminative capabilities from the model. The lack of comprehensive evaluation of SAMs limits understanding of their performance boundaries, which may hinder the design of future models. In this paper, we conduct a thorough evaluation of SAMs on 11 CD concepts across 2D and 3D images and videos in various visual modalities within natural, medical, and industrial scenes. We develop a unified evaluation framework for SAM and SAM~2 that supports manual, automatic, and intermediate self-prompting, aided by our specific prompt generation and interaction strategies. We further explore the potential of SAM~2 for in-context learning and introduce prompt robustness testing to simulate real-world imperfect prompts. Finally, we analyze the benefits and limitations of SAMs in understanding CD concepts and discuss their future development in segmentation tasks.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Al-Dhabyani, M
W. Al-Dhabyani, M. Gomaa, H. Khaled, and A. Fahmy. Dataset of breast ultrasound images. Data in Brief , 28:104863, 2020. 2, 8, 11
2020
-
[2]
L. W. Barsalou. Context-independent and context- dependent information in concepts. Memory & Cog- nition, 10(1):82–93, 1982. 3, 5
1982
-
[3]
Bergmann, M
P. Bergmann, M. Fauser, D. Sattlegger, and C. Steger. Mvtec ad–a comprehensive real-world dataset for un- supervised anomaly detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 9592–9600, 2019. 2, 8, 11, 13
2019
-
[4]
Bernal, F
J. Bernal, F. J. S´ anchez, G. Fern´ andez-Esparrach, D. Gil, C. Rodr ´ ıguez, and F. Vilari˜ no. Wm-dova maps for accurate polyp highlighting in colonoscopy: Valida- tion vs. saliency maps from physicians. Computerized Medical Imaging and Graphics , 43:99–111, 2015. 2, 8
2015
-
[5]
Bideau and E
P. Bideau and E. Learned-Miller. It’s moving! a proba- bilistic model for causal motion segmentation in moving camera videos. In Proceedings of European Conference on Computer Vision , pages 433–449, 2016. 2, 6, 12
2016
-
[6]
V. I. Butoi, J. J. G. Ortiz, T. Ma, M. R. Sabuncu, J. Guttag, and A. V. Dalca. Universeg: Universal med- ical image segmentation. In Proceedings of IEEE Inter- national Conference on Computer Vision, pages 21438– 21451, 2023. 1, 5, 10, 11, 19
2023
-
[7]
Caesar, J
H. Caesar, J. Uijlings, and V. Ferrari. Coco-stuff: Thing and stuff classes in context. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 1209–1218, 2018. 3
2018
-
[8]
Carass, S
A. Carass, S. Roy, A. Jog, J. L. Cuzzocreo, E. Magrath, A. Gherman, J. Button, J. Nguyen, F. Prados, C. H. Su- dre, et al. Longitudinal multiple sclerosis lesion segmen- tation: resource and challenge.NeuroImage, 148:77–102,
Show all 102 references
-
[9]
G. Chen, K. Han, and K.-Y. K. Wong. Tom-net: Learn- ing transparent object matting from a single image. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 9233–9241, 2018. 5
2018
-
[10]
G. Chen, L. Li, Y. Dai, J. Zhang, and M. H. Yap. Aau- net: An adaptive attention u-net for breast lesions seg- mentation in ultrasound images. IEEE Transactions on Medical Imaging, 42(5):1289–1300, 2023. 11
2023
-
[11]
T. Chen, L. Zhu, C. Ding, R. Cao, Y. Wang, Z. Li, L. Sun, P. Mao, and Y. Zang. Sam fails to segment anything?–sam-adapter: Adapting sam in underper- formed scenes: Camouflage, shadow, medical image seg- mentation, and more. arXiv preprint arXiv:2304.09148,
-
[12]
Z. Chen, Q. Xu, R. Cong, and Q. Huang. Global context-aware progressive aggregation network for salient object detection. In Proceedings of AAAI Con- ference on Artificial Intelligence , pages 10599–10606,
-
[13]
Cheng, H
X. Cheng, H. Xiong, D. Fan, Y. Zhong, M. Harandi, T. Drummond, and Z. Ge. Implicit motion handling for video camouflaged object detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 13854–13863, 2022. 12
2022
-
[14]
Cheng, H
X. Cheng, H. Xiong, D.-P. Fan, Y. Zhong, M. Harandi, T. Drummond, and Z. Ge. Implicit motion handling for video camouflaged object detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 13864–13873, 2022. 2, 6, 12
2022
-
[15]
C ¸ i¸ cek, A
¨O. C ¸ i¸ cek, A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ronneberger. 3d u-net: Learning dense volumetric segmentation from sparse annotation. In S. Ourselin, L. Joskowicz, M. R. Sabuncu, G. B. ¨Unal, and W. M. W. III, editors, Proceedings of International Conference on ...
2016
-
[16]
Codella, V
N. Codella, V. Rotemberg, P. Tschandl, M. E. Celebi, S. Dusza, D. Gutman, B. Helba, A. Kalloo, K. Liopy- ris, M. Marchetti, et al. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic). arXiv preprint arX...
2018 arXiv
-
[17]
R. Cong, H. Yang, Q. Jiang, W. Gao, H. Li, C. Wang, Y. Zhao, and S. Kwong. Bcs-net: Boundary, context, and semantic for automatic covid-19 lung infection seg- mentation from ct images. IEEE Transactions on In- strumentation and Measurement , 71:1–11, 2022. 5
2022
-
[18]
Cordts, M
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. En- zweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The cityscapes dataset for semantic urban scene under- standing. In Proceedings of IEEE Conference on Com- puter Vision and Pattern Recognition, pages 3213–3223,
-
[19]
Deng and X
H. Deng and X. Li. Anomaly detection via reverse distillation from one-class embedding. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 9727–9736, 2022. 11
2022
-
[20]
Fan, M.-M
D.-P. Fan, M.-M. Cheng, Y. Liu, T. Li, and A. Borji. Structure-measure: A new way to evaluate foreground maps. In Proceedings of IEEE International Conference on Computer Vision , pages 4548–4557, 2017. 8
2017
-
[21]
Fan, G.-P
D.-P. Fan, G.-P. Ji, M.-M. Cheng, and L. Shao. Con- cealed object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence , 44(10):6024–6042,
-
[22]
Fan, G.-P
D.-P. Fan, G.-P. Ji, G. Sun, M.-M. Cheng, J. Shen, and L. Shao. Camouflaged object detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 2777–2787, 2020. 2, 6, 11
2020
-
[23]
Fan, G.-P
D.-P. Fan, G.-P. Ji, P. Xu, M.-M. Cheng, C. Sakaridis, and L. Van Gool. Advances in deep concealed scene understanding. Visual Intelligence , 1(1):16, 2023. 2
2023
-
[24]
Fan, G.-P
D.-P. Fan, G.-P. Ji, T. Zhou, G. Chen, H. Fu, J. Shen, and L. Shao. Pranet: Parallel reverse attention net- work for polyp segmentation. In Proceedings of Inter- national Conference on Medical Image Computing and Computer-Assisted Intervention , pages 263–273, 2020. 5, 11
2020
-
[25]
D.-P. Fan, W. Wang, M.-M. Cheng, and J. Shen. Shift- ing more attention to video salient object detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 8554–8564, 2019. 2, 6, 12
2019
-
[26]
D.-P. Fan, T. Zhou, G.-P. Ji, Y. Zhou, G. Chen, H. Fu, J. Shen, and L. Shao. Inf-net: Automatic covid-19 lung infection segmentation from ct images. IEEE Transac- tions on Medical Imaging , 39(8):2626–2637, 2020. 2, 8, 11
2020
-
[27]
K. Fan, C. Wang, Y. Wang, C. Wang, R. Yi, and L. Ma. Rfenet: Towards reciprocal feature evolution for glass segmentation. In Proceedings of International Joint Conference on Artificial Intelligence , pages 717–725,
-
[28]
Hashemi, M
M. Hashemi, M. Akhbari, and C. Jutten. Delve into mul- tiple sclerosis (MS) lesion exploration: A modified at- tention u-net for MS lesion segmentation in brain MRI. 22 Xiaoqi Zhao 1 et al. Computers in Biology and Medicine , 145:105402, 2022. 12
2022
-
[29]
H. He, X. Li, G. Cheng, J. Shi, Y. Tong, G. Meng, V. Prinet, and L. Weng. Enhanced boundary learn- ing for glass-like object segmentation. In Proceedings of IEEE International Conference on Computer Vision , pages 15839–15848, 2021. 11
2021
-
[30]
J. Hu, Y. Yang, X. Guo, B. Peng, H. Huang, and T. Ma. Decor-net: A covid-19 lung infection segmentation net- work improved by emphasizing low-level features and decorrelating features. In IEEE International Sympo- sium on Biomedical Imaging , pages 1–5, 2023. 11
2023
-
[31]
D. Jha, P. H. Smedsrud, M. A. Riegler, P. Halvorsen, T. d. Lange, D. Johansen, and H. D. Johansen. Kvasir- seg: A segmented polyp dataset. In Proceedings of In- ternational Conference on Multimedia Modeling , pages 451–462, 2020. 2, 8
2020
-
[32]
Ji, Y.-C
G.-P. Ji, Y.-C. Chou, D.-P. Fan, G. Chen, H. Fu, D. Jha, and L. Shao. Progressively normalized self-attention network for video polyp segmentation. In Proceedings of International Conference on Medical Image Com- puting and Computer-Assisted Intervention, pages 142– 152, 2021....
2021
-
[33]
Ji, D.-P
G.-P. Ji, D.-P. Fan, P. Xu, M.-M. Cheng, B. Zhou, and L. Van Gool. Sam struggles in concealed scenes– empirical study on” segment anything”. arXiv preprint arXiv:2304.06022, 2023. 2, 3, 6
2023 arXiv
-
[34]
W. Ji, J. Li, Q. Bi, T. Liu, W. Li, and L. Cheng. Seg- ment anything is not always perfect: An investigation of sam on different real-world applications. Machine Intelligence Research, 21:617–630, 2024. 2, 6
2024
-
[35]
L. Ke, M. Ye, M. Danelljan, Y. Liu, Y.-W. Tai, C.-K. Tang, and F. Yu. Segment anything in high quality. In Proceedings of International Conference and Workshop on Neural Information Processing Systems , 2024. 1, 5
2024
-
[36]
D. P. Kingma. Adam: A method for stochastic opti- mization. In Proceedings of International Conference on Learning Representations, 2015. 20
2015
-
[37]
Kirillov, E
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.- Y. Lo, et al. Segment anything. In Proceedings of IEEE International Conference on Computer Vision , pages 4015–4026, 2023. 1, 19
2023
-
[38]
Lachmann and C
T. Lachmann and C. Van Leeuwen. Individual pattern representations are context independent, but their col- lectiverepresentation is context dependent. The Quar- terly Journal of Experimental Psychology Section A , 58:1265–1294, 2005. 5
2005
-
[39]
Lamdouar, C
H. Lamdouar, C. Yang, W. Xie, and A. Zisserman. Be- trayed by motion: Camouflaged object discovery via mo- tion segmentation. 2020. 2
2020
-
[40]
T.-N. Le, T. V. Nguyen, Z. Nie, M.-T. Tran, and A. Sug- imoto. Anabranch network for camouflaged object seg- mentation. Computer Vision and Image Understand- ing, 184:45–56, 2019. 2, 6, 11
2019
-
[41]
Li and Y
G. Li and Y. Yu. Visual saliency based on multiscale deep features. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 5455– 5463, 2015. 2, 6, 11
2015
-
[42]
Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille. The secrets of salient object segmentation. In Proceed- ings of IEEE Conference on Computer Vision and Pat- tern Recognition, pages 280–287, 2014. 2, 6, 11
2014
-
[43]
Lian and H
S. Lian and H. Li. Evaluation of segment anything model 2: The role of sam2 in the underwater environ- ment. arXiv preprint arXiv:2408.02924 , 2024. 2, 3, 6
2024 arXiv
-
[44]
S. Lian, Z. Zhang, H. Li, W. Li, L. T. Yang, S. Kwong, and R. Cong. Diving into underwater: Segment anything model guided underwater salient instance segmentation and a large-scale dataset. In Proceedings of Interna- tional Conference on Machine Learning , 2024. 2
2024
-
[45]
J. Lin, Z. He, and R. W. Lau. Rich context aggregation with reflection prior for glass surface detection. In Pro- ceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 13415–13424, 2021. 5
2021
-
[46]
N. Liu, K. Nan, W. Zhao, X. Yao, and J. Han. Learning complementary spatial–temporal transformer for video salient object detection. IEEE Transactions on Neu- ral Networks and Learning Systems, 35(8):10663–10673,
-
[47]
W. Liu, X. Shen, C.-M. Pun, and X. Cun. Explicit vi- sual prompting for low-level structure segmentations. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 19434–19445, 2023. 5
2023
-
[48]
Y. Liu, M. Zhu, H. Li, H. Chen, X. Wang, and C. Shen. Matcher: Segment anything with one shot using all-purpose feature matching. arXiv preprint arXiv:2305.13310, 2023. 1
2023 arXiv
-
[49]
X. Lu, Y. Cao, S. Liu, C. Long, Z. Chen, X. Zhou, Y. Yang, and C. Xiao. Video shadow detection via spatio-temporal interpolation consistency training. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 3116–3125, 2022. 2, 7, 11, 12
2022
-
[50]
X. Lu, Y. Cao, S. Liu, C. Long, Z. Chen, X. Zhou, Y. Yang, and C. Xiao. Video shadow detection via spatio-temporal interpolation consistency training. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 3106–3115, 2022. 12
2022
-
[51]
Z. Luo, N. Liu, W. Zhao, X. Yang, D. Zhang, D.-P. Fan, F. Khan, and J. Han. Vscode: General visual salient and camouflaged object detection with 2d prompt learning. In Proceedings of IEEE Conference on Computer Vi- sion and Pattern Recognition, pages 17169–17180, 2024. 5
2024
-
[52]
Y. Lv, J. Zhang, Y. Dai, A. Li, B. Liu, N. Barnes, and D.-P. Fan. Simultaneously localize, segment and rank the camouflaged objects. In Proceedings of IEEE Con- ference on Computer Vision and Pattern Recognition , pages 11591–11601, 2021. 2, 6, 11
2021
-
[53]
J. Ma, Y. He, F. Li, L. Han, C. You, and B. Wang. Segment anything in medical images. Nature Commu- nications, 15(1):654, 2024. 1
2024
-
[54]
Margolin, L
R. Margolin, L. Zelnik-Manor, and A. Tal. How to eval- uate foreground maps? In Proceedings of IEEE Con- ference on Computer Vision and Pattern Recognition , pages 248–255, 2014. 8
2014
-
[55]
Martial, D
C. Martial, D. Stawarczyk, and A. D’Argembeau. Neural correlates of context-independent and context- dependent self-knowledge. Brain and Cognition , 125:23–31, 2018. 5
2018
-
[56]
Mehta, A
R. Mehta, A. Filos, U. Baid, C. Sako, R. McKinley, M. Rebsamen, K. Datwyler, R. Meier, P. Radojewski, G. K. Murugesan, et al. Qu-brats: Miccai brats 2020 challenge on quantifying uncertainty in brain tumor segmentation-analysis of ranking scores and benchmark- ing results. arX...
2020 arXiv
-
[57]
Mishra, R
P. Mishra, R. Verk, D. Fornasier, C. Piciarelli, and G. L. Foresti. Vt-adl: A vision transformer network for im- age anomaly detection and localization. In Proceedings of International Symposium on Industrial Electronics , pages 01–06, 2021. 2, 8, 11 Comprehensive Evaluation o...
2021
-
[58]
Y. Pang, X. Zhao, T.-Z. Xiang, L. Zhang, and H. Lu. Zoomnext: A unified collaborative pyramid network for camouflaged object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024. 5, 6, 11, 12
2024
-
[59]
Perazzi, J
F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition , pages 724– 732, 2016. 2, 6, 12, 13
2016
-
[60]
X. Qin, H. Dai, X. Hu, D.-P. Fan, L. Shao, and L. Van Gool. Highly accurate dichotomous image seg- mentation. In ECCV, pages 38–56, 2022. 2
2022
-
[61]
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R¨ adle, C. Rolland, L. Gustafson, et al. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 1, 19
2024 arXiv
-
[62]
K. Roth, L. Pemula, J. Zepeda, B. Sch¨ olkopf, T. Brox, and P. V. Gehler. Towards total recall in industrial anomaly detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition , pages 14298–14308, 2022. 11
2022
-
[63]
J. Ruan, S. Xiang, M. Xie, T. Liu, and Y. Fu. Malunet: A multi-attention and light-weight unet for skin le- sion segmentation. In IEEE International Conference on Bioinformatics and Biomedicine , pages 1150–1156,
-
[64]
J. Ruan, M. Xie, J. Gao, T. Liu, and Y. Fu. Ege-unet: An efficient group enhanced unet for skin lesion segmen- tation. In Proceedings of International Conference on Medical Image Computing and Computer-Assisted In- tervention, pages 481–490, 2023. 11
2023
-
[65]
A. N. Rumaksari, S. Sumpeno, and A. D. Wibawa. Background subtraction using spatial mixture of gaus- sian model with dynamic shadow filtering. In Proceed- ings of International Seminar on Intelligent Technology and Its Applications , pages 296–301, 2017. 5
2017
-
[66]
Sarica, D
B. Sarica, D. Z. Seker, and B. Bayram. A dense residual u-net for multiple sclerosis lesions segmentation from multi-sequence 3d MR images. International Journal of Medical Informatics , 170:104965, 2023. 12
2023
-
[67]
D. She, Y. Zhang, Z. Zhang, H. Li, Z. Yan, and X. Sun. Eoformer: Edge-oriented transformer for brain tumor segmentation. In Proceedings of International Con- ference on Medical Image Computing and Computer- Assisted Intervention, pages 333–343, 2023. 12
2023
-
[68]
Silva, A
J. Silva, A. Histace, O. Romain, X. Dray, and B. Granado. Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer. International Journal of Computer Assisted Radiology and Surgery, 9:283–293, 2014. 2, 8
2014
-
[69]
Skurowski, H
P. Skurowski, H. Abdulameer, J. B laszczyk, T. Depta, A. Kornacki, and P. Kozie l. Animal camouflage analysis: Chameleon database. Unpublished Manuscript, 2018. 2
2018
-
[70]
J. Sun, K. Xu, Y. Pang, L. Zhang, H. Lu, G. P. Hancke, and R. W. H. Lau. Adaptive illumination mapping for shadow detection in raw images. InProceedings of IEEE International Conference on Computer Vision , pages 12663–12672, 2023. 11
2023
-
[71]
Tajbakhsh, S
N. Tajbakhsh, S. R. Gurudu, and J. Liang. Automated polyp detection in colonoscopy videos using shape and context information. IEEE Transactions on Medical Imaging, 35(2):630–644, 2015. 2, 8
2015
-
[72]
F. Tang, L. Wang, C. Ning, M. Xian, and J. Ding. Cmu- net: A strong convmixer-based medical ultrasound im- age segmentation network. In IEEE International Sym- posium on Biomedical Imaging , pages 1–5, 2023. 11
2023
-
[73]
Tang and B
L. Tang and B. Li. Evaluating sam2’s role in cam- ouflaged object detection: From sam to sam2. arXiv preprint arXiv:2407.21596, 2024. 2
2024 arXiv
-
[74]
Z. Tu, T. Xia, C. Li, X. Wang, Y. Ma, and J. Tang. Rgb-t image saliency detection via collaborative graph learning. IEEE Transactions on Multimedia, 22(1):160– 173, 2019. 2
2019
-
[75]
V´ azquez, J
D. V´ azquez, J. Bernal, F. J. S´ anchez, G. Fern´ andez- Esparrach, A. M. L´ opez, A. Romero, M. Drozdzal, and A. Courville. A benchmark for endoluminal scene seg- mentation of colonoscopy images. Journal of Healthcare Engineering, 2017(1):4037190, 2017. 2, 8
2017
-
[76]
T. F. Y. Vicente, M. Hoai, and D. Samaras. Leave- one-out kernel optimization for shadow detection. In Proceedings of IEEE International Conference on Com- puter Vision , pages 3388–3396, 2015. 8
2015
-
[77]
T. F. Y. Vicente, L. Hou, C.-P. Yu, M. Hoai, and D. Samaras. Large-scale training of shadow detectors with noisily-annotated shadow examples. In Proceedings of European Conference on Computer Vision , pages 816–832, 2016. 2, 7, 11
2016
-
[78]
J. Wang, X. Li, and J. Yang. Stacked conditional gen- erative adversarial networks for jointly learning shadow detection and shadow removal. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 1788–1797, 2018. 2, 7, 11
2018
-
[79]
L. Wang, H. Lu, Y. Wang, M. Feng, D. Wang, B. Yin, and X. Ruan. Learning to detect salient objects with image-level supervision. In Proceedings of IEEE Con- ference on Computer Vision and Pattern Recognition , pages 136–145, 2017. 2, 6, 11, 13
2017
-
[80]
M. Wang, Y. Feng, Y. Tang, T. Zhang, Y. Liang, and C. Lv. Global-local medical sam adaptor based on full adaption. arXiv preprint arXiv:2409.17486 , 2024. 1
2024
-
[81]
X. Wang, W. Wang, Y. Cao, C. Shen, and T. Huang. Im- ages speak in images: A generalist painter for in-context visual learning. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 6830– 6839, 2023. 1
2023
-
[82]
X. Wang, X. Zhang, Y. Cao, W. Wang, C. Shen, and T. Huang. Seggpt: Towards segmenting everything in context. In Proceedings of IEEE International Confer- ence on Computer Vision , pages 1130–1140, 2023. 1, 5, 10, 11, 19
2023
-
[83]
Y. Wang, R. Wang, X. Fan, T. Wang, and X. He. Pixels, regions, and objects: Multiple enhancement for salient object detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition , pages 10031–10040, 2023. 11
2023
-
[84]
J. Wei, Y. Hu, S. Cui, S. K. Zhou, and Z. Li. Weakpolyp: You only look bounding box for polyp segmentation. In Proceedings of International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 757–766, 2023. 11
2023
-
[85]
J. Wu, W. Ji, Y. Liu, H. Fu, M. Xu, Y. Xu, and Y. Jin. Medical sam adapter: Adapting segment any- thing model for medical image segmentation. arXiv preprint arXiv:2304.12620, 2023. 1
2023 arXiv
-
[86]
Y. Wu, Y. Liu, L. Zhang, M. Cheng, and B. Ren. EDN: salient object detection via extremely-downsampled net- work. IEEE Transactions on Image Processing , pages 3125–3136, 2022. 11
2022
-
[87]
E. Xie, W. Wang, W. Wang, M. Ding, C. Shen, and P. Luo. Segmenting transparent objects in the wild. 24 Xiaoqi Zhao 1 et al. In Proceedings of European Conference on Computer Vision, pages 696–711, 2020. 2, 7, 11
2020
-
[88]
H. Xing, S. Gao, Y. Wang, X. Wei, H. Tang, and W. Zhang. Go closer to see better: Camouflaged ob- ject detection via object area amplification and figure- ground conversion. IEEE Transactions on Circuits and Systems for Video Technology, 33(10):5444–5457, 2023. 6, 11
2023
-
[89]
Q. Yan, L. Xu, J. Shi, and J. Jia. Hierarchical saliency detection. In Proceedings of IEEE Conference on Com- puter Vision and Pattern Recognition, pages 1155–1162,
-
[90]
C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang. Saliency detection via graph-based manifold ranking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 3166–3173, 2013. 2, 6, 11
2013
-
[91]
H. Yang, T. Wang, X. Hu, and C. Fu. SILT: shadow- aware iterative label tuning for learning to detect shad- ows from noisy labels. In Proceedings of IEEE Interna- tional Conference on Computer Vision , pages 12641– 12652, 2023. 11
2023
-
[92]
X. Yuan, G. Cheng, K. Yan, Q. Zeng, and J. Han. Small object detection via coarse-to-fine proposal generation and imitation learning. In Proceedings of IEEE Inter- national Conference on Computer Vision , pages 6317– 6327, 2023. 11
2023
-
[93]
Zhang, D.-P
J. Zhang, D.-P. Fan, Y. Dai, X. Yu, Y. Zhong, N. Barnes, and L. Shao. Rgb-d saliency detection via cascaded mu- tual information minimization. In Proceedings of IEEE International Conference on Computer Vision , pages 4338–4347, 2021. 2
2021
-
[94]
Zhang, P
R. Zhang, P. Lai, X. Wan, D. Fan, F. Gao, X. Wu, and G. Li. Lesion-aware dynamic kernel for polyp segmen- tation. In Proceedings of International Conference on Medical Image Computing and Computer-Assisted In- tervention, pages 99–109, 2022. 11
2022
-
[95]
X. Zhao, H. Jia, Y. Pang, L. Lv, F. Tian, L. Zhang, W. Sun, and H. Lu. M 2snet: Multi-scale in multi-scale subtraction network for medical image segmentation. arXiv preprint arXiv:2303.10894 , 2023. 12
2023 arXiv
-
[96]
X. Zhao, H. Liang, P. Li, G. Sun, D. Zhao, R. Liang, and X. He. Motion-aware memory network for fast video salient object detection. IEEE Transactions on Image Processing, 33:709–721, 2024. 12
2024
-
[97]
X. Zhao, Y. Pang, Z. Chen, Q. Yu, L. Zhang, H. Liu, J. Zuo, and H. Lu. Towards automatic power battery de- tection: New challenge benchmark dataset and baseline. In Proceedings of IEEE Conference on Computer Vi- sion and Pattern Recognition, pages 22020–22029, 2024. 2, 8, 11
2024
-
[98]
X. Zhao, Y. Pang, W. Ji, B. Sheng, J. Zuo, L. Zhang, and H. Lu. Spider: A unified framework for context- dependent concept understanding. In Proceedings of International Conference on Machine Learning , pages 60906–60926, 2024. 1, 3, 5, 8, 10, 11, 19
2024
-
[99]
X. Zhao, Y. Pang, L. Zhang, H. Lu, and L. Zhang. To- wards diverse binary segmentation via a simple yet gen- eral gated network. International Journal of Computer Vision, 132:4157–4234, 2024. 5
2024
-
[100]
X. Zhao, L. Zhang, and H. Lu. Automatic polyp seg- mentation via multi-scale subtraction network. In Pro- ceedings of International Conference on Medical Image Computing and Computer-Assisted Intervention , pages 120–130, 2021. 5
2021
-
[101]
T. Zhou, Y. Zhang, Y. Zhou, Y. Wu, and C. Gong. Can sam segment polyps? arXiv preprint arXiv:2304.07583,
-
[102]
Y. Zou, J. Jeong, L. Pemula, D. Zhang, and O. Dabeer. Spot-the-difference self-supervised pre- training for anomaly detection and segmentation. In Proceedings of European Conference on Computer Vi- sion, pages 392–408, 2022. 2, 8
2022
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.