{"id":"977c18b1-177a-4e26-90cd-e8660e8d08db","arxiv_id":"2412.01240","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A unified evaluation of SAM and SAM 2 on 11 context-dependent concepts over 33 datasets shows box prompts dominate, SAM 2 lags SAM in some static-image settings, and both are prompt-sensitive.","lead":"This paper reports a large benchmark comparing SAM and SAM 2 on 11 context-dependent segmentation concepts across 33 datasets in natural, medical, and industrial scenes, using images, video, and 3D scans. It finds box prompts dominate, SAM 2 helps in video and 3D but sometimes hurts on static images, and both models are sensitive to imperfect prompts.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline performance boundaries are oracle-conditioned: everything-mode and point-mode results use GT overlap filtering and GT-guided clicks, so the Sec. 3.9 claims (box dominance, SAM 2 everything weakness) need a GT-free re-run before supporting the abstract's open-world framing.","rationale":"The reader's weakest assumption is exactly the oracle-conditioned protocol: GT-derived prompts and GT-filtered outputs are treated as the right way to measure SAMs' CD segmentation ability. My independent reading confirms this is the most load-bearing concern because it touches the majority of headline claims—box-prompt dominance, SAM 2's everything-mode deficit, and the 3D superiority claim all inherit GT information. The paper is transparent about the protocols, and the released code makes a GT-free re-run feasible, so this is a correctable overclaim rather than a fatal flaw. The separate Tab. 15a baseline duplication is a real local artifact but does not threaten the benchmark's overall structure as much as the protocol issue. With the current wording, CONDITIONAL remains the right verdict: the paper should either add GT-free evidence or rephrase the conclusions as an ideal-prompt upper-bound evaluation.","tokens_in":31751,"tokens_out":5448,"duration_ms":55878,"concrete_test":"Recompute Tabs. 3-9 with a GT-free everything-mode scorer: for each image, aggregate all predicted masks by max-pooling (or by SAM's own confidence ranking) and compute the reported metrics, without any >90% GT-overlap filtering. As a control, repeat the OFS with thresholds 80% and 95%. If SAM 2's everything-mode deficit relative to SAM shrinks or reverses, or if the box-vs-everything ordering changes, the Sec. 3.5/3.9 conclusions are artifacts of the oracle filter. A second half of the same check: score point mode with a single click at the GT centroid instead of the iterative IoU>=0.9 oracle loop; if box no longer dominates point in most concepts, the 'box prompts generally most advantageous' claim is protocol-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claims (Sec. 3.9 I, II, V) rest on numbers produced under ground-truth-conditioned protocols: automatic everything mode keeps only predicted entities with >90% overlap with the GT mask (Sec. 3.4.1); point mode adds corrective clicks at the largest GT error region until IoU reaches 0.9 (max 6 clicks) (Sec. 3.4.1); the 3D anchor slice is chosen as the one with the largest GT foreground (Sec. 3.4.4). Thus the reported 'performance boundaries' are upper envelopes under idealized prompt feedback, not the zero-shot or realistic interactive behavior the abstract and Sec. 1 imply. In particular, 'SAM 2 performs worse on everything' may reflect OFS's sensitivity to mask granularity rather than an intrinsic deficit: SAM 2's automatic masks can be more fragmented/instance-like, so fewer satisfy the >90% per-entity overlap test, suppressing recall even when the union covers the object. Likewise, the 'box prompt is generally most advantageous' conclusion compares GT boxes with oracle-corrected point clicks; with a single fixed click or realistic noisy clicks the ordering could differ. Section 3.8 perturbs GT prompts but does not remove GT conditioning, so it does not rescue the headline. This is not an internal contradiction—the protocols are clearly stated—but it means the central claim as written is broader than what the tables establish. Section 5's limitations list dataset constraints but do not flag this GT-conditioning.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a large-scale empirical evaluation of SAM and SAM 2 on eleven context-dependent (CD) segmentation concepts across 33 datasets, covering 2D images, videos, and 3D medical volumes in natural, medical, and industrial scenes. The authors propose a unified prompting framework that includes everything, box, point, mask, in-context learning, and prompt-robustness modes, and they distill the results into design insights for a future SAM 3. The headline conclusions are that box prompts are generally the most advantageous prompt type, that SAM 2 is not uniformly better than SAM and is worse on everything and point modes, that SAM 2 shows in-context-learning potential, and that bidirectional inference with multi-frame mask prompts enables SAM 2 to surpass specialized 3D medical segmentation models.","tokens_in":32033,"tokens_out":5102,"duration_ms":45598,"significance":"The benchmark is timely and potentially useful: it is the most comprehensive CD-concept evaluation of SAM and SAM 2 that I am aware of, it covers underexplored medical and industrial modalities, and it ships a unified evaluation toolkit and code repository, which supports reproducibility. The robustness analysis in Section 3.8 and the in-context-learning exploration in Section 3.4.2 go beyond earlier SAM evaluations and provide practically relevant information for deployment. If the headline findings were established by the reported data, they would give the community a reference map for prompt selection on CD concepts and a concrete evidence base for SAM 3 design; that is a meaningful contribution. However, as detailed below, several load-bearing conclusions outrun the ground-truth-conditioned protocols that produced the numbers, and one stated conclusion is contradicted by the paper's own tables.","major_comments":[{"comment":"The headline performance conclusions are derived under ground-truth-conditioned protocols rather than zero-shot or realistic interactive conditions. In automatic everything mode, only predicted entities whose overlap with the GT is greater than 90% are retained and merged (Sec. 3.4.1); in point mode, corrective clicks are placed at the largest GT-error region until IoU reaches 0.9 or six clicks are used (Sec. 3.4.1); and in 3D, the anchor slice is selected as the one with the largest GT foreground (Sec. 3.4.4). Consequently, claims such as \"box prompts are generally the most advantageous\" and \"SAM 2 performs worse on everything and point prompts\" describe behavior under oracle feedback, not the open-world or zero-shot behavior implied by the abstract and Section 1. The everything-mode comparison may also be confounded by the OFS granularity test: SAM 2's more fragmented automatic masks would be discarded by the >90% per-entity overlap rule even when their union is correct. Section 5's limitations do not disclose this GT-conditioning. Please either re-run automatic and point modes without GT-based filtering/correction (or with a realistic interaction model), re-evaluate the 3D anchor choice without GT, or restrict the conclusions to the oracle-conditioned setting and revise the abstract and Section 3.9 accordingly.","section":"§3.4.1 and §3.9, Items I, II, V"},{"comment":"The statement that \"SAM 2 (point) and SAM 2 (everything) are consistently weaker than their corresponding SAM variants\" is contradicted by the paper's own tables. For point mode, Table 4 reports COD10K F_beta of 0.864 for SAM 2 versus 0.823 for SAM, and Table 9 reports SAM 2 with higher Dice on COVID-19 (0.687 vs 0.352), BUSI (0.783 vs 0.694), ISIC-2018 (0.641 vs 0.504), and Polyp-Five (0.862 vs 0.641). For everything mode, the comparison is also uneven across tasks and datasets. The claim in Section 3.9 Item II therefore needs to be revised into a concept- and dataset-specific statement, or the tables must be corrected.","section":"§3.5 and Tables 4, 9"},{"comment":"The specialist baselines 3D U-Net and EoFormer are reported with identical Dice values across all four MRI modalities (Flair, T1ce, T1, T2); for example, 3D U-Net shows 0.900/0.807/0.792 repeated verbatim in every column. Identical multi-modality results are implausible and suggest the scores were transcribed from a modality-independent source or copied incorrectly. Since Section 3.7 and Section 3.9 Item V use this comparison to claim that SAM 2 surpasses specialized models such as DRU-Net and 3D U-Net, the baseline values must be verified against the original publications and corrected; otherwise the superiority claim is unsupported.","section":"§3.7, Table 15(a)"}],"minor_comments":[{"comment":"The OFS overlap threshold of 90% is introduced without any sensitivity analysis; a threshold sweep (for example, 50%, 75%, 90%) would clarify how much the everything-mode rankings depend on this arbitrary choice.","section":"§3.4.1"},{"comment":"The meaning of \"—\" differs across tables: in Table 7 it indicates zero prediction accuracy per the PBD protocol, while in Table 9 it means the specialized model does not support that lesion type. A unified caption-level explanation would improve readability.","section":"Tables 3–15"},{"comment":"The claim that SAM's image-based predictions \"fluctuate frame by frame\" is supported only by qualitative visualizations; a temporal consistency metric (for example, average inter-frame IoU of predicted masks) would make the comparison quantitative.","section":"§3.6"},{"comment":"References [13] and [14] appear to be the same paper (Implicit Motion Handling for Video Camouflaged Object Detection) with different page ranges; this looks like a duplicate citation.","section":"References [13, 14]"},{"comment":"The prompt-type icons (for example, \"/buromobelexperte\", \"/d⌢t-circle\", \"/border-s◎yle\") render as odd symbols in the text; please use standard notation or clearly typeset glyphs in the final version.","section":"Prompt icons throughout"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision. The GT-conditioned protocols are clearly described, so this is not a case of hidden methodology, but the abstract and Section 3.9 frame the results as open-world performance boundaries and Section 5 does not flag the conditioning; that gap is load-bearing. The contradiction between Section 3.5 and Tables 4/9 is straightforward and must be fixed, and Table 15(a) needs verification. The benchmark itself and the released code are valuable enough to warrant a revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my read of arXiv:2412.01240. It's a large, genuinely useful evaluation study, but the headline performance boundaries are oracle-conditioned. The everything-mode numbers keep only predicted entities with >90% overlap with the GT mask; point mode adds corrective clicks at the largest GT error region until IoU reaches 0.9; the 3D anchor slice is chosen as the one with the largest GT foreground. All of that is clearly stated in Sec. 3.4, so it's not hidden, but it means the abstract and Sec. 3.9 claims describe an upper envelope under idealized feedback, not zero-shot or realistic interactive behavior.\n\nWhat's genuinely new: the breadth is real—11 CD concepts, 33 datasets, natural/medical/industrial scenes, 2D/video/3D—plus three protocol elements I haven't seen in prior SAM evaluations: SAM 2 in-context-learning inference (Sec. 3.4.2), prompt robustness testing (Sec. 3.8), and bidirectional 3D inference (Sec. 3.4.4). The code, evaluation toolkits, and standard datasets are public. That is reproducible infrastructure and a solid reference map for where SAMs work and where they fail.\n\nSoft spots, in order of severity. First, the GT-conditioned protocols: the Sec. 3.9 conclusions—especially box dominance and SAM 2's everything-mode weakness—need a GT-free re-run before they can be framed as open-world deployment boundaries. I'd like to see a version with a single fixed click, no GT filtering, and a fixed anchor slice. Second, there's an internal contradiction: Sec. 3.5 and Sec. 3.9 say SAM 2 is consistently weaker than SAM with point and everything prompts, but Tab. 4 (COD10K point F-beta 0.864 vs 0.823) and Tab. 9 (LOS point Dice 0.687 vs 0.352 on COVID, 0.783 vs 0.694 on BUSI) show SAM 2 point beating SAM point. That's a load-bearing summary claim that doesn't match its own tables. Third, Tab. 15a lists identical baseline scores for 3D U-Net and EoFormer across all four BraTS modalities—suspicious enough to need a check. Fourth, Tab. 16 has internally odd deltas: some perturbations improve results (e.g., SAM 2 box P-AP on MVTec +28%), which conflicts with the 'consistent performance drops' narrative in the text. The Limitations section is candid about unimplemented outlooks but doesn't mention these protocol and table issues.\n\nWho this is for: practitioners who want a broad, if idealized, reference for deploying SAM/SAM 2 on saliency, camouflage, medical lesions, and industrial defects, and who want to know the failure cases (shadows, tiny targets, non-material concepts). It deserves a serious referee—the empirical map is valuable—but it needs revision: fix the summary contradictions, either re-run everything/point under GT-free protocols or reframe the claims as oracle-assisted upper bounds, and verify the baseline transcriptions.\n\nMy recommendation: send it to peer review with major revision requested. I'd also bring it to a reading group to dissect the protocols, because the paper is a good teaching case for why evaluation design shapes conclusions.","headline":"Broad, useful SAM/SAM2 evaluation map, but headline performance boundaries are GT-conditioned upper envelopes and two summary claims contradict the paper's own tables.","tokens_in":32730,"tokens_out":2317,"would_cite":true,"duration_ms":21548,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This evaluation maps SAM and SAM 2's performance boundaries on 11 context-dependent segmentation concepts, finding box prompts generally best and SAM 2's edge confined to video and 3D.","keywords":["Segment Anything Model","SAM 2","context-dependent concepts","prompt robustness","in-context learning","segmentation evaluation","bidirectional inference","benchmark"],"falsifier":"Rerun the headline settings on a held-out subset of these datasets with no oracle: prompt points and boxes from an off-the-shelf detector or saliency model, no overlap filtering, and no ground-truth-based anchor selection; if SAM 2's everything-mode and point-mode deficits and its 3D advantage shrink or reverse, the central conclusions are artifacts of ground-truth-conditioned prompting. Independently, re-extract the specialist baselines in Table 15a from the original papers, since identical scores across all four modalities would indicate transcription error.","tokens_in":1706,"feed_emoji":"🧩","tokens_out":1810,"duration_ms":82134,"temperature":0.7,"pith_summary":"This paper tries to establish where the Segment Anything Model (SAM) and its video-capable successor SAM 2 actually stand when the things to segment are context-dependent: visual saliency, camouflage, shadows, transparent objects, medical lesions, and industrial defects. Across 33 datasets, 11 concepts, and image, video, and 3D data, the authors build a unified prompting framework and produce a performance map. Their central findings are that box prompts are the most reliable prompt type, that SAM 2 does not uniformly beat SAM and is weaker on everything-mode and point-prompt tasks, and that SAM 2's temporal memory becomes a decisive advantage in video and 3D once multiple mask prompts and bidirectional inference are used. The value is a reference baseline for deploying SAMs on real-world context-dependent tasks and guidance for the next generation of the model.","feed_headline":"SAM 2 lags SAM on point and 'everything' prompts","feed_subtitle":"A 33-dataset study maps where both models fail on camouflage, shadows, defects, and medical lesions.","key_machinery":"The load-bearing mechanism is the unified evaluation framework itself: six prompting settings (everything, mask, box, point, in-context learning, and prompt robustness) applied to SAM and SAM 2 through three prompt-generation strategies. The prediction-based propagated prompt turns the previous frame's predicted mask into the next frame's prompt, letting an image-trained SAM behave as a video segmenter; the bidirectional inference strategy picks the slice with the largest ground-truth foreground mask as an anchor and propagates masks in both directions through a 3D volume; and the in-context learning mode feeds SAM 2 the first 20 training images and masks of a concept as exemplars. The everything-mode outputs are filtered by an overlap filtering strategy that keeps only predicted entities whose overlap with the ground truth exceeds 90%. These mechanisms together convert frozen, off-the-shelf SAMs into measurable systems across heterogeneous data without task-specific fine-tuning.","core_discovery":"The paper's central claim is that SAM and SAM 2 have clear, measurable performance boundaries on context-dependent concepts, and that those boundaries depend more on prompt type, data modality, and target material than on model scale. In static images, box prompts dominate on 10 of the 11 concepts and rescue performance on hard cases like camouflage and medical lesions. In everything mode and point-click mode, SAM 2 often performs worse than SAM, which the authors attribute to temporal memory and dynamic prompt attention adding unfavorable bias when no temporal context is available. In video and 3D, the picture flips: SAM 2 with prediction-based propagated prompts, multi-frame mask prompts, and a bidirectional inference strategy for volumetric data surpasses task-specific medical segmentors on brain tumor and multiple sclerosis lesion segmentation. The paper also claims SAM 2 shows real but incomplete in-context learning ability, and that both models are highly sensitive to imperfect prompts, degrading sharply under small perturbations.","pith_inferences":["If the oracle-conditioned prompts inflate performance, then the real-world gap between SAMs and specialized models on camouflage, shadows, and medical lesions is larger than the tables suggest; a fully automatic prompt-setting evaluation would test this directly.","The finding that SAM 2 underperforms SAM on everything mode in static images hints that the memory module or dynamic prompt attention introduces a temporal prior that is harmful without temporal input; inspecting memory-bank activity on single frames could localize the cause.","The box-prompt dominance across 10 of 11 concepts suggests SAMs act essentially as strong object proposal matchers: given spatial bounds they segment well, but their ability to discover context-dependent regions on their own is limited, so adding a saliency or anomaly proposal module as an automatic prompt generator could be a practical route toward a better next-generation model.","The 3D success with bidirectional multi-frame mask prompts implies that medical volumes can be treated as videos without architectural changes; testing on CT and MRI modalities beyond the ones reported would establish whether that advantage generalizes."],"forward_implications":["Box prompts should be the default interface when deploying SAMs on context-dependent image tasks; point and everything modes should be treated as unreliable outside well-separated objects.","SAM 2's temporal memory pays off only when there is temporal structure; for single-frame inference, the original SAM remains competitive or better, so upgrading to SAM 2 is not automatically an improvement.","In video and 3D medical segmentation, SAM 2 with mask prompts and multi-frame propagation can serve as a strong zero-shot baseline, competitive with or better than task-specific networks.","Evaluation practices should include prompt-robustness testing, because small perturbations to boxes, points, or masks move scores by percentages that change qualitative conclusions.","The in-context learning results suggest that a few exemplars can steer SAM 2 toward new concepts, though not yet to the level of existing unified models."],"supporting_citations":[{"why":"It supplies the original model whose pretrained weights and prompt interfaces the evaluation exercises.","marker":"[37]"},{"why":"It supplies the successor model whose temporal memory and mask-prompt video interface the evaluation probes.","marker":"[61]"},{"why":"It supplies the concept taxonomy and the unified-model results used as baselines in the in-context learning mode.","marker":"[98]"},{"why":"It supplies one of the generalist in-context segmentors compared in the in-context learning experiments.","marker":"[82]"},{"why":"It supplies a medical generalist segmentor used as a baseline for the lesion segmentation comparisons.","marker":"[6]"},{"why":"It supplies the power-battery X-ray dataset and location-and-count metrics that ground the findings on tiny targets.","marker":"[97]"},{"why":"It supplies the volumetric brain-tumor benchmark used to demonstrate the bidirectional inference advantage.","marker":"[56]"},{"why":"It supplies the multiple-sclerosis lesion benchmark used to claim performance beyond a strong specialized network.","marker":"[8]"},{"why":"It supplies the industrial anomaly dataset and the AUROC/PRO metrics used in the surface-defect failure claims.","marker":"[3]"}],"fun_headline_variants":["SAM 2 point and auto prompts underperform SAM","Box prompts beat everything mode on hard segmentation","SAM 2 excels in video, flops on static point prompts","Prompt type trumps model scale for SAM failures"],"cache_read_input_tokens":34560,"weakest_assumption_plain":"The headline numbers assume that ground-truth-derived prompts and ground-truth-filtered outputs are the right way to measure SAMs' context-dependent segmentation ability, and if those oracle conditions are removed, the reported performance boundaries may not describe realistic zero-shot or interactive use.","fun_headline_variants_meta":{"raw":{"variants":["SAM 2 point and auto prompts underperform SAM","Box prompts beat everything mode on hard segmentation","SAM 2 excels in video, flops on static point prompts","Prompt type trumps model scale for SAM failures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000302,"raw_usage":{"total_tokens":1781,"prompt_tokens":1031,"completion_tokens":750,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":687}},"tokens_in":647,"tokens_out":750,"duration_ms":7507,"temperature":1.0,"reasoning_tokens":687,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:33:35.328325+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the headline settings on a held-out subset of these datasets with no oracle: prompt points and boxes from an off-the-shelf detector or saliency model, no overlap filtering, and no ground-truth-based anchor selection; if SAM 2's everything-mode and point-mode deficits and its 3D advantage shrink or reverse, the central conclusions are artifacts of ground-truth-conditioned prompting. Independently, re-extract the specialist baselines in Table 15a from the original papers, since identical scores across all four modalities would indicate transcription error.","supporting_citations":[{"cited_title":"Kirillov, E","cited_arxiv_id":null,"evidence_quote":"It supplies the original model whose pretrained weights and prompt interfaces the evaluation exercises."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the concept taxonomy and the unified-model results used as baselines in the in-context learning mode."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies one of the generalist in-context segmentors compared in the in-context learning experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the power-battery X-ray dataset and location-and-count metrics that ground the findings on tiny targets."},{"cited_title":"QU-BraTS: MICCAI BraTS 2020 Challenge on Quantifying Uncertainty in Brain Tumor Segmentation - Analysis of Ranking Scores and Benchmarking Results","cited_arxiv_id":"2112.10074","evidence_quote":"It supplies the volumetric brain-tumor benchmark used to demonstrate the bidirectional inference advantage."}],"review_version":1}