Pith. sign in

REVIEW 2 major objections 2 minor 64 references

ARTEMIS evolves reliable temporal masks via a vision-language agent to segment polyps in videos from points, scribbles or few dense labels.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-26 18:09 UTC pith:D2ORURFY

load-bearing objection ARTEMIS adds a debate-and-judge VLM agent plus a prototype transport module to SAM2 propagation for weak-supervision polyp video segmentation, but the abstract supplies no numbers so the gains remain unverified. the 2 major comments →

arxiv 2606.20161 v1 pith:D2ORURFY submitted 2026-06-18 cs.CV

ARTEMIS: Agent-guided Reliability-aware Temporal Mask Evolution for Imperfectly Supervised Video Polyp Segmentation

classification cs.CV
keywords video polyp segmentationimperfect supervisiontemporal mask evolutionvision-language agentreliability-aware learningweakly supervised segmentationSAM2
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper presents ARTEMIS as a framework that learns dense, temporally consistent polyp masks from inexpensive supervision rather than full pixel labels. It begins with coarse masks generated by SAM2 from points or scribbles, then deploys a debate-and-judge vision-language agent to identify reliable anchor frames. These anchors are propagated bidirectionally to fill in unreliable frames, after which a segmenter is trained with reliability-guided reference selection, a prototype transport module, and a robust loss that down-weights noisy masks. The approach targets clinical video polyp segmentation where motion blur, weak contrast, and sparse guidance make standard pseudo-labeling unreliable. If the method works as described, it would allow accurate segmentation while using far less expensive annotation than dense labeling.

Core claim

ARTEMIS initializes coarse masks from weak supervision or dense anchors, employs a vision-language agent to select reliable temporal anchors under weak supervision, propagates those anchors bidirectionally with SAM2 to refine other frames, and trains the final segmenter with reliability-aware components including reference prototype transport and robust loss that down-weights noisy supervision instead of discarding samples. Experiments on SUN-SEG and CVC-ClinicDB-612 under scribble, point, and limited-label settings show state-of-the-art performance.

What carries the argument

The vision-language agent that debates and judges to select reliable temporal anchors from initial coarse masks, which are then bidirectionally propagated and used in reliability-aware robust learning.

Load-bearing premise

The vision-language agent can reliably identify and select high-quality temporal anchors from the initial coarse masks generated under weak supervision.

What would settle it

Running the full ARTEMIS pipeline versus a version without the agent on SUN-SEG or CVC-ClinicDB-612 under point or scribble supervision and finding no performance gain or a reversal of the reported state-of-the-art ranking.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Coarse masks from SAM2 can be refined over time instead of used directly, reducing boundary leakage and geometry degradation.
  • Reliability assessment allows the training process to down-weight noisy frames rather than discard difficult samples.
  • Bidirectional propagation from selected anchors improves temporal consistency across video frames.
  • The Reference Prototype Transport Module maintains target identity when moving from reference to target frames.
  • State-of-the-art results hold across scribble, point, and limited dense-label settings on the tested polyp video datasets.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same agent-plus-propagation loop could be tested on other medical video tasks that suffer from sparse labels, such as lesion tracking in endoscopy.
  • If the reliability scores prove stable, they might be used at inference time to flag frames for clinician review instead of accepting every output mask.
  • The framework's dependence on SAM2 suggests that swapping in newer mask generators could be a direct next step without changing the rest of the pipeline.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper proposes ARTEMIS, a unified framework for imperfectly supervised video polyp segmentation. It initializes coarse masks via SAM2 from points/scribbles or uses dense labels as anchors, employs a debate-and-judge vision-language agent to select reliable temporal anchors under weak supervision, propagates them bidirectionally with SAM2, and trains a segmenter via reliability-guided reference selection, a Reference Prototype Transport Module, and a reliability-aware robust loss that down-weights noisy samples. Experiments on SUN-SEG and CVC-ClinicDB-612 claim state-of-the-art performance under scribble, point, and limited-label settings.

Significance. If the experimental claims hold with proper validation, the work addresses a clinically relevant setting where dense annotations are costly. The combination of VLM-based reliability assessment, temporal propagation, and robust loss could offer a practical advance over direct pseudo-labeling from SAM2, particularly for handling boundary leakage and motion artifacts in polyp videos. Releasing code is a positive step toward reproducibility.

major comments (2)
  1. [Method (agent-guided reliability assessment and reliability-aware robust loss)] The central assumption that the vision-language agent can reliably judge mask quality from visual/textual cues without additional supervision (stated in the method description) is load-bearing for both anchor selection and loss weighting; the manuscript should provide an independent validation (e.g., human study or correlation with ground-truth boundary metrics) rather than relying solely on downstream segmentation gains.
  2. [Experiments] The abstract and experimental claims assert SOTA results on SUN-SEG and CVC-ClinicDB-612, yet the supplied description contains no quantitative metrics, ablation tables, or error analysis; the full manuscript must include these (with comparisons to recent weak-supervision baselines) to substantiate that the agent-driven components drive the reported improvements rather than the base SAM2 initialization.
minor comments (2)
  1. [Method] Notation for reliability scores and loss coefficients should be defined explicitly with equations to avoid ambiguity in how they interact with the Reference Prototype Transport Module.
  2. [Abstract] The abstract would be strengthened by including one or two key quantitative results (e.g., Dice or Jaccard improvements) alongside the SOTA claim.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback and the recognition of the clinical relevance of imperfectly supervised video polyp segmentation. We address each major comment below.

read point-by-point responses
  1. Referee: [Method (agent-guided reliability assessment and reliability-aware robust loss)] The central assumption that the vision-language agent can reliably judge mask quality from visual/textual cues without additional supervision (stated in the method description) is load-bearing for both anchor selection and loss weighting; the manuscript should provide an independent validation (e.g., human study or correlation with ground-truth boundary metrics) rather than relying solely on downstream segmentation gains.

    Authors: We agree that an independent validation of the agent's mask-quality judgments would strengthen the manuscript. The current version relies on downstream segmentation gains and component ablations to demonstrate utility. In revision we will add a correlation analysis: on the limited-label subset we will report agreement (e.g., Pearson correlation or rank correlation) between the agent's reliability scores and ground-truth boundary metrics (Dice, boundary F-score) computed against the dense labels. This provides direct evidence independent of the final segmenter performance. revision: partial

  2. Referee: [Experiments] The abstract and experimental claims assert SOTA results on SUN-SEG and CVC-ClinicDB-612, yet the supplied description contains no quantitative metrics, ablation tables, or error analysis; the full manuscript must include these (with comparisons to recent weak-supervision baselines) to substantiate that the agent-driven components drive the reported improvements rather than the base SAM2 initialization.

    Authors: The full manuscript contains Tables 1–2 with quantitative Dice, IoU and HD95 results under point, scribble and limited-label protocols on both datasets, plus direct comparisons against recent weak-supervision and SAM2-based baselines. Table 3 isolates the contribution of the debate-and-judge agent, Reference Prototype Transport Module and reliability-aware loss via controlled ablations. We will ensure these tables, together with any additional error analysis or statistical tests, are clearly presented and highlighted in the revised version so that the incremental benefit of the agent-driven components is evident. revision: yes

Circularity Check

0 steps flagged

No significant circularity identified

full rationale

The paper presents an empirical framework for imperfectly supervised video polyp segmentation, with the core claims resting on experimental SOTA results across external datasets (SUN-SEG, CVC-ClinicDB-612) under multiple weak-supervision regimes. The method description outlines an agent-guided pipeline and reliability-aware components, but the supplied text contains no equations, derivations, or self-citations that reduce any claimed prediction or result to its own inputs by construction. No self-definitional loops, fitted inputs renamed as predictions, or load-bearing self-citation chains are present. The experimental validation therefore remains independent of internal definitions.

Axiom & Free-Parameter Ledger

2 free parameters · 2 axioms · 0 invented entities

The central claim rests on the untested premise that the vision-language agent produces trustworthy reliability scores and that SAM2 propagation preserves geometry better than direct pseudo-labeling. No free parameters are explicitly listed in the abstract, but the reliability threshold and loss weighting coefficients are implicit fitted choices. No new physical entities are postulated.

free parameters (2)
  • reliability threshold for anchor selection
    Used by the agent to decide which masks become temporal anchors; value not stated in abstract.
  • loss weighting coefficients in reliability-aware robust loss
    Control how much noisy masks are down-weighted; chosen to optimize reported performance.
axioms (2)
  • domain assumption SAM2 can propagate masks bidirectionally while preserving polyp geometry better than direct pseudo-labeling from weak inputs.
    Invoked in the propagation step described in the abstract.
  • ad hoc to paper The vision-language agent can judge mask reliability from visual and textual cues without additional supervision.
    Central to the debate-and-judge mechanism.

pith-pipeline@v0.9.1-grok · 5839 in / 1558 out tokens · 18333 ms · 2026-06-26T18:09:08.821049+00:00 · methodology

0 comments
read the original abstract

Imperfectly supervised video polyp segmentation (VPS) aims to learn dense, temporally consistent masks from inexpensive supervision, including weak annotations (points, scribbles) and semi-supervision with few densely labeled frames. This setting is clinically valuable but challenging due to weak contrast, ambiguous boundaries, motion blur, and specular highlights, compounded by sparse pixel-level guidance. While SAM2 can generate dense masks from sparse inputs, direct pseudo-labeling often yields geometry-degraded masks with boundary leakage, underutilizes temporal consistency, and ignores reliability. To address these issues, we propose ARTEMIS, a unified framework for imperfectly supervised VPS driven by agent-guided reliability-aware temporal mask evolution. ARTEMIS initializes coarse masks from available supervision: SAM2 converts points/scribbles, while dense labels serve as reliable anchors. A debate-and-judge vision-language agent selects reliable temporal anchors under weak supervision, which are propagated bidirectionally with SAM2 to refine unreliable or unlabeled frames. Finally, ARTEMIS trains the segmenter using temporal reliability-aware robust learning, incorporating reliability-guided reference selection, a Reference Prototype Transport Module, and reliability-aware robust loss. These components assess mask reliability, evolve anchors over time, transport target identity across frames, and down-weight noisy supervision instead of discarding difficult samples. Experiments on SUN-SEG and CVC-ClinicDB-612 under scribble, point, and limited-label settings demonstrate that ARTEMIS achieves state-of-the-art performance. Code will be released at https://github.com/wangtong627/ARTEMIS.

Figures

Figures reproduced from arXiv: 2606.20161 by Guanyu Yang, Jinxing Zhou, Siwen Wang, Tong Wang, Yaolei Qi, Yuting He, Yutong Xie.

Figure 1
Figure 1. Figure 1: Challenges in VPS. (a) Low contrast between the foreground polyp within the green box and surrounding background; (b) motion blur; (c) scale variations; and (d) partial occlusion. colonoscopy. Compared with static image segmentation [9]– [14], VPS is more challenging because polyps often share similar color, texture, and boundaries with surrounding mu￾cosa. As shown in [PITH_FULL_IMAGE:figures/full_fig_p0… view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of imperfectly supervised VPS paradigms. Prior pipelines hard-filter pseudo labels, discarding useful samples and remaining noise-sensitive. ARTEMIS follows complete-then-learn: Stage 1 evolves reli￾able anchors, while Stage 2 uses reliability-guided reference selection, RPTM, and robust loss across weak/semi-supervised settings. solution for imperfectly supervised VPS, where pseudo labels must … view at source ↗
Figure 3
Figure 3. Figure 3: Qualitative comparison of sparse and dense prompts. Sparse Prompt uses yellow point prompts, while Dense Prompt uses green heavy noisy mask prompts on the same frames. Compared with sparse prompts, dense noisy masks provide stronger spatial priors for polyp localization. III. METHODOLOGY A. Methodology Overview This work follows a complete-then-learn pipeline. Stage 1 converts available supervision into co… view at source ↗
Figure 5
Figure 5. Figure 5: Overview of Stage 1: Agent-guided Bidirectional Mask Evolution. Given an input frame and SAM2 coarse mask, a debate-and-judge agent scores its anchor reliability. Candidate anchors are filtered by temporal NMS and propagated bidirectionally with SAM2 to produce evolved pseudo masks. C. Stage 1: Agent-guided Bidirectional Mask Evolution From observations to design. Given a video {It} T t=1 with It ∈ R 3×H×W… view at source ↗
Figure 6
Figure 6. Figure 6: Overview of Stage 2: Temporal Reliability-aware Robust Learning. ARTEMIS selects a reliable reference using mask reliability cues, transports reference identity tokens to target frames via RPTM, and applies a reliability-aware robust loss to suppress residual pseudo-label noise. where ϕv(Xt) ∈ R HfWf ×C is the value projection and Zt ∈ R N×C . This reference-conditioned pooling suppresses visually similar … view at source ↗
Figure 7
Figure 7. Figure 7 [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Qualitative comparison. Under the 1/8 Labeled Training Data setting, the two video clips, each with four frames, contain polyps highly similar to the background. ARTEMIS avoids the over-segmentation and under-segmentation observed in competing methods and identifies the polyp regions accurately [PITH_FULL_IMAGE:figures/full_fig_p010_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Comparison of Dice-threshold curves on the SUN-SEG-Hard-Seen across four supervision settings. From left to right, the subplots show results under Scribble Supervision, Point Supervision, 1/8 Labeled Training Data, and 1/16 Labeled Training Data. 0.00 0.20 0.40 0.60 0.80 1.00 Threshold 0.60 0.68 0.76 0.84 0.92 1.00 F Scribble Supervision 0.00 0.20 0.40 0.60 0.80 1.00 Threshold Point Supervision 0.00 0.20 0… view at source ↗
Figure 10
Figure 10. Figure 10: Comparison of Fβ-threshold curves on the SUN-SEG-Hard-Seen across four supervision settings. From left to right, the subplots show results under Scribble Supervision, Point Supervision, 1/8 Labeled Training Data, and 1/16 Labeled Training Data. TABLE IX ABLATION STUDY OF MASK EVOLUTION STRATEGIES IN STAGE 1 ON SUN-SEG-EASY-UNSEEN AND SUN-SEG-HARD-UNSEEN SETS. CM: COARSE MASK; FR: FORWARD REFINEMENT; BIR: … view at source ↗
Figure 12
Figure 12. Figure 12: Failure cases. Under Scribble Supervision, severe distractors, rapid motion, and uneven illumination may still affect fine boundary details and cause incomplete segmentation. Failure cases [PITH_FULL_IMAGE:figures/full_fig_p012_12.png] view at source ↗
Figure 11
Figure 11. Figure 11: Activation visualization of w/o RPTM and RPTM. Under Scribble Supervision, RPTM transports reliable reference-frame identity tokens to the current frame, enhancing polyp-region responses and suppressing distractors. TABLE XIV STAGE 2 PARAMETER ANALYSIS OF THE REFERENCE IDENTITY TOKEN NUMBER (N) IN RPTM ON SUN-SEG-EASY-UNSEEN AND SUN-SEG-HARD-UNSEEN SETS. N SUN-SEG-Easy-Unseen (%) SUN-SEG-Hard-Unseen (%) S… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

64 extracted references · 7 canonical work pages · 5 internal anchors

  1. [1]

    Global cancer statistics 2022: Globocan estimates of incidence and mortality worldwide for 36 cancers in 185 countries,

    F. Bray, M. Laversanne, H. Sung, J. Ferlay, R. L. Siegel, I. Soer- jomataram, and A. Jemal, “Global cancer statistics 2022: Globocan estimates of incidence and mortality worldwide for 36 cancers in 185 countries,”CA Cancer J. Clin., vol. 74, no. 3, pp. 229–263, 2024

  2. [2]

    Comparative validation of polyp detection methods in video colonoscopy: results from the miccai 2015 endoscopic vision challenge,

    J. Bernal, N. Tajkbaksh, F. J. Sanchez, B. J. Matuszewski, H. Chen, L. Yu, Q. Angermann, O. Romain, B. Rustad, I. Balasingham et al., “Comparative validation of polyp detection methods in video colonoscopy: results from the miccai 2015 endoscopic vision challenge,” IEEE Trans. Med. Imaging, vol. 36, no. 6, pp. 1231–1249, 2017

  3. [3]

    Context-aware convolutional neural network for grading of colorectal cancer histology images,

    M. Shaban, R. Awan, M. M. Fraz, A. Azam, Y .-W. Tsang, D. Snead, and N. M. Rajpoot, “Context-aware convolutional neural network for grading of colorectal cancer histology images,”IEEE Trans. Med. Imaging, vol. 39, no. 7, pp. 2395–2405, 2020

  4. [4]

    Progressively normalized self-attention network for video polyp seg- mentation,

    G.-P. Ji, Y .-C. Chou, D.-P. Fan, G. Chen, H. Fu, D. Jha, and L. Shao, “Progressively normalized self-attention network for video polyp seg- mentation,” inProc. Int. Conf. Med. Image Comput. Comput.-Assist. Interv. (MICCAI). Springer, 2021, pp. 142–152

  5. [5]

    Video polyp segmentation: A deep learning perspective,

    G.-P. Ji, G. Xiao, Y .-C. Chou, D.-P. Fan, K. Zhao, G. Chen, and L. Van Gool, “Video polyp segmentation: A deep learning perspective,” Mach. Intell. Res., vol. 19, no. 6, pp. 531–549, 2022

  6. [6]

    Sali: Short- term alignment and long-term interaction network for colonoscopy video polyp segmentation,

    Q. Hu, Z. Yi, Y . Zhou, F. Peng, M. Liu, Q. Li, and Z. Wang, “Sali: Short- term alignment and long-term interaction network for colonoscopy video polyp segmentation,” inProc. Int. Conf. Med. Image Comput. Comput.- Assist. Interv. (MICCAI). Springer, 2024, pp. 531–541

  7. [7]

    Stddnet: Harnessing mamba for video polyp segmentation via spatial-aligned temporal modeling and discrim- inative dynamic representation learning,

    G. Chen, H. Wu, and J. Qin, “Stddnet: Harnessing mamba for video polyp segmentation via spatial-aligned temporal modeling and discrim- inative dynamic representation learning,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025, pp. 21 364–21 373

  8. [8]

    Cmsa-net: Causal multi-scale aggregation with adaptive multi-source reference for video polyp segmentation,

    T. Wang, Y . Qi, S. Wang, I. Razzak, G. Yang, and Y . Xie, “Cmsa-net: Causal multi-scale aggregation with adaptive multi-source reference for video polyp segmentation,”arXiv preprint arXiv:2602.22821, 2026

  9. [9]

    Pranet: Parallel reverse attention network for polyp segmentation,

    D.-P. Fan, G.-P. Ji, T. Zhou, G. Chen, H. Fu, J. Shen, and L. Shao, “Pranet: Parallel reverse attention network for polyp segmentation,” in Proc. Int. Conf. Med. Image Comput. Comput.-Assist. Interv. (MICCAI). Springer, 2020, pp. 263–273

  10. [10]

    Polyp segmentation via semantic en- hanced perceptual network,

    T. Wang, X. Qi, and G. Yang, “Polyp segmentation via semantic en- hanced perceptual network,”IEEE Trans. Circuits Syst. Video Technol., 2024

  11. [11]

    Polypseg+: A lightweight context-aware network for real-time polyp segmentation,

    H. Wu, Z. Zhao, J. Zhong, W. Wang, Z. Wen, and J. Qin, “Polypseg+: A lightweight context-aware network for real-time polyp segmentation,” IEEE Trans. Cybern., vol. 53, no. 4, pp. 2610–2621, 2022

  12. [12]

    Epsegnet: Lightweight semantic recalibration and assembly for efficient polyp segmentation,

    H. Wu and Z. Zhao, “Epsegnet: Lightweight semantic recalibration and assembly for efficient polyp segmentation,”IEEE Trans. Neural Netw. Learn. Syst., vol. 36, no. 8, pp. 13 805–13 817, 2025

  13. [13]

    Ctnet: Contrastive transformer network for polyp segmentation,

    B. Xiao, J. Hu, W. Li, C.-M. Pun, and X. Bi, “Ctnet: Contrastive transformer network for polyp segmentation,”IEEE Trans. Cybern., vol. 54, no. 9, pp. 5040–5053, 2024

  14. [14]

    The devil is in the boundary: Boundary-enhanced polyp segmentation,

    Z. Liu, S. Zheng, X. Sun, Z. Zhu, Y . Zhao, X. Yang, and Y . Zhao, “The devil is in the boundary: Boundary-enhanced polyp segmentation,”IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 7, pp. 5414–5423, 2024

  15. [15]

    Structure-consistent weakly supervised salient object detection with local saliency coherence,

    S. Yu, B. Zhang, J. Xiao, and E. G. Lim, “Structure-consistent weakly supervised salient object detection with local saliency coherence,” in Proc. AAAI Conf. Artif. Intell. (AAAI), vol. 35, no. 4, 2021, pp. 3234– 3242

  16. [16]

    Weakly-supervised camouflaged object detection with scribble annotations,

    R. He, Q. Dong, J. Lin, and R. W. Lau, “Weakly-supervised camouflaged object detection with scribble annotations,” inProc. AAAI Conf. Artif. Intell. (AAAI), vol. 37, no. 1, 2023, pp. 781–789

  17. [17]

    Tree energy loss: Towards sparsely annotated semantic segmentation,

    Z. Liang, T. Wang, X. Zhang, J. Sun, and J. Shen, “Tree energy loss: Towards sparsely annotated semantic segmentation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022, pp. 16 907–16 916

  18. [18]

    Mean teachers are better role mod- els: Weight-averaged consistency targets improve semi-supervised deep learning results,

    A. Tarvainen and H. Valpola, “Mean teachers are better role mod- els: Weight-averaged consistency targets improve semi-supervised deep learning results,”Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017

  19. [19]

    Mcf: Mutual correction framework for semi-supervised medical image segmentation,

    Y . Wang, B. Xiao, X. Bi, W. Li, and X. Gao, “Mcf: Mutual correction framework for semi-supervised medical image segmentation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 15 651–15 660

  20. [20]

    Caussl: Causality- inspired semi-supervised learning for medical image segmentation,

    J. Miao, C. Chen, F. Liu, H. Wei, and P.-A. Heng, “Caussl: Causality- inspired semi-supervised learning for medical image segmentation,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 21 426– 21 437

  21. [21]

    Cross-view mutual learning for semi-supervised medical image segmentation,

    S. Wu, X. Wei, X. Chen, Y . Ren, J. He, and X. Pu, “Cross-view mutual learning for semi-supervised medical image segmentation,” inProc. ACM Int. Conf. Multimedia (ACMMM), 2024, pp. 9253–9261

  22. [22]

    Alternate diverse teaching for semi-supervised medical image segmentation,

    Z. Zhao, Z. Wang, L. Wang, D. Yu, Y . Yuan, and L. Zhou, “Alternate diverse teaching for semi-supervised medical image segmentation,” in Proc. Eur. Conf. Comput. Vis. (ECCV). Springer, 2024, pp. 227–243

  23. [23]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollar, and R. Girshick, “Segment anything,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 4015–4026

  24. [24]

    SAM 2: Segment Anything in Images and Videos

    N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Radle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V . Alwala, N. Carion, C.-Y . Wu, R. Girshick, P. Dollar, and C. Feichtenhofer, “SAM 2: Segment anything in images and videos,”arXiv preprint arXiv:2408.00714, 2024

  25. [25]

    Weakly-supervised concealed object segmentation with sam- based pseudo labeling and multi-scale feature grouping,

    C. He, K. Li, Y . Zhang, G. Xu, L. Tang, Y . Zhang, Z. Guo, and X. Li, “Weakly-supervised concealed object segmentation with sam- based pseudo labeling and multi-scale feature grouping,”Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, pp. 30 726–30 737, 2023

  26. [26]

    Weakpolyp-sam: Segment anything model-driven weakly-supervised polyp segmentation,

    Y . Zhao, T. Zhou, Y . Gu, Y . Zhou, Y . Zhang, Y . Wu, and H. Fu, “Weakpolyp-sam: Segment anything model-driven weakly-supervised polyp segmentation,”Knowl.-Based Syst., vol. 322, p. 113701, 2025

  27. [27]

    Segment concealed objects with incomplete supervision,

    C. He, K. Li, Y . Zhang, Z. Yang, Y . Pang, L. Tang, C. Fang, Y . Zhang, L. Kong, X. Liet al., “Segment concealed objects with incomplete supervision,”IEEE Trans. Pattern Anal. Mach. Intell., 2025

  28. [28]

    Sam-cod+: Sam-guided unified framework for weakly-supervised camouflaged object detection,

    H. Chen, P. Wei, G. Guo, and S. Gao, “Sam-cod+: Sam-guided unified framework for weakly-supervised camouflaged object detection,”IEEE Trans. Circuits Syst. Video Technol., vol. 35, no. 5, pp. 4635–4647, 2024

  29. [29]

    Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations

    J. Ge, J. Cao, X. Li, X. Zhu, C. Liu, B. Liu, C. Feng, and I. Patras, “D3ETOR: Debate-enhanced pseudo labeling and frequency-aware pro- gressive debiasing for weakly-supervised camouflaged object detection with scribble annotations,”arXiv preprint arXiv:2512.20260, 2026

  30. [30]

    Relax image-specific prompt re- quirement in sam: A single generic prompt for segmenting camouflaged objects,

    J. Hu, J. Lin, S. Gong, and W. Cai, “Relax image-specific prompt re- quirement in sam: A single generic prompt for segmenting camouflaged objects,” inProc. AAAI Conf. Artif. Intell. (AAAI), vol. 38, no. 11, 2024, pp. 12 511–12 518

  31. [31]

    Sam-adapter: Adapting segment anything in underperformed scenes,

    T. Chen, L. Zhu, C. Deng, R. Cao, Y . Wang, S. Zhang, Z. Li, L. Sun, Y . Zang, and P. Mao, “Sam-adapter: Adapting segment anything in underperformed scenes,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 3367–3375

  32. [32]

    St-sam: Sam-driven self- training framework for semi-supervised camouflaged object detection,

    X. Hu, F. Sun, J. Liu, F. Xu, and X. Zhang, “St-sam: Sam-driven self- training framework for semi-supervised camouflaged object detection,” inProc. ACM Int. Conf. Multimedia (ACMMM), 2025, pp. 8194–8203

  33. [33]

    Learnable prompting sam-induced knowledge distillation for semi- supervised medical image segmentation,

    K. Huang, T. Zhou, H. Fu, Y . Zhang, Y . Zhou, C. Gong, and D. Liang, “Learnable prompting sam-induced knowledge distillation for semi- supervised medical image segmentation,”IEEE Trans. Med. Imaging, vol. 44, no. 5, pp. 2295–2306, 2025

  34. [34]

    Asgnet: Adaptive spectrum guidance network for automatic polyp segmentation,

    Y . Sun, H. Zhang, J. Qian, J. Yang, and L. Luo, “Asgnet: Adaptive spectrum guidance network for automatic polyp segmentation,”IEEE Trans. Circuits Syst. Video Technol., 2026

  35. [35]

    Domain-interactive contrastive learning and prototype-guided self-training for cross-domain polyp segmentation,

    Z. Lu, Y . Zhang, Y . Zhou, Y . Wu, and T. Zhou, “Domain-interactive contrastive learning and prototype-guided self-training for cross-domain polyp segmentation,”IEEE Trans. Med. Imaging, 2024

  36. [36]

    Wavepolyp: Video polyp segmentation via hierarchical wavelet-based feature aggregation and inter-frame divergence perception,

    Y . Zhang, G. Chen, Y . He, H. Wu, and J. Qin, “Wavepolyp: Video polyp segmentation via hierarchical wavelet-based feature aggregation and inter-frame divergence perception,” inProc. Int. Conf. Learn. Represent. (ICLR), 2026

  37. [37]

    Hfsti-net: Hierarchical frequency-spatial-temporal interactions for video polyp segmentation,

    Y . He, G. Chen, Y . Zhang, H. Wu, and J. Qin, “Hfsti-net: Hierarchical frequency-spatial-temporal interactions for video polyp segmentation,” inProc. Int. Conf. Learn. Represent. (ICLR), 2026

  38. [38]

    An embedding-unleashing video polyp segmentation framework via region linking and scale alignment,

    Z. Fang, X. Guo, J. Lin, H. Wu, and J. Qin, “An embedding-unleashing video polyp segmentation framework via region linking and scale alignment,” inProc. AAAI Conf. Artif. Intell. (AAAI), vol. 38, no. 2, 2024, pp. 1744–1752

  39. [39]

    Leveraging hallucinations to reduce manual prompt dependency in promptable segmentation,

    J. Hu, J. Lin, J. Yan, and S. Gong, “Leveraging hallucinations to reduce manual prompt dependency in promptable segmentation,”Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 37, pp. 107 171–107 197, 2024

  40. [40]

    Group-wise learning for weakly supervised semantic segmentation,

    T. Zhou, L. Li, X. Li, C.-M. Feng, J. Li, and L. Shao, “Group-wise learning for weakly supervised semantic segmentation,”IEEE Trans. Image Process., vol. 31, pp. 799–811, 2021

  41. [41]

    Pixel-level domain adaptation: A new perspective for enhancing weakly supervised semantic segmentation,

    Y . Du, Z. Fu, and Q. Liu, “Pixel-level domain adaptation: A new perspective for enhancing weakly supervised semantic segmentation,” IEEE Trans. Image Process., vol. 33, pp. 4654–4669, 2024

  42. [42]

    Weakly supervised semantic segmentation via alternate self- dual teaching,

    D. Zhang, H. Li, W. Zeng, C. Fang, L. Cheng, M.-M. Cheng, and J. Han, “Weakly supervised semantic segmentation via alternate self- dual teaching,”IEEE Trans. Image Process., vol. 34, pp. 3086–3095, 2023

  43. [43]

    Weaktr: Exploring plain vision transformer for weakly-supervised semantic seg- mentation,

    L. Zhu, Y . Li, J. Fang, Y . Liu, X. Hao, W. Liu, and X. Wang, “Weaktr: Exploring plain vision transformer for weakly-supervised semantic seg- mentation,”IEEE Trans. Image Process., 2026. IEEE TRANSACTIONS ON IMAGE PROCESSING 14

  44. [44]

    Semi-supervised spatial temporal attention network for video polyp segmentation,

    X. Zhao, Z. Wu, S. Tan, D.-J. Fan, Z. Li, X. Wan, and G. Li, “Semi-supervised spatial temporal attention network for video polyp segmentation,” inProc. Int. Conf. Med. Image Comput. Comput.-Assist. Interv. (MICCAI). Springer, 2022, pp. 456–466

  45. [45]

    Boost the inference with co-training: A depth-guided mutual learning framework for semi- supervised medical polyp segmentation,

    Y . Li, Z. Zhu, Y . Zhang, Y . Chen, and Z. Yu, “Boost the inference with co-training: A depth-guided mutual learning framework for semi- supervised medical polyp segmentation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025, pp. 10 394–10 403

  46. [46]

    Vpsentry: Semi-supervised video polyp segmentation via sentry-guided long-term prototype fusion with correlation dynamic propagation,

    G. Chen, X. Luo, H. Wu, and J. Qin, “Vpsentry: Semi-supervised video polyp segmentation via sentry-guided long-term prototype fusion with correlation dynamic propagation,” inProc. AAAI Conf. Artif. Intell. (AAAI), vol. 40, no. 4, 2026, pp. 2850–2858

  47. [47]

    Uncertainty-aware cross-training for semi-supervised medical image segmentation,

    K. Huang, T. Zhou, H. Fu, Y . Zhang, Y . Zhou, and X.-J. Wu, “Uncertainty-aware cross-training for semi-supervised medical image segmentation,”IEEE Trans. Image Process., 2025

  48. [48]

    Boundary-aware prototype in semi-supervised medical image segmentation,

    Y . Wang, B. Xiao, X. Bi, W. Li, and X. Gao, “Boundary-aware prototype in semi-supervised medical image segmentation,”IEEE Trans. Image Process., vol. 33, pp. 5456–5467, 2024

  49. [49]

    Background matters: A cross-view bidirec- tional modeling framework for semi-supervised medical image segmen- tation,

    L. Cao, J. Li, and Y . Shi, “Background matters: A cross-view bidirec- tional modeling framework for semi-supervised medical image segmen- tation,”IEEE Trans. Image Process., 2025

  50. [50]

    Minet: Weakly-supervised camouflaged object detection through mutual interaction between region and edge cues,

    Y . Niu, L. Yang, R. Xu, Y . Li, and Y . Chen, “Minet: Weakly-supervised camouflaged object detection through mutual interaction between region and edge cues,” inProc. ACM Int. Conf. Multimedia (ACMMM), 2024, pp. 6316–6325

  51. [51]

    Progressive representation learning for weakly-supervised camouflaged object detection,

    S. Gao, Q. Guo, Y . Feng, C. Chen, X. Wei, Y . Wang, and W. Zhang, “Progressive representation learning for weakly-supervised camouflaged object detection,” inProc. ACM Int. Conf. Multimedia (ACMMM), 2025, pp. 2752–2761

  52. [52]

    Imagpose: A unified conditional framework for pose-guided person generation,

    F. Shen and J. Tang, “Imagpose: A unified conditional framework for pose-guided person generation,”Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 37, pp. 6246–6266, 2024

  53. [53]

    Advancing pose-guided image synthesis with progressive conditional diffusion models,

    F. Shen, H. Ye, J. Zhang, C. Wang, X. Han, and Y . Wei, “Advancing pose-guided image synthesis with progressive conditional diffusion models,” inProc. Int. Conf. Learn. Represent. (ICLR), vol. 2024, 2024, pp. 10 481–10 503

  54. [54]

    ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

    C.-M. Chan, W. Chen, Y . Su, J. Yu, W. Xue, S. Zhang, J. Fu, and Z. Liu, “Chateval: Towards better llm-based evaluators through multi- agent debate, 2023,”arXiv preprint arXiv:2308.07201, vol. 3, 2023

  55. [55]

    Encouraging divergent thinking in large language models through multi-agent debate,

    T. Liang, Z. He, W. Jiao, X. Wang, Y . Wang, R. Wang, Y . Yang, S. Shi, and Z. Tu, “Encouraging divergent thinking in large language models through multi-agent debate,” inProc. Conf. Empirical Methods Natural Lang. Process. (EMNLP), 2024, pp. 17 889–17 904

  56. [56]

    Scaling large language model-based multi- agent collaboration,

    C. Qian, Z. Xie, Y . Wang, W. Liu, K. Zhu, H. Xia, Y . Dang, Z. Du, W. Chen, C. Yanget al., “Scaling large language model-based multi- agent collaboration,” inProc. Int. Conf. Learn. Represent. (ICLR), vol. 2025, 2025, pp. 41 488–41 505

  57. [57]

    Superpixel-guided iterative learning from noisy labels for medical image segmentation,

    S. Li, Z. Gao, and X. He, “Superpixel-guided iterative learning from noisy labels for medical image segmentation,” inProc. Int. Conf. Med. Image Comput. Comput.-Assist. Interv. (MICCAI). Springer, 2021, pp. 525–535

  58. [58]

    Qwen2.5-VL Technical Report

    S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y . Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y . Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin, “Qwen2.5-vl technical report,”arXiv preprint arXiv:2502.13923, 2025

  59. [59]

    Mamba: Linear-Time Sequence Modeling with Selective State Spaces

    A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,”arXiv preprint arXiv:2312.00752, 2023

  60. [60]

    Pvt v2: Improved baselines with pyramid vision transformer,

    W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pvt v2: Improved baselines with pyramid vision transformer,” Comput. Vis. Media, vol. 8, pp. 415 – 424, 2021

  61. [61]

    WM-DOV A maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians,

    J. Bernal, F. J. Sanchez, G. Fernandez-Esparrach, D. Gil, C. Rodriguez, and F. Vilarino, “WM-DOV A maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians,”Comput. Med. Imaging Graph., vol. 43, pp. 99–111, 2015

  62. [62]

    Structure-measure: A new way to evaluate foreground maps,

    D.-P. Fan, M.-M. Cheng, Y . Liu, T. Li, and A. Borji, “Structure-measure: A new way to evaluate foreground maps,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2017, pp. 4548–4557

  63. [63]

    Enhanced-alignment measure for binary foreground map evaluation,

    D.-P. Fan, C. Gong, Y . Cao, B. Ren, M.-M. Cheng, and A. Borji, “Enhanced-alignment measure for binary foreground map evaluation,” arXiv preprint arXiv:1805.10421, 2018

  64. [64]

    How to evaluate fore- ground maps?

    R. Margolin, L. Zelnik-Manor, and A. Tal, “How to evaluate fore- ground maps?” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2014, pp. 248–255. Tong Wangis currently a Ph.D. candidate at the Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Ministry of Education), Southeast Universi...