Pith. sign in

REVIEW 4 major objections 4 minor 66 references

TAGS: 3D Tumor-Adaptive Guidance for SAM

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Fusing text and organ cues lifts 3D tumor Dice by up to 46.88%.

desk verdict Solid adaptation recipe with honest ablations, but the headline SOTA numbers rest on a lopsided benchmark protocol; the matched comparison against the 3D SAM Adapter is the credible part. read the letter →

arxiv 2505.17096 v2 pith:EDN6YFIA submitted 2025-05-21 eess.IV cs.CV

classification eess.IVcs.CV
keywords tumorsegmentation3DmedicalimagingfoundationmodeladaptationSAMCLIPmulti-promptfusionorganpromptinteractive
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a frozen 2D Segment Anything Model (SAM) can be turned into a strong 3D tumor segmenter by feeding it three kinds of guidance: an automatically generated organ mask, CLIP's text embeddings of tumor and healthy descriptions, and interactive click points. The proposed framework, TAGS, inserts lightweight alignment adapters between the stages of SAM's image encoder so text semantics refine spatial features before the mask decoder produces a volume. On kidney, liver, and pancreas tumor datasets, TAGS reports higher Dice and surface agreement than nnUNet, the 3D SAM Adapter, and several medical foundation models, using only 27.82 million tunable parameters. The central claim is that multi-prompt fusion, rather than building or retraining a 3D foundation model, is what closes the 2D-to-3D gap for tumor segmentation.

What carries the argument

The load-bearing mechanism is a stage-wise alignment adapter placed between the four stages of SAM's frozen image encoder. Each adapter is one fully connected layer with a residual connection scaled by λ=0.2, projecting intermediate visual features toward CLIP's text embedding space; a multi-modal alignment loss (focal plus dice on cosine similarity between adapter outputs and text features) pushes the encoder to attend to tumor-relevant semantics. Around this sit three prompt channels: an annotation-free organ prompt from TotalSegmentator (the third input channel), dual-category text prompts from CLIP with state-level and template-level descriptions, and point prompts at inference. The text encoder is omitted at inference, while the organ prompt and alignment adapters are kept.

What would settle it

Run nnUNet and the other baselines on the official KiTS21, LiTS, and MSD-Pancreas test sets (or standard five-fold cross-validation) with published hyperparameters and compare Dice/NSD with TAGS; a properly tuned nnUNet scoring near 80 Dice on kidney tumors instead of 50.24 would falsify the +46.88% claim, even if TAGS still beats the 3D SAM Adapter.

Watch

Extended reading notes

Core claim

TAGS claims to show that the 2D-to-3D gap in medical foundation models can be bridged by hierarchical multi-prompt fusion instead of by pre-training a new 3D model. The framework overlays a pseudo-organ mask from TotalSegmentator on the input volume, encodes dual-category text prompts ('tumor present/absent' at state and template levels) with CLIP, and aligns each of SAM's four encoder stages to those text embeddings through a single fully connected adapter with a scaled residual connection. At inference the text encoder is dropped, and one to three point prompts inside the tumor drive the frozen prompt encoder and mask decoder. Across KiTS21, LiTS, and MSD-Pancreas, TAGS reports Dice scores of 80.39–80.83 on kidney tumors, 59.69–66.23 on liver tumors, and 59.96–61.04 on pancreas tumors, surpassing all compared baselines and at least doubling nnUNet's kidney-tumor Dice in this protocol.

Load-bearing premise

The headline gains assume that the authors' custom 70/10/20 re-split of KiTS21, LiTS, and MSD-Pancreas, plus their training protocol for baselines, is the fair comparison; if official test sets or properly tuned baselines score much higher, the reported margins collapse.

Editorial extensions

If this is right

  • Click-based interactive segmentation can match or exceed fully supervised 3D segmentation on tumor tasks while keeping most of a 2D foundation model frozen.
  • Adding organ-level and text-level guidance is worth more than adding extra point prompts alone, since TAGS's gain over the 3D SAM Adapter (up to +13% Dice) comes at nearly the same parameter count.
  • Because only the organ's name is needed to build the text prompts, the recipe should transfer to new tumor types without retraining a medical text encoder.
  • The reported ICC above 90% across point prompt placements implies that clinical users do not need precisely placed clicks to get consistent segmentations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the authors leave implicit is that the organ mask may be doing much of the work: in the ablation, adding the organ prompt alone improves kidney Dice by 1.92 points while text alignment alone improves it by 1.89, so a direct test would be to degrade the organ mask quality and measure the drop in tumor Dice.
  • The custom 70/10/20 split makes the headline +46.88% protocol-dependent; a testable extension is to rerun the same comparisons on official test sets to see how much of the margin survives.
  • The framework inherits TotalSegmentator's failure modes, so a natural stress test is to apply TAGS to organs or imaging contrasts where TotalSegmentator produces poor masks and check whether tumor Dice degrades accordingly.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript proposes TAGS, a parameter-efficient adaptation of 2D SAM to 3D tumor segmentation. The method augments the 3D SAM Adapter with three components: a TotalSegmentator-derived organ mask added as a third input channel, CLIP text embeddings that are aligned to encoder features through lightweight stage-wise adapters, and 3D point prompts. TAGS is evaluated on KiTS21, LiTS, and MSD-Pancreas, with claims of state-of-the-art results (+46.88% over nnU-Net and at least +13% over medical foundation models). Ablations analyze the organ prompt, text prompt, multi-level alignment, point-prompt strategies, and text encoders, and the supplementary material includes fine-tuned comparisons with SAM-based baselines.

Significance. The framework is clearly described and novel in its combination of organ masks, CLIP-based text alignment, and 3D point prompts for SAM adaptation. The comparison against the 3D SAM Adapter under the same training/inference protocol is the strongest evidence and suggests real gains. The ablations are thorough, and the release of code and models is a practical strength. However, the headline state-of-the-art claims are not supported by the current benchmarking protocol: the classic baselines are under-trained on a nonstandard 70/10/20 split, several foundation models are evaluated zero-shot rather than fine-tuned, and TAGS receives interactive clicks and an organ mask that automatic baselines do not receive. With a corrected protocol and more cautious claims, the paper could make a useful contribution to 2D-to-3D medical foundation-model adaptation.

major comments (4)
  1. [Table 1 and Sec. 4.2] The claim of a +46.88% improvement over nnU-Net is not supported by the evidence. The nnU-Net baseline scores 50.24 Dice on kidney tumors and 36.46 Dice on pancreas tumors, which are far below established nnU-Net performance on these datasets; moreover, no standard deviations or number of runs are reported. Because all methods are trained only on a custom 70/10/20 split, the authors must either tune the classic baselines carefully on that split, or evaluate on the original challenge test sets, before claiming state-of-the-art. In addition, TAGS uses interactive point prompts and a TotalSegmentator organ mask as an extra input channel, whereas nnU-Net and the other classic baselines are fully automatic; the comparison should therefore be framed as interactive segmentation versus automatic segmentation, not as a general head-to-head state-of-the-art comparison.
  2. [Table 1 and Table 9] The main comparison with SAM-based foundation models is asymmetric. In Table 1, SAM-B, SAM-Med2D, SAM-Med3D, SegVol, and Universal Model are evaluated without fine-tuning on the target training split, while TAGS is trained on it. The supplementary fine-tuned results in Table 9 are a step in the right direction, but several fine-tuned models perform worse than their zero-shot counterparts and no fine-tuning details (epochs, learning rate, prompt strategy, checkpoints) are given, so it is unclear whether the baselines were tuned to a fair stopping point. The authors should report full fine-tuning details or explicitly describe Table 1 as a zero-shot comparison.
  3. [Abstract and Table 1] The abstract and Sec. 1 claim that TAGS surpasses 'other established medical FMs' by 'at least +13%'. This is internally inconsistent with Table 1: TAGS (3pts) reaches 61.04 Dice on pancreas tumors, which is below Universal Model's 61.28. The gain over the 3D SAM Adapter is also not 'at least +13%' in relative Dice terms (kidney: 80.83 vs 77.16, liver: 66.23 vs 58.96, pancreas: 61.04 vs 55.73). Please make all summary statistics match the table and state explicitly whether the reported percentages are absolute or relative.
  4. [Sec. 3.3, Sec. 3.4, and Eq. (3)-(4)] The text prompt is described as part of a 'multi-prompt fusion' that 'refines SAM's spatial attention', but at inference the text prompt is omitted and only the organ mask and point prompts are used. The text embeddings enter only through the training loss in Eq. (3), making the design a text-aligned training objective rather than an actual text-prompted inference mechanism. This is a legitimate approach, but the contribution statement and Fig. 2 should be revised to avoid implying that CLIP text features are injected at test time.
minor comments (4)
  1. [Sec. 3.1] There is a typo: 'SAM is a versatile, promptable segmentation model that that leverages' should read 'that leverages'.
  2. [Supplementary Table 8 and Sec. 3.2] The statement 'we used only tumor labels and excluded organ annotations from the process' is confusing because TotalSegmentator organ masks are a core input; it should specify that ground-truth organ labels from the datasets are excluded, not organ masks in general.
  3. [Table 1] No standard deviations, confidence intervals, or statistical significance tests are reported for the main results, despite the variability introduced by random point prompts and a single split. Reporting these would substantially strengthen the paper.
  4. [Sec. 4.2] The sentence 'For SAM-based models pre-trained on the three evaluation datasets, we perform direct inference' is ambiguous, since 'pre-trained on the datasets' could imply the baselines already saw the target data; please rephrase to clarify that direct inference means zero-shot use of released checkpoints without training on the target splits.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; the weak-baseline SOTA concern is a benchmarking issue, not a self-referential derivation.

full rationale

TAGS is an empirical adaptation method rather than a derivation from first principles, and its load-bearing components are externally grounded: organ prompts come from TotalSegmentator, text embeddings from CLIP, the image backbone from SAM, and the 3D adapter structure from Gong et al. [17], none of which are outputs of this paper. The tumor masks used for training and evaluation are dataset ground truths, not model predictions, and the 70/10/20 re-split is a protocol choice rather than a fitted quantity. The lambda=0.2 scaling is selected on the validation split and is ordinary hyperparameter tuning; it does not define the reported Dice numbers by construction. The ablation in Table 4, which directly uses aligned features for prediction, is an internal diagnostic rather than a claim that the final segmentation equals the alignment loss. The only self-citations ([32], [40], [61]) are ordinary prior-work references and are not load-bearing: no uniqueness theorem is invoked, no fitted parameter is renamed as a prediction, and no core ansatz is justified solely by the authors' own previous papers. The weak nnUNet baseline and custom evaluation split could undermine the external validity of the headline +46.88% claim, but that is a benchmarking and correctness issue, not circularity. No equation in the paper reduces a predicted quantity to its own input by construction.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on two external tools (TotalSegmentator and CLIP) and on the inherited 3D SAM Adapter architecture, plus several design choices such as the residual scaling factor and prompt counts. None of these are new entities; they are borrowed or tuned components.

free parameters (3)
  • Residual scaling factor lambda = 0.2
    Set by hand as 'optimal' in supplementary Section 2; used in Eq. 2 for all alignment adapters.
  • Inference point prompt count = 1 or 3
    Choice of 1 or 3 points; the 3-point setting is used for the headline +46.88% claim.
  • Training foreground:background patch sampling ratio = 2:1
    Chosen in Section 4.1; affects the sampled point prompts used to train the decoder.
assumptions (4)
  • domain assumption TotalSegmentator produces organ masks of sufficient quality on test volumes
    The organ mask is an input channel at inference (Section 3.2); Table 3 shows using ground-truth organ masks improves Dice, so errors in the external mask directly affect performance.
  • domain assumption CLIP text embeddings provide semantically useful guidance for 3D medical volumes despite CLIP being trained on 2D natural images
    The alignment loss in Section 3.3 trains adapters to match image features to CLIP text features; if CLIP carries no usable medical semantics, the text guidance would be noise.
  • domain assumption Tumors generally reside within specific organs
    The organ prompt design in Section 3.2 relies on this to localize tumors; it is stated in the text and used to justify the method.
  • domain assumption The 3D SAM Adapter backbone of Gong et al. is a valid base whose point-prompt decoder can be reused
    TAGS reuses the 3D SAM Adapter's prompt encoder, decoder, and 3D inflation without re-validating them; the paper takes this prior model's reliability as given.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TAGS: 3D Tumor-Adaptive Guidance for SAM." pith.science (2026). https://pith.science/paper/EDN6YFIA

@misc{pith2026250517096,
  author       = {Pith},
  title        = {Pith review of: TAGS: 3D Tumor-Adaptive Guidance for SAM},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EDN6YFIA}},
  note         = {Machine review of arXiv:2505.17096}
}
read the original abstract

Foundation models (FMs) such as CLIP and SAM have recently shown great promise in image segmentation tasks, yet their adaptation to 3D medical imaging-particularly for pathology detection and segmentation-remains underexplored. A critical challenge arises from the domain gap between natural images and medical volumes: existing FMs, pre-trained on 2D data, struggle to capture 3D anatomical context, limiting their utility in clinical applications like tumor segmentation. To address this, we propose an adaptation framework called TAGS: Tumor Adaptive Guidance for SAM, which unlocks 2D FMs for 3D medical tasks through multi-prompt fusion. By preserving most of the pre-trained weights, our approach enhances SAM's spatial feature extraction using CLIP's semantic insights and anatomy-specific prompts. Extensive experiments on three open-source tumor segmentation datasets prove that our model surpasses the state-of-the-art medical image segmentation models (+46.88% over nnUNet), interactive segmentation frameworks, and other established medical FMs, including SAM-Med2D, SAM-Med3D, SegVol, Universal, 3D-Adapter, and SAM-B (at least +13% over them). This highlights the robustness and adaptability of our proposed framework across diverse medical segmentation tasks.

Figures

Figures reproduced from arXiv: 2505.17096 by the authors.

Figure 1
Figure 1. Comparison of 4 approaches for SAM-driven volumet [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The overall framework of proposed model (TAGS). TAGS includes 3 different prompts: (a) annotation-free organ prompt; (b) [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Qualitative visualizations of TAGS and SAM-based approaches for kidney, liver, and pancreas tumor segmentation. Lesion areas [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (7 more)
Figure 5
Figure 5. Figure 5: Visualization of segmentation results on the MSD [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 4
Figure 4. Figure 4: Effectiveness analysis of the necessity of using multi [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 6
Figure 6. Figure 6: Illustration of different point prompt selection strategies. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Illustration of different model structures: (a) TAGS when [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Lists of two-level text prompt description that we used for feature alignment. [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Qualitative visualizations of TAGS and other volumetric segmentation benchmarks approaches for kidney, liver, and pancreas [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Performance differences between “random selection” and “edge selection” as the number of point prompts increase. [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

66 extracted references · 43 canonical work pages

  1. [1]

    Syn- thetic boost: Leveraging synthetic data for enhanced vision- language segmentation in echocardiography

    Rabin Adhikari, Manish Dhakal, Safal Thapaliya, Kanchan Poudel, Prasiddha Bhandari, and Bishesh Khanal. Syn- thetic boost: Leveraging synthetic data for enhanced vision- language segmentation in echocardiography. In Interna- tional Workshop on Advances in Simplifying Medical Ultra- sound, pages 89–99. Springer, 2023. 3

  2. [2]

    The medical segmentation decathlon.Nature communications, 13(1):4128, 2022

    Michela Antonelli, Annika Reinke, Spyridon Bakas, Key- van Farahani, Annette Kopp-Schneider, Bennett A Landman, Geert Litjens, Bjoern Menze, Olaf Ronneberger, Ronald M Summers, et al. The medical segmentation decathlon.Nature communications, 13(1):4128, 2022. 5, 1

  3. [3]

    The liver tumor segmentation benchmark (lits)

    Patrick Bilic, Patrick Christ, Hongwei Bran Li, Eugene V orontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, et al. The liver tumor segmentation benchmark (lits). Medical Image Analysis, 84:102680, 2023. 5, 1

  4. [4]

    Sam3d: Segment anything model in volumetric medical images

    Nhat-Tan Bui, Dinh-Hieu Hoang, Minh-Triet Tran, Gian- franco Doretto, Donald Adjeroh, Brijesh Patel, Arabinda Choudhary, and Ngan Le. Sam3d: Segment anything model in volumetric medical images. In 2024 IEEE International Symposium on Biomedical Imaging (ISBI), pages 1–4. IEEE,

  5. [5]

    Monai: An open-source framework for deep learning in healthcare

    M Jorge Cardoso, Wenqi Li, Richard Brown, Nic Ma, Eric Kerfoot, Yiheng Wang, Benjamin Murrey, Andriy Myro- nenko, Can Zhao, Dong Yang, et al. Monai: An open-source framework for deep learning in healthcare. arXiv preprint arXiv:2211.02701, 2022. 4

  6. [6]

    Ma-sam: Modality-agnostic sam adap- tation for 3d medical image segmentation

    Cheng Chen, Juzheng Miao, Dufan Wu, Aoxiao Zhong, Zhiling Yan, Sekeun Kim, Jiang Hu, Zhengliang Liu, Lichao Sun, Xiang Li, et al. Ma-sam: Modality-agnostic sam adap- tation for 3d medical image segmentation. Medical Image Analysis, 98:103310, 2024. 2, 3

  7. [7]

    Transunet: Rethinking the u-net architec- ture design for medical image segmentation through the lens of transformers

    Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie, Ehsan Adeli, Yan Wang, et al. Transunet: Rethinking the u-net architec- ture design for medical image segmentation through the lens of transformers. Medical Image Analysis, 97:103280, 2024. 2

  8. [9]

    Sam-med2d

    Junlong Cheng, Jin Ye, Zhongying Deng, Jianpin Chen, Tianbin Li, Haoyu Wang, Yanzhou Su, Ziyan Huang, Ji- long Chen, Lei Jiang, et al. Sam-med2d. arXiv preprint arXiv:2308.16184, 2023. 2

Show all 66 references
  1. [10]

    Orgunetr: Utilizing organ information and squeeze and exci- tation block for improved tumor segmentation.IEEE Access,

    Sanghyuk Roy Choi, Jungro Lee, and Minhyeok Lee. Orgunetr: Utilizing organ information and squeeze and exci- tation block for improved tumor segmentation.IEEE Access,

  2. [11]

    3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion

    ¨Ozg¨un C ¸ ic ¸ek, Ahmed Abdulkadir, Soeren S Lienkamp, Thomas Brox, and Olaf Ronneberger. 3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2016: 19th International Conference, At...

  3. [12]

    Clip-art: Contrastive pre-training for fine-grained art classification

    Marcos V Conde and Kerem Turgutlu. Clip-art: Contrastive pre-training for fine-grained art classification. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3956–3960, 2021. 3

  4. [13]

    Segment anything model (sam) for digital pathology: Assess zero- shot segmentation on whole slide imaging

    Ruining Deng, Can Cui, Quan Liu, Tianyuan Yao, Lu- cas W Remedios, Shunxing Bao, Bennett A Landman, Lee E Wheless, Lori A Coburn, Keith T Wilson, et al. Segment anything model (sam) for digital pathology: Assess zero- shot segmentation on whole slide imaging. arXiv preprint ar...

  5. [14]

    Bert: Pre-training of deep bidirectional trans- formers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional trans- formers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the asso- ciation for computational linguistics: human l...

  6. [15]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 3

  7. [16]

    Segvol: Universal and interactive volumetric medical image segmen- tation

    Yuxin Du, Fan Bai, Tiejun Huang, and Bo Zhao. Segvol: Universal and interactive volumetric medical image segmen- tation. arXiv preprint arXiv:2311.13385, 2023. 2, 5, 6

  8. [17]

    3dsam-adapter: Holistic adaptation of sam from 2d to 3d for promptable tumor segmentation

    Shizhan Gong, Yuan Zhong, Wenao Ma, Jinpeng Li, Zhao Wang, Jingyang Zhang, Pheng-Ann Heng, and Qi Dou. 3dsam-adapter: Holistic adaptation of sam from 2d to 3d for promptable tumor segmentation. Medical Image Analysis , 98:103324, 2024. 2, 3, 4, 5, 6, 1

  9. [18]

    A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero- shot detection of abnormalities

    Ibrahim Ethem Hamamci, Sezgin Er, Furkan Almas, Ayse Gulnihan Simsek, Sevval Nil Esirgun, Irem Dogan, Muhammed Furkan Dasdelen, Bastian Wittmann, Enis Sim- sar, Mehmet Simsar, et al. A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero-...

  10. [19]

    Unetr: Transformers for 3d med- ical image segmentation

    Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath, Dong Yang, Andriy Myronenko, Bennett Landman, Holger R Roth, and Daguang Xu. Unetr: Transformers for 3d med- ical image segmentation. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 574–58...

  11. [20]

    The state of the art in kidney and kidney tumor segmentation in contrast-enhanced ct imaging: Results of the kits19 challenge

    Nicholas Heller, Fabian Isensee, Klaus H Maier-Hein, Xi- aoshuai Hou, Chunmei Xie, Fengyi Li, Yang Nan, Guangrui Mu, Zhiyong Lin, Miofei Han, et al. The state of the art in kidney and kidney tumor segmentation in contrast-enhanced ct imaging: Results of the kits19 challenge. M...

  12. [21]

    When sam meets medical images: An investigation of seg- ment anything model (sam) on multi-phase liver tumor seg- mentation

    Chuanfei Hu, Tianyi Xia, Shenghong Ju, and Xinde Li. When sam meets medical images: An investigation of seg- ment anything model (sam) on multi-phase liver tumor seg- mentation. arXiv preprint arXiv:2304.08506, 2023. 1

  13. [22]

    Adapting visual-language models for generalizable anomaly detection in medical im- ages

    Chaoqin Huang, Aofan Jiang, Jinghao Feng, Ya Zhang, Xin- chao Wang, and Yanfeng Wang. Adapting visual-language models for generalizable anomaly detection in medical im- ages. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1...

  14. [23]

    nnu-net: a self-configuring method for deep learning-based biomedical image segmen- tation

    Fabian Isensee, Paul F Jaeger, Simon AA Kohl, Jens Pe- tersen, and Klaus H Maier-Hein. nnu-net: a self-configuring method for deep learning-based biomedical image segmen- tation. Nature methods, 18(2):203–211, 2021. 2, 5, 1

  15. [24]

    Winclip: Zero- /few-shot anomaly classification and segmentation

    Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, and Onkar Dabeer. Winclip: Zero- /few-shot anomaly classification and segmentation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19606–19616, 2023. 4

  16. [25]

    Jorge Jovicich, Silvester Czanner, Xiao Han, David Salat, Andre van der Kouwe, Brian Quinn, Jenni Pacheco, Mari- lyn Albert, Ronald Killiany, Deborah Blacker, et al. Mri- derived measurements of human subcortical, ventricular and intracranial brain volumes: reliability effects...

  17. [26]

    Segment any- thing

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 4015–4026, 202...

  18. [27]

    Medclip-sam: Bridging text and image to- wards universal medical image segmentation

    Taha Koleilat, Hojat Asgariandehkordi, Hassan Rivaz, and Yiming Xiao. Medclip-sam: Bridging text and image to- wards universal medical image segmentation. arXiv preprint arXiv:2403.20253, 2024. 2

  19. [28]

    3d ux-net: A large kernel volumetric convnet modernizing hierarchical transformer for medical image seg- mentation

    Ho Hin Lee, Shunxing Bao, Yuankai Huo, and Bennett A Landman. 3d ux-net: A large kernel volumetric convnet modernizing hierarchical transformer for medical image seg- mentation. arXiv preprint arXiv:2209.15076, 2022. 5, 1

  20. [29]

    Medlsam: Localize and segment anything model for 3d ct images

    Wenhui Lei, Wei Xu, Kang Li, Xiaofan Zhang, and Shaoting Zhang. Medlsam: Localize and segment anything model for 3d ct images. Medical Image Analysis, page 103370, 2024. 2

  21. [31]

    Clipsam: Clip and sam collabora- tion for zero-shot anomaly segmentation

    Shengze Li, Jianjian Cao, Peng Ye, Yuhan Ding, Chongjun Tu, and Tao Chen. Clipsam: Clip and sam collabora- tion for zero-shot anomaly segmentation. arXiv preprint arXiv:2401.12665, 2024. 2

  22. [32]

    Text-guided foundation model adaptation for long- tailed medical image classification

    Sirui Li, Li Lin, Yijin Huang, Pujin Cheng, and Xiaoying Tang. Text-guided foundation model adaptation for long- tailed medical image classification. In 2024 IEEE Interna- tional Symposium on Biomedical Imaging (ISBI), pages 1–5. IEEE, 2024. 1

  23. [33]

    Focal loss for dense object detection

    Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Doll´ar. Focal loss for dense object detection. In Pro- ceedings of the IEEE international conference on computer vision, pages 2980–2988, 2017. 4

  24. [34]

    Clip-driven universal model for organ segmentation and tumor detection

    Jie Liu, Yixiao Zhang, Jie-Neng Chen, Junfei Xiao, Yongyi Lu, Bennett A Landman, Yixuan Yuan, Alan Yuille, Yucheng Tang, and Zongwei Zhou. Clip-driven universal model for organ segmentation and tumor detection. In Proceedings of the IEEE/CVF International Conference on Compute...

  25. [35]

    Segment anything in medical images

    Jun Ma, Yuting He, Feifei Li, Lin Han, Chenyu You, and Bo Wang. Segment anything in medical images. Nature Communications, 15(1):654, 2024. 2

  26. [36]

    Crepe: Can vision-language foundation models reason compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 10910–10921, 2023

    Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna. Crepe: Can vision-language foundation models reason compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 10910–10921, 2023. 1

  27. [37]

    Segment anything model for medical image analysis: an experimental study

    Maciej A Mazurowski, Haoyu Dong, Hanxue Gu, Jichen Yang, Nicholas Konz, and Yixin Zhang. Segment anything model for medical image analysis: an experimental study. Medical Image Analysis, 89:102918, 2023. 2

  28. [38]

    V-net: Fully convolutional neural networks for volumetric medical image segmentation

    Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV), pages 565–571. Ieee, 2016. 1

  29. [39]

    A guide to combat har- monization of imaging biomarkers in multicenter studies

    Fanny Orlhac, Jakoba J Eertink, Anne-S ´egol`ene Cottereau, Jos´ee M Zijlstra, Catherine Thieblemont, Michel Meignan, Ronald Boellaard, and Ir `ene Buvat. A guide to combat har- monization of imaging biomarkers in multicenter studies. Journal of Nuclear Medicine, 63(2):172–179...

  30. [40]

    Optimizing synthetic data for enhanced pan- creatic tumor segmentation

    Linkai Peng, Zheyuan Zhang, Gorkem Durak, Frank H Miller, Alpay Medetalibeyoglu, Michael B Wallace, and Ulas Bagci. Optimizing synthetic data for enhanced pan- creatic tumor segmentation. In International Workshop on Personalized Incremental Learning in Medicine , pages 35–

  31. [41]

    Explor- ing transfer learning in medical image segmentation using vision-language models

    Kanchan Poudel, Manish Dhakal, Prasiddha Bhandari, Ra- bin Adhikari, Safal Thapaliya, and Bishesh Khanal. Explor- ing transfer learning in medical image segmentation using vision-language models. arXiv preprint arXiv:2308.07706,

  32. [42]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...

  33. [43]

    Sam 2: Segment anything in images and videos

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 2

  34. [44]

    U- net: Convolutional networks for biomedical image segmen- tation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, pa...

  35. [45]

    Laion-5b: An open large-scale dataset for training next generation image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural In- f...

  36. [46]

    Is sam 2 better than sam in medical image segmentation?arXiv preprint arXiv:2408.04212, 2024

    Sourya Sengupta, Satrajit Chakrabarty, and Ravi Soni. Is sam 2 better than sam in medical image segmentation?arXiv preprint arXiv:2408.04212, 2024. 2

  37. [47]

    Unetr++: delving into efficient and accurate 3d medical image segmentation

    Abdelrahman M Shaker, Muhammad Maaz, Hanoona Rasheed, Salman Khan, Ming-Hsuan Yang, and Fahad Shah- baz Khan. Unetr++: delving into efficient and accurate 3d medical image segmentation. IEEE Transactions on Medi- cal Imaging, 2024. 5

  38. [48]

    Self-supervised pre-training of swin trans- formers for 3d medical image analysis

    Yucheng Tang, Dong Yang, Wenqi Li, Holger R Roth, Bennett Landman, Daguang Xu, Vishwesh Nath, and Ali Hatamizadeh. Self-supervised pre-training of swin trans- formers for 3d medical image analysis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recogn...

  39. [49]

    Yfcc100m: The new data in multimedia research

    Bart Thomee, David A Shamma, Gerald Friedland, Ben- jamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016. 3

  40. [50]

    Attention is all you need

    A Vaswani. Attention is all you need. Advances in Neural Information Processing Systems, 2017. 1

  41. [51]

    Integrated treatment planning in percutaneous microwave ablation of lung tumors

    Haoyu Wang, Hongrui Yi, Jie Liu, and Lixu Gu. Integrated treatment planning in percutaneous microwave ablation of lung tumors. In 2022 44th Annual International Confer- ence of the IEEE Engineering in Medicine & Biology Society (EMBC), pages 4974–4977. IEEE, 2022. 1

  42. [52]

    Sam-med3d: Towards general-purpose seg- mentation models for volumetric medical images, 2024

    Haoyu Wang, Sizheng Guo, Jin Ye, Zhongying Deng, Jun- long Cheng, Tianbin Li, Jianpin Chen, Yanzhou Su, Ziyan Huang, Yiqing Shen, Bin Fu, Shaoting Zhang, Junjun He, and Yu Qiao. Sam-med3d: Towards general-purpose seg- mentation models for volumetric medical images, 2024. 2, 5, 6

  43. [53]

    Joint learning of 3d lesion segmentation and classifica- tion for explainable covid-19 diagnosis

    Xiaofei Wang, Lai Jiang, Liu Li, Mai Xu, Xin Deng, Lisong Dai, Xiangyang Xu, Tianyi Li, Yichen Guo, Zulin Wang, et al. Joint learning of 3d lesion segmentation and classifica- tion for explainable covid-19 diagnosis. IEEE transactions on medical imaging, 40(9):2463–2476, 2021. 1

  44. [54]

    Medclip: Contrastive learning from unpaired medical images and text, 2022

    Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. Medclip: Contrastive learning from unpaired medical images and text, 2022. 3, 7, 2

  45. [55]

    To- talsegmentator: robust segmentation of 104 anatomic struc- tures in ct images

    Jakob Wasserthal, Hanns-Christian Breit, Manfred T Meyer, Maurice Pradella, Daniel Hinck, Alexander W Sauter, Tobias Heye, Daniel T Boll, Joshy Cyriac, Shan Yang, et al. To- talsegmentator: robust segmentation of 104 anatomic struc- tures in ct images. Radiology: Artificial In...

  46. [56]

    Medical sam adapter: Adapting seg- ment anything model for medical image segmentation.arXiv preprint arXiv:2304.12620, 2023

    Junde Wu, Wei Ji, Yuanpei Liu, Huazhu Fu, Min Xu, Yanwu Xu, and Yueming Jin. Medical sam adapter: Adapting seg- ment anything model for medical image segmentation.arXiv preprint arXiv:2304.12620, 2023. 2, 3

  47. [57]

    Cotr: Efficiently bridging cnn and transformer for 3d medi- cal image segmentation

    Yutong Xie, Jianpeng Zhang, Chunhua Shen, and Yong Xia. Cotr: Efficiently bridging cnn and transformer for 3d medi- cal image segmentation. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th Inter- national Conference, Strasbourg, France, September...

  48. [58]

    Uniseg: A prompt-driven universal segmenta- tion model as well as a strong representation learner

    Yiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen, and Yong Xia. Uniseg: A prompt-driven universal segmenta- tion model as well as a strong representation learner. In In- ternational Conference on Medical Image Computing and Computer-Assisted Intervention , pages 508–518. Springer,

  49. [59]

    Florence: A new foundation model for computer vision

    Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021. 1

  50. [60]

    Scaling vision transformers

    Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lu- cas Beyer. Scaling vision transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12104–12113, 2022. 3

  51. [61]

    Large- scale multi-center ct and mri segmentation of pancreas with deep learning

    Zheyuan Zhang, Elif Keles, Gorkem Durak, Yavuz Tak- tak, Onkar Susladkar, Vandan Gorade, Debesh Jha, Asli C Ormeci, Alpay Medetalibeyoglu, Lanhong Yao, et al. Large- scale multi-center ct and mri segmentation of pancreas with deep learning. arXiv preprint arXiv:2405.12367, 2024. 1

  52. [62]

    Torr, and Li Zhang

    Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H.S. Torr, and Li Zhang. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proceedings of the IEEE/CVF Confe...

  53. [63]

    Zegclip: Towards adapting clip for zero-shot se- mantic segmentation

    Ziqin Zhou, Yinjie Lei, Bowen Zhang, Lingqiao Liu, and Yifan Liu. Zegclip: Towards adapting clip for zero-shot se- mantic segmentation. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 11175–11185, 2023. 3

  54. [64]

    Segment everything everywhere all at once

    Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee. Segment everything everywhere all at once. Advances in Neural Information Processing Systems, 36, 2024. 4 TAGS: 3D Tumor-Adaptive Guidance for SAM Supplementary...

  55. [65]

    The KiTS dataset [20] originates from the MICCAI 2021 Kid- ney and Kidney Tumor Segmentation Challenge

    Dataset Details Table 8 presents detailed information on each dataset. The KiTS dataset [20] originates from the MICCAI 2021 Kid- ney and Kidney Tumor Segmentation Challenge. The LiTS dataset [3] comes from the MICCAI 2017 Liver Tumor Seg- mentation Challenge, while the MSD-Pa...

  56. [66]

    Our pre-processing pipeline follows the approach in [17]

    Supplementary Implementation Details Data Processing. Our pre-processing pipeline follows the approach in [17]. We resample anisotropic images to the target spacing, followed by intensity clipping and normal- ization. For data augmentation, each sample has a 50% probability of...

  57. [67]

    As a supplement to Fig

    Additional Qualitative Evaluations Qualitative Visualization on Other Benchmarks. As a supplement to Fig. 3, we present a qualitative visualization comparison of TAGS with four classic volumetric segmen- tation benchmarks and CLIP-based benchmarks in Fig. 9. Both Fig. 3 and Fi...

  58. [68]

    sin- gle alignment adapter ablation experiment

    Additional Ablation Studies Model Structure of Multi-level Alignment Ablation. To clarify the distinctions in Sec. 4.3, we present visual rep- resentations of the model structures in Fig. 7 for the “sin- gle alignment adapter ablation experiment” referenced in Fig. 4, the “CLI...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.