REVIEW 4 major objections 8 minor 30 references
Adapting under 1% of SAM3 with LoRA beats zero-shot and fully fine-tuned medical SAMs on surgical concept segmentation while fitting on a single consumer GPU.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
LoRA on SAM3’s prompt encoder, detector, and tracker (0.98% of parameters) raises surgical concept-segmentation mIoU over zero-shot SAM3 and Medical SAM3 while fitting in ~9 GB GPU memory.
T0 review reviewed 2026-07-30 challenge →
load-bearing objection Solid single-GPU LoRA recipe with real gains over zero-shot SAM3; the “beats fully fine-tuned Medical SAM3” claim is the soft leg and needs fixing before anyone leans on it. the 4 major comments →
Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
A LoRA-adapted SAM3 that trains only the prompt encoder, DETR-style detector, and tracker (8.32 M parameters, 0.98% of the model) while freezing the vision backbone consistently outperforms zero-shot SAM3 and fully fine-tuned Medical SAM3 variants on prompt-driven surgical concept segmentation across CholecSeg8k, EndoVis18, and CaDISv2, with peak training memory of 9 GB on one consumer GPU and masks usable downstream for reconstruction and physical simulation.
What carries the argument
Low-Rank Adaptation (LoRA) injected only into the prompt encoder’s cross-modal layers, the detector’s attention projections, and the tracker’s attention layers, with the vision backbone fully frozen and a unified surgical concept vocabulary aligning text prompts across datasets.
Load-bearing premise
That the near-floor scores reported for fully fine-tuned Medical SAM3 under the authors’ prompts and protocol are a fair baseline rather than a mismatch in prompting, input mode, or evaluation setup.
What would settle it
Re-run Medical SAM3 2D/3D on the same held-out CholecSeg8k and EndoVis18 frames with the paper’s exact text concepts and scoring script; if its mIoU on classes such as abdominal wall, grasper, and liver rises to match or beat the LoRA model, the claim of outperforming full medical fine-tunes fails.
If this is right
- Surgical concept segmentation can be specialized on a single consumer GPU without full foundation-model fine-tuning.
- One shared LoRA weight plus a unified concept vocabulary can serve multiple procedures without per-dataset retraining.
- Class-wise masks from the adapted model can drive region-level material assignment in Gaussian-splatting reconstruction and MPM-style physics simulation.
- Preserving the frozen vision backbone while adapting prompt, detect, and track paths is presented as enough to close the surgical domain gap for this task.
Where Pith is reading between the lines
- If the frozen backbone already carries usable surgical appearance features, similar LoRA-only adaptation may transfer to other narrow clinical domains (endoscopy, interventional radiology) with little extra memory.
- The large reported gap versus Medical SAM3 invites a controlled re-evaluation protocol that isolates prompt format and 2D vs video input before treating full fine-tuning as categorically worse.
- Extending the same adapters from still frames to full surgical video sequences, as the authors flag for future work, is the natural next stress test of tracker LoRA stability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper adapts SAM3 to surgical concept segmentation via Low-Rank Adaptation (LoRA). Low-rank adapters (r=16, α=32) are injected into the prompt encoder, DETR-style detector, and tracker while the vision backbone is frozen, giving 8.32M trainable parameters (0.98% of 849M) and ~9 GB peak training memory on a single RTX 3090. A unified concept vocabulary maps native class labels of CholecSeg8k, EndoVis18, and CaDISv2 to shared text prompts, enabling one universal LoRA weight across datasets. Per-class mIoU results on CholecSeg8k (Table 1) and EndoVis18 (Table 2) show large gains over zero-shot SAM3 and near-zero/single-digit scores for Medical SAM3 (2D/3D); a downstream Gaussian-splatting reconstruction plus MPM physics simulation pipeline is demonstrated qualitatively.
Significance. If the results hold, the contribution is a useful and practical one: a concrete, reproducible recipe (r=16, α=32, AdamW, bf16, 10 epochs, single RTX 3090 at ~9 GB peak memory) for adapting SAM3 to surgical concept segmentation with 0.98% trainable parameters, evaluated on three public surgical benchmarks, with a promised open-source release. The demonstration that a frozen vision backbone plus adapters on the prompt encoder/detector/tracker suffices for large per-class gains over zero-shot SAM3 is a meaningful negative-space result for the surgical PEFT literature, and the single universal LoRA weight across datasets is a practical selling point. The downstream reconstruction/simulation demonstration, while qualitative, illustrates a plausible deployment path. The work is incremental in method terms (standard LoRA applied to SAM3) but timely, and the efficiency claims are directly measurable and falsifiable.
major comments (4)
- [§3.2, Tables 1–2] Tables 1–2, §1 and Abstract: the headline claim that the method 'outperforms fully fine-tuned Medical SAM3 under identical training settings' is not supported by the experiments as described. Medical SAM3 scores 0.00 on Abdominal Wall and 4.99/8.65 on Liver (Table 1) — near-empty outputs on the two largest, easiest structures — while zero-shot SAM3 scores 92.57 on Liver. This pattern is the signature of a prompt-interface or protocol mismatch, not of a fine-tuned medical foundation model. Two specific ambiguities must be resolved: (a) §3.2 attributes Medical SAM3's failure to 'full fine-tuning across 33 multi-modal datasets diluting specialized surgical features', which describes the off-the-shelf checkpoint run zero-shot on surgical data — yet §1 claims comparison 'under identical training settings'. Was Medical SAM3 actually fine-tuned on the same surgical training split with the same
- [§3.2 / Fig. 4] CaDISv2 is named as one of three evaluation benchmarks in the Abstract, §1 (contributions), and §3.1 (held-out videos 2, 12, 22), but the paper contains no quantitative table for it. §3.2 reports only a single aggregate number ('overall mIoU of 26.9% and mDice of 37.0%') with no per-class breakdown and no statement of which dataset this aggregate refers to; the only CaDISv2 evidence is the Fig. 4 radar chart, from which exact values cannot be read. A per-class results table for CaDISv2, matching Tables 1–2, is needed for the 'consistent across three benchmarks' claim.
- [§1, Tables 1–2] §1 states the method 'outperforms dataset-specific SurgTPGS', but SurgTPGS does not appear in Table 1, Table 2, or Fig. 4 — no comparison numbers are reported anywhere. Either include the comparison or remove the claim. Note also that SurgTPGS is a dataset-specific method while the proposed model uses one universal weight, so the comparison conditions should be stated explicitly.
- [Tables 1–2, §2.3] Tables 1–2 report single-run numbers with no error bars, multi-seed runs, or significance tests. Several claimed wins are modest (e.g., EndoVis18 Seq_9 instrument-clasper: 40.51 vs CAT-Seg 38.94; instrument-wrist: 39.18 vs 24.34 is larger, but kidney-parenchyma 73.53 vs CAT-Seg 71.98 is within typical seed noise for LoRA fine-tuning on small surgical splits). Given that the test sets comprise only 4, 2, and 3 videos respectively, at minimum multi-seed means ± std, or a paired per-frame test, is needed to support 'consistent' improvements. Additionally, no ablation on LoRA rank/insertion sites is provided even though §2.3 makes a specific architectural argument for adapting the prompt encoder, detector, and tracker while freezing the vision encoder; a small ablation (e.g., rank {4,16,64}, encoder-only vs decoder-only insertion) would substantiate that design choice.
minor comments (8)
- [Tables 1–2] Table 1 and Table 2 captions: 'Quantative' → 'Quantitative'; the captions state 'highlighted the first, second, and third' but do not say what visual convention (bold/underline/color) encodes each rank.
- [§3.1, Tables 1–2] The Evaluation Metrics paragraph (§3.1) promises both mIoU and Dice, but Tables 1–2 report only mIoU; Dice appears only in the §3.2 aggregate sentence and Fig. 4. Report both metrics in the tables or state why mIoU alone is shown.
- [§3.2, §2.3] §3.2, final paragraph: 'Zero-shot SAM 3' has a stray space; §2.3 'DETR decoder' is used while §2.2 calls the component a 'DETR-style detector' — unify terminology.
- [§2.3] The 80 GB figure for full fine-tuning (§1, §2.3) is asserted without citation or measurement; 'over 80% reduction' from 80 GB to 9 GB is arithmetically ~89%. Please give the source of the 80 GB number (own measurement? which checkpoint/batch size?) or qualify it.
- [§2.1] The unified vocabulary V is central to the 'universal LoRA weight' contribution, but the actual label→text mapping is never shown. Please include the full mapping table (e.g., as a supplement), since cross-dataset homonym classes (e.g., 'grasper' vs 'instrument-clasper') are exactly where this design could silently degrade.
- [§3.2] Baselines CLIP and SurgVLP are not segmentation models per se; a sentence explaining how segmentation outputs were obtained from them (e.g., which segmentation head or decoding procedure) would make Table 1–2 reproducible.
- [Fig. 4] Fig. 4: axes cover 'a curated subset' of categories — state the selection criterion in the caption, and consider including per-dataset aggregate values so readers can cross-check against Tables 1–2.
- [§3.3, Abstract] §3.3 (Application): the reconstruction/simulation pipeline is illustrative only. One sentence quantifying or at least describing the validation of the simulation output (beyond the qualitative Fig. 5) would temper the 'directly deployed' language in the Abstract, which currently overstates what is demonstrated.
Circularity Check
No circularity: standard supervised PEFT evaluation on held-out surgical splits; metrics are external mIoU/Dice, not quantities forced by the fit.
full rationale
This is an empirical computer-vision paper: LoRA adapters are trained on labeled surgical frames (unified concept vocabulary V) and scored with standard mIoU/Dice on held-out video sequences from CholecSeg8k, EndoVis18, and CaDISv2. The reported gains over zero-shot SAM3 and other baselines are ordinary out-of-sample comparisons; nothing in the method defines the evaluation metric to equal the training objective by construction, and there is no uniqueness theorem, self-cited ansatz, or fitted parameter renamed as a first-principles prediction. Self-citations (SurgTPGS, Endo-4DGS, EndoGSim) appear only as related work or optional downstream consumers of the masks, not as load-bearing premises that force the segmentation scores. Concerns about whether Medical SAM3 was fairly prompted or fine-tuned under identical settings are experimental-validity issues, not circularity. Derivation chain is self-contained against external benchmarks; steps list is empty by design.
Axiom & Free-Parameter Ledger
free parameters (4)
- LoRA rank r and scale α =
r=16, α=32
- LoRA insertion sites =
prompt encoder + detector + tracker only
- Optimization hyperparameters =
LR=1e-4, 10 epochs, etc.
- Unified concept vocabulary mapping V
axioms (5)
- domain assumption Low-rank updates W = W0 + BA suffice to adapt frozen Transformer linear maps for cross-domain surgical transfer (standard LoRA hypothesis).
- domain assumption SAM3 vision-encoder features generalize to endoscopic surgical frames well enough that the backbone can remain fully frozen.
- domain assumption Text concept prompts from a unified surgical vocabulary are a valid and comparable interface across CholecSeg8k, EndoVis18, and CaDISv2.
- ad hoc to paper Held-out video IDs listed in §3.1 are sufficiently representative test splits for the claimed generalization.
- domain assumption Standard segmentation metrics (mIoU, Dice) and confidence-weighted NMS fusion in SAM3 are appropriate success criteria for “surgical concept segmentation.”
invented entities (1)
-
Unified surgical concept vocabulary V (cross-dataset label→text map)
no independent evidence
Cite this review
Pith. "Pith review of Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation." pith.science (2026). https://pith.science/paper/RGWXZFIX
@misc{pith2026260723694,
author = {Pith},
title = {Pith review of: Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/RGWXZFIX}},
note = {Machine review of arXiv:2607.23694}
}
read the original abstract
Efficient surgical segmentation empowers clinical diagnosis, intraoperative monitoring, and downstream robotic pipelines for reconstruction and simulation. Although prompt-driven foundation models like Segment Anything Model 3 (SAM3) achieve strong segmentation performance on natural images, surgical data exhibits domain gaps against its pre-training data, resulting in degraded segmentation accuracy. Furthermore, existing medical SAM methods require full-parameter fine-tuning, incurring heavy computational consumption and low efficiency. To address these limitations, this work proposes a parameter-efficient Low-Rank Adaptation (LoRA) adaptation of SAM3 for surgical concept segmentation. We inject low-rank adapters into the prompt encoder, detector and tracker while fully freezing the vision backbone, which only optimizes 0.98% of the total model parameters and supports training on a single consumer GPU. Comprehensive experiments demonstrate that our method consistently outperforms zero-shot SAM3 and other mainstream baselines, and the generated segmentation results can be directly deployed to support downstream robotic surgical scene reconstruction and physical simulation pipelines.
Figures
Reference graph
Works this paper leans on
-
[1]
Allan,M.,Kondo,S.,Bodenstedt,S.,Leger,S.,Kadkhodamohammadi,R.,Luengo, I., Fuentes, F., Flouty, E., Mohammed, A., Pedersen, M., et al.: 2018 robotic scene segmentation challenge. arXiv preprint arXiv:2001.11190 (2020) Parameter-Efficient Adaptation of SAM3 for Surgical Segmentation 9 Fig. 5.From segmentation to downstream applications.Left:masks predicted ...
Pith/arXiv arXiv 2018
-
[2]
In: European conference on computer vision
Cao, H., Wang, Y., Chen, J., Jiang, D., Zhang, X., Tian, Q., Wang, M.: Swin- unet: Unet-like pure transformer for medical image segmentation. In: European conference on computer vision. pp. 205–218. Springer (2022)
2022
-
[3]
arXiv preprint arXiv:2511.16719 (2025)
Carion, N., Gustafson, L., Hu, Y.T., Debnath, S., Hu, R., Suris, D., Ryali, C., Alwala,K.V.,Khedr,H.,Huang,A.,etal.:Sam3:Segmentanythingwithconcepts. arXiv preprint arXiv:2511.16719 (2025)
Pith/arXiv arXiv 2025
-
[4]
In: European conference on computer vision
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End- to-end object detection with transformers. In: European conference on computer vision. pp. 213–229. Springer (2020)
2020
-
[5]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Cho,S.,Shin,H.,Hong,S.,Arnab,A.,Seo,P.H.,Kim,S.:Cat-seg:Costaggregation for open-vocabulary semantic segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4113–4123 (2024)
2024
-
[6]
Nature machine intelligence5(3), 220–235 (2023)
Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., Hu, S., Chen, Y., Chan, C.M., Chen, W., et al.: Parameter-efficient fine-tuning of large-scale pre-trained language models. Nature machine intelligence5(3), 220–235 (2023)
2023
-
[7]
arXiv preprint arXiv:2601.10880 (2026)
Ding, T., Song, C., Tu, J., Yan, Z., Shao, Y., Wang, Z., Shang, Y., Han, T., Tian, Y.: Medical sam3: A foundation model for universal prompt-driven medical image segmentation. arXiv preprint arXiv:2601.10880 (2026)
arXiv 2026
-
[8]
Medical Image Analysis71, 102053 (2021)
Grammatikopoulou, M., Flouty, E., Kadkhodamohammadi, A., Quellec, G., Chow, A., Nehme, J., Luengo, I., Stoyanov, D.: Cadis: Cataract dataset for surgical rgb- image segmentation. Medical Image Analysis71, 102053 (2021)
2021
-
[9]
IEEE Transactions on Biomedical Engineering69(3), 1173–1185 (2021)
Guan, H., Liu, M.: Domain adaptation for medical image analysis: a survey. IEEE Transactions on Biomedical Engineering69(3), 1173–1185 (2021)
2021
-
[10]
In: Proceedings of the IEEE/CVF winter conference on applications of computer vi- sion
Hatamizadeh, A., Tang, Y., Nath, V., Yang, D., Myronenko, A., Landman, B., Roth, H.R., Xu, D.: Unetr: Transformers for 3d medical image segmentation. In: Proceedings of the IEEE/CVF winter conference on applications of computer vi- sion. pp. 574–584 (2022)
2022
-
[11]
arXiv preprint arXiv:2012.12453 (2020)
Hong, W.Y., Kao, C.L., Kuo, Y.H., Wang, J.R., Chang, W.L., Shih, C.S.: Cholec- seg8k: a semantic segmentation dataset for laparoscopic cholecystectomy based on cholec80. arXiv preprint arXiv:2012.12453 (2020)
Pith/arXiv arXiv 2012
-
[12]
Iclr1(2), 3 (2022) 10 C
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al.: Lora: Low-rank adaptation of large language models. Iclr1(2), 3 (2022) 10 C. Liu et al
2022
-
[13]
ACM Transactions on Graphics (TOC)37(4), 1–14 (2018)
Hu, Y., Fang, Y., Ge, Z., Qu, Z., Zhu, Y., Pradhana, A., Jiang, C.: A moving least squares material point method with displacement discontinuity and two-way rigid body coupling. ACM Transactions on Graphics (TOC)37(4), 1–14 (2018)
2018
-
[14]
In: ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP)
Huang,H.,Lin,L.,Tong,R.,Hu,H.,Zhang,Q.,Iwamoto,Y.,Han,X.,Chen,Y.W., Wu, J.: Unet 3+: A full-scale connected unet for medical image segmentation. In: ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP). pp. 1055–1059. Ieee (2020)
2020
-
[15]
In: Medical Image Computing and Computer Assisted Inter- vention (MICCAI)
Huang,Y.,Bai,L.,Cui,B.,Yuan,K.,Wang,G.,Hoque,M.I.,Padoy,N.,Navab,N., Ren, H.: Surgtpgs: Semantic 3d surgical scene understanding with text promptable gaussian splatting. In: Medical Image Computing and Computer Assisted Inter- vention (MICCAI). pp. 584–594. Springer (2026)
2026
-
[16]
In: Medical Image Computing and Computer-Assisted Intervention (MICCAI)
Huang, Y., Cui, B., Bai, L., Guo, Z., Xu, M., Islam, M., Ren, H.: Endo-4dgs: Endoscopic monocular scene reconstruction with 4d gaussian splatting. In: Medical Image Computing and Computer-Assisted Intervention (MICCAI). pp. 197–207. Springer (2024)
2024
-
[17]
ACM Trans
Kerbl, B., Kopanas, G., Leimkühler, T., Drettakis, G.: 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph.42(4), 139–1 (2023)
2023
-
[18]
In: Proceedings of the IEEE/CVF international conference on computer vision
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., et al.: Segment anything. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 4015–4026 (2023)
2023
-
[19]
arXiv preprint arXiv:2605.16022 (2026)
Liu, C., Huang, Y., Bai, L., Cui, B., Ren, H.: Endogsim: Physics-aware 4d dynamic endoscopic scene simulations via mllm-guided gaussian splatting. arXiv preprint arXiv:2605.16022 (2026)
Pith/arXiv arXiv 2026
-
[20]
Nature communications15(1), 654 (2024)
Ma, J., He, Y., Li, F., Han, L., You, C., Wang, B.: Segment anything in medical images. Nature communications15(1), 654 (2024)
2024
-
[21]
arXiv preprint arXiv:2504.03600 (2025)
Ma, J., Yang, Z., Kim, S., Chen, B., Baharoon, M., Fallahpour, A., Asakereh, R., Lyu, H., Wang, B.: Medsam2: Segment anything in 3d medical images and videos. arXiv preprint arXiv:2504.03600 (2025)
Pith/arXiv arXiv 2025
-
[22]
In: 2016 fourth international confer- ence on 3D vision (3DV)
Milletari, F., Navab, N., Ahmadi, S.A.: V-net: Fully convolutional neural networks for volumetric medical image segmentation. In: 2016 fourth international confer- ence on 3D vision (3DV). pp. 565–571. Ieee (2016)
2016
-
[23]
In: International conference on machine learning
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. pp. 8748–8763. PmLR (2021)
2021
-
[24]
In: International Conference on Learning Representations
Ravi, N., Gabeur, V., Hu, Y.T., Hu, R., Ryali, C., Ma, T., Khedr, H., Rädle, R., Rolland, C., Gustafson, L., et al.: Sam 2: Segment anything in images and videos. In: International Conference on Learning Representations. vol. 2025, pp. 28085–28128 (2025)
2025
-
[25]
In: International Conference on Medical image computing and computer-assisted intervention
Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedi- cal image segmentation. In: International Conference on Medical image computing and computer-assisted intervention. pp. 234–241. Springer (2015)
2015
-
[26]
Radiology: Artificial Intelligence 5(5), e230024 (2023)
Wasserthal, J., Breit, H.C., Meyer, M.T., Pradella, M., Hinck, D., Sauter, A.W., Heye, T., Boll, D.T., Cyriac, J., Yang, S., et al.: Totalsegmentator: robust segmen- tation of 104 anatomic structures in ct images. Radiology: Artificial Intelligence 5(5), e230024 (2023)
2023
-
[27]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Wu, G., Yi, T., Fang, J., Xie, L., Zhang, X., Wei, W., Liu, W., Tian, Q., Wang, X.: 4d gaussian splatting for real-time dynamic scene rendering. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 20310– 20320 (2024) Parameter-Efficient Adaptation of SAM3 for Surgical Segmentation 11
2024
-
[28]
In: Proceedings of the Computer Vision and Pattern Recognition (CVPR)
Xie, T., Zong, Z., Qiu, Y., Li, X., Feng, Y., Yang, Y., Jiang, C.: Physgaussian: Physics-integrated 3d gaussians for generative dynamics. In: Proceedings of the Computer Vision and Pattern Recognition (CVPR). pp. 4389–4398 (2024)
2024
-
[29]
Medical Image Analysis105, 103644 (2025)
Yuan, K., Srivastav, V., Yu, T., Lavanchy, J.L., Marescaux, J., Mascagni, P., Navab, N., Padoy, N.: Learning multi-modal representations by watching hundreds of surgical video lectures. Medical Image Analysis105, 103644 (2025)
2025
-
[30]
arXiv preprint arXiv:2408.00874 (2024)
Zhu, J., Hamdi, A., Qi, Y., Jin, Y., Wu, J.: Medical sam 2: Segment medical images as video via segment anything model 2. arXiv preprint arXiv:2408.00874 (2024)
Pith/arXiv arXiv 2024
This paper was first reviewed by grok-4.5 on July 30, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.