REVIEW 4 major objections 4 minor 43 references
Parameter Efficient Fine-Tuning of Segment Anything Model for Biomedical Imaging
T0 review · 4 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Finetuning SAM for biomedical images is memory-bound by activations, not trainable parameters, so tuning only the late encoder blocks delivers most of full finetuning's benefit at a fraction of the memory cost.
desk verdict Solid empirical PEFT benchmark for SAM with a genuinely new late-PEFT recipe; the activation-bound claim is well measured in one training regime but overgeneralized without testing checkpointing or larger batches. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the measurement of memory consumption during training, decomposed into parameter gradients versus activation gradients, captured by the ratio of model parameters to sequence length: SAM's ViT-B has a ratio near 20,000, far below GPT-3 or LLaMA, meaning activations dominate. The paper's operational tools are 'late PEFT'—applying LoRA, QLoRA, or full unfreezing only to the last layers of the image encoder (about 50% of the blocks)—and the comparison of nine PEFT methods (LoRA, QLoRA, AdaptFormer, SSF, FacT, bias, layer-norm, attention, and free-encoder tuning) across six microscopy and six medical datasets.
What would settle it
Measure peak training memory for full finetuning versus late finetuning on ViT-L or ViT-H with gradient checkpointing enabled, or with a batch size larger than 1; if late finetuning no longer reduces memory significantly (or full finetuning fits in the same budget), the activation-bound diagnosis and the late-PEFT recommendation would be refuted for those settings.
Extended reading notes
Core claim
The central claim is that PEFT layer placement is more important than layer type for vision transformers. Concretely, the paper finds that finetuning SAM's ViT encoder is constrained by activation memory, not by the number of trainable parameters, and that introducing trainable parameters or unfreezing weights only in the late layers of the encoder produces nearly the same segmentation quality as full finetuning while cutting memory use. It also reports that PEFT does not beat full finetuning even with a single annotated training image, contrary to suggestions in earlier work, and that QLoRA helps mainly when the domain gap to the base model is small.
Load-bearing premise
The conclusion that finetuning is activation-bound rests on memory measurements from a single setup—ViT-B, batch size 1, the LIVECell dataset, and ordinary backpropagation without gradient checkpointing—and the paper generalizes it to all datasets, larger models, and other training settings.
Editorial extensions
If this is right
- Late finetuning of only the last ~50% of ViT blocks cuts training memory (e.g., from 52.1 GB to 45.1 GB for ViT-B with late FT) while staying close to full finetuning in segmentation quality.
- Standard parameter-saving PEFT methods like LoRA, AdaptFormer, SSF, and FacT yield only marginal memory reductions because they do not remove activation gradients.
- Freezing the entire image encoder is the only PEFT choice viable on CPU-scale hardware, while late PEFT is the recommended compromise for consumer GPUs.
- PEFT does not provide a quality advantage over full finetuning even in the one-image training regime, implying that gains from fewer trainable parameters do not materialize for SAM.
- QLoRA performs well only when starting from domain-specific baselines such as muSAM and MedicoSAM and degrades for default SAM, suggesting that quantization hurts more under large domain shifts.
Reading between the lines
- If the activation-bound conclusion generalizes to other ViT-based foundation models, future PEFT designs should target activation memory (freezing early blocks, activation checkpointing, or input downsampling) rather than parameter count.
- The memory comparison may shift under gradient checkpointing or larger batch sizes, so the late-PEFT recipe should be re-tested in those settings before being adopted as a universal rule.
- Extending the same recipe to SAM2 or video-based segmentation will require re-measuring the activation-to-parameter ratio, since longer sequences change the balance the paper identifies.
- A testable prediction is that combining late-layer adapters with early-block freezing will match the memory savings of late PEFT while retaining more tunable capacity at the late layers; the paper's results suggest this but do not explicitly test it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an empirical study of parameter-efficient fine-tuning (PEFT) methods applied to Segment Anything Model (SAM) and two domain-specific SAM variants (µSAM and MedicoSAM) for biomedical image segmentation. It compares additive and selective PEFT methods against full fine-tuning on six microscopy and six medical imaging datasets, reporting segmentation quality, trainable parameter counts, training times, and peak VRAM usage. The paper introduces 'late PEFT' variants in which only late encoder blocks are adapted, and argues that for vision transformers the placement of PEFT layers matters more for efficiency than the type of layer, because SAM fine-tuning is activation-bound rather than parameter-bound. It also proposes a resource-efficient finetuning recipe using one training image and one validation image, with code released.
Significance. If the activation-bound conclusion holds, this is a practically useful result for vision transformers: freezing early encoder blocks and tuning only late blocks can deliver most of the benefit of full fine-tuning at a fraction of the memory cost. The study is broader than prior work on PEFT for SAM, includes detailed ablations, and releases code, which are clear strengths. However, the central efficiency claim rests on a narrow set of measurements (one dataset, one batch size, no gradient checkpointing), and the quality comparisons lack repeated runs or statistical support. The paper's recommendations are therefore somewhat ahead of the evidence.
major comments (4)
- [Appendix A; Tables 5–7] The activation-bound conclusion and the late-PEFT recommendation are inferred from efficiency measurements on a single configuration: LIVECell, ViT-B, batch size 1, and standard backpropagation without gradient checkpointing. Under gradient checkpointing or larger batch sizes, activation memory ceases to be the dominant bottleneck, and the memory advantage of late PEFT (e.g., 9.7 GB vs 13.3 GB for Full FT in Table 8) can shrink or disappear; the paper nevertheless generalizes to ViT-L/ViT-H and to consumer-GPU recommendations. Please add at least one benchmark with gradient checkpointing and one with a larger batch size, or explicitly narrow the claimed scope of the activation-bound conclusion.
- [§3.1; Tables 1–4] All quality comparisons are reported from single runs per method and dataset, without standard deviations or significance tests, yet the conclusions rely on small differences among top methods (e.g., Table 1 AIS: Full FT 0.352, Attn Tune 0.352, LoRA 0.343; Table 3 Point: Full FT 0.511 vs Attn Tune 0.522). The statements that LoRA is the best overall and that late PEFT incurs only a small quality loss are not supportable at this precision without repeated seeds or paired significance testing. Please provide seed variability or otherwise temper the ranking claims.
- [Table 3; Fig. 7] Table 3 reports averages 'over all 9 datasets' but the caption of Fig. 7 states that the late finetuning experiments were only run on the core 6 medical datasets. If Late FT averages are computed over a different subset than the other methods, the comparison 'Late FT is almost as good as Full FT' is not valid. Please recompute all averages on the common subset or report per-dataset results for the late experiments.
- [Appendix B] The QLoRA evaluation at inference uses full-precision pretrained weights plus the learned LoRA adapters instead of the 4-bit quantized weights used during training. The reported QLoRA quality numbers therefore do not correspond to the actual inference-time quantization scheme, and the claim that QLoRA performs well for small domain gaps may not hold for the quantized deployment. Please report both the current approximate evaluation and an evaluation using the true quantized inference, or explicitly characterize the current numbers as an upper bound.
minor comments (4)
- [Table 1] Table 1 contains a duplicated 'AdaptFormer' row (0.334/0.291 AIS); one entry likely belongs to a different method or configuration and should be relabeled.
- [Appendix C] The authors note that the choice of the two manual training images is crucial; no sensitivity analysis over alternative image choices is provided, so the resource-efficient finetuning results should be interpreted with this caveat.
- [Abstract] The abstract states that the placement of PEFT layers is more important for efficiency than the type of layer; this is supported only in the memory dimension and not in segmentation quality, and the wording should be clarified.
- [Fig. 4; Section 3.1] The terms 'LoRA-C' and 'LoRA-A' are used in the main text and figures, but their definitions appear only in Appendix A; please define them at first use.
Circularity Check
No significant circularity: the central efficiency finding and late-PEFT recipe are direct measurements against external baselines (default SAM, CellSeg1, published datasets); self-citations to the authors' μSAM/MedicoSAM are supportive but not load-bearing.
full rationale
The derivation chain is empirical at every load-bearing step. Appendix A states the activation-bound hypothesis from a parameter-to-sequence-length heuristic, then tests it by direct measurement on LIVECell (batch size 1): 'unfreezing the first block of the image encoder reduced memory usage by approximately 500MB, while unfreezing the last block saved around 5GB compared to full fine-tuning.' Tabs. 5–7 measure VRAM, per-iteration time, and trainable parameters directly and show that PEFT type (LoRA 51.2 GB, QLoRA 51.1 GB, AdaptFormer 50.8 GB, FacT 51.2 GB, SSF 53.3 GB for ViT-B) barely moves memory while placement (Late FT 45.1 GB, Late LoRA 43.6 GB, Freeze Encoder 35.5 GB) does. The recommendation tiering (Freeze Encoder for CPU, Late FT for consumer GPU, Full FT for server GPU) is a summary of measured memory (Tab. 8: 5.8/9.7/13.3 GB) and measured quality (Figs. 3–4), not a quantity derived from the inputs by construction. Quality comparisons are made against external benchmarks: default SAM, the published CellSeg1 workflow, and 12 published datasets. The paper's reliance on the authors' own μSAM (Archit et al., 2025a) and MedicoSAM (Archit et al., 2025b) as base models, and on μSAM's earlier freeze-encoder result, is self-citation, but it is not load-bearing: the central placement-vs-type and activation-bound findings are established on default SAM (App. A, Tabs. 5–7), and the freeze-encoder quality claim is independently re-measured here (Tabs. 1–4), so no central claim reduces to an unverified self-citation. Stated limitations—efficiency reported 'only for a single dataset' (App. E.2), recommendations 'only tested for finetuning SAM' (App. E.3), and the single GPU configuration without gradient checkpointing flagged by the reviewer—are generality caveats about external validity, not circular reductions. The one adjacent concern is that the 50% late-PEFT setting was selected on the same datasets later used for the headline quality claim (Sec. 3.1 and Figs. 13–14); that is test-set-informed model selection, which could inflate reported quality, but the efficiency argument is structural (measured VRAM) and the quality numbers are measurements rather than predictions forced by construction, so per the review rules it does not constitute circularity. Score 2 reflects the presence of minor, non-load-bearing self-citation; the derivation itself is self-contained against external benchmarks.
Assumptions & free parameters
free parameters (6)
- LoRA rank =
32
- LoRA alpha =
1
- Learning rate =
1e-5
- AdaptFormer projection size =
64
- FacT rank =
16
- Late PEFT layer fraction =
50%
assumptions (6)
- domain assumption SAM fine-tuning memory usage is dominated by activation gradients, not trainable parameter count
- domain assumption Efficiency measurements on LIVECell generalize to other datasets
- domain assumption Segmentation quality trends observed with ViT-B transfer to ViT-L and ViT-H
- domain assumption Simulated interactive prompting (automatic prompts with 7-8 correction iterations) reflects real interactive use
- ad hoc to paper The two manually selected training images are representative of each dataset
- ad hoc to paper QLoRA inference with full-precision weights plus LoRA (instead of quantized weights) is a valid way to evaluate the finetuned model
Cite this review
Pith. "Pith review of Parameter Efficient Fine-Tuning of Segment Anything Model for Biomedical Imaging." pith.science (2026). https://pith.science/paper/YPJKJ7KC
@misc{pith2026250200418,
author = {Pith},
title = {Pith review of: Parameter Efficient Fine-Tuning of Segment Anything Model for Biomedical Imaging},
year = {2026},
howpublished = {\url{https://pith.science/paper/YPJKJ7KC}},
note = {Machine review of arXiv:2502.00418}
}
read the original abstract
Segmentation is an important analysis task for biomedical images, enabling the study of individual organelles, cells or organs. Deep learning has massively improved segmentation methods, but challenges remain in generalization to new conditions, requiring costly data annotation. Vision foundation models, such as Segment Anything Model (SAM), address this issue through improved generalization. However, these models still require finetuning on annotated data, although with less annotations, to achieve optimal results for new conditions. As a downside, they require more computational resources. This makes parameter-efficient finetuning (PEFT) relevant. We contribute the first comprehensive study of PEFT for SAM applied to biomedical images. We find that the placement of PEFT layers is more important for efficiency than the type of layer for vision transformers and we provide a recipe for resource-efficient finetuning. Our code is publicly available at https://github.com/computational-cell-analytics/peft-sam.
Figures
Figures from the paper (16 more)
Reference graph
Works this paper leans on
-
[1]
Segment anything for microscopy.Nature Methods, 22(3):579–591, February 2025a
Anwai Archit, Luca Freckmann, Sushmita Nair, Nabeel Khalid, Paul Hilt, Vikas Ra- jashekar, Marei Freitag, Carolin Teuber, Genevieve Buckley, Sebastian von Haaren, Sag- nik Gupta, Andreas Dengel, Sheraz Ahmed, and Constantin Pape. Segment anything for microscopy.Nature Methods, 22(3):579–591, February 2025a. ISSN 1548-7105. doi: 10.1038/s41592-024-02580-4....
-
[3]
doi: 10.1609/ aaai.v38i10.28978
ISSN 2159-5399. doi: 10.1609/ aaai.v38i10.28978. URLhttp://doi.org/10.1609/aaai.v38i10.28978. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Zieg...
-
[5]
Curran Associates Inc. ISBN 9781713829546. URLhttps://proceedings.neurips.cc/paper/2020/file/ 81f7acabd411274fcf65ce2070ed568a-Paper.pdf. Juan C. Caicedo, Allen Goodman, Kyle W. Karhohs, Beth A. Cimini, Jeanelle Acker- man, Marzieh Haghighi, CherKeng Heng, Tim Becker, Minh Doan, Claire McQuin, Mohammad Rohban, Shantanu Singh, and Anne E. Carpenter. Nucleu...
work page 2020
-
[6]
Yanlin Wu, Zhihong Wang, Xiongfeng Yang, Hong Kang, Along He, and Tao Li
URLhttp://doi.org/10.1007/978-3-031-72684-2_6. Yanlin Wu, Zhihong Wang, Xiongfeng Yang, Hong Kang, Along He, and Tao Li. Trans-sam: Transfer segment anything model to medical image segmentation with parameter-efficient fine-tuning.Knowledge-Based Systems, 310:112909,
-
[9]
URLhttps://doi. org/10.48550/arXiv.1810.04805. Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale.ICLR,
-
[10]
URLhttp://doi.org/10.1038/s41597-024-03634-0
1038/s41597-024-03634-0. URLhttp://doi.org/10.1038/s41597-024-03634-0. Hanxue Gu, Haoyu Dong, Jichen Yang, and Maciej A. Mazurowski. How to build the best medical image segmentation algorithm using foundation models: a comprehensive empirical study with segment anything model.Machine Learning for Biomedical Imaging, 3(May 2025):88–120, May
-
[11]
ISSN 2052-4463. doi:
-
[14]
doi: 10.1038/s41597-024-03844-6
ISSN 2052-4463. doi: 10.1038/s41597-024-03844-6. URL http://doi.org/10.1038/s41597-024-03844-6. Fabian Isensee, Paul F. Jaeger, Simon A. A. Kohl, Jens Petersen, and Klaus H. Maier- Hein. nnu-net: a self-configuring method for deep learning-based biomedical image segmentation.Nature Methods, 18(2):203–211, December
Show all 43 references
-
[15]
doi: 10.1038/s41592-020-01008-z
ISSN 1548-7105. doi: 10.1038/s41592-020-01008-z. URLhttp://doi.org/10.1038/s41592-020-01008-z. Uriah Israel, Markus Marks, Rohit Dilip, Qilin Li, Changhua Yu, Emily Laubscher, Shenyi Li, Morgan Schwartz, Elora Pradhan, Ada Ates, et al. A foundation model for cell 12 PEFT-SAM s...
-
[16]
11.17.567630
URLhttps://doi.org/10.1101/2023. 11.17.567630. Malte Jensen, Andreas Clemmensen, Jacob Gorm Hansen, Julie van Krimpen Mortensen, Emil N. Christensen, Andreas Kjaer, and Rasmus Sejersten Ripa. 3d whole body pre- clinical micro-ct database of subcutaneous tumors in mice with ann...
2023 doi
-
[17]
doi: 10.1038/ s41597-024-03814-y
ISSN 2052-4463. doi: 10.1038/ s41597-024-03814-y. URLhttp://doi.org/10.1038/s41597-024-03814-y. Shibo Jie and Zhi-Hong Deng. Fact: Factor-tuning for lightweight adaptation on vi- sion transformer.Proceedings of the AAAI Conference on Artificial Intelligence, 37 (1):1060–1068, June
-
[18]
doi: 10.1609/aaai.v37i1.25187
ISSN 2159-5399. doi: 10.1609/aaai.v37i1.25187. URL http://doi.org/10.1609/aaai.v37i1.25187. Florian Jug and Bioicons Contributors. Bioicons - free biology icons,
-
[19]
doi: 10.1038/s41597-022-01388-1
ISSN 2052-4463. doi: 10.1038/s41597-022-01388-1. URL http://doi.org/10.1038/s41597-022-01388-1. Juliet W. Lefferts, Suzanne Kroes, Matthew B. Smith, Paul J. Niem¨ oller, Natascha D. A. Nieuwenhuijze, Heleen N. Sonneveld van Kooten, Cornelis K. van der Ent, Jeffrey M. Beekman, ...
-
[20]
doi: 10.1038/s42003-024-05966-4
ISSN 2399-3642. doi: 10.1038/s42003-024-05966-4. URLhttps://doi.org/ 10.1038/s42003-024-05966-4. Dongze Lian, Daquan Zhou, Jiashi Feng, and Xinchao Wang. Scaling & shifting your features: A new baseline for efficient model tuning. InAdvances in Neural Information Processing Sy...
-
[21]
Jun Ma, Yuting He, Feifei Li, Lin Han, Chenyu You, and Bo Wang
URLhttps://papers.neurips.cc/paper_files/ paper/2022/file/00bb4e415ef117f2dee2fc3b778d806d-Paper-Conference.pdf. Jun Ma, Yuting He, Feifei Li, Lin Han, Chenyu You, and Bo Wang. Segment anything in medical images.Nature Communications, 15(1):654,
2022
-
[22]
doi: 10.1038/s41592-019-0658-6
ISSN 1548-7105. doi: 10.1038/s41592-019-0658-6. URLhttps://doi.org/10.1038/s41592-019-0658-6. Marius Pachitariu and Carsen Stringer. Cellpose 2.0: how to train your own model. Nature Methods, 19(12):1634–1641, December
-
[23]
doi: 10.1038/ s41592-022-01663-4
ISSN 1548-7105. doi: 10.1038/ s41592-022-01663-4. URLhttps://doi.org/10.1038/s41592-022-01663-4. Constantin Pape, Roman Remme, Adrian Wolny, Sylvia Olberg, Steffen Wolf, Lorenzo Cerrone, Mirko Cortese, Severina Klaus, Bojana Lucic, Stephanie Ullrich, Maria Anders- ¨Osswein, St...
-
[24]
George Pu, Anirudh Jain, Jihan Yin, and Russell Kaplan
URLhttps://doi.org/10.1002/bies.202000257. George Pu, Anirudh Jain, Jihan Yin, and Russell Kaplan. Empirical analysis of the strengths and weaknesses of peft techniques for llms,
- [27]
-
[29]
URLhttp://doi.org/10.1007/978-3-031-20053-3_29. 14 PEFT-SAM Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, ...
-
[30]
URLhttps://doi.org/10.48550/arXiv.2307. 09288. Hernando M. Vergara, Constantin Pape, Kimberly I. Meechan, Valentyna Zinchenko, Chris- tel Genoud, Adrian A. Wanner, Kevin Nzumbi Mutemi, Benjamin Titze, Rachel M. Templin, Paola Y. Bertucci, Oleg Simakov, Wiebke D¨ urichen, Pedro...
-
[31]
doi: 10.1016/j.cell.2021.07.017
ISSN 0092-8674. doi: 10.1016/j.cell.2021.07.017. URLhttps://doi.org/10.1016/j.cell.2021.07.017. Athul Vijayan, Tejasvinee Atul Mody, Qin Yu, Adrian Wolny, Lorenzo Cerrone, Soeren Strauss, Miltos Tsiantis, Richard S. Smith, Fred A. Hamprecht, Anna Kreshuk, and Kay Schneitz. A d...
2021 doi
-
[32]
doi: 10.1242/dev.202800
ISSN 0950-1991. doi: 10.1242/dev.202800. URLhttps://doi.org/10.1242/dev.202800. Jakob Wasserthal, Hanns-Christian Breit, Manfred T Meyer, Maurice Pradella, Daniel Hinck, Alexander W Sauter, Tobias Heye, Daniel T Boll, Joshy Cyriac, Shan Yang, et al. Totalsegmentator: robust se...
1991 doi
-
[33]
Xiaobao Wei, Jiajun Cao, Yizhu Jin, Ming Lu, Guangyu Wang, and Shanghang Zhang.I- MedSAM: Implicit Medical Image Segmentation with Segment Anything, page 90–107
URLhttps://doi.org/10.1148/ryai.230024. Xiaobao Wei, Jiajun Cao, Yizhu Jin, Ming Lu, Guangyu Wang, and Shanghang Zhang.I- MedSAM: Implicit Medical Image Segmentation with Segment Anything, page 90–107. Springer Nature Switzerland, November
-
[34]
1016/j.knosys.2024.112909
URLhttps://doi.org/10. 1016/j.knosys.2024.112909. Yi Xin, Siqi Luo, Haodi Zhou, Junlong Du, Xiaohong Liu, Yue Fan, Qing Li, and Yuntao Du. Parameter-efficient fine-tuning for pre-trained vision models: A survey,
2024
-
[35]
15 Teuber Archit Pape Lingling Xu, Haoran Xie, Si-Zhao Joe Qin, Xiaohui Tao, and Fu Lee Wang
URL https://doi.org/10.48550/arXiv.2402.02242. 15 Teuber Archit Pape Lingling Xu, Haoran Xie, Si-Zhao Joe Qin, Xiaohui Tao, and Fu Lee Wang. Parameter- efficient fine-tuning methods for pretrained language models: A critical review and as- sessment,
- [36]
-
[37]
Theodore Zhao, Yu Gu, Jianwei Yang, Naoto Usuyama, Ho Hin Lee, Sid Kiblawi, Tristan Naumann, Jianfeng Gao, Angela Crabtree, Jacob Abel, et al
URLhttps://doi.org/10.48550/arXiv.2304.13785. Theodore Zhao, Yu Gu, Jianwei Yang, Naoto Usuyama, Ho Hin Lee, Sid Kiblawi, Tristan Naumann, Jianfeng Gao, Angela Crabtree, Jacob Abel, et al. A foundation model for joint segmentation, detection and recognition of biomedical objec...
-
[38]
Peilin Zhou, Bo Du, and Yongchao Xu
URLhttps://doi.org/10.1038/s41592-024-02499-w. Peilin Zhou, Bo Du, and Yongchao Xu. Cellseg1: Robust cell segmentation with one training image,
-
[39]
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee
URLhttps://doi.org/10.48550/arXiv.2412.01410. Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee. Segment everything everywhere all at once. InPro- ceedings of the 37th International Conference on Neural Informat...
-
[40]
URLhttps: //openreview.net/forum?id=UHBrWeFWlL
Curran Associates Inc. URLhttps: //openreview.net/forum?id=UHBrWeFWlL. Appendix A. Efficiency of PEFT For the optimal use of PEFT methods in vision transformers, it is important to understand the main contributors to the memory footprint during training. Unlike typical LLM tra...
2020
-
[41]
LoRA Classic
suggest that once memory constraints are taken care of, per- formance can be further optimized by introducing additional trainable parameters through more low-rank adapters. Rather than limiting LoRA to the query and value matrices of the attention blocks, which is a common pr...
2022
-
[42]
has shown that 17 Teuber Archit Pape the choice of image is crucial in this setting. Once these two images are selected, we use 1 image for training and the other exclusively for validation, and test our trained model on the entire test set, for consistency with other experime...
2019
-
[43]
Note that we perform all experiments for segmentation quality with the smallest model, ViT-B, because it was shown e.g
dataset, a large dataset with annotations for cell segmentation in phase-contrast microscopy. Note that we perform all experiments for segmentation quality with the smallest model, ViT-B, because it was shown e.g. in (Gu et al., 2025; Archit et al., 2025b) that using it does n...
2025
-
[2000]
doi: 10.2214/ajr.174.1.1740071
ISSN 1546-3141. doi: 10.2214/ajr.174.1.1740071. URLhttp://doi.org/10.2214/ajr.174. 1.1740071. Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby, Jakob Verbeek, and Herv´ e J´ egou. Three Things Everyone Should Know About Vision Transformers, page 497–515. Springer Nature Switzerland,
- [2018]
-
[2019]
doi: 10.1038/s41592-019-0612-7
ISSN 1548-7105. doi: 10.1038/s41592-019-0612-7. URLhttps: //doi.org/10.1038/s41592-019-0612-7. Cheng Chen, Juzheng Miao, Dufan Wu, Aoxiao Zhong, Zhiling Yan, Sekeun Kim, Jiang Hu, Zhengliang Liu, Lichao Sun, Xiang Li, et al. Ma-sam: Modality-agnostic sam adaptation for 3d medi...
-
[2020]
10 PEFT-SAM Han Cai, Chuang Gan, Ligeng Zhu Massachusetts Institute of Technology, and Song Han Massachusetts Institute of Technology
URLhttps://proceedings.neurips.cc/paper_files/ paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf. 10 PEFT-SAM Han Cai, Chuang Gan, Ligeng Zhu Massachusetts Institute of Technology, and Song Han Massachusetts Institute of Technology. Tinytl: reduce memory, not paramete...
2020
-
[2021]
URLhttps://doi.org/10.1038/s41592-021-01249-6
doi: 10.1038/s41592-021-01249-6. URLhttps://doi.org/10.1038/s41592-021-01249-6. Database Center for Life Science (DBCLS). Togotv - life science video portal,
-
[2022]
Ryan Conrad and Kedar Narayan
URLhttps://proceedings.neurips.cc/paper_files/paper/2022/ file/69e2f49ab0837b71b0e0cb7c555990f8-Paper-Conference.pdf. Ryan Conrad and Kedar Narayan. Instance segmentation of mitochondria in electron mi- croscopy images with a generalist deep learning model trained on a diverse...
2022
-
[2023]
doi: 10.1016/j.cels.2022.12.006
ISSN 2405-4712. doi: 10.1016/j.cels.2022.12.006. URL https://doi.org/10.1016/j.cels.2022.12.006. Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. Qlora: Effi- cient finetuning of quantized llms,
2022 doi
-
[2024]
Samyadeep Basu, Shell Hu, Daniela Massiceti, and Soheil Feizi
URLhttps://doi.org/10.48550/arXiv.2404.13506. Samyadeep Basu, Shell Hu, Daniela Massiceti, and Soheil Feizi. Strong baselines for parameter-efficient few-shot fine-tuning.Proceedings of the AAAI Conference on Ar- tificial Intelligence, 38(10):11024–11031, March
-
[2025]
doi: 10.59275/j.melba.2025-86a6
ISSN 2766-905X. doi: 10.59275/j.melba.2025-86a6. URL http://doi.org/10.59275/j.melba.2025-86a6. Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath, Dong Yang, Andriy Myronenko, Bennett Landman, Holger R. Roth, and Daguang Xu. Unetr: Transformers for 3d medical image segmentation. In...
2025 doi
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.