REVIEW 4 major objections 4 minor 23 references
SegBook: A Simple Baseline and Cookbook for Volumetric Medical Image Segmentation
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Fine-tuning a full-body CT pretrained model yields larger Dice gains on small and large downstream datasets than on medium-sized ones, a 'bottleneck effect' the authors attribute to dataset size.
desk verdict Useful benchmark, but the headline 'bottleneck effect' does not survive scrutiny; the dataset collection and modality-transfer results are the real contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
STU-Net, a scalable U-Net variant with base/large/huge configurations, pretrained in fully supervised fashion on TotalSegmentator (1204 CT volumes, 104 anatomical structures). Fine-tuning initializes the segmentation head randomly, uses 10x learning rate for the head, and follows nnU-Net's automatic preprocessing and patch-based training. The benchmark's dataset-size buckets (boundaries at 100 and 400 cases) are the analytical tool that exposes the bottleneck effect.
What would settle it
Recompute the fine-tuning gain on a size-balanced subset of the 87 datasets where each bucket contains the same proportion of modalities and target types (e.g., only CT structure tasks), and check whether the U-shaped gain pattern persists. If the medium bucket no longer shows a dip, the bottleneck effect is an artifact of confounding.
Extended reading notes
Core claim
The central claim is that the benefit of supervised pretraining on full-body CT is not monotonic in downstream dataset size. Across 87 datasets, average Dice gain from fine-tuning STU-Net is roughly 3% for datasets with fewer than 100 cases and more than 400 cases, but only about 1% for datasets in between (100–400). The authors call this a 'bottleneck effect'. They also show positive transfer across modalities (CT to MRI, PET, ultrasound) and across target types (structure to lesion), with larger models generally gaining more.
Load-bearing premise
The bottleneck effect assumes that the size buckets (fewer than 100, 100–400, more than 400 cases) isolate the effect of dataset size; if medium-sized datasets in this collection happen to contain harder tasks (more lesions, more MRI), the observed U-shape could be caused by task composition rather than size.
Editorial extensions
If this is right
- Full-body CT pretraining can serve as a general initialization for MRI, PET, and ultrasound segmentation, reducing the need for task-specific pretraining.
- For datasets with 100–400 cases, fine-tuning yields only modest gains, so practitioners may need other strategies (longer fine-tuning, different regularization) to unlock transfer benefits.
- Larger STU-Net models show greater absolute gains from pretraining, suggesting model capacity helps when transferring to diverse tasks.
- The 87-dataset benchmark provides a public testbed for comparing future transfer learning methods in volumetric segmentation.
Reading between the lines
- The bottleneck effect might be a composition artifact: medium-sized datasets in this collection may be enriched for hard targets (lesions, MRI), which would make the U-shape a property of task difficulty rather than dataset size; a matched analysis controlling for modality and target would settle this.
- If the bottleneck is real, it suggests a 'transfer gap': small datasets overfit and benefit from the pretrained prior, large datasets allow the model to learn the new task fully, while intermediate datasets are too large to rely on the prior but too small to override it cleanly.
- The finding that structure pretraining helps lesion segmentation hints that anatomical priors capture general image features; testing whether the gain holds for lesions in non-CT modalities would sharpen this claim.
- The benchmark's public release could be used to test whether the U-shaped gain persists across other architectures and pretraining strategies, such as self-supervised pretraining.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SegBook, a benchmark of 87 public volumetric medical image segmentation datasets, and uses STU-Net (pre-trained on TotalSegmentator full-body CT) as a representative model to study transfer learning across dataset sizes, modalities, targets, and model scales. The authors report three main findings: (1) a non-linear "bottleneck effect" in which fine-tuning gains are larger for small and large datasets than for medium ones; (2) effective transfer from CT to other modalities such as MRI, including for targets unseen during pre-training; and (3) adaptability of CT structural pre-training to both structure and lesion segmentation targets. The paper also provides a task-specific nnU-Net baseline and compares STU-Net at base, large, and huge scales with and without pre-training.
Significance. If the bottleneck effect is real, the paper would document a practically important non-linearity in how supervised pre-training benefits downstream volumetric segmentation, with direct implications for dataset collection and fine-tuning strategy. The benchmark itself, spanning 87 datasets with diverse modalities, targets, and sizes, is a substantial resource, and the open release of the evaluation protocol and results would support future comparisons. The paper also provides a useful empirical demonstration that full-body CT pre-training transfers across modalities and to lesion targets, which is valuable evidence for the community. However, the headline bottleneck claim currently rests on uncontrolled bin averages and lacks statistical support, so the benchmark's ultimate value depends on strengthening that analysis.
major comments (4)
- [Section 3.1, Table 2] The central bottleneck claim is not statistically supported. Table 2 reports mean Dice gains per dataset-size bucket without any error bars, confidence intervals, or significance tests, and the text in Section 3.1 interprets differences of roughly 1.5-2% versus 2.2-3.7% as a clear non-linear pattern. Given that these are averages over heterogeneous datasets, the claim needs at least per-dataset improvement values, a variance estimate, and a test (e.g., a permutation or bootstrap test) to establish that the medium-bin deficit is not within sampling noise.
- [Section 2.1, Figure 3] The dataset-size boundaries of 100 and 400 cases are chosen without justification or sensitivity analysis. Since the non-monotonic pattern in Table 2 depends entirely on which datasets fall into each bin, the authors should show that the bottleneck effect persists under alternative thresholds (e.g., different small/medium/large cutoffs) and ideally treat dataset size as a continuous covariate rather than a hand-binned factor.
- [Section 3.1, Table 2 and Table 5] The medium-bin average gain may be confounded by task composition. Table 5 shows that the medium bin contains a mix of CT and MRI, structure and lesion targets, and tasks of varying difficulty; the paper never disaggregates the per-dataset gains by modality, target type, label count, or baseline difficulty. Because the bottleneck effect is non-monotonic, it is load-bearing that the medium-bin deficit survives when these factors are controlled. The authors should provide a stratified analysis or a regression that includes these covariates.
- [Section 3.1, Summary paragraph] The phrase "significantly higher performance gains" in the Summary is too strong given the absence of any statistical significance testing. Please rephrase to reflect the observed magnitude and the need for inferential validation, or add the missing analyses.
minor comments (4)
- [Section 2.1, Table 1] The column headers contain typos: "SeenSturctute" should be "SeenStructure" and "Lesion&Sturcture" should be "Lesion&Structure".
- [Section 3.3.2] In the sentence about US transfer, the text reads "with ∆(huge) of 0.84, 2.47, and 3.40 in STU-Net-base, STU-Net-large and STU-Net, respectively"; this should say "STU-Net-base, STU-Net-large, and STU-Net-huge" for clarity.
- [Section 4.1] The conclusion restates findings (1)-(3) but does not mention the limitations of the binning analysis or the absence of statistical tests; a sentence acknowledging these caveats would improve balance.
- [Section 2.3] The transfer-learning setup uses a 10x learning rate for the segmentation head, but the paper does not state whether the same hyperparameters were used for training from scratch and for all model scales; please clarify to rule out a hyperparameter confound.
Circularity Check
No significant circularity: the paper's claims are empirical measurements on 87 external public datasets; STU-Net self-citation is not load-bearing.
full rationale
This paper is an empirical benchmark, not a derivation. The central claims — the 'bottleneck effect' and modality/target transferability — are descriptive summaries of measured Dice improvements when fine-tuning STU-Net pretrained on TotalSegmentator across 87 public downstream datasets. There is no equation in which the predicted quantity is defined as, or fitted from, the input quantity. The pretrained model (STU-Net) is developed by overlapping authors, but it is not used as a source of theoretical justification; rather, it is evaluated against external public datasets and a from-scratch nnU-Net baseline. The choice of 100/400-case bin boundaries is arbitrary and could affect the interpretation of the bottleneck effect, but this is a methodological/confounding concern, not circularity: the binned Dice deltas are reported measurements, not quantities constructed to equal the inputs by definition. No self-citation is invoked to forbid alternatives or to justify a uniqueness claim. Therefore no circular step meets the evidentiary standard required by the review rules.
Assumptions & free parameters
free parameters (2)
- dataset scale boundaries =
100 and 400 cases
- head learning rate multiplier =
10x
assumptions (4)
- domain assumption DSC is a comparable metric across heterogeneous tasks and object sizes.
- domain assumption The 87 collected datasets are a representative sample for evaluating transfer in volumetric medical segmentation.
- domain assumption The nnU-Net automatic pipeline provides a fair and equally appropriate baseline for every downstream dataset.
- domain assumption TotalSegmentator supervised pre-training represents full-body CT pre-training in general.
Cite this review
Pith. "Pith review of SegBook: A Simple Baseline and Cookbook for Volumetric Medical Image Segmentation." pith.science (2026). https://pith.science/paper/OMALYYQ3
@misc{pith2026241114525,
author = {Pith},
title = {Pith review of: SegBook: A Simple Baseline and Cookbook for Volumetric Medical Image Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/OMALYYQ3}},
note = {Machine review of arXiv:2411.14525}
}
read the original abstract
Computed Tomography (CT) is one of the most popular modalities for medical imaging. By far, CT images have contributed to the largest publicly available datasets for volumetric medical segmentation tasks, covering full-body anatomical structures. Large amounts of full-body CT images provide the opportunity to pre-train powerful models, e.g., STU-Net pre-trained in a supervised fashion, to segment numerous anatomical structures. However, it remains unclear in which conditions these pre-trained models can be transferred to various downstream medical segmentation tasks, particularly segmenting the other modalities and diverse targets. To address this problem, a large-scale benchmark for comprehensive evaluation is crucial for finding these conditions. Thus, we collected 87 public datasets varying in modality, target, and sample size to evaluate the transfer ability of full-body CT pre-trained models. We then employed a representative model, STU-Net with multiple model scales, to conduct transfer learning across modalities and targets. Our experimental results show that (1) there may be a bottleneck effect concerning the dataset size in fine-tuning, with more improvement on both small- and large-scale datasets than medium-size ones. (2) Models pre-trained on full-body CT demonstrate effective modality transfer, adapting well to other modalities such as MRI. (3) Pre-training on the full-body CT not only supports strong performance in structure detection but also shows efficacy in lesion detection, showcasing adaptability across target tasks. We hope that this large-scale open evaluation of transfer learning can direct future research in volumetric medical image segmentation.
Figures
Reference graph
Works this paper leans on
-
[1]
Medical Image Analysis, 83:102628,
Crossmoda 2021 challenge: Benchmark of cross-modality domain adaptation techniques for vestibular schwannoma and cochlea segmentation. Medical Image Analysis, 83:102628,
work page 2021
-
[6]
Y . Deng, C. Wang, Y . Hui, Q. Li, J. Li, S. Luo, M. Sun, Q. Quan, S. Yang, Y . Hao, et al. Ctspine1k: a large-scale dataset for spinal vertebrae segmentation in computed tomography. arXiv preprint arXiv:2105.14711,
-
[8]
M. R. Hernandez Petzsche, E. de la Rosa, U. Hanning, R. Wiest, W. Valenzuela, M. Reyes, M. Meyer, S.-L. Liew, F. Kofler, I. Ezhov, et al. Isles 2022: A multi-center magnetic resonance imaging stroke lesion segmentation dataset. Scientific data, 9(1):762,
work page 2022
- [9]
-
[10]
A. F. Kazerooni, N. Khalili, X. Liu, D. Haldar, Z. Jiang, S. M. Anwar, J. Albrecht, M. Adewole, U. Anazodo, H. Anderson, S. Bagheri, U. Baid, T. Bergquist, A. J. Borja, E. Calabrese, V . Chung, G.-M. Conte, F. Dako, J. Eddy, I. Ezhov, A. Familiar, K. Farahani, S. Haldar, J. E. Iglesias, A. Janas, E. Johansen, B. V . Jones, F. Kofler, D. LaBella, H. A. Lai...
work page 2023
-
[12]
Z. Lambert, C. Petitjean, B. Dubray, and S. Kuan. Segthor: Segmentation of thoracic organs at risk in ct images. In 2020 Tenth International Conference on Image Processing Theory, Tools and Applications (IPTA), pages 1–6. IEEE,
work page 2020
-
[13]
X. Li, G. Luo, K. Wang, H. Wang, J. Liu, X. Liang, J. Jiang, Z. Song, C. Zheng, H. Chi, et al. The state-of-the-art 3d anisotropic intracranial hemorrhage segmentation on non-contrast head ct: The instance challenge. arXiv preprint arXiv:2301.03281,
-
[14]
G. Luo, K. Wang, J. Liu, S. Li, X. Liang, X. Li, S. Gan, W. Wang, S. Dong, W. Wang, et al. Efficient automatic segmentation for multi-level pulmonary arteries: The parse challenge. arXiv preprint arXiv:2304.03708, 2023a. X. Luo and X. Zhuang. X-metric: An n-dimensional information-theoretic framework for groupwise registration and deep combined computing....
Show all 23 references
-
[15]
X. Luo, J. Fu, Y . Zhong, S. Liu, B. Han, M. Astaraki, S. Bendazzoli, I. Toma-Dasu, Y . Ye, Z. Chen, et al. Segrap2023: A benchmark of organs-at-risk and gross tumor volume segmentation for radiotherapy planning of nasopharyngeal carcinoma. arXiv preprint arXiv:2312.09576, 202...
-
[16]
J. Ma, Y . Zhang, S. Gu, C. Ge, S. Ma, A. Young, C. Zhu, K. Meng, X. Yang, Z. Huang, et al. Unleashing the strengths of unlabeled data in pan-cancer abdominal organ quantification: the flare22 challenge. arXiv preprint arXiv:2308.05862,
-
[17]
Menze, A
12 B. Menze, A. Jakab, S. Bauer, J. Kalpathy-Cramer, K. Farahani, J. Kirby, Y . Burren, N. Porz, J. Slotboom, R. Wiest, L. Lanczi, E. Gerstner, M.-A. Weber, T. Arbel, B. Avants, N. Ayache, P. Buendia, L. Collins, N. Cordier, J. Corso, A. Criminisi, T. Das, H. Delingette, C. De...
-
[18]
A. W. Moawad, A. Janas, U. Baid, D. Ramakrishnan, L. Jekel, K. Krantchev, H. Moy, R. Saluja, K. Osenberg, K. Wilms, M. Kaur, A. Avesta, G. C. Pedersen, N. Maleki, M. Salimi, S. Merkaj, M. von Reppert, N. Tillmans, J. Lost, K. Bousabarah, W. Holler, M. Lin, M. Westerhoff, R. Ma...
2023
-
[19]
M. L. N. Q. L. C. Organizers. Mediastinal lymph node quantification (lnq): Segmentation of heteroge- neous ct data. LNQ 2023 Grand Challenge, 2023a. URL https://lnq2023.grand-challenge. org/lnq2023/. Accessed: 2024-11-06. S.-U. C. . Organizers. Smile-uhura challenge 2023: Smal...
2023
-
[20]
Pedrosa, G
J. Pedrosa, G. Aresta, C. Ferreira, M. Rodrigues, P. Leitão, A. S. Carvalho, J. Rebelo, E. Negrão, I. Ramos, A. Cunha, et al. Lndb: a lung nodule database on computed tomography. arXiv preprint arXiv:1911.08434,
1911 arXiv
-
[21]
Podobnik, P
G. Podobnik, P. Strojan, P. Peterlin, B. Ibragimov, and T. Vrtovec. Han-seg: The head and neck organ-at-risk ct and mr segmentation dataset. Medical physics, 50(3):1917–1927,
1917
-
[23]
DOI: 10.7937/QSTF- ST65
URL https://doi.org/10.7937/QSTF-ST65. DOI: 10.7937/QSTF- ST65. S. Wang, C. Qin, C. Wang, K. Wang, H. Wang, C. Chen, C. Ouyang, X. Kuang, C. Dai, Y . Mo, Z. Shi, C. Dai, X. Chen, H. Wang, and W. Bai. The extreme cardiac mri analysis challenge under respiratory motion (cmrxmotion),
-
[2015]
2015.zF0vlOPv
URL http://doi.org/10.7937/K9/TCIA. 2015.zF0vlOPv. DOI: 10.7937/K9/TCIA.2015.zF0vlOPv. A. Carass, S. Roy, A. Jog, J. L. Cuzzocreo, E. Magrath, A. Gherman, J. Button, J. Nguyen, F. Prados, C. H. Sudre, et al. Longitudinal multiple sclerosis lesion segmentation: resource and cha...
2015 doi
-
[2019]
LaBella, M
11 D. LaBella, M. Adewole, M. Alonso-Basanta, T. Altes, S. M. Anwar, U. Baid, T. Bergquist, R. Bhalerao, S. Chen, V . Chung, G.-M. Conte, F. Dako, J. Eddy, I. Ezhov, D. Godfrey, F. Hilal, A. Familiar, K. Farahani, J. E. Iglesias, Z. Jiang, E. Johanson, A. F. Kazerooni, C. Kent...
2023
-
[2021]
doi: 10.1016/j.media.2020.101821
ISSN 1361-8415. doi: 10.1016/j.media.2020.101821. URL https://www. sciencedirect.com/science/article/pii/S1361841520301857. N. Heller, F. Isensee, D. Trofimova, R. Tejpaul, Z. Zhao, H. Chen, L. Wang, A. Golts, D. Khapun, D. Shats, et al. The kits21 challenge: Automatic segment...
2020
-
[2022]
U. Baid, S. Ghodasara, S. Mohan, M. Bilello, E. Calabrese, E. Colak, K. Farahani, J. Kalpathy- Cramer, F. C. Kitamura, S. Pati, et al. The rsna-asnr-miccai brats 2021 benchmark on brain tumor segmentation and radiogenomic classification. arXiv preprint arXiv:2107.02314,
2021 arXiv
-
[2023]
Bloch, A
N. Bloch, A. Madabhushi, H. Huisman, J. Freymann, J. Kirby, M. Grauer, A. Enquobahrie, C. Jaffe, L. Clarke, and K. Farahani. Nci-isbi 2013 challenge: Automated segmentation of prostate structures. The Cancer Imaging Archive,
2013
-
[2024]
Styner, J
M. Styner, J. Lee, B. Chin, M. Chin, O. Commowick, H. Tran, S. Markovic-Plese, V . Jewells, and S. Warfield. 3d segmentation in the clinic: A grand challenge ii: Ms lesion segmentation. MIDAS journal, 2008:1–6,
2008
-
[8415]
doi: https://doi.org/10.1016/j.media.2022.102628. M. Adewole, J. D. Rudie, A. Gbadamosi, O. Toyobo, C. Raymond, D. Zhang, O. Omidiji, R. Akinola, M. A. Suwaid, A. Emegoakor, N. Ojo, K. Aguh, C. Kalaiwo, G. Babatunde, A. Ogunleye, Y . Gbadamosi, K. Iorpagher, E. Calabrese, M. A...
2022
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.