Pith. sign in

REVIEW 3 major objections 5 minor 36 references

Robust Renal Mass Segmentation on CT: A Validation Study of an AI-Based Framework

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A CT kidney-and-abnormality segmentation model trained only on public datasets can match human-level agreement at home and outperform existing public segmenters on external cohorts.

desk verdict A genuinely useful public model with a large external validation, but the kidney-Dice claim is partially undercut by an unanalyzed hilum protocol mismatch. read the letter →

arxiv 2505.07573 v2 pith:X2IDAUJV submitted 2025-05-12 cs.CV cs.AI

classification cs.CVcs.AI
keywords kidneysegmentationabnormalitynnU-NetcomputedtomographyexternalvalidationDicesimilaritycoefficientsubgroupanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that a CT kidney-and-abnormality segmentation model trained exclusively on publicly available data can match human-level agreement on its home-institution test set and still perform well on external cohorts from other institutions. If true, this matters because kidney volume and lesion size—currently often assessed by eye—could be measured automatically and reproducibly, supporting staging, treatment decisions, and active surveillance. The authors report that their model, built on an nnU-Net backbone, reaches Dice scores comparable to an independent human observer on the local private test set and significantly outperforms two established reference models on all tested datasets. Subgroup analyses across patient sex, age, CT contrast phase, and tumour histologic subtype show small performance differences. The model weights and code are released for reuse.

What carries the argument

The load-bearing machinery is a 3D nnU-Net v2 model with Residual Encoder L presets, a self-configuring neural architecture for medical image segmentation, trained as a five-fold ensemble with the full-resolution configuration selected by cross-validation. Before inference and training, a pre-processing step uses the TotalSegmentator anatomical segmenter to locate the lower lung lobes and the urinary bladder and crops a region of interest around the kidneys, reducing compute and standardizing fields of view. A post-processing step removes predicted abnormalities not connected to the kidney unless they exceed 100 cm³ and keeps attached abnormalities with axial diameter above 3 mm, using kidney-volume and RECIST-inspired heuristics. The training target merges cyst and tumour labels from the public challenge into a single 'abnormality' class, while deliberately leaving the two datasets' differing hilum-inclusion protocols untouched.

What would settle it

Re-annotate a random subset of the external test scans under both annotation protocols—hilum included versus excluded, and cyst/tumour split versus merged—and recompute Dice and HD95 against each reference; if the scores diverge by more than the reported inter-observer spread, the claimed external generalizability is largely an artefact of protocol agreement rather than anatomical accuracy. A simpler first check is to run the released model on non-contrast CTs, where the paper already concedes failure, and measure the Dice drop.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that a kidney-abnormality segmentation model built from the nnU-Net framework and trained only on the public kidney-tumour challenge dataset plus a Dutch public kidney-abnormality dataset generalizes to unseen CT scans. The reported numbers are Dice 0.95–0.96 for the kidney region and 0.66 for the abnormality region on the local 50-scan private test set, 0.95 and 0.86 on a 28-scan public clear-cell carcinoma cohort, and 0.84 and 0.83 on a 1510-scan German surgical cohort, with the model outperforming the TotalSegmentator and BAMF reference models on every comparison at corrected significance. The paper further claims that the model's performance lies within human inter-observer variability on the local test set and that subgroup analyses reveal no large biases by sex, age, contrast phase, or tumour subtype.

Load-bearing premise

The merged training labels define one coherent segmentation target even though the two source datasets disagree about whether the renal hilum is part of the kidney and whether cysts should be a separate class, and the authors re-annotated neither.

Editorial extensions

If this is right

  • If the central claim holds, a segmentation model with human-comparable kidney Dice can be built without any proprietary training data, lowering the barrier for reproducible model development and external validation.
  • Automated kidney and abnormality volumes could replace subjective visual estimates in staging, nephrometry scoring, and longitudinal monitoring such as active surveillance.
  • The released weights and code let other centres test the model on their own CT protocols; the paper explicitly notes it fails on non-contrast scans, so retraining with non-contrast examples is the likely remedy.
  • Significant superiority over two public reference models suggests that organ-specific fine-tuning on merged public datasets is a practical route when general-purpose segmenters miss kidney abnormalities.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the authors leave implicit: the reported Dice values are measured against reference masks that follow each site's local protocol, so on the challenge dataset's hilum-inclusive protocol the kidney scores could shift; cross-dataset numbers are protocol-relative rather than anatomy-absolute.
  • Because the public clear-cell carcinoma cohort's reference masks were AI-generated and radiologist-corrected, part of the high external Dice may reflect agreement with an automated annotation style rather than pure anatomical truth; a fully manual re-annotation subset would separate these contributions.
  • Detection precision at a 0.5 IoU threshold on the home test set is markedly lower than on the large surgical cohort, suggesting small or subtle abnormalities in a less-selected population remain a weak point; a reader study could quantify the clinical cost of false negatives.
  • A direct testable extension is to retrain the identical pipeline on versions of the challenge data with hilum included or excluded and compare Dice under both reference protocols, quantifying the label-mismatch bias explicitly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript presents Renal-Net, a kidney and kidney-abnormality segmentation framework based on nnU-Net v2 (ResEnc-L, full-resolution, 5-fold ensemble), trained exclusively on two public datasets: KiTS23 and the Radboudumc kidney abnormality dataset. The pipeline uses TotalSegmentator-based ROI cropping and a rule-based post-processing step. The model is evaluated on a private Radboudumc test set (50 cases), the public TCGA-KIRC set (28 cases), and a large private Charité test set (1,510 scans). Performance is reported as Dice and HD95 for the kidney, kidney+abnormality, and abnormality regions, with comparisons against TotalSegmentator and the BAMF model, detection metrics, subgroup analyses, and qualitative case reviews. The authors claim that the model generalizes to external data, outperforms existing state-of-the-art models across all tested datasets, and reaches human-level performance on the Radboudumc test set, and they release code and weights.

Significance. If the validation withstands scrutiny, this is a practically valuable contribution: it shows that a model trained only on public data can reach high Dice values on large external cohorts, and it provides one of the largest external validations of a kidney/abnormality segmentation model (1,510 clinical CT scans). The explicit comparison with human inter-observer variability, the release of code and weights via GitHub, Zenodo, and grand-challenge.org, and the reporting of detection and subgroup results are all strengths. The main risk to the central claims is the inconsistency in the kidney-label protocol between the two training datasets, which directly affects the interpretation of kidney-region Dice on the Radboudumc test set and, consequently, the human-level and superiority claims.

major comments (3)
  1. [Section 2.1.3, Figure 2, Tables A.4-A.6] The merged kidney label is not a well-defined training target because KiTS23 includes the renal hilum in the kidney mask (489 scans) while the Radboudumc training set excludes it (215 scans), and the authors state that they deliberately did not re-annotate either dataset. The Radboudumc private test set follows the hilum-excluded protocol. For a model that follows the KiTS majority convention, agreement with this reference is capped at Dice = 2V/(2V+h), where V is parenchyma volume and h is non-fat hilar tissue volume; with h/V around 0.1 this cap is about 0.95, which is exactly the range reported for the Radboudumc B30 (0.95 ± 0.04) and B20 (0.96 ± 0.01) kidney Dice. The reported kidney Dice therefore does not by itself establish anatomically correct hilum handling, and it is not a clean basis for the human-level comparison or for the superiority claim over TotalSegmentator and BAMF on the kidney region. Since the paper motivates kidney volume as a biomarker, a systematic hilum over-segmentation bias would have clinical consequences. Please quantify the hilum effect, for example by reporting metrics with the hilum region excluded, by re-annotating a subset under a single protocol, or by evaluating against a hilum-inclusive reference, and temper the Discussion's hypothesis that exposure to annotation variability improves robustness unless it is supported by an ablation or sensitivity analysis.
  2. [Section 3.2, Section 4.2.3, Table A.6] The statement that the proposed model 'significantly outperforms' the BAMF model on the Charité test set rests on Mann-Whitney tests over 1,510 scans, but the absolute differences in mean Dice are extremely small (kidney 0.84 vs 0.83; kidney+abnormality 0.89 vs 0.89; abnormality 0.83 vs 0.81). With this sample size, statistical significance can be reached for differences that are not practically meaningful. Please report effect sizes, paired difference distributions, or 95% confidence intervals, and revise the abstract's claim of outperforming state-of-the-art models 'across all tested datasets' to reflect the magnitude of the advantage on the Charité set.
  3. [Section 4.4, Figures 5 and 6] The subgroup analyses are presented as evidence of robustness, but no statistical comparisons are made between subgroups, and the TCGA-KIRC analysis uses only 28 scans split into four age groups and two sex groups. The boxplots alone do not support the conclusion of 'consistent high performance' across subgroups. Please either add statistical testing or explicitly state that the subgroup results are descriptive only, and avoid the strong robustness language in the abstract until this is supported.
minor comments (5)
  1. [Appendix A, Table A.4] The BAMF row for the kidney region appears to contain a formatting error: '0.90 ± 0.02 0.90 ± 0.04' is ambiguous and the HD95 values for B20 and B30 are missing or misplaced. Please correct the table.
  2. [Abstract and Section 2.3] The abstract and title use 'Renal Mass Segmentation' while the body consistently uses 'Kidney Abnormality Segmentation'; please align the terminology throughout.
  3. [Footnote 1, Abstract, Section 2.3] Footnote 1 states that all URLs will be made public after acceptance, while the abstract and text say the algorithm and code are publicly accessible at the given GitHub and Zenodo links. Please clarify the current availability status and provide the exact DOIs or version identifiers.
  4. [Section 3.1.4] The single-kidney evaluation rule for the Charité set, in which the predicted region overlapping the reference is selected and a zero score is assigned when there is no overlap, should be discussed as a potential source of favorable bias; consider a sensitivity analysis that evaluates both predicted kidney regions.
  5. [Figures 3-6] The boxplots state that the whiskers show the full distribution while outliers are omitted; this is contradictory and should be clarified in the figure captions.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the model is trained on public datasets and evaluated on held-out and external test sets, with no test-set-fitted parameters and no load-bearing self-citation chain.

full rationale

The paper's derivation chain is self-contained in the sense required by the circularity analysis. Renal-Net is trained with nnU-Net on KiTS23 plus the public Radboudumc kidney abnormality dataset, and its reported Dice and HD95 values are computed on test sets that were not used to fit any model parameter or postprocessing threshold. The postprocessing choices (3 mm axial diameter and 100 cm3 volume exception) are justified from RECIST and published average kidney volumes, not from the test data. Comparisons against TotalSegmentator and the BAMF model use externally trained reference systems, and the BAMF model is deliberately excluded from the TCGA-KIRC comparison because TCGA-KIRC was in its training set, which is an anticircular design choice. The Radboudumc private test set shares a protocol with one training set, but that is a distributional overlap rather than a definitional or fitted-input circularity: no quantity from that test set enters the training objective, model selection, or threshold setting. The only self-citations, such as [30] for the human inter-observer variability values, are empirical measurements from prior work and are not used as an unverified uniqueness theorem or as premises that smuggle in the target conclusion. The hilum annotation-protocol mismatch between KiTS and Radboudumc labels is a real label-consistency and external-validity concern, but it is a correctness risk rather than a circularity, because the paper does not define its target metric in terms of the model's output or fit the reported result through that inconsistency.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper contributes no new mathematical derivations or entities. It leans on existing frameworks (nnU-Net, TotalSegmentator) and public annotations. The central empirical claims rest on dataset compatibility, tool reliability, and reference annotation quality, plus several hand-set post-processing thresholds.

free parameters (3)
  • Post-processing volume threshold for unattached predicted abnormalities = 100 cm3
    In Section 2.3.2, predicted abnormality regions not attached to the kidney are removed unless volume exceeds 100 cm3, chosen from ultrasound-based average kidney volume (146/134 cm3). This hand-chosen threshold changes detection and segmentation outputs.
  • Minimum axial diameter for attached predicted abnormalities = 3 mm
    Section 2.3.2: only abnormalities attached to the kidney with axial diameter greater than 3 mm are retained, justified by RECIST (5/10 mm) but 3 mm itself is an arbitrary choice affecting false positives.
  • nnU-Net configuration (ResEnc-L, full-resolution, 5-fold ensemble)
    Section 2.3: the full-resolution ResEnc L model was selected because cross-validation indicated it performed best; low-resolution and cascade variants were trained but not used. This is model selection on training data, not a test-fit constant.
assumptions (4)
  • domain assumption KiTS23 and Radboudumc kidney and abnormality annotations can be merged into one coherent training target without re-annotating the hilum difference.
    Section 2.1.3 states the annotation protocols differ on hilum inclusion and cyst/tumor labels and the authors chose not to re-annotate; the central validation assumes this does not bias learning.
  • domain assumption TotalSegmentator reliably detects lower lung lobes and bladder in all training and test CT scans, so the ROI crop never removes relevant kidney tissue.
    Section 2.3.1 uses TotalSegmentator to crop between these landmarks; a localization failure would eliminate kidneys or abnormalities from the input.
  • domain assumption The Charite reference annotations (one kidney per scan, by one medical student and two radiologists) are accurate and sufficiently complete for evaluating segmentation.
    Section 2.2.3 describes the annotation process; no inter-observer variability or quality control is reported for this 1,510-scan set, the largest evidence source.
  • domain assumption nnU-Net self-configuration and 5-fold cross-validation yield an unbiased model for external deployment.
    Sections 2.3 and 3 rely on nnU-Net automated hyperparameter choices and ensemble; the paper does not analyze how configuration choices affect external generalization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Robust Renal Mass Segmentation on CT: A Validation Study of an AI-Based Framework." pith.science (2026). https://pith.science/paper/X2IDAUJV

@misc{pith2026250507573,
  author       = {Pith},
  title        = {Pith review of: Robust Renal Mass Segmentation on CT: A Validation Study of an AI-Based Framework},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/X2IDAUJV}},
  note         = {Machine review of arXiv:2505.07573}
}
read the original abstract

Renal mass segmentation has important potential to enhance the clinical workflow, especially in settings requiring quantitative assessments. Kidney volume could serve as an important biomarker for renal diseases, with changes in volume correlating directly with kidney function. Currently, clinical practice often relies on subjective visual assessment for evaluating kidney size and kidney lesions, including tumors and cysts, which are typically staged based on diameter, volume, and anatomical location. To support a more objective and reproducible approach, this research aims to develop a robust, thoroughly validated renal mass segmentation algorithm, named Renal-Net. We employ publicly available training datasets and leverage the state-of-the-art medical image segmentation framework nnU-Net. Validation is conducted using both proprietary and public test datasets, with segmentation performance quantified by Dice coefficient and the 95th percentile Hausdorff distance. Furthermore, we analyze robustness across subgroups based on patient sex, age, CT contrast phases, and tumor histologic subtypes. Our findings demonstrate that our segmentation algorithm, trained exclusively on publicly available data, generalizes effectively to external test sets and outperforms existing state-of-the-art models across all tested datasets. Subgroup analyses reveal consistent high performance, indicating strong robustness and reliability. The developed algorithm and associated code are publicly accessible at https://github.com/DIAGNijmegen/oncology-kidney-abnormality-segmentation.

Figures

Figures reproduced from arXiv: 2505.07573 by the authors.

Figure 1
Figure 1. Overview of the research. The proposed AI-based segmentation model is trained [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Visualization of the annotation protocols of the two training datasets. On the [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Segmentation results on the healthy cohort of the Radboudumc test set expressed [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Segmentation results expressed in Dice (left column) and Hausdorff distance [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: Subgroup analysis performed on the TCGA-KIRC test set. Subgroups that are [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: Subgroup analysis performed on the Charit´e Universit¨atsmedizin Berlin test [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: Presented are two cases (top and bottom) from the Radboudumc test set and [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]
Figure 8
Figure 8. Figure 8: Presented are two cases (top and bottom) from the Charit´e Universit¨atsmedizin [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

36 extracted references · 18 canonical work pages

  1. [1]

    Cirillo, S

    L. Cirillo, S. Innocenti, F. Becherucci, Global epidemiology of kidney cancer, Nephrology Dialysis Transplantation 39 (2024) 920–928. doi:10. 1093/ndt/gfae036

  2. [2]

    R. J. Motzer, E. Jonasch, N. Agarwal, A. Alva, M. Baine, K. Becker- mann, M. I. Carlo, T. K. Choueiri, B. A. Costello, I. H. Derweesh, A. De- sai, Y. Ged, S. George, J. L. Gore, N. Haas, S. L. Hancock, P. Kapur, C. Kyriakopoulos, E. T. Lam, P. N. Lara, C. Lau, B. Lewis, D. C. Mad- off, B. Manley, M. D. Michaelson, A. Mortazavi, L. Nandagopal, E. R. Plimac...

  3. [3]

    Filippiadis, G

    D. Filippiadis, G. Mauri, P. Marra, G. Charalampopoulos, N. Gennaro, F. De Cobelli, Percutaneous ablation techniques for renal cell carcinoma: Current status and future trends, International Journal of Hyperthermia 36 (2019) 21–30. doi: 10.1080/02656736.2019.1647352

  4. [4]

    J. P. Mershon, M. N. Tuong, N. S. Schenkman, Thermal ablation of the small renal mass: A critical analysis of current literature, Minerva Uro- logica e Nefrologica 72 (2020). doi: 10.23736/S0393-2249.19.03572-0

  5. [5]

    Capitanio, F

    U. Capitanio, F. Montorsi, Renal cancer, The Lancet 387 (2016) 894–

  6. [6]

    Kutikov, B

    A. Kutikov, B. L. Egleston, Y.-N. Wong, R. G. Uzzo, Evaluating Over- all Survival and Competing Risks of Death in Patients With Localized Renal Cell Carcinoma Using a Comprehensive Nomogram, Journal of Clinical Oncology 28 (2010) 311–317. doi: 10.1200/JCO.2009.22.4816

  7. [7]

    Kutikov, R

    A. Kutikov, R. G. Uzzo, The R.E.N.A.L. Nephrometry Score: A Com- prehensive Standardized System for Quantitating Renal Tumor Size, Location and Depth, The Journal of Urology 182 (2009) 844–853. doi:10.1016/j.juro.2009.05.035

  8. [8]

    P. Gaur, W. Gedroyc, P. Hill, ADPKD—what the radiologist should know, The British Journal of Radiology 92 (2019) 20190078. doi: 10. 1259/bjr.20190078. 24

Show all 36 references
  1. [9]

    Joskowicz, D

    L. Joskowicz, D. Cohen, N. Caplan, J. Sosna, Inter-observer variability of manual contour delineation of structures in CT, European Radiology 29 (2019) 1391–1399. doi: 10.1007/s00330-018-5695-5

  2. [10]

    C. R. Meyer, T. D. Johnson, G. McLennan, D. R. Aberle, E. A. Kazerooni, H. MacMahon, B. F. Mullan, D. F. Yankelevitz, E. J. Van Beek, S. G. Armato, M. F. McNitt-Gray, A. P. Reeves, D. Gur, C. I. Henschke, E. A. Hoffman, P. H. Bland, G. Laderach, R. Pais, D. Qing, C. Piker, J. ...

  3. [11]

    Bilic, P

    P. Bilic, P. Christ, H. B. Li, E. Vorontsov, A. Ben-Cohen, G. Kaissis, A. Szeskin, C. Jacobs, G. E. H. Mamani, G. Chartrand, F. Loh¨ ofer, J. W. Holch, W. Sommer, F. Hofmann, A. Hostettler, N. Lev-Cohain, M. Drozdzal, M. M. Amitai, R. Vivanti, J. Sosna, I. Ezhov, A. Sekuboy- i...

  4. [12]

    Staal, M

    J. Staal, M. Abramoff, M. Niemeijer, M. Viergever, B. van Ginneken, Ridge-based vessel segmentation in color images of the retina, IEEE Transactions on Medical Imaging 23 (2004) 501–509. doi: 10.1109/TMI. 2004.825627. 25

  5. [13]

    Litjens, R

    G. Litjens, R. Toth, W. van de Ven, C. Hoeks, S. Kerkstra, B. van Ginneken, G. Vincent, G. Guillard, N. Birbeck, J. Zhang, R. Strand, F. Malmberg, Y. Ou, C. Davatzikos, M. Kirschner, F. Jung, J. Yuan, W. Qiu, Q. Gao, P. E. Edwards, B. Maan, F. van der Heijden, S. Ghose, J. Mit...

  6. [14]

    Ronneberger, P

    O. Ronneberger, P. Fischer, T. Brox, U-Net: Convolutional Networks for Biomedical Image Segmentation, 2015. doi: 10.48550/arXiv.1505. 04597. arXiv:1505.04597

  7. [15]

    Isensee, P

    F. Isensee, P. F. Jaeger, S. A. A. Kohl, J. Petersen, K. H. Maier-Hein, nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation, Nature Methods 18 (2021) 203–211. doi: 10.1038/ s41592-020-01008-z

  8. [16]

    Wasserthal, H.-C

    J. Wasserthal, H.-C. Breit, M. T. Meyer, M. Pradella, D. Hinck, A. W. Sauter, T. Heye, D. T. Boll, J. Cyriac, S. Yang, M. Bach, M. Segeroth, TotalSegmentator: Robust Segmentation of 104 Anatomic Structures in CT Images, Radiology: Artificial Intelligence 5 (2023) e230024. doi:...

  9. [17]

    E. Yang, C. K. Kim, Y. Guan, B.-B. Koo, J.-H. Kim, 3D multi- scale residual fully convolutional neural network for segmentation of extremely large-sized kidney tumor, Computer Methods and Programs in Biomedicine 215 (2022) 106616. doi: 10.1016/j.cmpb.2022.106616

  10. [18]

    Rombolotti, F

    M. Rombolotti, F. Sangalli, D. Cerullo, A. Remuzzi, E. Lanzarone, Au- tomatic cyst and kidney segmentation in autosomal dominant polycys- tic kidney disease: Comparison of U-Net based methods, Computers in Biology and Medicine 146 (2022) 105431. doi: 10.1016/j.compbiomed. 2022.105431

  11. [19]

    Korfiatis, A

    P. Korfiatis, A. Denic, M. E. Edwards, A. V. Gregory, D. E. Wright, A. Mullan, J. Augustine, A. D. Rule, T. L. Kline, Automated Seg- mentation of Kidney Cortex and Medulla in CT Images: A Multisite Evaluation Study, Journal of the American Society of Nephrology 33 (2022) 420. ...

  12. [20]

    R. Khan, C. Chen, A. Zaman, J. Wu, H. Mai, L. Su, Y. Kang, B. Huang, RenalSegNet: Automated segmentation of renal tumor, veins, and ar- teries in contrast-enhanced CT scans, Complex & Intelligent Systems 11 (2025) 131. doi: 10.1007/s40747-024-01751-2

  13. [21]

    S. Pang, A. Du, M. A. Orgun, Z. Yu, Y. Wang, Y. Wang, G. Liu, CTu- morGAN: A unified framework for automatic computed tomography tumor segmentation, European Journal of Nuclear Medicine and Molec- ular Imaging 47 (2020) 2248–2268. doi: 10.1007/s00259-020-04781-3

  14. [22]

    M. J. J. de Grauw, E. Th. Scholten, E. J. Smit, M. J. C. M. Rutten, M. Prokop, B. van Ginneken, A. Hering, The ULS23 challenge: A base- line model and benchmark dataset for 3D universal lesion segmentation in computed tomography, Medical Image Analysis 102 (2025) 103525. doi:1...

  15. [23]

    Heller, F

    N. Heller, F. Isensee, K. H. Maier-Hein, X. Hou, C. Xie, F. Li, Y. Nan, G. Mu, Z. Lin, M. Han, G. Yao, Y. Gao, Y. Zhang, Y. Wang, F. Hou, J. Yang, G. Xiong, J. Tian, C. Zhong, J. Ma, J. Rickman, J. Dean, B. Stai, R. Tejpaul, M. Oestreich, P. Blake, H. Kaluzniak, S. Raza, J. Ro...

  16. [24]

    Heller, F

    N. Heller, F. Isensee, D. Trofimova, R. Tejpaul, Z. Zhao, H. Chen, L. Wang, A. Golts, D. Khapun, D. Shats, Y. Shoshan, F. Gilboa- Solomon, Y. George, X. Yang, J. Zhang, J. Zhang, Y. Xia, M. Wu, Z. Liu, E. Walczak, S. McSweeney, R. Vasdev, C. Hornung, R. Solaiman, J. Schoephoer...

  17. [25]

    G. K. Murugesan, D. McCrumb, M. Aboian, T. Verma, R. Soni, F. Memon, K. Farahani, L. Pei, U. Wagner, A. Y. Fedorov, D. Clu- nie, S. Moore, J. V. Oss, AI-Generated Annotations Dataset for Di- verse Cancer Radiology Collections in NCI Image Data Commons, 2024. doi:10.48550/arXiv...

  18. [26]

    Myronenko, D

    A. Myronenko, D. Yang, Y. He, D. Xu, Automated 3D Segmen- tation of Kidneys and Tumors in MICCAI KiTS 2023 Challenge, in: N. Heller, A. Wood, F. Isensee, T. R¨ adsch, R. Teipaul, N. Pa- panikolopoulos, C. Weight (Eds.), Kidney and Kidney Tumor Segmenta- tion, Springer Nature S...

  19. [27]

    MONAI, MONAI: Medical Open Network for AI, Zenodo, 2024

    C. MONAI, MONAI: Medical Open Network for AI, Zenodo, 2024. doi:10.5281/zenodo.13942962

  20. [28]

    Isensee, T

    F. Isensee, T. Wald, C. Ulrich, M. Baumgartner, S. Roy, K. Maier-Hein, P. F. Jaeger, nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation, 2024. arXiv:2404.09556

  21. [29]

    G. E. Humpire-Mamani, L. Builtjes, C. Jacobs, B. Van Ginneken, M. Prokop, E. T. Scholten, Dataset for: Kidney abnormality segmenta- tion in thorax-abdomen CT scans, 2023. doi:10.5281/ZENODO.8014289

  22. [30]

    G. E. H. Mamani, N. Lessmann, E. T. Scholten, M. Prokop, C. Jacobs, B. van Ginneken, Kidney abnormality segmentation in thorax-abdomen CT scans, 2023. arXiv:2309.03383

  23. [31]

    O. Akin, P. Elnajjar, M. Heller, R. Jarosz, B. J. Erickson, S. Kirk, Y. Lee, M. W. Linehan, R. Gautam, R. Vikram, K. M. Garcia, C. Roche, E. Bonaccio, J. Filippini, The Cancer Genome Atlas Kidney Renal Clear Cell Carcinoma Collection (TCGA-KIRC), 2016. doi: 10.7937/ K9/TCIA.20...

  24. [32]

    S. A. Emamian, M. B. Nielsen, J. F. Pedersen, L. Ytte, Kidney dimen- sions at sonography: correlation with age, sex, and habitus in 665 adult volunteers., AJR. American journal of roentgenology 160 (1993) 83–86

  25. [33]

    E. A. Eisenhauer, P. Therasse, J. Bogaerts, L. H. Schwartz, D. Sargent, R. Ford, J. Dancey, S. Arbuck, S. Gwyther, M. Mooney, L. Rubinstein, 28 L. Shankar, L. Dodd, R. Kaplan, D. Lacombe, J. Verweij, New response evaluation criteria in solid tumours: Revised RECIST guideline (...

  26. [34]

    Maier-Hein, A

    L. Maier-Hein, A. Reinke, P. Godau, M. D. Tizabi, F. Buettner, E. Christodoulou, B. Glocker, F. Isensee, J. Kleesiek, M. Kozubek, M. Reyes, M. A. Riegler, M. Wiesenfarth, A. E. Kavur, C. H. Su- dre, M. Baumgartner, M. Eisenmann, D. Heckmann-N¨ otzel, T. R¨ adsch, L. Acion, M. ...

  27. [90]

    doi:10.6004/jnccn.2022.0001

  28. [906]

    doi:10.1016/S0140-6736(15)00046-X

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.