Pith. sign in

REVIEW 5 major objections 6 minor 3 references

TissUnet: Improved Extracranial Tissue and Cranium Segmentation for Children through Adulthood

T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read TissUnet is a deep learning model that segments skull bone, subcutaneous fat, and muscle from routine three-dimensional T1-weighted brain MRI, with or without contrast enhancement, across the human lifespan and in the presence of brain…

desk verdict Useful engineering with a real gap, but the pediatric/tumor accuracy claims rest on circular AI-CT validation and a 10-case adult-only manual study. read the letter →

arxiv 2506.05660 v2 pith:VAGR2LW3 submitted 2025-06-06 cs.CV cs.AI

classification cs.CVcs.AI
keywords whole-headsegmentationMRIdeeplearningartificialintelligencepediatricbraintumorskullthicknessbodycompositionnnU-Net
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

TissUnet takes a standard T1-weighted brain MRI and automatically outlines three extracranial tissues: skull bone, subcutaneous fat, and muscle. The paper's central claim is that this segmentation is accurate, fast, and reproducible enough to support large-scale studies in healthy children and adults as well as in patients with brain tumors, a setting where previous tools were untested or failed. Against CT-derived labels in healthy adults the model reaches a median Dice of 0.79, and against expert manual annotations it reaches 0.83 in healthy subjects and 0.81 in tumor cases, outperforming the prior state of the art. If correct, TissUnet would let researchers quantify craniofacial morphology, treatment effects on body composition, and cardiometabolic risk from MRI data that are already being collected for other purposes.

What carries the argument

The load-bearing machinery is a nnU-Net v2—a self-configuring U-Net architecture—trained on pseudo-labels: CT scans from the multi-center SynthRAD2023 dataset were segmented by TotalSegmentator into skull, subcutaneous fat, and temporalis muscle masks, and those masks were propagated to co-registered T1-weighted MRIs to generate training ground truth without manual annotation. To handle heterogeneity across publicly shared MRI data, the pipeline adds a brain-mask-guided region-of-interest cropping step that standardizes the extracranial field of view and reduces the impact of defacing algorithms and scanner differences. A separate component estimates skull thickness by detecting the orbital roof with a DenseNet landmark model and measuring the median of 100 tangents per 1-mm axial slice over sixteen slices, with no dependence on CT Hounsfield-unit thresholds.

What would settle it

Have two or more expert radiologists manually segment skull, subcutaneous fat, and muscle on a diverse set of at least 50 pediatric and 50 brain-tumor T1-weighted MRIs spanning the age range, then compare TissUnet’s Dice and volume bias against these manual labels per tissue and subgroup; if subcutaneous fat Dice falls below roughly 0.5 in any subgroup, or if fat volume shows a systematic bias that tracks with age or tumor location, the CT-pseudo-label training strategy is the likely cause.

Watch

Extended reading notes

Core claim

The central discovery is that a single nnU-Net v2–based model, TissUnet, can segment skull, subcutaneous fat, and muscle from standard T1-weighted brain MRI with clinically useful accuracy across pediatric and adult populations and in the presence of intracranial pathology. The authors show this by training on 155 paired MRI-CT scans from the SynthRAD2023 dataset, using TotalSegmentator’s CT segmentations propagated to co-registered MRI as pseudo ground truth, and then validating on nine external datasets. Against AI-CT labels on 37 healthy adults, TissUnet reaches a median Dice of 0.79; against expert manual annotations it reaches 0.83 in healthy subjects and 0.81 in brain tumor cases, compared with 0.73 and 0.60 for GRACE, the prior state of the art. In a blind review of 45 cases, all TissUnet segmentations were rated acceptable while 84% of GRACE outputs required revision, and an acceptability study of 108 MRIs ended at an 89% acceptance rate after adjudication. TissUnet also produced skull-thickness estimates closer to CT reference than four established MRI-based methods, and its temporalis muscle volume showed a significant inverse association with blood cholesterol in 888 adolescents.

Load-bearing premise

The load-bearing premise is that pseudo-labels generated by TotalSegmentator on co-registered CT scans, after propagation to MRI, are accurate enough to serve as training ground truth for skull, fat, and muscle across ages and pathologies, and that the small independent manual check (one expert on ten cases) does not hide systematic errors in those labels.

Editorial extensions

If this is right

  • Large retrospective cohorts can be analyzed for extracranial tissue volumes without manual annotation, enabling studies of craniofacial morphology, treatment toxicity such as sarcopenia in pediatric brain tumor survivors, and cardiometabolic risk from existing T1w MRI.
  • Skull thickness can be measured from routine MRI with accuracy close to CT, supporting cranial growth tracking and surgical planning without radiation exposure.
  • The model works with both contrast-enhanced and non-contrast T1-weighted sequences, so it can be applied to routine clinical and research scans.
  • TissUnet’s robustness to defacing means anonymized and publicly shared MRI datasets become usable for extracranial quantification.
  • Automated temporalis muscle volumetry associates with blood cholesterol in adolescents, suggesting a path to imaging-based cardiometabolic risk markers in pediatric populations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the training labels come from CT pseudo-segmentations propagated through registration, any systematic error in TotalSegmentator’s skull, fat, or muscle labels—especially in pediatric skulls or tumor-distorted anatomy—will be learned by TissUnet; the validation against only ten manually annotated cases is too small to bound such biases, particularly for subcutaneous fat where Dice is lowest
  • The brain-mask ROI cropping that removes the anterior face standardizes volumes but excludes facial and upper-anterior tissues, so downstream facial morphology studies may need a different region of interest.
  • The cholesterol association is cross-sectional and from a single cohort; the predictive value of temporalis volume against established metrics such as BMI should be tested prospectively across diverse populations.
  • A direct extension would be to retrain or fine-tune on manually segmented pediatric and fat-saturated MRI to test whether the pseudo-label approach carries over to those sequences, which the paper itself flags as uncertain.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The manuscript presents TissUnet, an nnU-Net-based model for segmenting skull, subcutaneous fat, and muscle from T1-weighted brain MRI, with or without contrast. Training uses 155 paired MRI-CT scans from SynthRAD2023, with pseudo-labels generated by TotalSegmentator on CT and propagated to co-registered MRI. Validation is reported across nine external datasets: median Dice of 0.79 vs AI-CT-derived labels on 37 healthy adults (CERMEP), 0.83 (healthy) and 0.81 (tumor) vs expert manual annotations on 10 cases, 89% acceptability on 108 MRIs, and a blinded comparative review with 100% acceptable ratings for TissUnet vs 16% for GRACE. The paper also proposes a skull thickness estimation pipeline and demonstrates an association between temporalis muscle volume and blood cholesterol in the ABCD cohort.

Significance. If the accuracy and generalizability claims hold, TissUnet would be a practically valuable tool for opportunistic quantification of extracranial tissues from routine brain MRI, with plausible applications in craniofacial growth assessment, treatment toxicity monitoring, and cardiometabolic risk studies. The manuscript has notable strengths: the model weights and code are publicly available, the evaluation spans multiple public datasets and age groups, the rotation ablation study provides a useful robustness check, and the ABCD application demonstrates a concrete downstream use. However, the strength of the central claim is currently limited by the circular AI-CT validation and the very small manual annotation study, so the contribution is promising but not yet convincingly established.

major comments (5)
  1. [§2.1, §2.3, Supplementary Methods 4, Table 1] The AI-CT validation on CERMEP (Table 1, N=37) uses reference labels generated by TotalSegmentator, which is the same tool that produced the training pseudo-labels (Supplementary Methods 4). This comparison therefore measures agreement between TissUnet and its own teacher rather than with an independent anatomical truth, making the reported Dice of 0.79 circular for the purpose of establishing accuracy. Please re-label this experiment as an agreement study against the pseudo-label generator, or provide an independent CT-based or manually annotated reference, and adjust the abstract and Key Results accordingly.
  2. [§2.3, Table 2] The only non-circular quantitative validation is the expert manual annotation study, but it includes just 10 cases (5 healthy adults from CERMEP and 5 adults with glioblastoma from ACRIN), with no pediatric cases and only a single expert annotator. The manuscript's central claim of accuracy 'for children through adulthood' and in tumor cases is therefore not quantitatively supported by per-class Dice in those populations; the pediatric and tumor evidence rests on subjective acceptability ratings and a single-rater blinded review. Please add quantitative validation with independent references stratified by age group and pathology, or substantially temper the generalizability claims.
  3. [Table 2, fat row] In the manual annotation study, TissUnet's median Dice for subcutaneous fat in healthy subjects is 0.59, while GRACE achieves 0.73, meaning the previous state-of-the-art is actually better on this tissue class. This is directly relevant to the abstract's cardiometabolic-risk application, which depends on reliable fat quantification. The manuscript should explicitly report and discuss this per-class reversal rather than presenting only overall Dice superiority; as written, Table 2 is inconsistent with the blanket statement that TissUnet outperforms GRACE.
  4. [Discussion, first paragraph] The Discussion states 'The model achieved median Dice scores of 0.81 and 0.83 in healthy and brain tumor cohorts, respectively,' which reverses the Results (0.83 healthy, 0.81 tumor, per §3.1 and Table 2). This numeric inconsistency appears also in the Key Results and should be corrected, and the full text should be checked for any other swapped cohort-specific values.
  5. [§3.1, Figure 1E, §2.3] The blinded comparative review was conducted by B.H.K., a co-author and board-certified radiation oncologist, with N=45 MRIs. While the blinding is a strength, the evaluation is single-rater and not independent of the study team, and the reported 100% acceptable rate for TissUnet should be interpreted with this in mind. Please present this result as a non-independent, single-rater assessment and avoid the unqualified claim of 'excellent performance' without acknowledging the potential for reviewer bias.
minor comments (6)
  1. [Abstract and §3.1] The abstract reports 'N=108 MRIs' for acceptability testing, but §3.1 reports N=289 acceptable, N=34 unacceptable, and N=1 bad image. These counts appear to be per-mask (3 tissue classes × 108 = 324), not per-MRI; please clarify the unit of analysis and ensure consistency between the abstract and Results.
  2. [Figure 1E caption] The caption states 'N=54 MRI' but then says 'All cases (N=45, 100%)' and 'GRACE, 7 cases (16%)... 38 cases (84%)'; the total in the second set is 45, not 54. Please reconcile the N reported in the caption.
  3. [Table 4] The table header uses 'cm²' for volumetric differences, but volumes should be in cm³. Please correct the unit notation in Table 4 and the accompanying text.
  4. [Introduction and Discussion] There is a typo 'standart' in the Introduction, and the Discussion refers to 'SPM125' once while the rest of the text uses 'SPM25'; please correct these errors.
  5. [Throughout] The model name is written inconsistently as 'TissUnet' and 'TissUNet' (e.g., Table 2, Figure 1E). Please standardize the name.
  6. [Supplementary Methods 5] The acceptability inter-rater agreement values in Supplementary Table S1 include some very low coefficients for skull (e.g., R1-R2 = 0.200 in brain tumor cases); the manuscript should comment on why skull agreement is so much lower than fat and muscle, since skull is the highest-Dice class in the quantitative evaluations.

Circularity Check

1 steps flagged · score 5.0 of 10

AI-CT validation is circular: TissUnet is trained on TotalSegmentator pseudo-labels and then 'validated' against TotalSegmentator-derived labels; independent manual evidence covers only 10 adult cases.

  1. fitted input called prediction [Section 2.1 (Training), Section 2.3 (Evaluation), Results 3.1, Table 1]
    "we used segmentations generated by a previously validated CT-based AI algorithm, TotalSegmentator (Wasserthal et al., 2023), as initial ground truth labels that were then propagated to the co-registered T1w MRI (Supplementary Methods 7). ... we compared TissUnet-predicted segmentations of the skull, fat, and muscle to reference segmentations generated from CT using TotalSegmentator ... on CERMEP dataset (N=37, Figure 2B)."

    TissUnet is trained to reproduce TotalSegmentator's skull/fat/muscle labels on SynthRAD MRI-CT pairs. The CERMEP 'AI-CT as ground-truth' reference is generated by the same TotalSegmentator algorithm. Therefore the reported Dice of 0.79 measures agreement between TissUnet and its own label teacher, not agreement with independent anatomical truth. Any systematic error in TotalSegmentator, especially for subcutaneous fat, pediatric skulls, or tumor-distorted anatomy, is inherited by TissUnet and is invisible in this metric. The only non-circular quantitative check, expert manual annotations (Table 2), covers just 10 adults (5 healthy, 5 tumor), so the pediatric and broad-pathology accuracy claims lean substantially on this circular comparison.

full rationale

The paper's central claim is supported by several validation arms. The AI-CT validation on CERMEP is circular because the reference labels are produced by the same TotalSegmentator tool that generated the training pseudo-labels, making the 0.79 Dice a measure of agreement with the teacher rather than anatomical truth. The independent manual annotation study provides non-circular evidence, but it includes only 10 adult cases, so the quantitative pediatric and tumor-pathology claims rest partly on the circular AI-CT comparison and on subjective acceptability/blinded review. The cholesterol association is exploratory and does not add circularity. One minor self-citation (Zapaishchykova et al. 2023 for skull-thickness landmark detection) is not load-bearing to the main segmentation claim. Overall, partial circularity in one of the four validation arms warrants a score of 5.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The central claim rests mainly on the accuracy of AI-generated training labels, CT-MRI alignment assumptions, and the trustworthiness of a single expert's annotations on 10 cases. The fitted values (Box-Cox lambda, HU threshold, skull-thickness aggregation choices) affect only secondary analyses. No invented entities are proposed; the brain-mask ROI pipeline is an algorithmic procedure, not an entity.

free parameters (3)
  • Box-Cox lambda for cholesterol transformation = 0.2
    Chosen to normalize the cholesterol distribution (Section 2.6); affects the exploratory regression, not the segmentation claim.
  • CT skull reference HU threshold = 471 HU default; 300-800 tested
    Adopted from Delso 2015 for the skull thickness reference; the paper reports sensitivity across windows (Supplementary Methods 3), with one subject excluded at 800 HU.
  • Skull thickness aggregation parameters = 16 slices, 100 tangents, central 95%, 10 mm offset
    Hand-chosen pipeline settings (Supplementary Methods 2) that define the reported skull thickness estimates.
assumptions (6)
  • domain assumption TotalSegmentator CT labels are accurate enough to serve as training pseudo-labels for skull, fat, and muscle
    Invoked in Section 2.1; the cited general validation does not cover pediatric skulls or tumor-distorted anatomy.
  • domain assumption Co-registered MRI-CT pairs are aligned sufficiently for label propagation
    Section 2.1 and Supplementary Methods 4: CT labels transferred to T1w MRI without reported residual-error analysis.
  • domain assumption Rigid registration to age-dependent NIHPD atlases and 1 mm isotropic resampling preserve anatomical size
    Section 2.2; the volumetric comparisons assume this registration does not distort volumes.
  • domain assumption Single-expert manual annotations on 10 cases are an unbiased accuracy reference
    Section 2.3; one neuroradiologist (H.S.) annotated 10 cases; no inter-observer reliability is reported for the reference.
  • standard math nnU-Net v2 automated configuration and augmentation suit this three-class task
    Section 2.1 and Supplementary Methods 7; framework-level assumption standard in this literature.
  • domain assumption HD-BET brain masks are accurate across ages and pathologies
    Supplementary Methods 6; the ROI cropping and defacing normalization depend on mask quality.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TissUnet: Improved Extracranial Tissue and Cranium Segmentation for Children through Adulthood." pith.science (2026). https://pith.science/paper/VAGR2LW3

@misc{pith2026250605660,
  author       = {Pith},
  title        = {Pith review of: TissUnet: Improved Extracranial Tissue and Cranium Segmentation for Children through Adulthood},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VAGR2LW3}},
  note         = {Machine review of arXiv:2506.05660}
}
read the original abstract

Extracranial tissues visible on brain magnetic resonance imaging (MRI) may hold significant value for characterizing health conditions and clinical decision-making, yet they are rarely quantified. Current tools have not been widely validated, particularly in settings of developing brains or underlying pathology. We present TissUnet, a deep learning model that segments skull bone, subcutaneous fat, and muscle from routine three-dimensional T1-weighted MRI, with or without contrast enhancement. The model was trained on 155 paired MRI-computed tomography (CT) scans and validated across nine datasets covering a wide age range and including individuals with brain tumors. In comparison to AI-CT-derived labels from 37 MRI-CT pairs, TissUnet achieved a median Dice coefficient of 0.79 [IQR: 0.77-0.81] in a healthy adult cohort. In a second validation using expert manual annotations, median Dice was 0.83 [IQR: 0.83-0.84] in healthy individuals and 0.81 [IQR: 0.78-0.83] in tumor cases, outperforming previous state-of-the-art method. Acceptability testing resulted in an 89% acceptance rate after adjudication by a tie-breaker(N=108 MRIs), and TissUnet demonstrated excellent performance in the blinded comparative review (N=45 MRIs), including both healthy and tumor cases in pediatric populations. TissUnet enables fast, accurate, and reproducible segmentation of extracranial tissues, supporting large-scale studies on craniofacial morphology, treatment effects, and cardiometabolic risk using standard brain T1w MRI.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

3 extracted references · 1 canonical work pages

  1. [1]

    ACRIN-FMISO-BRAIN. (n.d.). The Cancer Imaging Archive (TCIA). Retrieved March 19, 2025, from https://www.cancerimagingarchive.net/collection/acrin-fmiso-brain/ Casey, B. J., Cannonier, T., Conley, M. I., Cohen, A. O., Barch, D. M., Heitzeg, M. M., Soules, M. E., Teslovich, T., Dellarco, D. V., Garavan, H., Orr, C. A., Wager, T. D., Banich, M. T., Speer, N...

  2. [2]

    https://doi.org/10.3390/cancers16020415 Zapaishchykova, A., Tak, D., Boyd, A., Ye, Z., Aerts, H. J. W. L., & Kann, B. H. (2023). SegmentationReview: A Slicer3D extension for fast review of AI-generated segmentations. Software Impacts, 17, 100536. https://doi.org/10.1016/j.simpa.2023.100536

  3. [91]

    E., Long, X., Paniukov, D., Bagshawe, M., & Lebel, C

    https://doi.org/10.1186/s13550-021-00830-6 Reynolds, J. E., Long, X., Paniukov, D., Bagshawe, M., & Lebel, C. (2020). Calgary Preschool magnetic resonance imaging (MRI) dataset. Data in Brief, 29, 105224. https://doi.org/10.1016/j.dib.2020.105224 Rivkin, M. J., Ball, W. S., Wang, D.-J., McCracken, J. T., Brandt, M., Fletcher, J., McKinstry, R., Evans, A.,...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.