{"id":"aa8e664a-b5a7-4efc-bdc3-78812dbde613","arxiv_id":"2501.16879","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The paper presents holiAtlas, a multimodal, multiscale, densely labeled brain MRI atlas built from 75 subjects with 350 substructure labels.","lead":"Researchers built a new digital map of the human brain that combines three MRI contrasts and 350 labeled anatomical regions at a fine scale. The atlas is freely available and could help scientists measure small brain structures that standard atlases blur together.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The atlas's accuracy claim rests on unquantified label validity, and the most fragile link is the WMn-dependent set: synthetic WMn from §2.2 is used in §2.4 to train/apply thalamic nuclei segmentation, with only visual QC.","rationale":"I read the paper as a resource-description and construction-methodology contribution. The central claim is that holiAtlas is a valid multimodal, multiscale, densely labelled structural atlas; that claim requires the labels to be accurate enough to serve as a reference. The paper provides extensive process documentation: 75 HCP subjects, seven software protocols, semi-automatic MAP-based relabeling, manual ITK-SNAP correction, group-wise template construction, and majority-vote labels. These are real and valuable supports, and the public release of the atlas and scripts allows independent use. However, the accuracy of the final labels is never quantitatively compared with an independent ground truth, expert manual delineations, or a well-validated atlas. The most load-bearing instance of this gap is the WMn chain: WMn is synthesized rather than acquired, and the thalamic nuclei — among the finest and most clinically interesting labels — are derived from that synthetic contrast. A failure of WMn synthesis would not merely blur the average image; it would systematically distort the segmentations used to build the labels, and visual QC cannot detect subtle but consistent errors in tiny nuclei. This is an internal-validity concern, not a disagreement with community practice; even if automatic methods are generally near inter-rater variability, that claim is cited for other tools and is not demonstrated here for the fused protocol. The reader's weakest assumption identified the same issue, and the conditional verdict is appropriate. The concrete test I propose would settle whether the synthetic WMn proxy is fit for the segmentation task by directly comparing real- and synthetic-based nuclei segmentations on held-out subjects. If the test passed, the concern would be substantially retired; if it failed, the atlas's WMn-dependent labels would need to be relabeled from real WMn or explicitly described as provisional. Either way, the current conditional status remains the right verdict until such evidence is supplied.","tokens_in":14912,"tokens_out":4325,"duration_ms":39729,"concrete_test":"Use the DeepMultiBrain dataset, which contains real T1w, T2w, and WMn images, in a held-out evaluation. Train the T1w/T2w-to-WMn synthesis network exactly as described in §2.2 on a subset (e.g., 40 subjects) and synthesize WMn for the remaining 15. Run the THOMAS-derived nuclei segmentation of §2.4 on both the real and the synthetic WMn of the held-out subjects, after registering both to the atlas space, and compute per-nucleus Dice and volume bias between the real- and synthetic-based segmentations. If the mean Dice across the 12/13 thalamic labels is clearly below 0.7, or if small nuclei such as the habenular nucleus and mammillothalamic tract show very low overlap, the synthetic WMn proxy cannot support the atlas's WMn-dependent labels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that holiAtlas is a valid densely labelled multimodal reference atlas at 0.125 mm3. For this to hold, the 350 labels produced by the protocol fusion must be anatomically accurate. The paper supports this with careful protocol engineering, semi-automatic correction, and expert visual QC, but provides no quantitative accuracy validation. The weakest point is the WMn-dependent label chain. In §2.2, WMn is not acquired for the 75 HCP subjects: it is synthesized from T1w and T2w using a network trained on a private dataset, and the paper states only that quality was 'validated by our experts' with no metric. In §2.4, a THOMAS-trained network is used to segment 12 or 13 thalamic nuclei per hemisphere from these synthetic WMn images. Any systematic synthesis artifact — intensity bias, over-smoothing, or hallucinated contrast in deep gray matter — is inherited by the nuclear boundaries and, after majority voting across subjects, by the atlas labels. This matters most for small structures such as the habenular nucleus, mammillothalamic tract, and subthalamic nucleus, where visual QC is least reliable. The paper itself acknowledges the concern: 'the fact that it is based on the fusion of automatic segmentations may raise doubts about its accuracy.' Without independent quantitative evidence that the synthetic WMn is equivalent to real WMn for this segmentation task, the validity of all WMn-derived substructure labels in the atlas is unestablished.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces holiAtlas, a multimodal (T1w, T2w, WMn) structural brain MRI atlas built from 75 healthy HCP subjects. The authors resample the original 0.7 mm HCP images to a 0.5 mm isotropic grid (0.125 mm^3 voxels), synthesize WMn images from T1w/T2w using a network trained on a private dataset, register all subjects to a group-wise template with FireANTS, and fuse outputs of seven automatic segmentation tools (vol2Brain, hypothalamus_seg, BrainVISA, FreeSurfer, pBrain, HIPS, CERES) with semi-automatic correction and manual QC. The final protocol has 350 substructure labels, with coarser groupings into 54 structures, 9 tissues, and 1 organ. The atlas and labels are publicly released. The central claims are that the atlas is ultra-high resolution, multimodal, and densely labelled, and that it can serve as a reference for brain segmentation and disease studies.","tokens_in":15285,"tokens_out":3096,"duration_ms":26523,"significance":"If the construction is sound, holiAtlas is a potentially useful community resource: it combines multiple modalities, offers hierarchical label organization, and is publicly available. The authors also make the construction scripts available via FireANTS. However, the significance is tempered by two load-bearing weaknesses: the 'ultra-high resolution' claim is based on interpolation of native 0.7 mm data rather than higher-resolution acquisition, and the WMn modality is synthetic, with no quantitative validation of its fidelity for the subsequent segmentation tasks. The atlas artifact itself is a contribution, but the paper's framing overstates its resolution and the anatomical validity of the WMn-derived labels is not established.","major_comments":[{"comment":"The '0.125 mm3 resolution' claim is misleading. The HCP T1w and T2w images are acquired at 0.7 mm isotropic voxels (§2.1), and §2.2 states they were affine-registered to MNI152 space at 0.5 mm voxels. This is interpolation/upsampling, not a gain in true spatial resolution. The effective resolution of the T1w/T2w templates remains acquisition-limited by the original 0.7 mm sampling (and by the 1 mm working resolution of the synthesis network used for WMn). The title and abstract's 'ultra-high resolution' should be rephrased to 'resampled to 0.125 mm^3 voxels' or '0.5 mm grid' whenever the source data are not natively acquired at that resolution.","section":"Abstract; §2.1; §2.2; §4"},{"comment":"The WMn images used to construct the atlas are synthetic, generated by a network trained on a private 55-subject dataset, and the paper only states that the output 'was validated by our experts' with no quantitative metric. These synthetic WMn images are then used in §2.4 to train and apply a deep network for thalamic nuclei segmentation (based on THOMAS data). Any systematic synthesis artifact—intensity bias, over-smoothing, or hallucinated contrast in deep gray matter—will propagate directly into the nuclear boundaries and, after majority voting, into the atlas labels. This is particularly concerning for the small structures explicitly mentioned as error-prone (Mammillothalamic Tract, Habenular Nucleus). The authors must provide quantitative evidence (e.g., Dice agreement between synthetic WMn segmentation and real WMn segmentation, or comparison against manual labels) or substantially soften the claims about WMn-dependent substructure accuracy.","section":"§2.2; §2.4; §4"},{"comment":"No quantitative evaluation of the atlas labels is provided. The paper states in §4 that 'the inclusion of human expertise ensures the accuracy and reliability of the final atlas' and in the limitations paragraph that 'the fact that it is based on the fusion of automatic segmentations may raise doubts about its accuracy,' but no overlap measures, volume comparisons with independent manual segmentations, or inter-rater statistics are reported. Since the atlas is proposed as a reference, the absence of any quantitative label validation is a load-bearing gap. At minimum, the authors should report overlap of the final atlas labels with one or more independent manual delineations (even on a subset of structures) or clearly state that label accuracy rests solely on expert visual QC.","section":"§3; §4"}],"minor_comments":[{"comment":"The phrase '0.125 mm3 resolution' appears repeatedly; since mm^3 is a volume unit, this should be '0.125 mm^3 voxel volume' or '0.5 mm isotropic resolution' to avoid confusion.","section":"Abstract; throughout"},{"comment":"There is a typo: 'A throughout QC was done' should read 'A thorough QC was done.'","section":"§2.4, Hippocampus subfield integration"},{"comment":"The reference list has several minor issues: 'Mazziota' should be 'Mazziotta' in the Discussion; reference 38 is empty; and 'compressive list' should be 'comprehensive list' in §1.","section":"§1 and References"},{"comment":"The description of the 'intermediate space' label is vague: it is defined as 'the WM region connecting all the nuclei,' but it is not clear how this region was derived or whether it was manually edited. Please clarify the label definition and its QC process.","section":"§2.4, Thalamus subfield integration"},{"comment":"The template construction mentions median-based sharpening and majority voting but does not specify the number of subjects contributing to each voxel or how missing labels (e.g., structures not present in all subjects) were handled. A sentence on the coverage or left/right symmetry of the labels would improve reproducibility.","section":"§2.5"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a useful public atlas resource, but the central claims need substantial revision. The resolution claim should be corrected to reflect the actual acquisition and resampling, and the synthetic WMn chain requires quantitative validation. The authors are also the developers of several of the segmentation tools used in the fusion (vol2Brain, pBrain, HIPS, CERES), which is not in itself a problem but reinforces the need for independent validation of the final labels. The lack of any quantitative label evaluation makes it difficult to assess the atlas's reliability for downstream applications."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the real contribution is a publicly released multiscale atlas built by carefully integrating seven existing segmentation protocols — and that resource is genuinely useful. But 'ultra-high resolution' is resampling of 0.7 mm native data to a 0.5 mm grid, the WMn contrast is synthetic, and the 350 labels have no quantitative accuracy validation. Use it as a reference and prior; don't treat it as validated ground truth.\n\nWhat's new: the integration itself. Seven protocols (vol2Brain, hypothalamus_seg, BrainVISA sulci, FreeSurfer brainstem/amygdala, pBrain, HIPS, CERES) fused into one coherent hierarchy — 350 substructures, 54 structures, 9 tissues, whole ICV — on a T1/T2/WMn template from 75 HCP subjects, all released under a CC license with label definitions and template scripts. The pipeline is documented in enough detail to critique: strided decomposition for the 1 mm tools, tissue-level correction with a UNet trained on 12 manually corrected cases, spatial-intensity MAP relabeling, manual QC with ITK-SNAP at every stage. The paper earns credit for publishing the artifact and for stating its own limitations plainly — it admits the fusion-of-automatic-segmentations accuracy concern in the Discussion, the young-cohort issue, and that the resolution gain mostly helps small structures.\n\nSoft spots, in proportion. The resolution claim is the mildest: 0.125 mm3 is a resampled grid, not acquisition resolution, though the sharpened averaging does give crisper templates than upsampled MNI152. More serious is the synthetic WMn chain: the synthesis network was trained on a private Bordeaux dataset, 'validated by our experts' with no metric, and it feeds the THOMAS-trained segmentation of 12-13 thalamic nuclei per hemisphere. Systematic synthesis bias would be inherited by the nuclear boundaries, and the paper itself notes manual correction concentrated on the two smallest structures (habenula, mammillothalamic tract). There is no quantitative evaluation of the final labels against manual tracing or any independent protocol — the strongest form of this concern is the paper's own sentence. These are real weaknesses for an atlas meant to serve as substructure-level ground truth, but they are disclosed, not hidden.\n\nWho it's for: anyone needing a freely available sub-1 mm multiscale reference for segmentation priors, education, or normative modeling — provided they read the Methods before believing the small deep-GM labels. I'd send it to peer review. The resource deserves referee time, and a revision that rewords the resolution claims and adds even a limited quantitative label validation would materially strengthen it.","headline":"A real, publicly released multiscale atlas built by careful integration of seven segmentation protocols — useful as a reference, but its 'ultra-high resolution' claim is resampling and its 350 labels have no quantitative accuracy validation.","tokens_in":15878,"tokens_out":4065,"would_cite":true,"duration_ms":33651,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"holiAtlas is a 0.125 mm3 multimodal brain atlas that fuses seven segmentation protocols into 350 substructure labels covering the whole intracranial cavity.","keywords":["brain atlas","MRI","multimodal imaging","ultra-high resolution","brain segmentation","multiscale labelling","thalamic nuclei","hippocampal subfields"],"falsifier":"Expert neuroanatomists manually label the substructures of the holiBrain protocol in, say, 20 subjects not used to build the atlas, and the labels are propagated from the atlas to those subjects; if the overlap between the atlas-propagated labels and the manual labels is low for small structures such as thalamic nuclei or hippocampal subfields, the claim that the 350-label atlas is a valid reference would be undermined.","tokens_in":14754,"feed_emoji":"🧠","tokens_out":9547,"duration_ms":76448,"temperature":0.7,"pith_summary":"This paper introduces holiAtlas, a reference atlas of the healthy human brain built from MRI scans of 75 young adults. The atlas combines T1-weighted, T2-weighted, and white-matter-nulled contrasts at a nominal 0.125 mm3 resolution, and assigns every voxel in the intracranial cavity a label. At its finest scale it carries 350 distinct labels, produced by merging seven existing automatic segmentation protocols into one consistent hierarchy and then correcting the result with semi-automatic and manual quality control. The authors argue that this dense, multiscale, multimodal reference will make small structures such as thalamic nuclei, hippocampal subfields, and cerebellar lobules measurable in ways that 1 mm3 atlases cannot support, which may help detect neurological disease earlier.","feed_headline":"350-label brain atlas built from seven MRI protocols at 0.125 mm3","feed_subtitle":"Multimodal reference map of the healthy brain promises substructure-level measures for early disease detection.","key_machinery":"The load-bearing object is the holiBrain protocol, a hierarchical label system that integrates seven independent delineation protocols into one consistent dense labelling. The integration proceeds by overlaying each new protocol onto the vol2Brain whole-brain segmentation and resolving mismatches through a Bayesian MAP relabelling step that combines smoothed spatial probability maps with intensity likelihoods. The atlas templates themselves are produced by symmetric group-wise normalization with FireANTS, a GPU-accelerated diffeomorphic registration method, iterated ten times; the final images are the median of Laplacian-sharpened volumes and the final labels come from majority voting across the 75 registered label maps.","core_discovery":"The central claim is that a holistic, densely labelled atlas of the human brain can be constructed at 0.125 mm3 resolution by fusing complementary parcellation protocols into a single consistent label system. The authors demonstrate the construction by nonlinearly registering and averaging images from 75 healthy subjects, applying seven delineation protocols, and reconciling their overlapping or conflicting labels with a spatial-intensity diffusion process that relabels voxels according to Gaussian spatial priors and intensity likelihoods. The result is a hierarchical atlas with 350 substructure labels, grouped into 54 structures, 9 tissue classes, and 1 whole-intracranial-cavity label, together with average T1w, T2w, and synthetic WMn templates. The paper's claim is that this atlas is anatomically accurate enough to serve as a reference for ultra-high-resolution segmentation and substructure-level volumetric analysis.","pith_inferences":["If the synthetic WMn contrast is a valid proxy for real white-matter-nulled acquisitions, the WMn template and the thalamic labels derived from it may transfer to other datasets; this transferability is testable by comparing against real WMn scans from a separate cohort, which the paper does not do.","The atlas's accuracy is only checked visually and by the fusion process itself; a quantitative evaluation against manual expert segmentations or histology-derived references would settle how much of the 350-label detail is real anatomy rather than algorithm agreement.","The 75-subject atlas spans only ages 22 to 35, so applying it to pediatric or elderly brains will encounter registration and label bias; building age-specific versions from the same pipeline is a natural next step the paper leaves implicit."],"forward_implications":["If the atlas is accurate, structural analyses can move from organ-level or structure-level volumes to substructure-level volumes, for example measuring atrophy of specific hippocampal subfields or thalamic nuclei rather than whole structures.","The multiscale hierarchy means a single atlas can support analyses at the organ, tissue, structure, and substructure levels without relabelling or coordinate changes.","Because the atlas is multimodal, segmentation methods trained on it can use T1w, T2w, and WMn contrasts simultaneously, potentially improving performance where only one contrast is available.","The public release of templates and label definitions gives other groups a common coordinate system for comparing results at 0.125 mm3 resolution."],"supporting_citations":[{"why":"Supplies the 75 healthy young adults' T1w and T2w images from which the atlas is averaged.","marker":"HCP1200 dataset"},{"why":"Provides the 135-label whole-brain segmentation used as the base layer of the protocol.","marker":"(Manjón et al., 2022)"},{"why":"Supplies the hypothalamus substructure segmentation integrated on top of the base labels.","marker":"(Billot et al, 2020)"},{"why":"Provides the sulci recognition and labelling pipeline used to assign cortical sulcus labels.","marker":"(Mangin et al., 2004)"},{"why":"Supplies the brainstem subdivision into midbrain, pons, medulla, and superior cerebellar peduncle.","marker":"(Iglesias et al., 2015a)"},{"why":"Supplies the amygdala nuclear subdivisions used in the protocol.","marker":"(Iglesias et al., 2015b)"},{"why":"Supplies the substantia nigra, red nucleus, and subthalamic nucleus segmentations.","marker":"(Manjón et al, 2020)"},{"why":"Supplies the hippocampal subfield segmentation (HIPS) with five subfields plus added fimbria and HATA.","marker":"(Romero et al, 2017)"},{"why":"Supplies the cerebellum lobule segmentation (CERES) used for the 12 lobules and cerebellar white matter.","marker":"(Romero et al., 2017)"},{"why":"Provides the GPU-accelerated symmetric registration used for group-wise template generation.","marker":"(Jena et al., 2025)"}],"fun_headline_variants":["Multimodal holistic brain atlas: 350 labels at 0.125 mm³","350-label brain atlas from 7 MRI protocols at 0.125 mm³","Ultra-high-res brain atlas with 350 dense labels from 7 protocols","Holistic brain atlas from 75 subjects: 350 labels, 7 protocols","Fusing 7 MRI protocols into one holistic 350-label brain atlas"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The atlas is only as accurate as the fusion of seven automatic segmentation tools, whose outputs were corrected by visual inspection and manual editing rather than validated quantitatively.","fun_headline_variants_meta":{"raw":{"variants":["Multimodal holistic brain atlas: 350 labels at 0.125 mm³","350-label brain atlas from 7 MRI protocols at 0.125 mm³","Ultra-high-res brain atlas with 350 dense labels from 7 protocols","Holistic brain atlas from 75 subjects: 350 labels, 7 protocols","Fusing 7 MRI protocols into one holistic 350-label brain atlas"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001784,"raw_usage":{"total_tokens":7019,"prompt_tokens":920,"completion_tokens":6099,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":5995}},"tokens_in":536,"tokens_out":6099,"duration_ms":35336,"temperature":1.0,"reasoning_tokens":5995,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T05:56:11.401467+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Expert neuroanatomists manually label the substructures of the holiBrain protocol in, say, 20 subjects not used to build the atlas, and the labels are propagated from the atlas to those subjects; if the overlap between the atlas-propagated labels and the manual labels is low for small structures such as thalamic nuclei or hippocampal subfields, the claim that the 350-label atlas is a valid reference would be undermined.","supporting_citations":[{"cited_title":"Dalca, Jonathan D","cited_arxiv_id":null,"evidence_quote":"Supplies the hypothalamus substructure segmentation integrated on top of the base labels."},{"cited_title":"F., Riviere, D., Cachia, A., Duchesnay, E., Cointepas, Y., Papadopoulos- Orfanos, D","cited_arxiv_id":null,"evidence_quote":"Provides the sulci recognition and labelling pipeline used to assign cortical sulcus labels."}],"review_version":1}