{"id":"e015d1b6-e6e4-4508-972b-209494aca9a2","arxiv_id":"2507.23245","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A multi-parametric, multi-stage fiber clustering atlas automatically identifies CN II, III, V, and VII/VIII pathways with wDice overlap comparable to expert manual tracing.","lead":"This paper builds a multi-nerve diffusion MRI tractography atlas that automatically maps the pathways of five pairs of cranial nerves using fiber clustering on 50 healthy subjects. The authors report high spatial overlap with expert manual tracing on new subjects and on patients with pituitary adenoma and craniopharyngioma, suggesting a faster preoperative mapping tool for skull base surgery.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section VIII says 'All data used in this study were simulated' while Sections III-A and IV validate on real HCP, MDM, PA, and CP data; the central multi-site validation claim is internally unsupported as written.","rationale":"The reader's weakest-assumption analysis correctly flags shared expert priors in the manual ground truth and atlas labels. I agree that this limits what wDice means. However, before that question can even be asked, the paper must establish that the validation data exist. Section VIII explicitly says all data were simulated, which contradicts the methods and results sections. This is not a stylistic concern; it determines whether the central empirical claim has any evidentiary basis. I am not accusing the authors of fabrication: the contradiction may be a boilerplate error, and the promised public repository allows a direct check. Because the contradiction is resolvable and the reader's other concerns remain, I do not think unconditional rejection is warranted. The current manuscript should not be accepted or treated as an established result until the provenance question is settled; hence the verdict should remain conditional, with the data-provenance resolution as a binding condition. If the check shows the data are simulated, the verdict should escalate to reject. This is an explicit data-support passage under the review rule, so it must be weighed in the verdict.","tokens_in":17140,"tokens_out":8560,"duration_ms":104422,"concrete_test":"Clone https://github.com/IPIS-XieLei/CNsAtlas and run the published pipeline on a small number of publicly available HCP subjects with the IDs listed in Fig. 3 (e.g., #100206, #103414) and, if accessible, MDM subjects #001-#005. Confirm whether the code reads real DWI/T1 files from these datasets and whether the automated identification reproduces wDice values near 0.74 (HCP) and 0.78 (MDM). In parallel, request from the authors a per-dataset provenance table (subject IDs, acquisition site, scanner, ethics approval) for the two PA patients and one CP patient. If the code cannot process the claimed real subjects, or if the provenance table is unavailable, the validation claim is not established; if the code does process them and reproduces the numbers, the 'simulated' statement is an error that must be corrected and the remaining concern reduces to ground-truth independence.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing premise of the paper is that the reported wDice values (0.7448 on HCP, 0.7827 on MDM) and the patient demonstrations were computed on real dMRI data from multiple acquisition sites. The text contradicts that premise in Section VIII (Data and Code Availability): after naming HCP, MDM, Xuanwu Hospital PA patients, and a CP patient, it states 'All data used in this study were simulated and did not require an ethical statement.' Section III-A specifies 110 HCP subjects, 20 MDM traveling subjects, and patients with real scanner parameters; Section IV describes manual ROI/ROA ground truth on those subjects; Tables I-II and Figs. 3-5 report outcomes. Simulation and real acquisition cannot both be true. If the simulation sentence is literal, the central claim is not an empirical result; if it is an error, the manuscript must be corrected before any validity assessment. This internal inconsistency, not the shared-expert-prior issue, is the first gate the central claim must pass. The public code and atlas repository makes the contradiction resolvable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a comprehensive diffusion tractography atlas for automatically mapping multiple cranial nerve pairs (CN II, CN III, CN V, and CN VII/VIII) from diffusion MRI. The atlas is built from 50 HCP subjects using pair-specific UKF tractography parameters and a two-stage fiber clustering pipeline, then applied to new subjects by registering their tractography into atlas space and assigning streamlines to atlas clusters. Validation is reported on 60 HCP subjects, 20 MDM subjects, two pituitary adenoma patients, and one craniopharyngioma patient, with mean wDice values of 0.7448 (HCP) and 0.7827 (MDM) against expert manual ROI/ROA-based identification. The authors claim this is the first comprehensive multi-pair CN tractography atlas and that the automated method achieves high spatial correspondence with expert manual annotations.","tokens_in":17475,"tokens_out":3513,"duration_ms":40862,"significance":"If the results are valid, the atlas would be a useful contribution: it addresses a real clinical need for automated, multi-pair CN mapping, and the authors state that code and atlas data are publicly available, which supports reproducibility and community use. The multi-stage clustering strategy and pair-specific tractography parameters are reasonable technical ideas. However, the manuscript's central empirical claim is currently undermined by a direct internal contradiction about whether the data were real or simulated, and the validation design relies on the same expert priors used to label the atlas. Resolving these issues is essential before the quantitative results can be interpreted as evidence of generalizable performance.","major_comments":[{"comment":"Section VIII states 'All data used in this study were simulated and did not require an ethical statement.' This directly contradicts Section III-A, which specifies real acquisition parameters for 110 HCP cases, 20 MDM traveling subjects, two PA patients, and one CP patient, and Section IV, which describes manual ROI/ROA ground-truth selection on those subjects. The reported wDice values, identification rates, and patient figures (Figs. 3-5) are meaningless if the underlying data were simulated, and if the sentence is an error, the manuscript cannot be assessed until it is corrected with an appropriate ethics/data statement. This internal inconsistency is the first gate the central claim must pass.","section":"Section VIII (Data and Code Availability), vs. Section III-A, Section IV-A"},{"comment":"The atlas clusters were labeled by expert manual annotation using the authors' own anatomical criteria (Section III-D2), and the validation ground truth is expert manual ROI/ROA selection following the same group's protocol from Ref. [8] (Section IV-A). Therefore, the wDice values measure agreement between the automated atlas and the same expert priors that defined the atlas, not agreement with an independent ground truth. This weakens the claim of 'ideal colocalisation' and the implication that the method can replace expert judgment. The authors should provide an independent validation, such as a second rater from a different institution, manual-manual variability, or comparison with an independent anatomical standard.","section":"Section III-D2 and Section IV-A"},{"comment":"The clustering hyperparameters (K=6000, then K=300, standard deviations 2.0 and 1.0, two iterations) are selected 'by considering' qualitative false-positive screening on the atlas-building data, and the text reports no quantitative selection criterion or sensitivity analysis. Since these choices directly determine which streamlines are retained as CN clusters, the generalizability claim for new subjects and new acquisition sites depends on these choices not being overfit to the 50 HCP training subjects. A sensitivity analysis or a validation of these parameter choices on held-out subjects should be reported.","section":"Section III-D2 (Multi-stage Fiber Clustering)"},{"comment":"The paper uses the threshold 'wDice ≥ 0.72' from Cousineau et al. [34] as evidence of high spatial overlap. However, Ref. [34] concerns test-retest reproducibility of healthy white matter fascicles, not agreement between automated CN mapping and expert manual identification. The threshold is not directly applicable to the authors' validation setting, and without a manual-manual wDice baseline for the same CNs on the same subjects, the reported values cannot be interpreted as 'high spatial correspondence.' The authors should either report manual-manual variability or justify the threshold in the CN context.","section":"Section V-B and Table II"}],"minor_comments":[{"comment":"There is a typo in the Related Work section: 'owever' should be 'However.'","section":"Section II-B"},{"comment":"The text refers to 'Section III-E' for tractography parameters, but the parameters are actually given in Section III-C. Please correct the cross-reference.","section":"Section VI (Discussion)"},{"comment":"The text states that approximately 50,000 fibers result for each CN pair and then says merging the five pairs gives approximately 200,000 fibers. Five times 50,000 is 250,000; if the merged count is correct, the per-pair estimate or the number of pairs should be clarified.","section":"Section III-C"},{"comment":"The paper states '5 pairs of CNs' but enumerates CN II, CN III, CN V, and CN VII/VIII, with VII/VIII described as a combined pair. The counting should be made explicit (e.g., whether CN VII and CN VIII are counted separately) to avoid confusion.","section":"Abstract and throughout"}],"recommendation":"major_revision","confidential_remarks":"The 'all data simulated' statement in Section VIII appears likely to be a copy-paste error from a simulation-based template, but it is exactly the kind of issue that must be resolved before review can proceed. If the data are real, the authors must supply the required ethics statements and correct the availability section; if they are simulated, the entire experimental section needs to be rewritten. I also note that the validation design, with atlas labels and ground-truth annotations from the same group, is a circularity risk that should be addressed even after the data-contradiction issue is fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nTwo things to know about arXiv:2507.23245. The resource is genuinely useful, and Section VIII currently says \"All data used in this study were simulated\" — which directly contradicts the entire empirical section, where HCP, MDM, and patient acquisitions are described in detail and used to compute wDice. That sentence has to be resolved before any serious validity assessment of the reported numbers.\n\nWhat's actually new: a single atlas covering five CN pairs (II, III, V, VII/VIII) simultaneously, with per-nerve UKF tracking parameters and a two-stage clustering cascade. That is a real step beyond the single-pair atlases in refs 17–20. The validation across HCP, MDM, and two patient cohorts is the right kind of test, and the wDice values (0.7448 on HCP, 0.7827 on MDM) are plausible for this problem. The code and atlas are promised on GitHub, which makes the work checkable.\n\nThe soft spots are real but second-order, except for the data sentence. First, the ground truth is manual ROI/ROA selection using a protocol from the same group (ref [8]) that also labelled the atlas clusters. So wDice measures agreement between two expert-guided methods, not independent accuracy. That's a moderate circularity, common in this literature, but worth stating. Second, clustering parameters appear chosen by screening on the atlas data; validation on held-out HCP subjects partly mitigates this. Third, the Discussion claims a \"100% identification rate for manually verifiable CNs,\" which Table I contradicts — several rows show auto failures (e.g., CN V-R on MDM: 5/9).\n\nIf the data-availability sentence is a literal error, this is a solid resource paper for anyone doing skull base tractography. If it is not an error, the central claim is unsupported. Given the public code and atlas, I'd give the authors a chance to correct it and then send it to peer review.","headline":"A genuinely useful multi-pair cranial nerve atlas, but Section VIII says all data were simulated while the whole paper reports real HCP, MDM, and patient experiments; that contradiction must be resolved before the wDice numbers mean anything.","tokens_in":18013,"tokens_out":2740,"would_cite":false,"duration_ms":30952,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single diffusion tractography atlas can automatically map eight cranial-nerve bundles spanning five nerve pairs, matching expert manual tracing across multiple imaging sites.","keywords":["cranial nerves","diffusion MRI tractography","multi-stage fiber clustering","multi-parametric tractography","UKF tractography","tractography atlas","skull base surgery","optic nerve"],"falsifier":"Run the atlas on a cohort where the true nerve locations are known independently—from intraoperative photographs, cadaveric dissection, or ultra-high-resolution 7T/5T diffusion images—and compare the eight automatically mapped bundles against those locations. If the mean overlap with that independent gold standard falls below the 0.72 threshold, or if the automated CN III bundles systematically miss or falsely cross the brainstem decussation that the authors acknowledge current tractography cannot resolve, the claim that the atlas achieves high spatial correspondence would be refuted.","tokens_in":16932,"feed_emoji":"🧠","tokens_out":6260,"duration_ms":61062,"temperature":0.7,"pith_summary":"This paper claims that a single diffusion tractography atlas can automatically map eight fiber bundles belonging to five pairs of cranial nerves—optic (CN II), oculomotor (CN III), trigeminal (CN V), and facial-vestibulocochlear (CN VII/VIII)—in a new person's brain without expert placement of regions of interest. The atlas is built from roughly one million streamlines generated from 50 HCP subjects using per-nerve tractography parameters, followed by a two-stage clustering cascade that reduces anatomically implausible fibers. On 60 held-out HCP subjects and 20 multi-shell MDM subjects, automated identification reached mean weighted Dice scores of 0.7448 and 0.7827 against expert manual tracing, above the 0.72 threshold commonly used for fiber-overlap agreement. It also identified nerves around tumors in two pituitary adenoma patients and one craniopharyngioma patient. If correct, this would replace a two-hour manual tracing workflow with an automated pipeline of under twenty minutes.","feed_headline":"Single atlas maps five cranial nerve pairs without manual tracing","feed_subtitle":"Automated diffusion tractography matches expert manual annotation across HCP, multi-shell, and tumor-patient data.","key_machinery":"The load-bearing object is the multi-parametric, multi-stage diffusion tractography atlas: a set of 74 expert-labeled fiber clusters in a common space that collectively define eight cranial-nerve bundles. It is built with two-tensor UKF tractography using parameter sets tuned per nerve pair (seeding/stopping FA, Qm, Ql), unbiased groupwise registration of approximately one million streamlines, initial spectral clustering at K=6,000 with two iterations at standard deviation 2.0, ROI screening, and a second enhanced clustering at K=200 with standard deviation 1.0. In a new subject, the same tracking and registration steps are run, streamlines are assigned to the nearest atlas clusters with the same outlier criterion, and the labeled clusters output the final CN pathways.","core_discovery":"The central discovery the authors are trying to establish is that a reusable, multi-nerve diffusion tractography atlas can reproduce expert manual identification of cranial nerve pathways across acquisition sites. They generate the atlas by merging and registering streamlines from 50 HCP subjects, running an initial spectral clustering into 6,000 clusters, retaining 106 candidate clusters, then refining them by enhanced clustering into 200 clusters and expert-selecting 74 that represent the eight CN bundles (CN II decussating and nondecussating, CN III left/right, CN V left/right, CN VII/VIII left/right). A new subject's streamlines are registered into the atlas space and assigned to the nearest atlas clusters, producing automated bundles without any manual ROI placement. The authors report that these automated bundles agree with expert manual ROI/ROA tracts at mean wDice 0.7448 (HCP) and 0.7827 (MDM), and that the atlas identifies nerves in patients where manual selection failed, which they take as evidence of robustness.","pith_inferences":["Editorial inference: because the atlas clusters and the manual ground truth were both defined by expert raters using the same CN anatomical judgment, the reported wDice may measure inter-method agreement more than independent anatomical accuracy; the real test would compare against intraoperative or ultra-high-resolution ground truth.","Editorial inference: the per-nerve parameter tuning that makes the atlas work suggests that extending it to the remaining seven CN pairs will require additional custom tracking protocols and possibly finer clusters near the brainstem, where current dMRI cannot resolve decussating fibers.","Editorial inference: the authors' own comparison with volumetric segmentation shows the two approaches fail differently—atlas streamlines elongate anatomically while voxel segmentation produces fewer extracranial false positives—so a combined pipeline could plausibly beat either alone.","Editorial inference: the atlas's success on manual-failure cases is promising but not yet proof of superiority; those recovered tracts need independent verification before being used in surgery."],"forward_implications":["A clinical user can obtain CN II, III, V, and VII/VIII pathways in under 20 minutes instead of the more than 2 hours typical of manual ROI-based tractography.","The same atlas transfers to data acquired at different sites and resolutions, including 1.25 mm HCP, 1.5 mm multi-shell, and lower-resolution clinical tumor scans.","In subjects where manual ROI identification failed entirely (e.g., 0/6 left CN III on HCP), the atlas still identified most of those tracts, suggesting it can recover pathways a manual protocol misses.","For skull base tumors, the atlas can show the spatial relation between a compressed nerve and the lesion, as demonstrated around the optic nerve in pituitary adenoma and craniopharyngioma patients.","The overlap scores are comparable to or above those of single-pair CN atlases, so mapping five pairs at once does not cost per-nerve accuracy."],"supporting_citations":[{"why":"Supplies the two-tensor Unscented Kalman Filter tractography method used for all CN tracking.","marker":"[14]"},{"why":"Provides the single-pair CN II atlas approach and the manual ROI/ROA style evaluation the paper extends to multiple nerves.","marker":"[17]"},{"why":"Provides the data-driven fiber clustering approach for CN III that the multi-stage clustering builds on.","marker":"[18]"},{"why":"Provides the CN V tractography atlas and shows why per-nerve tracking parameters matter.","marker":"[19]"},{"why":"Provides the CN VII/VIII automated identification framework that the comprehensive atlas generalizes.","marker":"[20]"},{"why":"Shows that uniform tractography parameters fail across CNs, motivating the multi-parametric strategy.","marker":"[21]"},{"why":"Supplies the fiber clustering and pairwise-distance segmentation methodology used in the spectral clustering stage.","marker":"[24]"},{"why":"Supplies the unbiased groupwise tractography registration that puts all subjects' streamlines in a common space.","marker":"[33]"},{"why":"Defines the wDice spatial-overlap metric and the 0.72 agreement threshold used as the evaluation benchmark.","marker":"[34]"}],"fun_headline_variants":["One atlas auto-maps five cranial nerve pairs","Automated atlas: five CN pairs, no manual tracing","Diffusion tractography atlas maps five CN pairs","Atlas replicates expert CN mapping without tracing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole evaluation rests on treating expert manual ROI/ROA tracing as an independent, correct ground truth for where the cranial nerves run, even though the same expert anatomical judgment and protocol also selected and labeled the atlas clusters.","fun_headline_variants_meta":{"raw":{"variants":["One atlas auto-maps five cranial nerve pairs","Automated atlas: five CN pairs, no manual tracing","Diffusion tractography atlas maps five CN pairs","Atlas replicates expert CN mapping without tracing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000786,"raw_usage":{"total_tokens":3534,"prompt_tokens":1077,"completion_tokens":2457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":693,"completion_tokens_details":{"reasoning_tokens":2397}},"tokens_in":693,"tokens_out":2457,"duration_ms":21327,"temperature":1.0,"reasoning_tokens":2397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:54:49.718637+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the atlas on a cohort where the true nerve locations are known independently—from intraoperative photographs, cadaveric dissection, or ultra-high-resolution 7T/5T diffusion images—and compare the eight automatically mapped bundles against those locations. If the mean overlap with that independent gold standard falls below the 0.72 threshold, or if the automated CN III bundles systematically miss or falsely cross the brainstem decussation that the authors acknowledge current tractography cannot resolve, the claim that the atlas achieves high spatial correspondence would be refuted.","supporting_citations":[{"cited_title":"Filtered multitensor tractography,","cited_arxiv_id":null,"evidence_quote":"Supplies the two-tensor Unscented Kalman Filter tractography method used for all CN tracking."},{"cited_title":"Automated identification of the retinogeniculate visual pathway using a high-dimensional tractography atlas,","cited_arxiv_id":null,"evidence_quote":"Provides the single-pair CN II atlas approach and the manual ROI/ROA style evaluation the paper extends to multiple nerves."},{"cited_title":"Automatic oculomotor nerve identification based on data- driven fiber clustering,","cited_arxiv_id":null,"evidence_quote":"Provides the data-driven fiber clustering approach for CN III that the multi-stage clustering builds on."},{"cited_title":"Cre- ation of a novel trigeminal tractography atlas for automated trigeminal nerve identification,","cited_arxiv_id":null,"evidence_quote":"Provides the CN V tractography atlas and shows why per-nerve tracking parameters matter."},{"cited_title":"Automated facial–vestibulocochlear nerve com- plex identification based on data-driven tractography clustering,","cited_arxiv_id":null,"evidence_quote":"Provides the CN VII/VIII automated identification framework that the comprehensive atlas generalizes."},{"cited_title":"Anatomical assessment of trigeminal nerve tractography using diffusion mri: A comparison of acquisition b-values and single- and multi-fiber tracking strategies,","cited_arxiv_id":null,"evidence_quote":"Shows that uniform tractography parameters fail across CNs, motivating the multi-parametric strategy."},{"cited_title":"Automatic tractography segmentation using a high-dimensional white matter atlas,","cited_arxiv_id":null,"evidence_quote":"Supplies the fiber clustering and pairwise-distance segmentation methodology used in the spectral clustering stage."},{"cited_title":"Unbiased groupwise registration of white matter tractography,","cited_arxiv_id":null,"evidence_quote":"Supplies the unbiased groupwise tractography registration that puts all subjects' streamlines in a common space."},{"cited_title":"A test-retest study on parkinson’s ppmi dataset yields statistically significant white matter fascicles,","cited_arxiv_id":null,"evidence_quote":"Defines the wDice spatial-overlap metric and the 0.72 agreement threshold used as the evaluation benchmark."}],"review_version":1}