{"id":"6a542e16-157e-42a4-9846-38822eb20ca7","arxiv_id":"2501.06229","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An open-source manually annotated 3D MRI vocal tract database for 10 French speakers enables benchmarking showing 3D U-Nets with transfer learning segment vocal tract airspace at Dice around 0.90.","lead":"This paper releases a new open-source database of 53 manually annotated 3D vocal tract MRI volumes from 10 French speakers, and benchmarks four deep learning segmentation models against it. The best models, a 3D U-Net and a transfer-learned 3D U-Net, reach about 0.90 Dice score, and transfer learning matches the full 3D U-Net with less than half the training data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark's accuracy numbers are not reference-independent: the STAPLE test truth includes the same voice-science team that made the training labels, and the paper concedes STAPLE can reinforce systematic errors; this is the weakest load-bearing assumption for the central claim.","rationale":"I agree with the reader's weakest assumption: the STAPLE reference is the load-bearing element, and the same-team annotator overlap is the specific mechanism by which it can fail. This is not a claim of bad faith; it is a structural evaluation risk. The dataset release, code repository, and reproducible training pipeline are genuine strengths, and the 3D U-Net results may well survive re-evaluation. But the central benchmark claim is stated as numerical accuracy, and that accuracy cannot be trusted until the reference is shown to be independent of the training-label style. The HD contradiction in Results is an additional sign that the quantitative section needs revision; the subject counts in Tables 1A/1B also appear to sum to 46 training volumes and 8 distinct subject IDs rather than the stated 45 volumes and 10 speakers, and these should be checked. None of these issues invalidate the dataset itself, so a conditional verdict with requested revision remains appropriate; my read does not move the reader's verdict.","tokens_in":16913,"tokens_out":11063,"duration_ms":102852,"concrete_test":"Recompute the benchmark using three test references from the released data: (a) the full three-annotator STAPLE, (b) STAPLE from only the two independent BME graduate-student annotators, and (c) the voice-science annotator alone. Report per-model Dice and HD against each reference. If removing the voice-science annotator from STAPLE (b vs a) changes any model's mean Dice by more than roughly 0.01–0.02, or changes the model ranking, the accuracy claim is reference-dependent and must be restated with independent-reference numbers. A stronger variant: have a new annotator with no role in training-label creation independently segment the eight test volumes and recompute the metrics against that reference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 3D U-Net and transfer-learned 3D U-Net reach Dice 0.896 and are usable for vocal-tract segmentation rests on the STAPLE consensus being an unbiased ground truth. That condition is insecure. In Methods (Data Annotation and Evaluation & Metrics), the eight test volumes are evaluated against a STAPLE fusion of three annotators, and one of those three is the same voice-science team (two graduate students plus expert vocologist D. Meyer) that produced the training labels. The models were trained to imitate that team's segmentation conventions, so including a member of that team in the test reference can inflate apparent accuracy even if the two BME graduate-student annotators are independent. The paper's own Discussion concedes that 'STAPLE can be sensitive to systematic errors; if annotators make consistent mistakes, the algorithm may reinforce rather than correct them.' Near teeth (MRI signal voids) and glottal landmarks the task is genuinely ambiguous, so convention-dependent errors are plausible. The ambiguity is not hypothetical: the printed HD results are internally contradictory, giving the transfer-learned model an average HD of both 15 ± 24.9 and 3.95 ± 5.2 in the same Results section. Until the reference is shown to be independent of the training-label style and the HD numbers are corrected, the headline Dice values overstate what is established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an open-source dataset of 53 manually annotated 3D vocal tract MRI volumes from 10 French speakers and benchmarks four deep learning segmentation architectures (2D slice-by-slice U-Net, 3D U-Net, 3D U-Net with transfer learning from lung CT, and 3D transformer U-NetR). The central empirical claim is that the 3D U-Net and the transfer-learned 3D U-Net achieve the highest Dice coefficients (0.896 ± 0.05 and 0.896 ± 0.04, respectively) on an eight-volume STAPLE-based test set, with transfer learning matching the 3D U-Net while using fewer than half the training volumes. The authors release the manual annotations and the model code, which supports reproducibility.","tokens_in":17189,"tokens_out":7232,"duration_ms":64511,"significance":"If the findings hold after correcting the evaluation issues, the paper provides a valuable community resource: 53 manually annotated 3D vocal tract volumes with STAPLE consensus and public training code. The empirical comparison of 2D versus 3D convolution versus transformer architectures is directly relevant to the speech-MRI community, and the transfer-learning result is practically meaningful for small-data settings. The authors are to be credited for releasing the data and code. However, the benchmark's numerical claims are currently compromised by the STAPLE reference dependence and by contradictory Hausdorff distance reporting; the dataset contribution is independent of these issues, but the accuracy claims need revision.","major_comments":[{"comment":"The test reference used for benchmarking is not independent of the training-label style. In the Methods section, the STAPLE reference for the eight test volumes is built from three annotators, and 'segmentations from a third annotator were those provided by the voice-science team as described above'—the same team (two voice-science graduate students plus D. Meyer) that produced the training labels. Because the models are trained to reproduce that team's segmentation conventions, including a member of that team in the test consensus can inflate apparent accuracy even if the two BME graduate-student annotators are independent. The Discussion itself concedes that 'STAPLE can be sensitive to systematic errors; if annotators make consistent mistakes, the algorithm may reinforce rather than correct them.' Given the known ambiguity near teeth (MRI signal voids) and glottal landmarks, this is not a hypothetical concern. Please re-evaluate the test set against a STAPLE consensus of only the two independent annotators, or report pairwise inter-annotator Dice and Hausdorff values to demonstrate that the shared team's segmentations do not dominate the consensus.","section":"Methods, Data Annotation and Evaluation & Metrics; Discussion"},{"comment":"The Hausdorff distance results are internally contradictory. The text states that the 2D slice-by-slice U-Net achieved an average HD of 11.3 ± 5.4, 'lower than the 3D U-Net (14 ± 28) and the transfer learning 3D U-Net (15 ± 24.9)', and then immediately states that 'The 3D U-Net transfer learning model showed the best HD distance with an average of (3.95 + 5.2)'. Both statements cannot be correct, and the discrepancy directly affects the conclusion about which models have lower boundary variability and which is 'best' in HD. Please correct the numbers in Table 2 and the text, and re-derive any conclusions about model ranking and variability after the correction.","section":"Results, Table 2 paragraph"}],"minor_comments":[{"comment":"The caption states that 'The four sounds are labeled' but then lists five labels (/f/, /l/, 'UP', /k/, /a/); please correct the count, for example to 'four sounds and one voiceless posture'.","section":"Results, Figure 4 caption"},{"comment":"The phoneme /kõn/ in the Abstract is written as 'kon' in Table 1A and as /k/ in Figure 4; please use a consistent phonetic transcription throughout.","section":"Abstract and Table 1A"},{"comment":"Please clarify whether the 20 volumes used for the transfer-learning model are a subset of the 45 training volumes and how they were selected; the current text only states that 'only 20 samples from the French speaker dataset were utilized.'","section":"Methods, Data Annotation"},{"comment":"Please report the final selected hyperparameters for each architecture (number of epochs, steps per epoch, learning rate, dropout, and frozen layers for the transfer-learning model) in the text or a supplementary table; listing only the search ranges is insufficient for reproducibility.","section":"Methods, Implementation and hyper-parameter tuning"},{"comment":"The large HD standard deviations (e.g., 14 ± 28 and 15 ± 24.9) suggest the presence of strong outliers; consider reporting the median and interquartile range in addition to mean ± SD, or identify the outlier volumes, to give a more robust summary of boundary errors.","section":"Results, Table 2"},{"comment":"References 40 and 45 cite the same STAPLE paper, and references 35 and 53 are also duplicates; please consolidate duplicate references.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The dataset and code release are the main contribution and fit the journal's scope. The evaluation methodology needs revision before the benchmark claims can be accepted. The authors should also verify that the Figshare and GitHub links are live and include a version or DOI. The paper relies heavily on the authors' prior transfer-learning work (ref. 54); the incremental contribution of the 3D transfer-learning benchmark over that work should be made explicit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the dataset is the contribution, and it is a real one. Fifty-three manually annotated 3D vocal tract volumes from 10 French speakers, with code and data online. That fills a gap; existing open datasets are mostly 2D mid-sagittal or small German samples. The annotation pipeline (two grad students plus an expert vocologist) is reasonable. So the paper deserves attention.\n\nThe benchmark is a fair application of known architectures, and the transfer-learning result is plausible: a 3D U-Net pre-trained on lung CT reaches parity with a from-scratch 3D U-Net using 20 vs 45 training volumes. Qualitatively the figures support that. I have no problem with the self-citation; the transfer learning protocol is from their prior work and it's cited.\n\nNow the soft spots, in proportion. The numbers as printed are not internally consistent. In the Results, the 2D U-Net is said to have HD 11.3 ± 5.4, 'lower than' the 3D U-Net (14 ± 28) and transfer learning (15 ± 24.9), and then two sentences later the transfer learning model 'showed the best HD distance with an average of (3.95 + 5.2).' Those cannot both be true. The table is not in the text, so I can't check which value is correct, but a reader cannot trust either until this is fixed. Also, no significance tests are reported; the Dice differences between top models are within one standard deviation, so the 'superior performance' language is stronger than the evidence.\n\nThe deeper concern is the STAPLE reference. For the eight test volumes, three annotators contributed, and one of them is the same voice-science team that produced the training labels. The models were trained to imitate that team's conventions, so including their labels in the test reference inflates apparent accuracy for all models, particularly the ones that overfit those conventions. The paper's own Discussion concedes STAPLE can reinforce systematic errors. Near teeth and glottal landmarks, the task is genuinely ambiguous. This doesn't sink the dataset, but it means the Dice values are not reference-independent; they are 'agreement with this annotation style' more than 'accuracy.' The fix is to re-evaluate on a reference built without the training-label team, or at least report the per-annotator Dice for the two independent BME students.\n\nBottom line: the dataset is worth having, and the transfer-learning observation is worth testing further. But the quantitative results section, as written, is not up to standard. A serious editor should send this to peer review, because the resource is valuable and the flaws are correctable. The authors need to fix the HD numbers, add significance testing or soften the superiority claims, and address the STAPLE dependence.","headline":"A genuinely useful open-source vocal tract MRI dataset with a benchmark that needs a corrected evaluation before the numbers can be trusted.","tokens_in":17703,"tokens_out":2101,"would_cite":true,"duration_ms":19360,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new open-source dataset of 53 manually annotated 3D vocal tract MRI volumes from 10 French speakers, benchmarked across four deep learning architectures, shows that 3D U-Nets—including a transfer-learned model trained on only 20…","keywords":["vocal tract segmentation","3D MRI","deep learning","U-Net","transfer learning","STAPLE","open-source dataset","speech MRI"],"falsifier":"Measure the pairwise Dice agreement among the three annotators on the eight test volumes; if the annotators' mutual agreement is comparable to or lower than the models' 0.896 Dice against the STAPLE consensus, the reported scores reflect the consensus construction rather than true anatomical accuracy. A second check: re-annotate the test volumes with an independent team and compare the resulting STAPLE consensus segmentations; large disagreement would invalidate the reference.","tokens_in":16741,"feed_emoji":"🗣️","tokens_out":7630,"duration_ms":57660,"temperature":0.7,"pith_summary":"This paper establishes that a small, open-source set of manually annotated 3D vocal tract MRI volumes is enough to train automatic segmentation models for speech research. The authors annotated 53 volumes from 10 French speakers covering 21 phonemes and 3 voiceless tasks, and benchmarked four deep learning architectures against a STAPLE consensus of three human annotators. The 3D U-Net and a transfer-learned 3D U-Net both reached an average Dice coefficient of 0.896, with the transfer-learned version using fewer than half the training volumes (20 vs 45). The 2D slice-by-slice U-Net scored lower (0.823), and the transformer-based 3D U-NetR produced frequent non-anatomical segmentations. If these results hold, researchers can automatically extract vocal tract geometry from MRI at scale, enabling larger studies of speech, singing, and voice disorders.","feed_headline":"3D U-Nets auto-segment vocal tract MRI at 0.90 Dice","feed_subtitle":"Transfer learning from lung CT matches a fully trained 3D U-Net using only 20 of 53 annotated volumes.","key_machinery":"The load-bearing mechanism is the transfer-learned 3D U-Net: a three-dimensional convolutional encoder–decoder architecture pre-trained to segment lungs in chest CT, then fine-tuned on 20 annotated vocal tract volumes with the early layers frozen. This setup lets the model reuse low-level biomedical image features from CT while learning upper-airway anatomy from a small target dataset. The evaluation machinery is the STAPLE algorithm, which fuses three independent manual segmentations into a probabilistic reference that mitigates inter-annotator variability when scoring test volumes.","core_discovery":"The central discovery is that 3D convolutional U-Nets, one trained from scratch on 45 annotated volumes and one pre-trained on lung CT and fine-tuned on 20 volumes, produce vocal tract segmentations that agree with a multi-annotator STAPLE reference at an average Dice coefficient of 0.896 (±0.05 and ±0.04). Transfer learning from an unrelated medical imaging domain (lung CT) gives the same accuracy with less than half the training data, while the 2D slice-by-slice U-Net and the transformer-based U-NetR lag behind in accuracy and anatomical plausibility. The paper also releases the dataset and code, with eight test volumes accompanied by three independent manual segmentations and their STAPLE consensus.","pith_inferences":["If the transfer-learning result generalizes across scanners and languages, then new vocal tract segmentation studies could be launched with just a handful of annotated volumes, making large-cohort articulatory research far cheaper.","The frequent mis-segmentation near teeth and the hard palate points toward a concrete testable extension: combining these models with zero-echo-time MRI or bone priors should reduce the 'islands of airspace' errors the authors describe.","The dataset's French phoneme inventory, paired with the released code, could support cross-linguistic comparisons of vocal tract geometry that English-only datasets cannot, and could serve as a pretraining resource for other upper-airway segmentation tasks.","A direct follow-up measurement the paper does not report: the pairwise Dice agreement among the three annotators on the test volumes. That number would show how much of the models' 0.896 score is limited by human rater variability."],"forward_implications":["Automatic segmentation of 3D vocal tract MRI is achievable with only about 20 manually annotated volumes when starting from a lung CT pre-trained model.","The released dataset provides the first open-source labeled 3D vocal tract volumes from French speakers, extending speech MRI resources beyond English.","Researchers can use the trained models to compute 3D vocal tract area functions and mid-sagittal shape changes from new MRI data without manual tracing.","The comparison suggests 3D convolutional networks are better suited than transformer U-Nets for this segmentation task with limited training data.","The models' consistent errors near teeth and the glottis define the specific anatomical regions where further methodological work is needed."],"supporting_citations":[{"why":"Supplies the 10-speaker French MRI database from which the 53 annotated volumes were drawn.","marker":"[13]"},{"why":"Provides the STAPLE algorithm used to fuse three manual segmentations into the reference ground truth.","marker":"[40]"},{"why":"Defines the base U-Net architecture that all four benchmark models derive from.","marker":"[49]"},{"why":"Extends U-Net to 3D volumes; the 3D U-Net and transfer-learning model are built on it.","marker":"[50]"},{"why":"Introduces the transformer-based UNETR architecture used as the fourth benchmark model.","marker":"[51]"},{"why":"Prior transfer-learning approach on speech MRI that motivates the small-data fine-tuning strategy.","marker":"[54]"},{"why":"Provides the OSIC lung CT dataset used for pre-training the transfer-learning model.","marker":"[41]"},{"why":"Supplies the lung segmentations in the CT volumes that serve as pre-training labels.","marker":"[42]"},{"why":"Prior 2D U-Net airway segmentation work that the slice-by-slice baseline extends.","marker":"[31]"}],"fun_headline_variants":["Transfer learning from lung CT halves MRI training data for vocal tract segmentation","3D U-Net with lung CT pretraining needs only 20 MRI volumes","Vocal tract MRI segmentation: 3D CNN beats 2D and transformer nets","Open MRI vocal tract dataset enables deep learning segmentation at Dice 0.90","3D CNNs outperform 2D and transformers for vocal tract MRI segmentation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results stand on the assumption that the STAPLE consensus of three annotators is an accurate reference for the vocal tract airspace, even though the annotation task is genuinely ambiguous near teeth and glottal landmarks and one of the annotators belongs to the same team that produced the training labels.","fun_headline_variants_meta":{"raw":{"variants":["Transfer learning from lung CT halves MRI training data for vocal tract segmentation","3D U-Net with lung CT pretraining needs only 20 MRI volumes","Vocal tract MRI segmentation: 3D CNN beats 2D and transformer nets","Open MRI vocal tract dataset enables deep learning segmentation at Dice 0.90","3D CNNs outperform 2D and transformers for vocal tract MRI segmentation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001184,"raw_usage":{"total_tokens":4794,"prompt_tokens":755,"completion_tokens":4039,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":371,"completion_tokens_details":{"reasoning_tokens":3937}},"tokens_in":371,"tokens_out":4039,"duration_ms":25892,"temperature":1.0,"reasoning_tokens":3937,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:38:33.962967+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the pairwise Dice agreement among the three annotators on the eight test volumes; if the annotators' mutual agreement is comparable to or lower than the models' 0.896 Dice against the STAPLE consensus, the reported scores reflect the consensus construction rather than true anatomical accuracy. A second check: re-annotate the test volumes with an independent team and compare the resulting STAPLE consensus segmentations; large disagreement would invalidate the reference.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior 2D U-Net airway segmentation work that the slice-by-slice baseline extends."},{"cited_title":"G., Sutton, B","cited_arxiv_id":null,"evidence_quote":"Supplies the 10-speaker French MRI database from which the 53 annotated volumes were drawn."},{"cited_title":"Local limit theorems in relatively hyperbolic groups II : the non-spectrally degenerate case","cited_arxiv_id":"2004.13986","evidence_quote":"Defines the base U-Net architecture that all four benchmark models derive from."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the transformer-based UNETR architecture used as the fourth benchmark model."},{"cited_title":"K., Tong, Y., Torigian, D","cited_arxiv_id":null,"evidence_quote":"Provides the OSIC lung CT dataset used for pre-training the transfer-learning model."}],"review_version":1}