{"id":"ed5ccb83-37b9-4541-ae98-e73f494a6b7d","arxiv_id":"1908.01594","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An attention U-Net with transfer learning automatically segments knee menisci in 3D UTE cones MRI as accurately as radiologists, and derived relaxation times agree closely with manual segmentation.","lead":"This paper tests an attention U-Net deep learning model, pretrained on natural images, to automatically outline knee menisci in specialized 3D UTE magnetic resonance images. The automatic outlines matched radiologist outlines about as well as the two radiologists matched each other, and yielded similar T1, T1rho, and T2* relaxation measurements.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-difference relaxometry claim is built on t-tests that treat 191 correlated slices as independent; with only 15 menisci, the tests are underpowered and cannot establish equivalence.","rationale":"I focused on the statistical analysis of the relaxometry comparison because it directly supports the practical conclusion that automatic segmentations can replace manual ROIs for T1/T1ρ/T2* measurement. The reader's weakest assumption concerns the validity of the manual ground truth; that is a meaningful concern, but even if the ground truth were a perfect gold standard, the paper's 'no differences' evidence would still be compromised by the correlated-slice analysis. The paper reports p-values from two-sided t-tests on 191 2D ROIs from 15 menisci. Because slices within a meniscus share anatomy and relaxation values, they are not independent; the effective sample size is roughly 15. Low power means the absence of a significant difference cannot be read as equivalence. A proper equivalence test with a clinically meaningful margin, or a mixed-effects model, is needed. The same logic applies to the Dice-based equivalence claim: the authors compare point estimates but do not formally test equivalence. My concern does not invalidate the feasibility contribution; it suggests the paper's central claim is stronger than the evidence supports as analyzed. For a conditional acceptance, the authors should be asked to reanalyze with appropriate statistical methods. Therefore I recommend UNCHANGED relative to the reader's CONDITIONAL verdict.","tokens_in":14821,"tokens_out":5915,"duration_ms":54439,"concrete_test":"Reanalyze the relaxometry comparison using the 15 menisci as the unit of analysis: average the per-slice T1, T1ρ, and T2* values within each meniscus, then perform paired t-tests between automatic and manual segmentations. Also run a two one-sided tests (TOST) equivalence procedure with a prespecified margin, e.g., ±5% of the manual mean, and construct 90% confidence intervals for the mean differences. If any confidence interval exceeds the equivalence bounds, the 'no differences' claim is not supported. For the Dice equivalence, compute per-menisci Dice scores and test whether the mean CNN-vs-radiologist Dice is equivalent to the mean radiologist-vs-radiologist Dice using TOST with a margin such as ±0.05.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that automatic segmentations yield 'no associated differences' in T1, T1ρ, and T2* rests on paired t-tests performed on 191 2D images drawn from just 15 test menisci. Because adjacent slices within a meniscus are spatially correlated, the effective sample size is at most 15, not 191. With n=15, the t-tests have low power; a non-significant p-value is not evidence of equivalence. The paper never reports an equivalence test with a prespecified margin (e.g., TOST), nor does it account for within-meniscus correlation (e.g., averaging per meniscus or a linear mixed model). The same issue affects the segmentation-equivalence claim: the Dice comparison between CNN and radiologist versus radiologist inter-observer is purely descriptive, with no formal equivalence test. Thus the evidence for both halves of the headline claim is systematically overconfident. The paper itself acknowledges the subjectivity of the reference image selection ('strictly subjective way') in the Discussion, which compounds the concern because the manual ROIs are the only anchor for relaxometry accuracy; however, the statistical unit-of-analysis flaw is more directly fatal and can be tested from the reported data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a transfer-learning-based 2D attention U-Net for segmenting knee menisci in 3D UTE cones MR images and for computing T1, T1ρ, and T2* from the resulting masks. Sixty-one subjects were imaged; two radiologists independently outlined menisci on subtracted AdiabT1ρ-weighted images; the data were split into training/validation/test sets (36/10/15), with 191 test slices from 15 menisci. The authors report Dice scores of 0.860 and 0.833 for the two CNNs versus radiologist inter-observer Dice of 0.820, high correlations (0.90-0.97) between manual and automatic relaxometry values, and no statistically significant differences in average T1, T1ρ, and T2* between manual and automatic segmentations. They conclude that the automatic segmentations are equivalent to the radiologists' inter-observer variability and can replace manual ROIs for meniscal relaxometry.","tokens_in":15010,"tokens_out":5153,"duration_ms":49515,"significance":"If the claims hold, this would be a useful practical result: it is, to the authors' knowledge, the first demonstration of fully automatic meniscal segmentation and relaxometry on 3D UTE cones data. The study has real strengths: a held-out test set, validation-based model selection, two independent radiologist reference standards, separate reporting for medial and lateral menisci, AUC and false-positive checks, and Bland-Altman plots. The segmentation performance (Dice around 0.83-0.86) is in line with prior knee menisci segmentation work. However, the load-bearing statistical claims of equivalence in both segmentation and relaxometry are not supported by the analyses as reported.","major_comments":[{"comment":"The claim of 'no associated differences' in T1, T1ρ, and T2* rests on t-tests applied to 191 2D images derived from only 15 test menisci. Adjacent 2D slices from the same meniscus are spatially correlated, so the effective sample size is at most 15, not 191. With n=15, the tests are underpowered, and a non-significant p-value does not establish equivalence. The manuscript should report per-meniscus or per-subject averages and use an equivalence test (e.g., TOST with a pre-specified margin) or a linear mixed-effects model that accounts for within-meniscus correlation. Without this, the abstract's statement that there were 'no associated differences' is not warranted.","section":"Methods, 'Model development and performance evaluation'; Results, Table 2"},{"comment":"The segmentation-equivalence claim ('equivalent to the inter-observer variability of two radiologists') is based on descriptive comparison of average Dice scores (0.860 and 0.833 versus 0.820). No formal non-inferiority or equivalence test is reported for the CNN-versus-radiologist agreement relative to the radiologist inter-observer agreement. The only significance test mentioned for Dice (CNN1 vs CNN2 0.882 versus Rad1 vs Rad2 0.820, p<0.001) is not described in the Methods and appears to treat per-slice Dice values as independent. A formal test, or a clearly worded claim of 'comparable' rather than 'equivalent', is needed.","section":"Results, Table 1; Discussion"},{"comment":"The reference standard is based on ROIs drawn on subtracted AdiabT1ρ-weighted images selected 'based on subjective assessment', and the Discussion acknowledges the selection was 'strictly subjective'. Because the CNNs are trained to reproduce these specific ROIs, systematic errors in the manual outlines or in the choice of subtraction image are propagated into the automatic segmentations and the relaxometry comparison. The authors should assess the sensitivity of their conclusions to the choice of subtraction image, or at minimum discuss how this limitation bears on the claim that the method can 'replace' manual ROIs.","section":"Methods, 'UTE imaging and data collection'; Discussion"}],"minor_comments":[{"comment":"The phrase 'no associated differences' should be rephrased to 'no statistically significant differences were detected' unless equivalence testing with a prespecified margin is provided.","section":"Abstract and Conclusions"},{"comment":"The test set is described as '191 2D images from 15 menisci' while the data split describes 15 subjects; clarify whether '15 menisci' refers to 15 subjects or 15 individual menisci, and state how many medial/lateral menisci are included.","section":"Methods, 'Model development and performance evaluation'"},{"comment":"There is a formatting error in the 95% CI for the CNN2 row: '0.885 0.904' is missing a dash and should read '0.885-0.904'.","section":"Table 1"},{"comment":"The caption says 'NIR value' but the text uses 'NAFP' (number of adiabatic full passages); please correct this inconsistency.","section":"Figure 1 caption"},{"comment":"The text refers to 'cartridge compartments' where 'cartilage compartments' is intended; also, in the Methods there is a duplicated 'T1' in 'per adiabatic T1T1ρ preparation.'","section":"Discussion"},{"comment":"The '1D convolutional block' added in front of the VGG19 blocks is described as consisting of 1x1 convolutional filters; calling it a 1D convolutional block is confusing and should be clarified.","section":"Methods, network description"},{"comment":"No code or data availability statement is provided, which limits reproducibility of the transfer-learning implementation.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about pseudoreplication is valid and is the main reason for major revision: the equivalence claims are not supported by the reported t-tests on 191 correlated slices from 15 menisci. The authors should be asked to reanalyze at the level of independent menisci/subjects and to perform formal equivalence testing. The paper is otherwise a reasonable feasibility study with useful descriptive results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid feasibility study. The authors show that a 2D attention U-Net, initialized with VGG19 weights, can segment menisci on 3D UTE cones MRI about as well as a radiologist, and that the resulting ROIs give essentially the same T1, T1rho, and T2* values as manual tracing. I believe that practical claim. The weakest part is the statistical framing of “equivalence” — the paper leans on t-tests over 191 2D slices from only 15 menisci, and those slices are correlated. The stress-test note is right: the effective sample size is far smaller than 191, so the non-significant p-values don’t establish equivalence, and the tests are likely underpowered. That’s a real flaw, but it’s fixable and doesn’t overturn the core result.\n\nWhat’s genuinely new: this is the first menisci segmentation and relaxometry work on 3D UTE cones data, a sequence that matters for short-T2 tissues. The authors use two radiologists’ ROIs, a proper train/validation/test split, and they evaluate downstream relaxometry, not just Dice. The correlations of 0.90–0.97 and the tiny mean differences (about 1 ms for T1rho, under 0.5 ms for T2*) are the most convincing evidence. They also compare cross-observer performance sensibly and acknowledge the subjectivity of the reference image selection.\n\nWhere it’s soft: beyond the statistical unit-of-analysis issue, the attention mechanism and transfer learning are never ablated, so we don’t know if they beat a plain U-Net. That’s a moderate omission for a methods-focused paper, though not a fatal one. No code or data is released, which limits reproducibility but is still common in this field. And the ground truth is the radiologists’ manual ROIs, so the CNN inherits whatever bias those ROIs carry — fine for automating the current workflow, but not a claim about true meniscal boundaries.\n\nOverall, the paper is a competent incremental contribution with a clear clinical motivation. The main result is believable and useful for the UTE and osteoarthritis communities. I’d send it to peer review and would ask for per-meniscus statistics (mixed model or TOST), a baseline U-Net comparison, and ideally a clear statement about effective sample size. It’s not a home run, but it’s a fair candidate for publication after revision.","headline":"A useful, credible feasibility study for automatic menisci segmentation and relaxometry in 3D UTE cones MRI, but the 'equivalence' claim is statistically overconfident and should be tightened before publication.","tokens_in":15637,"tokens_out":3230,"would_cite":true,"duration_ms":34315,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 2D attention U-Net with transfer learning can outline knee menisci on 3D UTE cones MR images at the same level as two human radiologists, and the automatic outlines yield T1, T1ρ, and T2* values statistically indistinguishable from…","keywords":["knee meniscus segmentation","ultrashort echo time MRI","UTE cones","attention U-Net","transfer learning","T1rho relaxometry","T2* relaxometry","osteoarthritis biomarkers"],"falsifier":"Build a consensus ground truth for the same or new UTE cones volumes from a panel of several radiologists (or from higher-resolution histology where available), then compute Dice scores and relaxometry differences for the two CNNs against that consensus. If the CNNs' Dice against consensus falls well below the radiologists' average Dice against consensus, or if the automatic T1, T1ρ, or T2* values differ from the consensus-based values by more than the 1.95%, 2.56-3.03%, and 4.59-6.15% average relative errors reported here, the equivalence claim would fail.","tokens_in":14572,"feed_emoji":"🦵","tokens_out":7024,"duration_ms":70213,"temperature":0.7,"pith_summary":"The paper is trying to establish that a fully automatic deep-learning pipeline can replace hand-drawn regions of interest for quantitative knee meniscus MRI on 3D ultrashort echo time (UTE) cones images. It trains two 2D attention U-Nets, each on the outlines of one of two radiologists, using transfer learning from a network pretrained on natural images. On held-out subjects the automatic segmentations reached Dice scores of 0.860 and 0.833, slightly above the 0.820 Dice score between the two radiologists themselves. The T1, T1ρ, and T2* averages computed from automatic and manual outlines showed no statistically significant differences and correlated at 0.90 to 0.97. If correct, this makes meniscal relaxometry feasible without time-consuming manual tracing, supporting osteoarthritis assessment.","feed_headline":"Deep learning outlines knee menisci as reliably as radiologists","feed_subtitle":"A 2D attention U-Net matches manual T1, T1ρ, and T2* reads on UTE knee MRI.","key_machinery":"The central object is a 2D attention U-Net, a U-Net whose skip connections are gated by attention layers so that encoder features far from the initial meniscus localization are suppressed. The first two encoder blocks are initialized from the corresponding blocks of a deep network pretrained on a large natural-image dataset, and a trainable 1×1 convolution block in front converts the normalized grayscale subtracted UTE images into the three-channel input format that pretrained network expects. Training uses a Dice-score loss directly, which suits the menisci's small pixel fraction. The resulting segmentations are fed into nonlinear least-squares fitting of the UTE cones data to compute T1, T1ρ, and T2* maps.","core_discovery":"The central claim is that a 2D attention U-Net with transfer learning, trained on radiologist-outlined regions of interest in subtracted AdiabT1ρ-weighted UTE cones slices, segments the knee menisci at a level equivalent to inter-observer variability. The models achieved Dice scores of 0.860 and 0.833 against the radiologist whose outlines trained them, while the two radiologists' manual segmentations agreed with each other at 0.820. The average T1, T1ρ, and T2* relaxation times computed from the automatic segmentations were not statistically different from those computed from manual outlines, with Pearson correlations between 0.90 and 0.97 and average relative errors of roughly 2% for T1, 3% for T1ρ, and 5-6% for T2*. The two CNNs also agreed with each other more strongly than the radiologists did (Dice 0.882 versus 0.820), which the paper reads as evidence that the learned segmentations are driven by image content rather than by annotator idiosyncrasy.","pith_inferences":["A further step, not tested here, would be to evaluate the same models against a consensus ground truth from many annotators; since the two CNNs agree more than the radiologists do, the models may be averaging out annotator noise, but any systematic boundary bias shared by the radiologists would still be baked into both the labels and the evaluation.","The error pattern—smallest for T1 and largest for T2*—suggests T2* relaxometry is the most sensitive to boundary disagreements; an OA study using T2* as its primary biomarker may need tighter segmentation agreement than this pipeline currently guarantees on atypical slices.","Because the method is trained on 2D slices, through-plane meniscus curvature is ignored; a likely extension, consistent with earlier work the paper cites, is to add a 3D refinement network, and a testable prediction is that this would push held-out Dice further above the radiologist baseline.","The same automatic masks could also be used for morphometric measures such as meniscal volume or extrusion on UTE data, an extension the paper mentions as future whole-joint work but does not validate here."],"forward_implications":["Automatic segmentations can replace manual tracing for meniscal T1, T1ρ, and T2* relaxometry on similar 3T UTE cones scans without statistically changing the average values.","The segmentation accuracy is at or above radiologist inter-observer level: a model trained on one radiologist's outlines agrees with the other radiologist about as well as the two radiologists agree with each other.","The CNNs produced no false positives on slices without menisci, supporting use of the model in fully automated slice-filtering and quantification pipelines.","The method could enable efficient extraction of osteoarthritis-relevant biomarkers, including relaxation times and ROI areas, from whole-knee UTE acquisitions.","Because the two CNNs agreed with each other more strongly than the two radiologists did, automatic segmentations may reduce inter-observer variability in meniscal measurements."],"supporting_citations":[{"why":"Supplies the base U-Net architecture that the attention mechanism and transfer learning modify.","marker":"13"},{"why":"Supplies the pretrained deep network whose first convolutional blocks initialize the encoder for transfer learning.","marker":"27"},{"why":"Supplies the large natural-image dataset used for pretraining those transferred blocks.","marker":"29"},{"why":"Supplies the attention-gated skip connections that let the network focus on small meniscus regions.","marker":"37"},{"why":"Provides the prior 2D U-Net meniscus and cartilage segmentation with relaxometry that this paper's Dice and correlation results are compared against.","marker":"18"},{"why":"Provides a prior CNN menisci segmentation baseline that this work compares with and is measured against in Dice terms.","marker":"19"},{"why":"Provides a prior whole-joint knee anatomy segmentation CNN and an earlier transfer-learning-from-MR approach that this paper extends.","marker":"17"},{"why":"Supplies the 3D adiabatic T1ρ-prepared UTE cones sequence used to generate the T1ρ-weighted images and the subtracted images the radiologists outlined.","marker":"8"},{"why":"Supplies the 3D UTE cones AFI and VFA method used to compute whole-knee T1 maps.","marker":"9"},{"why":"Establishes the short T2/T2* of meniscus that motivates UTE imaging and the T1ρ mapping approach used here.","marker":"26"}],"fun_headline_variants":["AI matches radiologists in knee meniscus MRI segmentation","Attention U-Net segments knee menisci on par with experts","Deep learning equals radiologist reliability for meniscus MRI","UTE MRI menisci: AI segmentation rivals human experts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the two radiologists' manually drawn outlines, made on subjectively selected subtracted images, are the correct meniscus boundary; because those outlines are both the training labels and the evaluation standard, any systematic error in them is learned and reproduced by the network.","fun_headline_variants_meta":{"raw":{"variants":["AI matches radiologists in knee meniscus MRI segmentation","Attention U-Net segments knee menisci on par with experts","Deep learning equals radiologist reliability for meniscus MRI","UTE MRI menisci: AI segmentation rivals human experts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000207,"raw_usage":{"total_tokens":1488,"prompt_tokens":1122,"completion_tokens":366,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":738,"completion_tokens_details":{"reasoning_tokens":299}},"tokens_in":738,"tokens_out":366,"duration_ms":3699,"temperature":1.0,"reasoning_tokens":299,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:08:26.702668+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a consensus ground truth for the same or new UTE cones volumes from a panel of several radiologists (or from higher-resolution histology where available), then compute Dice scores and relaxometry differences for the two CNNs against that consensus. If the CNNs' Dice against consensus falls well below the radiologists' average Dice against consensus, or if the automatic T1, T1ρ, or T2* values differ from the consensus-based values by more than the 1.95%, 2.56-3.03%, and 4.59-6.15% average relative errors reported here, the equivalence claim would fail.","supporting_citations":[{"cited_title":"U-net: Convolutional networks for biomedical image segmentation","cited_arxiv_id":null,"evidence_quote":"Supplies the base U-Net architecture that the attention mechanism and transfer learning modify."},{"cited_title":"Very deep convolutional networks for large-scale image recognition","cited_arxiv_id":null,"evidence_quote":"Supplies the pretrained deep network whose first convolutional blocks initialize the encoder for transfer learning."},{"cited_title":"Imagenet: A large-scale hierarchical image database","cited_arxiv_id":null,"evidence_quote":"Supplies the large natural-image dataset used for pretraining those transferred blocks."},{"cited_title":"Attention U-Net: Learning Where to Look for the Pancreas","cited_arxiv_id":null,"evidence_quote":"Supplies the attention-gated skip connections that let the network focus on small meniscus regions."},{"cited_title":"Use of 2D U-Net Convolutional Neural Networks for Automated Cartilage and Meniscus Segmentation of Knee MR Imaging Data to Determine Relaxometry and Morphometry","cited_arxiv_id":null,"evidence_quote":"Provides the prior 2D U-Net meniscus and cartilage segmentation with relaxometry that this paper's Dice and correlation results are compared against."},{"cited_title":"Knee menisci segmentation using convolutional neural networks: data from the Osteoarthritis Initiative","cited_arxiv_id":null,"evidence_quote":"Provides a prior CNN menisci segmentation baseline that this work compares with and is measured against in Dice terms."},{"cited_title":"Deep convolutional neural network for segmentation of knee joint anatomy","cited_arxiv_id":null,"evidence_quote":"Provides a prior whole-joint knee anatomy segmentation CNN and an earlier transfer-learning-from-MR approach that this paper extends."},{"cited_title":"3D adiabatic T1ρ prepared ultrashort echo time cones sequence for whole knee imaging","cited_arxiv_id":null,"evidence_quote":"Supplies the 3D adiabatic T1ρ-prepared UTE cones sequence used to generate the T1ρ-weighted images and the subtracted images the radiologists outlined."},{"cited_title":"Whole knee joint T1 values measured in vivo at 3T by combined 3D ultrashort echo time cones actual flip angle and variable flip angle methods","cited_arxiv_id":null,"evidence_quote":"Supplies the 3D UTE cones AFI and VFA method used to compute whole-knee T1 maps."},{"cited_title":"Ultrashort TE T1ρ (UTE T1ρ) imaging of the Achilles tendon and meniscus","cited_arxiv_id":null,"evidence_quote":"Establishes the short T2/T2* of meniscus that motivates UTE imaging and the T1ρ mapping approach used here."}],"review_version":1}