{"id":"3cfa5a74-5380-4525-9201-ab250c3f9508","arxiv_id":"2412.06717","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A Swin Transformer ensemble detects Bankart shoulder lesions on non-contrast and contrast MRIs with AUCs of 0.87 and 0.90 against surgical ground truth, though the non-contrast test set contains only 6 positives.","lead":"This paper trains a deep learning model to detect Bankart lesions, a common shoulder cartilage tear, on regular MRIs and on MRIs with injected contrast dye, using arthroscopic surgery findings as the ground truth. It reports roughly 85 percent accuracy and AUCs of 0.87 and 0.90, but the non-contrast MRI result rests on only six positive test cases and needs external validation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of non-invasive parity rests on a standard-MRI sensitivity estimate from only 6 positive test cases, so the headline comparison to radiologists reading MRAs is not statistically established.","rationale":"The reader's verdict is CONDITIONAL, and my analysis supports keeping that verdict. The reader's weakest-assumption identification centers on the radiologist baseline: routine clinical reports were not written for a focused Bankart-detection task, so the 16.7% standard-MRI sensitivity may understate true radiologist performance. That is a legitimate and important concern, and it applies directly to the 'surpasses radiologists' language. However, I find an even more load-bearing problem upstream: the standard-MRI model's sensitivity is computed from 6 positives, and the specificity from 65 negatives, while the MRA radiologist comparator uses 17 positives. The headline claim that a non-invasive standard MRI matches or surpasses radiologists reading MRAs reduces to comparing 5/6 against 14/17. Those proportions are not statistically distinguishable, and the confidence interval around 5/6 is enormous. This fragility is compounded by the threshold-selection procedure, which used a validation set with only about 4 positive standard-MRI cases. The AUC of 0.87 is also reported without a numeric confidence interval, despite the bootstrap shading in Figure 4. Taken together, the evidence supports the paper's feasibility claim but not the stronger comparative claim. The paper itself acknowledges the single-center, small-data limitations and calls for external validation, which is appropriate. I therefore agree with the CONDITIONAL verdict: the central contribution is plausible and worth reporting, but the headline claim about non-invasive parity should be framed as preliminary and should be accompanied by explicit uncertainty intervals. My concrete test would provide the missing statistical quantification and would tell us whether the comparative claim survives even a fair, controlled reader study. The independent strengths of the paper—surgical ground truth, multi-view ensembling, and a held-out test set—are real, but they do not resolve the small-denominator fragility of the key comparison.","tokens_in":8888,"tokens_out":3171,"duration_ms":36370,"concrete_test":"Compute exact Clopper-Pearson 95% confidence intervals for the Table 2 sensitivity and specificity entries and bootstrap 95% confidence intervals for the standard-MRI AUC (0.87) using the 71 test MRIs. Then perform an exact permutation or Fisher-style test comparing the standard-MRI model sensitivity/specificity pair (5/6 and 56/65) against the MRA radiology-report pair (14/17 and 25/29).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 4.2 claim that the standard-MRI model 'matched or surpassed' radiologists interpreting MRAs. The entire basis for that claim in Table 2 is a standard-MRI sensitivity of 83.33% (5/6) compared with the MRA radiology-report sensitivity of 82.35% (14/17). With only 6 positive standard-MRI test cases, a single missed or gained prediction moves sensitivity by roughly 17 percentage points, and the exact 95% confidence interval for 5/6 spans from about 36% to 99%. The two proportions are statistically indistinguishable, and the reported AUC of 0.87 is accompanied in Figure 4 by shaded bootstrap intervals, but no numeric intervals are given and the standard-MRI test set has only 6 positives and 65 negatives. In addition, the operating threshold for standard MRIs (0.71) was selected on a validation set containing only about 4 positive standard-MRI cases (10% of 40), so the threshold and the resulting sensitivity/specificity pair are fragile. The reader correctly notes that the radiologist baseline comes from routine clinical reports rather than a focused Bankart-reading task; that is a real validity threat, but it is downstream of the more basic sample-size problem. Even with a perfectly controlled radiologist comparator, 6 positive cases cannot support a claim that a non-invasive standard MRI matches or surpasses an invasive MRA read. Thus the central load-bearing assumption is not merely the fairness of the baseline, but the statistical stability of the model's standard-MRI performance estimate itself.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a deep learning pipeline for detecting Bankart lesions on both standard (non-contrast) shoulder MRI and MR arthrography (MRA). Separate Swin Transformer models are fine-tuned for each modality using arthroscopy-derived labels as ground truth, and multi-view ensembles of sagittal, axial, and coronal slice-level predictions are evaluated on a 20% hold-out test set (71 standard MRIs and 46 MRAs). The authors report AUCs of 0.87 and 0.90, with sensitivity/specificity around 83%/86% for standard MRI and 82%/86% for MRA. The central claim is that the standard-MRI model matches or surpasses radiologists interpreting MRAs, based on a comparison with retrospective clinical radiology reports. The paper is framed as a feasibility study and calls for external validation.","tokens_in":9168,"tokens_out":2587,"duration_ms":27668,"significance":"If the central claim were firmly established, the work would be clinically valuable because it could reduce reliance on invasive MRAs for a common shoulder pathology. The study has notable methodological strengths: ground truth comes from intraoperative arthroscopic findings, inter-rater agreement on a labeling subset is perfect (Fleiss's kappa = 1.0), the model is pre-trained on a relevant knee MRI dataset, and predictions are ensembled across multiple views. The MRA result (AUC 0.90 on 46 cases) is plausible and suggests the pipeline is functional. However, the headline comparison between standard MRI and radiologist-read MRA is not statistically supportable from the reported data, and the radiologist baseline is not derived from a controlled reading task. As presented, the significance of the paper is therefore limited to a promising proof-of-concept rather than a demonstrated clinical equivalence.","major_comments":[{"comment":"The claim that the standard-MRI model 'matched or surpassed' radiologists interpreting MRAs rests on a sensitivity of 83.33% (5/6) in the standard-MRI test set. With only six positive cases, the 95% confidence interval for this proportion spans roughly 36% to 99%, and a single changed prediction moves sensitivity by about 17 percentage points. The difference between 5/6 and the radiologist MRA sensitivity of 14/17 is not statistically distinguishable. The manuscript must report confidence intervals for all sensitivity, specificity, and accuracy figures, and either temper the equivalence claim or present it with explicit acknowledgment of this instability.","section":"4.2, Table 2"},{"comment":"The radiologist comparator is derived from original clinical radiology reports written for routine care, not from a focused evaluation of the anterior-inferior labrum. The reported standard-MRI sensitivity of 16.7% (1/6) is far below the 52-55% cited from the literature, which suggests the reports may not have targeted Bankart lesion detection. Without a reader study in which radiologists are asked to explicitly assess for Bankart lesions under the same conditions as the model, the 'surpasses radiologists' comparison is not a valid head-to-head evaluation. This issue affects the central claim and should be addressed by either performing a reader study or substantially reframing the conclusion.","section":"4.2, Table 2"},{"comment":"The decision threshold for the standard-MRI ensemble (0.71) was selected on a validation set containing only about 4 positive cases (10% of 40). The resulting sensitivity/specificity pair is therefore fragile, and the reported operating point may not reflect the model's true performance at a clinically meaningful threshold. The authors should report the full ROC curve with numeric confidence intervals, and discuss how the operating point would change under alternative threshold-selection strategies or a larger validation set.","section":"3.1, 4.2"},{"comment":"The ROC curves in Figure 4 include shaded 95% confidence intervals calculated by bootstrapping, but no numeric interval values are given anywhere in the text or tables. Since the test set for standard MRI has only 6 positive cases, the AUC of 0.87 is likely accompanied by a very wide interval. The manuscript should report the numeric bootstrap intervals for both AUCs and for the threshold-dependent metrics in Table 2.","section":"4.1, Figure 4"}],"minor_comments":[{"comment":"There is an inconsistency in the patient count: Section 2.1 states 546 patients with 586 MRIs, while Section 2.2 and Table 1 report 558 patients. The abstract also says 558 patients. This should be reconciled.","section":"2.1, 2.2, Table 1"},{"comment":"The inter-rater reliability statement reports a Fleiss's kappa of 1.0 on a 20-MRI subset. This is unusually perfect and may merit a brief explanation of the labeling procedure or the amount of discussion among raters.","section":"2.1"},{"comment":"The phrase 'significantly exceeding radiologist sensitivity' is used without a statistical test. Given the small sample, it would be more accurate to say 'higher' or 'nominally higher' unless a formal comparison is provided.","section":"4.2"},{"comment":"The first column header 'Total MRIs1' appears to have a superscript issue, and the meaning of the superscript is not explained. Consider simplifying the header.","section":"Table 1"},{"comment":"The augmentation procedure is described as 'ten-fold augmentation of training samples,' but it is not clear whether this increases the effective number of training epochs or the dataset size. A brief clarification would help reproducibility.","section":"2.3, 3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a conference-style feasibility study and clearly states its limitations. The main concern is that the abstract and conclusions overstate the evidence for non-invasive diagnostic equivalence based on a 6-positive-case test set and an unvalidated radiologist comparator. This is fixable within the manuscript's scope by adding confidence intervals, reframing the claims, and explicitly discussing the fragility of the threshold-dependent metrics. I would encourage the authors to consider repositioning the contribution as a feasibility study of deep learning for Bankart detection on standard MRI rather than a demonstration of superiority over radiologists."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent feasibility study that does one genuinely new thing—applying deep learning to Bankart lesions specifically, with intraoperative ground truth—but its headline claim is not supported by the number of positive cases behind it. The standard-MRI \"matched or surpassed radiologists on MRA\" conclusion rests on 5 true positives out of 6. One flipped prediction moves sensitivity by roughly 17 points; the 95% CI for 5/6 spans from about 36% to 99%. So the central comparison is statistically indistinguishable from chance at that sample size.\n\nWhat the paper does well: surgical ground truth is the right reference standard, the two-surgeon labeling protocol with kappa=1.0 is a solid touch, and the multi-view Swin ensemble is a sensible, if standard, architecture. The MRA result (AUC 0.90 on 46 cases) is plausible and would be a useful data point on its own. The authors also deserve credit for being explicit in the conclusion that the standard-MRI cohort is highly imbalanced and that stability across splits needs future study; they don't hide the fragility.\n\nSoft spots, in rough order of importance. First, the abstract and Section 4.2 overstate. \"Matched or surpassed radiologists interpreting MRAs\" is not established with 6 positives, regardless of how fair the radiologist baseline is. Second, the radiologist comparator is routine clinical reports, which were not written for a focused Bankart-reading task; that's a real validity threat for the specific comparison, though downstream of the sample-size issue. Third, the bootstrapped 95% CIs are shown as shaded bands but no numbers are given, which makes the CI claim hard to verify at reading time. Fourth, the dataset description is internally inconsistent: 546 patients in one place, 558 in another. Minor but should be fixed. I'd also note the threshold on standard MRI (0.71) was tuned on a validation set with about 4 positives; that's another source of optimism.\n\nNet: the core MRA feasibility result is believable; the non-invasive parity claim is a hypothesis, not a result. This is the kind of paper a serious referee should see—it is a legitimate new application with real clinical motivation—but it needs revision: reframe the standard-MRI claim, report numeric CIs, and ideally add external validation.\n\nI'd send it to peer review with a request for major revision. I would not cite the standard-MRI parity claim in my own work; I might cite the MRA model if the numbers hold up after revision.","headline":"A legitimate first DL application to Bankart lesions with solid surgical ground truth, but the headline non-invasive parity claim rests on only 6 positive cases and needs reframing.","tokens_in":9711,"tokens_out":1865,"would_cite":false,"duration_ms":18524,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a deep-learning ensemble can detect Bankart lesions on non-contrast shoulder MRIs with accuracy comparable to radiologists reading contrast-enhanced MR arthrograms, potentially reducing the need for invasive imaging.","keywords":["Bankart lesion","glenoid labral tear","deep learning","standard MRI","MR arthrography","Swin Transformer","multi-view ensemble","computer-aided diagnosis"],"falsifier":"Re-run the evaluation with radiologists explicitly asked to report on the anterior-inferior glenoid labrum for the same standard MRIs, blinded to arthroscopy results; if their focused-reading sensitivity matches or exceeds the model's 83.3 percent, the paper's central comparative claim fails. Alternatively, an external multi-site test set with more than six Bankart-positive standard MRIs could settle whether the 83.3 percent sensitivity (5 of 6) is real or a small-sample artifact.","tokens_in":8692,"feed_emoji":"🦴","tokens_out":8865,"duration_ms":76079,"temperature":0.7,"pith_summary":"This paper tries to establish that deep learning can diagnose Bankart lesions—tears of the anterior-inferior glenoid labrum—directly from standard, non-contrast shoulder MRIs, a setting where radiologists often miss the finding and patients are instead sent for invasive MR arthrography (MRA). Using arthroscopy as the gold standard, the authors trained separate Swin Transformer models on 335 standard MRIs and 251 MRAs, then ensembled sagittal, axial, and coronal predictions. The ensemble reached AUCs of 0.87 on standard MRIs and 0.90 on MRAs, with accuracy around 85 percent. The paper's headline finding is that on standard MRIs the model's 83.3 percent sensitivity far exceeded the 16.7 percent sensitivity of the original radiology reports, and its overall performance matched or surpassed radiologists interpreting MRAs. If true, this would give clinicians a non-invasive route to a diagnosis that today usually requires contrast injection or surgery.","feed_headline":"Deep learning reads Bankart tears on non-contrast shoulder MRI","feed_subtitle":"Ensemble hits 0.87 AUC on standard MRIs, matching radiologists on MR arthrograms and possibly cutting invasive imaging.","key_machinery":"The load-bearing mechanism is a Swin Transformer, a vision-transformer architecture, pretrained on a public knee MRI dataset and fine-tuned separately on standard shoulder MRIs and MRAs. Each 3D MRI is treated as a series of 2D slices; slice-level features are aggregated by max pooling into a per-scan vector, and separate models for the sagittal, axial, and coronal views are trained and their output probabilities averaged. The decision threshold is set on the validation set at the point where sensitivity and specificity are equal, and ground-truth labels come from intraoperative arthroscopy findings, the gold standard for Bankart lesions.","core_discovery":"The central claim is that a multi-view deep-learning ensemble can detect Bankart lesions on both standard MRIs and MR arthrograms with diagnostic performance comparable to or better than radiologists, and specifically that the standard-MRI model rivals radiologist performance on MRAs. On a hold-out set of 71 standard MRIs (6 with Bankart lesions), the model achieved an AUC of 0.87, accuracy 85.9 percent, sensitivity 83.3 percent, and specificity 86.2 percent; on 46 MRAs (17 with lesions), it achieved an AUC of 0.90, accuracy 84.8 percent, sensitivity 82.4 percent, and specificity 86.2 percent. The paper reports that its standard-MRI sensitivity of 83.3 percent is far above the 16.7 percent sensitivity of the original radiology reports on the same scans and is within the 74–96 percent range reported in the literature for radiologists reading MRAs, which motivates the conclusion that non-invasive MRI plus deep learning could substitute for arthrography.","pith_inferences":["If externally validated, the clinical pathway could shift toward standard MRI as the first-line imaging test for suspected Bankart lesions, reserving arthrography for cases where the model is uncertain or surgery is already planned.","The radiologist comparison is probably conservative in one way and optimistic in another: the reports were written for routine care rather than a focused anterior-inferior labrum assessment, so the 16.7 percent sensitivity may understate radiologists' true skill; conversely, the model was trained and tested on the same institution's equipment and protocols, so its edge could shrink on outside data","The test set contains only six Bankart-positive standard MRIs, so the reported 83.3 percent sensitivity corresponds to 5 correct detections; a reader should treat the sensitivity estimate as provisional until a larger, multi-site cohort is evaluated.","Because MRA and standard MRI cohorts differ in age, sex, and tear prevalence, the two models are not directly comparable; a model trained on one population may need recalibration before deployment in a different clinic."],"forward_implications":["A deep-learning screen on standard MRI could identify Bankart lesions without requiring contrast injection, avoiding the pain, cost, and rare complications of MR arthrography.","In the paper's test set, the standard-MRI model caught 5 of 6 Bankart lesions while the original radiology reports caught only 1 of 6; if replicated, this would address the main weakness of non-contrast shoulder MRI.","The multi-view ensemble outperformed every single-view model, so combining sagittal, axial, and coronal information appears important for detecting these subtle tears.","The MRA model's specificity of 86.2 percent sits slightly below the 91–98 percent range reported for radiologists, which the authors argue is clinically tolerable because imaging is interpreted alongside patient history."],"supporting_citations":[{"why":"Supplies the Swin Transformer architecture used as the image feature extractor.","marker":"[26]"},{"why":"Provides the public knee MRI dataset that the model is pretrained on before fine-tuning.","marker":"[27]"},{"why":"Gives the literature baseline for radiologist sensitivity (52–55%) on standard MRIs that the model's 83.3% sensitivity is compared against.","marker":"[9]"},{"why":"Reports non-contrast MRI diagnostic performance used as the standard-MRI literature comparator.","marker":"[22]"},{"why":"Supplies the literature range (74–96% sensitivity) for radiologists on MR arthrograms that the model's MRA performance is compared with.","marker":"[13]"}],"fun_headline_variants":["AI matches radiologists on MR arthrograms using only plain shoulder MRI","Deep learning spots Bankart tears on non-contrast MRI at 0.87 AUC","Non-invasive Bankart diagnosis: DL rivals MRI arthrogram performance","Swin transformer ensemble detects shoulder labral tears without contrast","Plain MRI plus AI matches MRA-level diagnosis of Bankart tears"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that the model surpasses radiologists rests on using routine clinical radiology reports as the radiologist baseline—reports not written for the specific task of detecting anterior-inferior labral tears—and on a test set with only six Bankart-positive standard MRIs, so both the comparison and the sensitivity estimate are fragile.","fun_headline_variants_meta":{"raw":{"variants":["AI matches radiologists on MR arthrograms using only plain shoulder MRI","Deep learning spots Bankart tears on non-contrast MRI at 0.87 AUC","Non-invasive Bankart diagnosis: DL rivals MRI arthrogram performance","Swin transformer ensemble detects shoulder labral tears without contrast","Plain MRI plus AI matches MRA-level diagnosis of Bankart tears"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000552,"raw_usage":{"total_tokens":2709,"prompt_tokens":1101,"completion_tokens":1608,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":717,"completion_tokens_details":{"reasoning_tokens":1514}},"tokens_in":717,"tokens_out":1608,"duration_ms":14319,"temperature":1.0,"reasoning_tokens":1514,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:19:25.828077+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the evaluation with radiologists explicitly asked to report on the anterior-inferior glenoid labrum for the same standard MRIs, blinded to arthroscopy results; if their focused-reading sensitivity matches or exceeds the model's 83.3 percent, the paper's central comparative claim fails. Alternatively, an external multi-site test set with more than six Bankart-positive standard MRIs could settle whether the 83.3 percent sensitivity (5 of 6) is real or a small-sample artifact.","supporting_citations":[{"cited_title":"Deep-learning-assisted diagnosis for knee magnetic resonance imaging: Development and retrospective validation of MRNet,","cited_arxiv_id":null,"evidence_quote":"Provides the public knee MRI dataset that the model is pretrained on before fine-tuning."},{"cited_title":"Assessment of the rotator cuff and glenoid labrum using an extremity MR system: MR results compared to surgical findings from a multi-center study,","cited_arxiv_id":null,"evidence_quote":"Gives the literature baseline for radiologist sensitivity (52–55%) on standard MRIs that the model's 83.3% sensitivity is compared against."},{"cited_title":"Non-contrast magnetic resonance imaging for diagnosing shoulder injuries,","cited_arxiv_id":null,"evidence_quote":"Reports non-contrast MRI diagnostic performance used as the standard-MRI literature comparator."},{"cited_title":"Accuracy of MR arthrog- raphy in the detection of posterior glenoid labral injuries of the shoulder,","cited_arxiv_id":null,"evidence_quote":"Supplies the literature range (74–96% sensitivity) for radiologists on MR arthrograms that the model's MRA performance is compared with."}],"review_version":1}