{"id":"41752178-e33d-4786-b2fc-595de1f4c4ac","arxiv_id":"2509.06830","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A DINOv2 vision transformer pre-trained on 200M unlabeled CT/MRI slices matches or beats prior foundation models and radiology residents on a new 19-task benchmark.","lead":"Researchers trained a medical AI model, Curia, on 200 million CT and MRI images from one hospital's routine scans, without using any labels. On a new 19-task test, it matched or beat existing medical AI systems and radiology residents, and it transferred knowledge between CT and MRI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'meets or surpasses radiologists' claim is not yet auditable: Fig. 1e rests on four residents, unreported test-subset sizes, and an undisclosed cross-task aggregation of AUC/accuracy/c-index.","rationale":"The reader's weakest_assumption was single-center pre-training data, but the paper itself acknowledges that limitation and partially mitigates it with multi-center public benchmarks. The most load-bearing part of the central claim is the abstract's assertion that Curia 'meets or surpasses the performance of radiologists.' That assertion rests on a human-evaluation procedure that is under-specified at exactly the points needed to audit it: subset sizes, per-reader scores, aggregation across incompatible metrics, and whether reader variability is modeled. This is a reporting gap rather than a demonstrated error, so it does not overturn the paper; it does justify the CONDITIONAL verdict and a request for clarification before the claim is taken at face value. The reader's rationale did note the radiologist comparison was under-specified, but their formal 'weakest assumption' was single-center data, hence partial agreement.","tokens_in":28056,"tokens_out":11067,"duration_ms":134368,"concrete_test":"Publish a supplementary table with per-task number of test images read by each of the four residents, per-reader metric values, and the exact formula used to pool AUC, balanced accuracy, c-index, and r² into Fig. 1e. Recompute the comparison with a mixed-effects model (task and reader as random effects) on the same test subsets; if the pooled advantage over the resident average is no longer significant or reverses, revise the abstract to 'comparable to resident radiologists on selected tasks'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim includes clinical superiority. The only support is Fig. 1e and §4.4. The text reports a mean over 4 final-year residents across 14 tasks, but never states: per-task test-set sizes, per-reader scores, which metric is averaged or how heterogeneous metrics (AUC, balanced accuracy, c-index, r²) are normalized, or whether the paired bootstrap test compares Curia to each resident or to the resident average with reader random effects. Without this, 'meets or surpasses resident radiologists' could be an artifact of the averaging scheme and a small convenience sample. Because this sentence is in the abstract and is the clinical headline, it is load-bearing. This is an auditability gap, not evidence of fraud; the underlying Curia-vs-FM comparisons may still be valid.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Curia, a family of Vision Transformer foundation models (ViT-B and ViT-L) pre-trained with DINOv2 self-supervision on roughly 200 million 2D CT and MRI slices from 150,000 exams (130 TB) collected at a single private hospital between 2019 and 2022. The authors also release CuriaBench, a 19-task external benchmark covering organ recognition, registration, segmentation, classification, regression, and survival prediction across CT and MRI. Evaluation uses frozen backbones with lightweight heads and compares against MedImageInsight, BiomedCLIP, and, in places, DINOv2 and a cancer imaging foundation model. Claims include strong few-shot and cross-modality generalization, competitive or superior performance relative to resident radiologists across 14 tasks, and clinically useful emergent properties. The base model weights are released on Hugging Face.","tokens_in":28271,"tokens_out":5489,"duration_ms":62507,"significance":"If the central claims hold, this is a valuable contribution: it demonstrates that self-supervised pre-training on a large, routine clinical imaging corpus can yield a transferable radiology backbone, and CuriaBench is a potentially useful public benchmark. The statistical protocol is a strength: 5 training runs, 1000-sample bootstrapping, paired tests, and external public datasets. Release of model weights improves reproducibility. However, the headline claim of meeting or surpassing resident radiologists is not currently auditable, several main-text claims are contradicted by the paper's own supplementary tables, and a few data-reporting errors undermine confidence in the statistical presentation. These issues are fixable but require substantive revision.","major_comments":[{"comment":"The headline claim that Curia 'meets or surpasses the performance of radiologists' is not auditable. The text reports a mean over four final-year residents across 14 tasks but omits per-task test-subset sizes, per-reader scores, the metric used when averaging heterogeneous metrics (AUC, balanced accuracy, c-index, r²), and whether the paired bootstrap test compares Curia to each resident or to the resident average. Please provide a full per-task radiologist evaluation table and the exact protocol; without this, the abstract claim is unsupported.","section":"§4.4, Fig. 1e"},{"comment":"The claim that Curia 'outperformed or matched the performance of other models across all image registration tasks' and 'outperformed others on all organ-specific metrics' is contradicted by Table D6. For XCAT CT→CT, BiomedCLIP has mean DSC 81.74 vs Curia-B 81.30 and Curia-L 80.12. For XCAT CT→MR, DINOv2 Large has higher liver DSC (86.22 vs 86.12 for Curia-B) and higher spleen DSC (71.90 vs 70.09 for Curia-B). Similarly, the Introduction's 'consistently and significantly outperforms existing foundation models' is contradicted by §2.4, where MedImageInsight significantly outperforms Curia-L on Abdominal Trauma (93.14 vs 87.10, P<0.001) and scores higher on Myocardial Infarction (94.08 vs 89.16). These claims must be qualified.","section":"§2.1, Table D6"},{"comment":"The text reports Curia-L AUROC 87.74 for the Alzheimer's disease task, but Table E23 reports 84.90 (95% CI 74.51–93.78). This is a direct numerical inconsistency. It must be resolved, since it affects the reported competitiveness in the neurodegenerative benchmark. The same table also lists Curia-B as 87.83, matching the text for Curia-B only.","section":"§2.5, Table E23"},{"comment":"The Kidney Cancer Survival benchmark uses tumor volumes segmented by the authors' own segmentation FM [41], then 'the FM' is used with a Cox layer to predict time-to-event. Please specify whether the survival features are extracted from Curia or from the same segmentation FM, and describe the steps taken to prevent information leakage between mask generation and risk prediction. With n=183 patients, the c-index comparisons are also underpowered; the validation-set threshold selection should be described in more detail.","section":"§4.5.2"},{"comment":"Two supplementary tables report inverted confidence intervals: Subarticular Stenosis Curia-B is 87.81 with lower 95% CI 89.09 and upper 95% CI 86.57 (Table E17), and Abdominal Trauma Curia-B is 82.63 with lower 95% CI 83.98 and upper 95% CI 81.31 (Table E20). These errors call into question the reliability of the other CI tables. Please correct them and re-check all bootstrap CI computations.","section":"Appendix E, Tables E17 and E20"}],"minor_comments":[{"comment":"The r² score is reported as '75.54' and '69.41' etc. Since r² is normally bounded by 1, these values should be expressed as percentages (e.g., 75.54%) or as 0.7554 to avoid confusion.","section":"§2.1, Table E10"},{"comment":"The heading 'Curia achieves leading performance in musculoskeletal disease assessment' overstates the results: for foraminal narrowing and spinal canal stenosis, Curia-L is statistically equivalent to MedImageInsight (P=0.87 and P=0.243, respectively). Consider rewording to 'comparable or better'.","section":"§2.3"},{"comment":"The radiologist comparison is only described textually; the figure itself should display the per-task values with confidence intervals, or a dedicated table should be added, so that the 'average of four residents' result can be inspected.","section":"Fig. 1e"},{"comment":"Please specify the level of the residents (e.g., final-year, but in which country's training system), whether they were supervised or board-certified, and how the 14 radiologist tasks were selected from the 19 benchmark tasks.","section":"§4.4"},{"comment":"The comparison with Harvard OncoFM is not apples-to-apples because OncoFM-finetuned updates the encoder, whereas Curia uses a frozen backbone. This is noted in the table caption, but it should also be prominently stated in the main text.","section":"Table E11"},{"comment":"The caption mentions 'Harvard-RT', while the text and references refer to 'Harvard Onco-FM'. Please unify the nomenclature.","section":"Fig. 5a caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's main risk is the un-auditable radiologist comparison in the abstract; this needs to be backed by detailed data. There are also several overclaims relative to the paper's own tables, and at least two numerical inconsistencies. I believe the authors can fix these within the scope of a revision, and the underlying benchmark and model may be valuable. The single-center pre-training corpus and the use of the authors' own segmentation model in the survival benchmark should be transparently discussed; the latter, in particular, deserves a careful leakage check before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things up front. Curia is a real step for radiology foundation models: DINOv2-style self-supervised learning on 200M CT/MRI slices from one hospital's PACS, with base weights released and a 19-task external benchmark. The most striking new result is the CT-to-MRI transfer: Curia drops only about 9 points on organ classification while MedImageInsight and BioMedCLIP drop 35-72 points. That is the kind of empirical claim worth arguing about.\n\nThe benchmark design is mostly solid. Five runs per model, 1,000-sample bootstrap CIs, paired tests, and downstream evaluation on public datasets. The scaling curves, registration, and segmentation comparisons are careful, and the paper gives credit for the DINOv2 recipe. The survival analysis against T-stage is a nice addition, even if the cohort is small.\n\nThe soft spots are real but not all equal. The abstract's 'meets or surpasses radiologists' is load-bearing and not currently auditable: four Paris residents, unreported test-subset sizes per task, and no description of how AUC, balanced accuracy, c-index, and R2 are aggregated into a single curve. The stress-test note is right that this could be an artifact of the averaging scheme. That should be fixed before the claim is taken at face value.\n\nSecond, Section 2.1 says Curia 'outperformed others on all organ-specific metrics' in XCAT CT-to-MR registration. Table D6 contradicts it: DINOv2 Large has a higher mean DSC than Curia-B and higher liver and right-kidney DSC than both Curia variants. The claim is overstated.\n\nThird, two appendix tables (E17, E20) have inverted confidence bounds for Curia-B. That is likely a transcription error, but it makes the appendix less trustworthy than it should be.\n\nFourth, the pre-training corpus is single-center and not released. The authors acknowledge this, and external evaluation mitigates it, but the 'largest corpus' claim cannot be independently audited. There is also mild circularity in the kidney-survival task, where lesions are segmented by the authors' own segmentation model before features are extracted.\n\nNone of this invalidates the core result: a frozen SSL model trained on routine clinical images transfers better than existing medical foundation models on a broad benchmark. The paper deserves serious refereeing, with the radiologist section rewritten, the registration claim corrected, and the appendix tables cleaned.","headline":"Serious radiology foundation model with a strong benchmark and a genuinely interesting cross-modal transfer result, but the 'beats residents' headline needs better support before I would repeat it.","tokens_in":28915,"tokens_out":2945,"would_cite":true,"duration_ms":32257,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Self-supervised pre-training on 200 million unlabeled CT and MRI slices from one hospital yields a radiology foundation model that matches or exceeds resident radiologists and prior foundation models across a 19-task benchmark.","keywords":["foundation model","radiology","self-supervised learning","CT","MRI","cross-modal generalization","few-shot learning","CuriaBench"],"falsifier":"Evaluate Curia-L with the paper's frozen-backbone linear-probe protocol on CT examinations collected from two or three hospitals with different scanner vendors and patient populations, using the same 54-class organ-recognition task. If accuracy falls toward the from-scratch ViT level or the CT-to-MRI transfer drop grows well beyond the reported 9.17 percentage points, the single-center pre-training claim of broad generalizability would be refuted.","tokens_in":27942,"feed_emoji":"🩻","tokens_out":8053,"duration_ms":92469,"temperature":0.7,"pith_summary":"The paper claims that a single self-supervised model trained on unlabeled routine CT and MRI scans from one hospital can serve as a general radiology backbone, replacing the usual one-task-one-model approach. That claim is supported by a new 19-task benchmark on which Curia, without fine-tuning its backbone, matches or beats two existing medical foundation models and, on most tasks, final-year resident radiologists. The authors identify two emergent properties that would matter clinically if confirmed: features transfer across CT and MRI without paired data, and near-maximal performance can be reached from very few labeled examples. They also show that frozen imaging features predict kidney-cancer survival better than the standard local tumor stage. If true, the result implies that large unlabeled hospital archives, not curated annotated datasets, are the main ingredient for broad radiology AI.","feed_headline":"Radiology AI trained on 200M unlabeled scans matches residents","feed_subtitle":"Single self-supervised model tops two medical foundation models across 19 tasks and transfers between CT and MRI.","key_machinery":"DINOv2 self-supervised pre-training on a vision transformer: a teacher–student objective that combines class-token feature alignment, patch-level masked prediction, and auxiliary regularization losses, trained with strong cropping augmentations on 200M unlabeled CT/MRI slices. This is the mechanism that produces the frozen features reused by all downstream heads; the paper argues that the scale and routine-clinical character of the corpus, not task supervision, is what gives the features their generality and their CT-MRI alignment.","core_discovery":"Curia is a pair of vision transformers (86M and 300M parameters) pre-trained with the DINOv2 algorithm on roughly 200 million 2D CT and MRI images drawn from the entire multi-year imaging output of a single hospital. The authors' central claim is that this unlabeled, single-center corpus is sufficient to learn a transferable representation of radiological anatomy and pathology. They evaluate the model by freezing the backbone and training only lightweight prediction heads — linear, cross-attention, regression, Cox survival, and SAM-based segmentation heads — on 19 public tasks spanning organ recognition, oncology, musculoskeletal disease, emergencies, neurodegeneration, and infection. Curia","pith_inferences":["The paper's single-center pretraining makes vendor and protocol shift the main untested risk; a direct follow-up is to pre-train on a second hospital's archive and measure whether the CuriaBench gains shrink or grow.","The reported cross-modal transfer implies a practical bootstrapping recipe the authors do not test: use CT-trained heads to auto-label MRI volumes, then train MRI-specific heads with those pseudo-labels to reduce annotation cost.","Because all benchmarks use public datasets with different preprocessing, holding head architecture and preprocessing fixed across models would isolate how much of the advantage comes from pretraining data versus the head design; the paper uses a grid search of learning rates but does not ablate head choices on the baselines.","The 2D-slice design leaves volumetric context on the table; the scaling curves suggest a native 3D extension of the same recipe, with more data or longer training, is the most direct path to further gains."],"forward_implications":["A single radiology foundation model can replace many task-specific models: Curia reaches leading results on 19 tasks spanning organ recognition, oncology, trauma, infection, musculoskeletal, and neurodegenerative imaging using only lightweight heads on frozen features.","Low-data regimes become practical: near-maximal anatomical accuracy from tens of labeled examples per class, and useful malignancy AUC from 50 or more examples, so annotation-hungry tasks can be bootstrapped quickly.","Cross-modality transfer is usable: a linear head trained on CT organ recognition transfers to MRI with only a 9-point drop, and CT-to-MRI registration improves over existing foundation models, so labels in one modality may help another.","Imaging features carry prognostic signal: kidney-cancer survival prediction from baseline CT features exceeds the local T-stage baseline, suggesting imaging-derived biomarkers can complement current staging.","Frozen-backbone evaluation is sufficient to stage progress: because Curia is compared without fine-tuning the backbone, the benchmark results reflect the quality of the learned representation itself."],"supporting_citations":[{"why":"Supplies the self-supervised teacher–student pre-training objective used to train Curia.","marker":"[6]"},{"why":"Supplies the vision transformer architecture used for Curia-B and Curia-L.","marker":"[14]"},{"why":"General-domain medical embedding model used as the main comparison baseline on every CuriaBench task.","marker":"[15]"},{"why":"Biomedical image–text foundation model used as the second comparison baseline throughout the benchmark.","marker":"[17]"},{"why":"Provides the prompted-segmentation evaluation protocol and the RadSAM baseline for measuring Curia's segmentation quality.","marker":"[21]"},{"why":"Provides the lung-nodule malignancy task split and the specialized cancer-model score Curia is compared against.","marker":"[25]"},{"why":"Supplies the imaging and clinical records for the kidney-cancer survival cohort.","marker":"[28]"},{"why":"Supplies organ masks and labels used to build the CT organ-recognition benchmark.","marker":"[32]"},{"why":"Supplies the lung-nodule data used in the malignancy classification task.","marker":"[37]"}],"fun_headline_variants":["Radiology AI trained on 200M scans rivals residents","Curia: one AI model reads CT and MRI, matches radiologists on 19 tasks","200M unlabeled scans yield radiology AI that beats prior foundation models","Hospital-scale imaging trains radiology AI with emergent cross-modality skills","Curia: open-weights radiology model tops residents and other FMs"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The model's pre-training corpus is the imaging output of a single hospital, and the paper assumes that one center's mix of scanners, protocols, and patients is varied enough to represent radiology as a whole; if it is not, the benchmark results will overstate how well Curia generalizes outside that center.","fun_headline_variants_meta":{"raw":{"variants":["Radiology AI trained on 200M scans rivals residents","Curia: one AI model reads CT and MRI, matches radiologists on 19 tasks","200M unlabeled scans yield radiology AI that beats prior foundation models","Hospital-scale imaging trains radiology AI with emergent cross-modality skills","Curia: open-weights radiology model tops residents and other FMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1084,"prompt_tokens":703,"completion_tokens":381,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":447,"completion_tokens_details":{"reasoning_tokens":283}},"tokens_in":447,"tokens_out":381,"duration_ms":4766,"temperature":1.0,"reasoning_tokens":283,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T23:01:03.239733+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate Curia-L with the paper's frozen-backbone linear-probe protocol on CT examinations collected from two or three hospitals with different scanner vendors and patient populations, using the same 54-class organ-recognition task. If accuracy falls toward the from-scratch ViT level or the CT-to-MRI transfer drop grows well beyond the reported 9.17 percentage points, the single-center pre-training claim of broad generalizability would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the self-supervised teacher–student pre-training objective used to train Curia."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prompted-segmentation evaluation protocol and the RadSAM baseline for measuring Curia's segmentation quality."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the lung-nodule malignancy task split and the specialized cancer-model score Curia is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the imaging and clinical records for the kidney-cancer survival cohort."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the lung-nodule data used in the malignancy classification task."}],"review_version":1}