{"id":"e93c4e32-f783-4ab1-b647-e7787f604baf","arxiv_id":"2412.12629","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A proprietary AI model for abdomen-pelvis CT detects 21 conditions with an average external AUC of 0.923, according to a retrospective multicenter study.","lead":"This paper reports a large external validation of a2z-1, a commercial AI model that screens abdomen-pelvis CT scans for 21 time-sensitive conditions. The model achieved an average AUC of 0.923 on two independent health-system datasets, and the authors argue it can act as a second reader to catch findings that radiologists miss.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Report-derived labels are the pivotal assumption; without per-condition labeler error rates or independent adjudication, the reported average AUC of 0.923 is not yet tied to true diagnostic performance.","rationale":"The reader's conditional verdict is appropriate: the evaluation is externally designed but not externally verifiable. I agree with the reader's weakest assumption on ground-truth labels. The reported labeler precision (88.7%) and the paper's own false-positive analysis indicate the reference standard is imperfect and its errors correlate with model scores. Because the bias direction is unknown, the average AUC is not established. This does not justify rejection—the model may be as good as claimed—but it justifies conditional acceptance pending independent adjudication or per-condition labeler error reporting. The requested check directly targets the weakest load-bearing premise.","tokens_in":12172,"tokens_out":5894,"duration_ms":58571,"concrete_test":"Stratified random sample from the 5,444 external studies: draw at least 100 label-positive and 100 label-negative cases for each of the 21 conditions (or all positives if fewer than 100). Have two independent radiologists, blinded to model output and original labels, adjudicate presence/absence of each condition using the CT images plus full clinical context. Recompute per-condition AUC against the adjudicated labels and compare with the reported values. If the mean AUC shifts by more than 0.02, or any critical condition (small bowel obstruction, acute pancreatitis) changes by more than 0.02, the central generalizability claim is not robust to reference-label error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim—external AUC 0.923 across 21 conditions—stands or falls on the 5,444 external reference labels extracted automatically from radiology reports. The only validation offered is the authors' own comparison of the NLP labeler to report text on an unspecified sample; Section 2.1 reports aggregate accuracy 99.4%, F1 92.6%, precision 88.7%, recall 99.05%, but gives no sample size or per-condition breakdown. Aggregate precision of 88.7% means roughly one in nine positive labels is a false positive overall, and for low-prevalence conditions the positive-label noise can be larger relative to true positives. More importantly, Section 2.5 shows that label errors are concentrated among high-confidence model predictions: several 'false positives' were actually findings present in the report but missed by the labeler, and some labels marked positive were absent from the report. This demonstrates the label noise is not independent of model output, so the direction and magnitude of bias in the per-condition AUCs is unknown. Without independent, blinded radiologist adjudication of a stratified sample, the 0.923 average cannot be distinguished from an artifact of report-to-label misclassification. The QA/missed-finding claim in Section 2.5 is also based on selected, unblinded case review rather than a systematic reference standard, so it cannot rescue the main metric.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports an external validation of a2z-1, a deep-learning model for detecting 21 abdominal and pelvic findings on CT. The authors claim an average AUC of 0.923 across 21 conditions on 5,444 external studies from two health systems, with consistent performance across sites, patient demographics, and imaging protocols. They also report that high-confidence model predictions identified findings missed by original radiologist reports, suggesting a quality-assurance role. The study is retrospective and uses automated extraction of ground-truth labels from radiology reports, with a small manual validation subset reported as 99.4% accurate, F1 92.6%, precision 88.7%, recall 99.05%. The paper includes subgroup analyses by age, sex, scan area, contrast type, slice thickness, and scanner manufacturer, and describes a three-tiered confidence categorization ('likely', 'possible', 'unlikely') with thresholds tuned on internal validation data.","tokens_in":12402,"tokens_out":4153,"duration_ms":37830,"significance":"If the reported performance holds, this would be a substantial contribution: a single AI system demonstrating high and consistent discrimination across 21 time-sensitive abdominal conditions on truly external data would be of considerable clinical interest, and the proposed confidence-tier workflow is a practical idea. The paper's strengths are the large external cohort, the breadth of conditions, and the explicit attempt to analyze label noise and missed findings. However, the central claim rests on the accuracy of automatically extracted report labels, and the manuscript's own analyses in Sections 2.5 and 2.6 show that label errors are concentrated among high-confidence model predictions, meaning the label noise is not independent of model output. Without per-condition labeler validation or independent blinded adjudication, the reported AUCs are not yet tied to true diagnostic performance. The subgroup and missed-findings claims also lack statistical support. The study is therefore promising but currently under-supported; it could become a strong validation report after targeted revisions.","major_comments":[{"comment":"The ground-truth labels for the external validation set are derived from automated extraction from radiology reports, and the only validation reported is the authors' own comparison of the labeler to report text on an unspecified sample, giving aggregate accuracy 99.4%, F1 92.6%, precision 88.7%, recall 99.05%. The sample size and per-condition breakdown are not reported, and the aggregate precision of 88.7% implies roughly one in nine positive labels is a false positive overall. More importantly, Section 2.5 documents cases where the model's high-confidence 'false positives' were actually findings present in the report but missed by the labeler, and Section 2.6 documents labels marked positive when the report did not mention the finding. This demonstrates that label noise is not independent of model output, so the direction and magnitude of bias in the per-condition AUCs is unknown. Without per-condition labeler error rates or an independent, blinded radiologist adjudication of a stratified sample, the reported average AUC of 0.923 cannot be distinguished from a report-to-label misclassification artifact. This is the load-bearing assumption for the paper's central claim.","section":"Section 2.1, Evaluation Details"},{"comment":"All AUC values are reported as point estimates without confidence intervals, and per-condition sample sizes are not provided either in the text or in Figure 5. For low-prevalence conditions such as aortic dissection or retroperitoneal hemorrhage, the uncertainty in AUC can be substantial, and the claim of 'consistent performance across sites' (e.g., small bowel obstruction AUCs of 0.979, 0.981, 0.947) is not supported by any statistical comparison or interval estimate. The absence of confidence intervals also affects the comparison between internal and external validation and the statements about conditions with 'improved' or 'decreased' performance across sites. Reporting CIs (e.g., DeLong or bootstrap) and per-condition sample sizes is essential for a validation study.","section":"Section 2.2 and Figure 5"},{"comment":"The subgroup analysis claims consistent performance across age groups, sex, scan areas, contrast types, slice thicknesses, and scanner manufacturers, but no confidence intervals or hypothesis tests are provided for any subgroup AUC. For example, the statement that 'multi-regional scans show a dip to around 0.87' compared with 0.92 for abdominal scans may reflect noise or confounding by disease mix, and without interval estimates or tests, the 'consistent performance' claim is not quantitatively supported. The paper should provide CIs for subgroup AUCs and, where appropriate, tests of interaction or equivalence, and should adjust for case-mix differences across subgroups.","section":"Section 2.3, Table 1 and Figure 6"},{"comment":"The analysis of high-confidence model predictions (false positives and false negatives) is based on selected, unblinded manual review by the model developers themselves, with no systematic sampling protocol, no prespecified adjudication criteria, no independent radiologist readers, and no denominator describing how many cases were reviewed. The manuscript reports anecdotes such as a missed retroperitoneal hemorrhage and a subtle pancreatitis, and states that 'labeling errors were less than 1%' without defining the denominator. The abstract's claim that a2z-1 'identified overlooked findings' is therefore not quantitatively supported. To support the quality-assurance claim, the authors need a structured reader study or adjudication process with a defined sample, blinded independent readers, and prespecified definitions of 'missed finding.'","section":"Sections 2.5 and 2.6"}],"minor_comments":[{"comment":"The abstract reports an average AUC of 0.931 for 'large-scale retrospective analysis' and 0.923 for external validation; the main text states 0.923 but does not clearly define what the 0.931 corresponds to (presumably internal validation). Please clarify the relationship between these numbers and define 'average AUC' (unweighted across conditions? weighted by sample size?).","section":"Abstract and Section 2.1"},{"comment":"Figure references are inaccurate: Figure 1 and Figure 3 are case examples, not performance figures, yet the text cites 'acute pancreatitis (Figure 1)' and 'peritoneal free air (Figure 3)' when describing AUC results. Please renumber or re-reference the correct figures.","section":"Section 2.1 and figure captions"},{"comment":"There is a typo: 'review an CTs' should be 'review a CT'. Figure 2's caption also contains 'abdominen-pelvis'.","section":"Section 1, Introduction"},{"comment":"The confidence-category thresholds were tuned on internal validation data, but the manuscript does not report how the target precisions of 0.8 and 0.4 were verified or their stability on the external sites. Please provide the achieved precision and recall for each category on both internal and external data.","section":"Section 2.4"},{"comment":"The sentence 'the proportion of labeling errors were less than 1%' is ambiguous; state explicitly what the numerator and denominator are (e.g., proportion of all external studies, or proportion of high-confidence false positives), and report the proportion separately for each condition.","section":"Section 2.5"},{"comment":"The manuscript provides no technical description of the a2z-1 model architecture, training data, or inference pipeline, and no statement of code or data availability. For a validation study, at least a summary of the model's development dataset and characteristics (e.g., scanner types, annotation process) is needed to assess generalizability.","section":"General"},{"comment":"The paper does not state whether institutional review board approval or a waiver was obtained for the retrospective use of the external datasets, nor does it mention patient consent or data de-identification. This information is required for publication in a medical imaging journal.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is authored by employees of a2z Radiology AI, and the external validation is performed by the model's developers, which creates a conflict-of-interest concern even though the data are external to training. The absence of independent adjudication of ground-truth labels is the most serious scientific issue, but it is fixable with additional validation. The paper also lacks confidence intervals, per-condition sample sizes, and a structured reader study, all of which are standard for a validation paper of this scope. I would encourage the editor to require that any revision include independent (or at least blinded) adjudication of a stratified sample of labels, and that the authors either provide per-condition labeler performance or re-estimate AUCs with label-noise sensitivity analysis. The paper is not yet suitable for acceptance, but the underlying study design is sound enough for a major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, this is a serious attempt at external validation, not a toy: 5,444 studies from two health systems, 21 conditions, and the authors report per-condition AUCs and subgroup breakdowns. Second, the entire edifice rests on labels extracted automatically from radiology reports, and the paper does not give you enough to trust those labels at the level required by the headline claim.\n\nWhat is new: the breadth. Most prior work is single-disease or internal-only. Showing a single model holding roughly 0.92 average AUC across two external sites for small bowel obstruction, pancreatitis, and similar conditions is a useful data point for the field, even without architectural novelty. The paper also deserves credit for being honest in sections 2.5 and 2.6: it shows actual cases where the labeler was wrong, acknowledges similar-pathology misclassifications, and admits that false negatives skew toward subtle cases. That candor is real.\n\nThe soft spots are concentrated exactly where the stress-test puts them. The labeler validation is described as a sample of reports covering every condition, with aggregate accuracy 99.4%, precision 88.7%, recall 99.05%, but no sample size, no per-condition breakdown, and no independent adjudication. Aggregate precision of 88.7% means roughly one in nine positive labels is a false positive on average; for lower-prevalence conditions it can be worse. The paper's own examples show label errors are not random: several 'false positives' were findings the labeler missed, so errors are correlated with model output. That makes the direction of bias in the per-condition AUCs unknown. The confidence-category thresholds were tuned on internal validation, which is not itself a problem, but the manual reviews in the missed-finding analysis were done by the model developers, unblinded. No confidence intervals appear anywhere, and the subgroup comparisons have no statistical tests. These are real gaps, not manufactured ones.\n\nThat said, the central validation claim is not self-contradictory, and the authors disclose the main weakness rather than burying it. The paper just does not yet convert a plausible result into an established one.\n\nWho gets value: any researcher tracking clinical validation of abdominal CT AI, especially people designing label-extraction pipelines, because this is a live example of how much weight a report labeler can be asked to carry. It deserves peer review, but a referee should insist on per-condition labeler error rates, confidence intervals, and an independent blinded read of a stratified sample. Without those, the headline AUC is not yet tied to true diagnostic performance.","headline":"Useful external-validation data on 21 conditions, but the report-derived ground truth is not yet proven accurate enough to trust the headline AUC.","tokens_in":12934,"tokens_out":2161,"would_cite":false,"duration_ms":19688,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that a2z-1, a single AI model, maintains high diagnostic accuracy across 21 abdomen-pelvis CT conditions in external validation, and that its high-confidence predictions can uncover clinically significant findings…","keywords":["abdomen-pelvis CT","multi-disease detection","external validation","artificial intelligence","radiology","AUC","quality assurance","deep learning"],"falsifier":"Have two board-certified radiologists independently re-read a stratified sample of the external validation studies covering all 21 conditions, resolve discrepancies by consensus, and recompute condition-specific AUCs against that reference standard; if the average AUC drops materially below 0.923, the automated labels used for ground truth were inflating performance.","tokens_in":11985,"feed_emoji":"🩻","tokens_out":6391,"duration_ms":49656,"temperature":0.7,"pith_summary":"This paper reports that a2z-1, a deep-learning model for abdomen-pelvis CT, keeps its diagnostic accuracy when tested on data from two health systems that did not contribute to its training. The average area under the ROC curve across 21 time-sensitive findings is 0.923, with particularly high scores for small bowel obstruction and acute pancreatitis. The authors also show that accuracy is stable across patient age, sex, scanner manufacturer, contrast protocol, and slice thickness. Finally, a review of high-confidence model outputs found cases where the model flagged clinically significant findings that the original radiology report missed, which suggests a role as a quality-assurance second reader. If these results hold, a single broad AI system could support radiologists across different institutions without per-site retraining.","feed_headline":"A single AI model detects 21 abdominal CT findings across hospitals","feed_subtitle":"AUC 0.923 on outside data; model flagged findings missed in original reports.","key_machinery":"The load-bearing object is a2z-1 itself, a deep neural network that takes abdomen-pelvis CT volumes as input and outputs probabilities for 21 clinically actionable findings. The evaluation machinery is the external-validation protocol: internal validation on 9223 studies selected the model, and two held-out health systems provided 5444 studies for generalization testing. Ground truth was generated by automated extraction of labels from signed radiologist reports, a process the authors validated on a sample at 99.4% accuracy. The confidence-based three-tier categorization, with thresholds set on internal data to target 0.8 precision for 'likely' and 0.4 precision for 'possible,' is what allows the model to be positioned as a workflow triage and quality-assurance tool.","core_discovery":"The central claim is that a2z-1 generalizes: on an external validation set of 5444 studies from 4907 patients at two independent health systems, it achieves an average AUC of 0.923 across 21 predefined actionable conditions. Per-condition AUCs range from 0.855 for colitis to 0.970 for unruptured aortic aneurysm, with small bowel obstruction at 0.958 and acute pancreatitis at 0.961. The paper further claims that the model's performance is consistent across demographic subgroups and imaging protocols, and that a confidence-based categorization into 'likely,' 'possible,' and 'unlikely' lets the model flag high-confidence findings for immediate attention. Manual review of discordant cases indicates that some apparent false positives were actually correct detections of findings missed or understated in the original radiologist reports, including subtle pancreatitis and cholecystitis. In the authors' telling, this makes a2z-1 the first AI model to demonstrate broad, externally validated performance across the abdomen-pelvis CT spectrum.","pith_inferences":["A natural next step, not tested here, is a prospective study where radiologists see a2z-1 outputs at read time; the claim that it catches missed findings would only become actionable if it changes management or outcomes.","The model's tendency to flag conditions in the same pathological spectrum as the report diagnosis (for example, colitis where the report says diverticulitis) suggests that exact-match AUC understates its clinically useful performance; evaluation metrics that reward anatomically or pathophysiologically related predictions would better reflect its value.","Because the authors are also the model's developers and the label validation was self-performed, independent replication by a third party on independently curated datasets would substantially strengthen confidence in the reported 0.923 average."],"forward_implications":["A single AI system could be deployed across hospitals and scanner manufacturers without retraining, based on the consistent external AUCs.","The model could serve as a second reader that resurfaces findings missed or hedged in the original report, based on the manual review cases.","The three-tier confidence system enables different workflow uses: immediate alerts for 'likely,' secondary review for 'possible,' and low attention for 'unlikely.'","Subgroup consistency implies the model would not systematically penalize particular age groups, sexes, or contrast protocols, supporting equitable deployment."],"supporting_citations":[{"why":"Establishes that up to 14% of abdomen-pelvis CT reports undergo clinically important changes on second review, motivating the quality-assurance use case.","marker":"Lauritzen et al., 2016"},{"why":"Documents inter-reader and intra-reader disagreement rates of 26% and 32%, supporting the need for a consistent second reader.","marker":"Abujudeh et al., 2010"},{"why":"Attributes 42% of radiological errors to failures of perception, the gap the model's missed-finding analysis targets.","marker":"Kim and Mansfield, 2014"},{"why":"Provides the comparison generalist foundation model for abdominal CT, which the paper argues has limited external validation and non-comparable metrics.","marker":"Blankemeier et al., 2024"},{"why":"Shows prior evidence that AI assistance reduces missed findings on emergency CT, serving as the precedent for a2z-1's second-reader claims.","marker":"Rueckel et al., 2021"}],"fun_headline_variants":["AI model detects 21 abdominal CT findings across hospitals","External validation: AI detects 21 CT findings with AUC 0.923","a2z-1 AI: 21 CT findings, AUC 0.923 on external data","AI flags missed abdominal findings in CT scans","External validation: a2z-1 detects 21 abdomen-pelvis CT findings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that the labels automatically extracted from radiology reports are accurate enough to serve as ground truth, a premise the paper supports with a 99.4% accuracy check that lacks a reported sample size or per-condition breakdown.","fun_headline_variants_meta":{"raw":{"variants":["AI model detects 21 abdominal CT findings across hospitals","External validation: AI detects 21 CT findings with AUC 0.923","a2z-1 AI: 21 CT findings, AUC 0.923 on external data","AI flags missed abdominal findings in CT scans","External validation: a2z-1 detects 21 abdomen-pelvis CT findings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000547,"raw_usage":{"total_tokens":2600,"prompt_tokens":916,"completion_tokens":1684,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":1588}},"tokens_in":532,"tokens_out":1684,"duration_ms":11110,"temperature":1.0,"reasoning_tokens":1588,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:52:07.547736+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two board-certified radiologists independently re-read a stratified sample of the external validation studies covering all 21 conditions, resolve discrepancies by consensus, and recompute condition-specific AUCs against that reference standard; if the average AUC drops materially below 0.923, the automated labels used for ground truth were inflating performance.","supporting_citations":[],"review_version":1}