{"id":"b69b751d-343f-4053-961a-1e9d4fa1a6fd","arxiv_id":"2508.13626","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of 46 public abdominal CT datasets finds substantial case overlap and geographic skew, threatening the real-world applicability of AI models.","lead":"A systematic review of 46 public abdominal CT datasets finds 59.1% case reuse and a 75.3% share from North America and Europe, raising concerns for AI model generalizability. The study also flags domain shift and selection bias in 63% and 57% of larger datasets, suggesting a need for more diverse, multi-institutional data collection.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central percentages depend on dataset-selection and outcome definitions that the abstract does not establish; full methods needed before the 59.1%/75.3% claims can be assessed.","rationale":"The reader's verdict of UNVERDICTED is appropriate because the abstract alone does not permit verification of the review's methodology. My stress-test identifies the same load-bearing weakness: the two headline percentages are aggregates over a dataset list whose completeness and coding definitions are not shown. The concrete test would settle whether the numbers are robust by comparing the reported dataset list against a reproducible search and by auditing the case-reuse definition on a sample. I do not see evidence of internal inconsistency or fraud; the concern is purely about unverifiable sampling and definitions. Therefore the correct action is to leave the verdict unchanged pending full-text review. Agreeing with the reader's weakest assumption is honest; the most important risk is not the statistics themselves but the undisclosed selection and coding choices that could change them.","tokens_in":605,"tokens_out":1738,"duration_ms":20679,"concrete_test":"Retrieve the full-text methods. Re-run the dataset identification using the stated databases, search strings, and inclusion criteria from the stated cutoff date. Generate the candidate pool, apply inclusion/exclusion, and recompute the 46-dataset set and the two reported percentages. Additionally, take a random sample of 10 dataset pairs flagged as 'case reuse' and independently verify overlap using DICOM metadata (patient ID, study UID, series UID). If the corrected geographic percentage moves more than 5 percentage points or the case-reuse rate changes under an alternate but plausible definition, the central claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's headline quantitative claims are aggregates over a set of 46 datasets whose completeness and coding rules are not described in the abstract. The 59.1% case-reuse rate depends entirely on how 'reuse' is defined—patient-level duplicate scans, overlapping image series, or cropped/derived versions of the same case—and the 75.3% geographic skew depends on how 'Western' is assigned (institution country, funding source, or patient origin). Without the search protocol, inclusion/exclusion criteria, and codebook, these numbers are not independently checkable. The reader's weakest assumption—that the 46 datasets are a complete and unbiased representation—is indeed load-bearing; if even a few large non-Western datasets (e.g., recent public releases from Asia or Africa) were missed, the reported skew would shift. This is not an internal inconsistency, but it is a verification gap: the central claims are only as strong as the unreported sampling frame and coding decisions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This systematic review claims to examine 46 publicly available abdominal CT datasets (50,256 studies) and reports two headline quantitative findings: a 59.1% case-reuse rate across datasets and a 75.3% geographic skew toward North America and Europe. For the 19 datasets with at least 100 cases, the authors report high-risk bias categories of domain shift (63%) and selection bias (57%). The abstract also proposes dataset improvement strategies, including multi-institutional collaboration, standardized protocols, and inclusion of diverse populations. Because the full text was not available for review, the assessment is limited to the abstract, which contains no details on search strategy, inclusion criteria, coding definitions, or statistical methods.","tokens_in":827,"tokens_out":1779,"duration_ms":19977,"significance":"If the reported figures are accurate, this review addresses an important and under-reported problem in medical imaging AI: redundancy and geographic bias in public abdominal CT datasets directly threaten model generalizability and equity. The topic is timely, and the quantitative claims—59.1% case reuse and 75.3% Western skew—would be valuable evidence for dataset developers and downstream users. The paper also offers constructive recommendations. However, the significance is currently conditional on the existence of a rigorous, reproducible methodology; the abstract alone provides no evidence of such methodology, so the contribution cannot yet be assessed as a scientific result. I credit the authors for undertaking a systematic review with explicit numerical outcomes, but the lack of verifiable methods in the available text means the claims are currently unsubstantiated.","major_comments":[{"comment":"The abstract reports concrete figures (46 datasets, 50,256 studies, 59.1% case reuse, 75.3% geographic skew) but provides no search protocol, inclusion/exclusion criteria, or method for identifying eligible datasets. The reader's weakest assumption—that the 46 datasets form a complete and unbiased representation—is load-bearing. Without a reproducible search strategy and a PRISMA-style flow diagram, the percentages could shift materially if even a few large non-Western datasets were missed or misclassified. This is a verification gap that must be closed before the central claims can be accepted.","section":"Abstract, first paragraph"},{"comment":"The 59.1% case-reuse rate depends entirely on the operational definition of 'reuse.' Ambiguities include whether reuse means the same patient appearing in multiple datasets, duplicate image series, or derived/cropped versions of the same underlying study. The abstract does not define the unit of analysis (patient, study, series, image) or the threshold for considering two cases 'reused.' Without a codebook and inter-rater reliability assessment, the redundancy claim is not independently checkable.","section":"Abstract, 'case reuse' definition"},{"comment":"The bias assessment is limited to 19 datasets with >=100 cases, but the abstract does not define the assessment framework for 'domain shift' or 'selection bias.' Were these categories scored by a standardized tool (e.g., PROBAST or QUADAS-2), by expert judgment, or by quantitative metrics? The 63% and 57% rates are meaningless without a clear rubric and evidence that the assessments were reproducible. This is a central methodological component and must be specified.","section":"Abstract, second paragraph (bias assessment)"},{"comment":"The 75.3% 'from North America and Europe' claim requires a definition of geographic attribution: institution country, funding source, patient origin, or dataset publication venue? The term 'Western/geographic skew' conflates geography with socioeconomic classification. If attribution is based on institution country, the number will differ from patient-origin-based attribution. The abstract provides no such definition, so the claim's precision is unclear.","section":"Abstract, geographic skew"}],"minor_comments":[{"comment":"The abstract does not mention a protocol registration or adherence to systematic review guidelines such as PRISMA. Including this information would improve transparency. Also, the phrase 'high-risk categories' could be clarified as 'high risk of bias' to avoid ambiguity about what is at risk.","section":"Abstract, general"},{"comment":"The term 'Western/geographic skew' is informal. Suggest using specific geographic regions (e.g., 'North America and Europe') or a clearly defined socioeconomic classification, and avoid conflating 'Western' with a geographic measure.","section":"Abstract, terminology"}],"recommendation":"uncertain","confidential_remarks":"This is an abstract-only review. The major concerns about missing methods could be fully addressed by the full text if it contains a detailed search strategy, inclusion criteria, coding definitions, and bias assessment rubric. I recommend obtaining the full manuscript before making an editorial decision. The abstract's central claims are plausible but unverifiable from the available material. No concerns about citation patterns or novelty were evident from the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The interesting thing here is the attempt to put hard numbers on how redundant and how geographically skewed the public abdominal CT dataset landscape is. The combination of case-reuse detection and bias categorization at this scale is a new compilation, not just another narrative review. If the numbers hold up, the 59.1% case-reuse and 75.3% Western skew are the kind of concrete facts the field needs. A review that maps the landscape and names domain shift and selection bias as the two most prevalent high-risk categories is worth having, and the proposed remedies are standard but sensible.\n\nThe soft spot is not the idea; it is the verifiability of the central claims. The abstract gives no search protocol, no inclusion/exclusion criteria, no codebook for what counts as 'case reuse' or as 'Western,' and no description of how the 46 datasets were selected. The stress-test note is right: if recent large non-Western datasets were missed, the skew percentage shifts. Same for case reuse; different definitions would move the number a lot. That makes the sampling frame load-bearing, and the abstract does not establish it. This is not an internal inconsistency, just a verification gap.\n\nI cannot judge whether the 46 datasets are complete or whether the coding is honest because the full text was not available. The abstract reads like a legitimate systematic review, and the authors seem to be direct about the aim being generalizability. If the full methods are as rigorous as the abstract implies, this deserves a serious referee. But from what I can see, the headline numbers should be treated as unverified until the methods are published.\n\nRecommendation: send to peer review if the full text supplies the missing protocol and codebook; desk-reject only if it does not. For now, treat the percentages as provisional.","headline":"A potentially valuable quantitative review of abdominal CT dataset bias, but the headline percentages are not assessable from the abstract alone.","tokens_in":1284,"tokens_out":3355,"would_cite":false,"duration_ms":28304,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematic review of 46 publicly available abdominal CT datasets finds 59.1% case reuse and 75.3% geographic skew toward North America and Europe, with domain shift and selection bias prevalent in larger datasets, undermining AI genera","keywords":["abdominal CT","public datasets","dataset bias","domain shift","selection bias","medical imaging AI","systematic review","geographic representation"],"falsifier":"Independently re-run the dataset search with a broader set of inclusion criteria and compute the overlap of imaging studies across all publicly available abdominal CT datasets using image-level similarity matching; if the true case-reuse rate falls well below 59.1% and the geographic balance shifts materially away from a Western majority, the central percentages of this review would not hold.","tokens_in":2229,"feed_emoji":"🩻","tokens_out":2061,"duration_ms":21746,"temperature":0.7,"pith_summary":"This paper tries to establish that the publicly available abdominal CT datasets used to train and evaluate AI medical imaging models are not as robust as the volume of data suggests. By reviewing 46 datasets totaling over 50,000 studies, it shows that more than half of the cases are reused across datasets and that three-quarters come from North America and Europe. On the 19 largest datasets, the most common high-risk biases are domain shift and selection bias. If correct, these imbalances mean AI models may fail when deployed to hospitals that differ demographically, geographically, or technologically from the training data sources. The paper argues for coordinated dataset improvement through multi-institutional collaboration, standardized protocols, and deliberate inclusion of diverse populations and imaging equipment.","feed_headline":"Public abdominal CT datasets: 59% redundant, 75% Western","feed_subtitle":"A 46-dataset review links these imbalances to domain shift and selection bias that undermine AI generalizability.","key_machinery":"The central mechanism is a systematic-review protocol applied to the dataset itself: a structured search to identify 46 public abdominal CT datasets, a redundancy check to quantify case reuse across datasets, a geographic-origin classification, and a bias-risk assessment restricted to the 19 datasets with at least 100 cases. These steps produce the headline percentages and the identification of domain shift and selection bias as the dominant high-risk categories.","core_discovery":"The paper reports a systematic review of 46 publicly available abdominal CT datasets totaling 50,256 studies, claiming that these datasets are substantially redundant (59.1% of cases are reused across datasets) and geographically skewed (75.3% originate from North America and Europe). For the 19 datasets with at least 100 cases, the most common high-risk biases are domain shift (63%) and selection bias (57%), both of which threaten the generalizability of AI models trained on them, especially to resource-limited healthcare environments.","pith_inferences":["The same redundancy and geographic-skew patterns likely affect public datasets for other anatomical regions and imaging modalities, though this review only quantifies abdominal CT.","A testable extension would be to measure per-dataset patient-level diversity (age, sex, body habitus, comorbid conditions) and correlate it with the reported bias categories, since the review does not provide such patient-level breakdowns.","The review's emphasis on resource-limited settings suggests a concrete remedy: curating datasets that oversample underrepresented populations and older scanner technology, then benchmarking model degradation across these subgroups.","If the 59.1% case-reuse figure holds, it implies that many published AI performance numbers for abdominal CT are partly memorization of a small pool of unique patients, which would inflate confidence in those systems when deployed to new institutions."],"forward_implications":["AI models trained on existing public abdominal CT datasets may perform poorly in hospitals outside North America and Europe, especially in resource-limited settings.","Benchmark results reported on these datasets likely overstate real-world performance because the same patients' images appear across multiple training and test sets.","Dataset developers should prioritize multi-institutional data collection covering diverse populations, scanner manufacturers, and imaging protocols rather than adding more cases from the same sources.","Reporting standards for public medical datasets should include explicit statements about patient overlap and geographic composition so downstream users can interpret model evaluations correctly."],"supporting_citations":[],"fun_headline_variants":["Abdominal CT datasets: 59% redundant, 75% from West, AI risk","CT dataset review: 59% case reuse, 75% Western skew","Public CT datasets: 59% redundant, 75% Western, bias high","Abdominal CT AI: 59% data reuse, 75% Western origin","CT sets review: 59% overlap, 75% Western, domain shift risk"],"cache_read_input_tokens":1152,"weakest_assumption_plain":"The set of 46 datasets this review selected is a complete and unbiased representation of all publicly available abdominal CT datasets; if the search missed or misclassified any substantial number of datasets, the reported percentages would change.","fun_headline_variants_meta":{"raw":{"variants":["Abdominal CT datasets: 59% redundant, 75% from West, AI risk","CT dataset review: 59% case reuse, 75% Western skew","Public CT datasets: 59% redundant, 75% Western, bias high","Abdominal CT AI: 59% data reuse, 75% Western origin","CT sets review: 59% overlap, 75% Western, domain shift risk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000849,"raw_usage":{"total_tokens":3489,"prompt_tokens":661,"completion_tokens":2828,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":405,"completion_tokens_details":{"reasoning_tokens":2720}},"tokens_in":405,"tokens_out":2828,"duration_ms":19684,"temperature":1.0,"reasoning_tokens":2720,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:56:06.606315+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-run the dataset search with a broader set of inclusion criteria and compute the overlap of imaging studies across all publicly available abdominal CT datasets using image-level similarity matching; if the true case-reuse rate falls well below 59.1% and the geographic balance shifts materially away from a Western majority, the central percentages of this review would not hold.","supporting_citations":[],"review_version":1}