{"id":"888fe71d-daec-405c-9330-af6d5d8b2938","arxiv_id":"2507.22939","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"PARROT is a new open corpus of 2,658 fictional expert-written radiology reports in 13 languages, plus a study showing humans distinguish them from AI text at only 53.9% accuracy.","lead":"PARROT is an open dataset of 2,658 fictional radiology reports, written by 76 radiologists in 21 countries and 13 languages, with ICD-10 labels and English translations. It matters because it gives medical AI researchers a privacy-safe multilingual benchmark where English-only datasets like MIMIC-CXR cannot go.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Representativeness of fictional PARROT reports is unvalidated: the human/AI study shows indistinguishability from GPT-o1, not from real clinical reports, so the central utility claim lacks direct support.","rationale":"The reader's weakest assumption is that volunteer radiologists writing fictional cases produce text representative of real reporting practice; my analysis converges on the same point. This is the most load-bearing concern because the dataset's stated purpose is to serve as a benchmark for clinical NLP across languages and regions. If the text distribution is systematically different from real reports, models evaluated on PARROT may not transfer to practice, undermining the central claim regardless of dataset size. The human/AI discrimination study is often cited as evidence of authenticity, but it only tests discriminability from AI-generated text, not from real clinical reports; both classes could be equally unrepresentative. The paper's own Limitations section concedes the absence of real imaging/outcomes and possible selection bias, but it does not directly test textual representativeness. This is an addressable gap, not a fatal flaw: a distributional comparison against real reports could validate or refute the assumption. The secondary concern about the 'largest' superlative is also real but less central; even if another open multilingual corpus were larger, PARROT would still be a useful resource, whereas unrepresentative text would undercut its main value proposition. Therefore my read does not change the reader's verdict of conditional acceptance, which already flags this assumption and asks for additional validation.","tokens_in":10996,"tokens_out":7169,"duration_ms":74755,"concrete_test":"Run a distributional validation: obtain a sample of real, de-identified radiology reports from a subset of PARROT contributing institutions (with ethics approval), matched by language, modality, and body region. Train a per-language classifier (e.g., logistic regression on TF-IDF features or a fine-tuned language model) to distinguish PARROT from real reports using cross-validation. If held-out AUROC substantially exceeds 0.8 in most languages, the fictional reports are systematically distinguishable and the representativeness assumption fails; if AUROC is near 0.5, the concern is mitigated. As a lighter check, compare simple linguistic features (report length, section structure, negation/hedging rates, abbreviation frequency) between the two sets with a two-sample test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Conclusion, §4.2) is that PARROT 'enables development and validation of NLP applications across linguistic, geographic, and clinical boundaries.' That claim requires the 2,658 fictional reports to be representative of real radiology reporting in each language/region. The evidence offered is the human/AI discrimination study (§3.3): participants achieved only 53.9% accuracy distinguishing PARROT reports from GPT-o1 output. But this establishes only that fictional expert reports and GPT-o1 text are similarly plausible; it does not establish that either resembles actual clinical reports. Both could be systematically different from real practice in overlapping ways (e.g., more structured, complete, and free of the abbreviations and inconsistencies common in busy clinical workflows). The authors explicitly instructed contributors (§2.2) to produce 'plausible but non-specific clinical scenarios' with 'typical incidental findings' and 'realistic frequencies,' which may create a dataset that is more canonical than real reports. The Limitations section (§4.1) acknowledges no link to real imaging/outcomes and recruitment/selection bias, but does not test whether the text distribution itself matches real-world reporting. Without a comparison against real de-identified reports from the same institutions/languages, the utility benchmark claim is unsupported. A secondary issue is the unsupported 'largest' superlative: no systematic comparison with other open multilingual corpora is provided, but this is less central to usefulness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PARROT, a dataset of 2,658 fictional radiology reports written by 76 radiologists from 21 countries in 13 languages, with metadata including imaging modality, anatomical region, clinical context, English translations for non-English reports, and ICD-10 code assignments. The authors also report a human-versus-AI discrimination study in which 154 participants (radiologists, other healthcare professionals, and non-healthcare professionals) judged whether individual reports were human-authored or generated by GPT-o1, achieving an overall accuracy of 53.9% with radiologists performing better (56.9%). The paper claims that PARROT is the largest openly available multilingual radiology-report dataset and that it enables development and validation of NLP applications across linguistic, geographic, and clinical boundaries without privacy constraints.","tokens_in":11221,"tokens_out":9692,"duration_ms":87894,"significance":"If the claims hold, PARROT would be a useful resource for multilingual NLP in radiology, addressing the English-centric bias of widely used datasets such as MIMIC-CXR. The fictional nature of the reports and the decision to release the data openly (under CC-BY-NC-SA) are practical strengths that avoid privacy barriers, and the human/AI discrimination study addresses a timely question about the detectability of AI-generated medical text. However, the manuscript overreaches in several respects: the 'largest' superlative is not backed by a systematic comparison with other open multilingual radiology-corpus resources, the descriptive statistics contain internal inconsistencies, the discrimination study is under-specified, and the central assumption that fictional reports are representative of real clinical reporting is not validated. The resource itself has potential, but the current presentation needs revision before the claims can be accepted.","major_comments":[{"comment":"The claim that PARROT 'represents the largest openly available multilingual radiology-report dataset' is not supported by any systematic comparison with existing resources. Section 4 contrasts only MIMIC-IV and MIMIC-CXR, which are English-only, and does not enumerate or compare counts with other open multilingual radiology-report corpora. Without a comparative table listing existing datasets, their languages, report counts, and licensing, the superlative is unsubstantiated. Please either provide that comparison or temper the claim (e.g., 'largest to our knowledge') with supporting evidence.","section":"Abstract and §4.2"},{"comment":"Table 1 lists body-area counts that sum to 3,397, whereas the dataset contains 2,658 reports. The text and abstract report chest (19.9%), abdomen (18.6%), head (17.3%), and pelvis (14.1%) as if these were proportions of reports, but these percentages are consistent with 677/3397, 631/3397, 588/3397, and 480/3397, respectively, indicating the percentages are calculated over body-area labels rather than reports. If a report can be assigned multiple body areas, this must be stated explicitly and report-level versus label-level statistics should be presented separately; if each report has exactly one body area, the table sums are erroneous. As written, this is an internal inconsistency in the core descriptive results.","section":"§3.2 and Table 1"},{"comment":"The sentence 'Nations from the Global South—Argentina (71; 2.7 %), China (100; 3.8 %), and Mexico (75; 2.8 %)—together represent roughly 19 % of the collection' is arithmetically incorrect: the three named countries contribute 246 reports, which is 9.3% of 2,658, not 19%. If the intended 19% figure includes additional countries beyond the three named, those countries should be listed; otherwise the percentage should be corrected. Given that geographic diversity is a central selling point, this overstatement must be fixed.","section":"§3.1"},{"comment":"The human-versus-AI differentiation study is not described in sufficient detail to be reproducible or interpretable. Section 2.4 does not specify the prompt, model version, or generation settings used for GPT-o1; the number of AI-generated reports; how the five human reports and five AI reports were selected and matched for each participant (e.g., by language, modality, and anatomical region); or whether participants saw reports with metadata and translations. Section 2.5 states that differences were assessed with a χ2 test, but Section 3.3 reports a logistic regression with participant-level estimates; the unit of analysis (participant vs. report), the handling of repeated observations (each participant judged ten reports), and the clustering of observations are not described. These gaps make the headline 53.9% accuracy and the radiologist-superiority finding difficult to evaluate.","section":"§2.4, §2.5, §3.3"},{"comment":"The central utility claim that PARROT 'serves as a valuable benchmark for creating and testing AI applications capable of functioning across diverse healthcare systems and languages' (§4.2) presupposes that the fictional reports are representative of real clinical reporting in each language and region. The paper does not provide direct evidence for this. The discrimination study (§3.3) shows only that PARROT reports are hard to distinguish from GPT-o1 output; it does not show that either text resembles real clinical reports. The instructions in §2.2 to create 'plausible but non-specific clinical scenarios' with 'typical incidental findings' and 'realistic frequencies' could systematically exclude the abbreviations, telegraphic phrasing, and inconsistencies common in real-world reports. The Limitations section (§4.1) acknowledges the lack of linkage to real imaging and outcomes and selection bias, but it does not test whether the corpus's textual distribution matches real reports. Please either add a validation study comparing lexical and structural statistics of PARROT reports against de-identified real reports from the same institutions, or explicitly state that the dataset is not a proxy for real clinical text and adjust the claims accordingly.","section":"§2.2 and §4.1"}],"minor_comments":[{"comment":"The discrimination study used reports in English, German, Italian, French, Greek, and Polish, while the dataset contains 13 languages; the paper should clarify why these six were chosen and whether the reported 53.9% accuracy is averaged across all six languages or weighted by participant counts.","section":"§2.4"},{"comment":"The row 'NA 0.0 U09' is confusing because U09 (post-COVID-19 condition) is a valid ICD-10 code; if it is classified as 'NA' this needs explanation, and the chapter percentages should sum to 100% with no unexplained 'NA' category.","section":"Table 3"},{"comment":"Both Figure 2 and Figure 3 appear to include a panel C titled 'Effect sizes from logistic regression model'; please check whether this is a duplication and ensure each figure presents a unique set of panels.","section":"Figures 2 and 3"},{"comment":"The description of ICD-10 coding says codes were assigned 'by the contributor or by B. Le Guellec with assistance from the o3-mini-high language model (OpenAI)'; for reproducibility, please specify the model version and date and describe how disagreements were resolved beyond 'dialogue with contributors.'","section":"§2.2"},{"comment":"The statement that 'evaluation metrics developed for one region often fail to generalize across reports from different countries' cites a paper about metric inconsistencies; please ensure the citation supports the claim about regional generalization.","section":"§1"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself appears useful and the open-release approach is commendable, but the manuscript's central claims are currently overstated or not fully supported. The internal inconsistencies in descriptive statistics and the under-specified discrimination study are fixable in revision. I did not find indications of questionable research practices; the main concerns are evidentiary and presentational. The paper would be suitable for publication after the authors provide a systematic comparison for the 'largest' claim, correct the arithmetic and reporting of body-area statistics, substantially expand the Methods for the human/AI study, and either validate or temper the representativeness claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nPARROT is a real contribution: a curated, open-licensed corpus of 2,658 fictional radiology reports written by 76 practicing radiologists in 13 languages from 21 countries, with metadata, ICD-10 codes, and English translations for non-English reports. That genuinely fills a gap—existing open radiology text resources like MIMIC-CXR are English-only, and they don't allow the kind of unrestricted sharing PARROT enables. The authors also ran a human-vs-AI discrimination test showing that 154 participants, including radiologists, could barely tell PARROT reports from GPT-o1 output (53.9% overall; radiologists 56.9%). That's a nice sanity check and it gives the field a data point.\n\nThe soft spots are real but mostly fixable. The 'largest openly available multilingual radiology-report dataset' claim is asserted without a systematic comparison. I don't know of a bigger open multilingual corpus either, but the paper should show its work—list candidate corpora and counts. More substantively, the paper's utility claim rests on the idea that fictional reports written by volunteers capture real-world reporting practice. The authors instruct contributors to write 'plausible but non-specific' cases with 'typical incidental findings,' which likely produces cleaner, more canonical text than busy clinical workflows generate. The discrimination study only shows that PARROT text is as plausible as GPT-o1 text; it doesn't compare either to real clinical reports. The Limitations section acknowledges no link to real imaging and recruitment bias, but it doesn't test distributional similarity to authentic reports. That's a gap, not a fatal flaw, because the resource can still be useful for NLP benchmarking even if it's not perfectly representative—but the claim as stated goes further than the evidence.\n\nThe human/AI study also needs procedural detail: how AI reports were generated and sampled, how many per language, how the 10 reports per participant were selected, and the exact wording of the task. Without those, the 53.9% is hard to interpret.\n\nWho is this for? Anyone building or testing multilingual medical NLP tools, especially for low-resource languages. It's a dataset paper, not a mechanism paper, so the value is in the resource plus the descriptive statistics. I'd encourage you to look at the repo (I couldn't inspect it from here) and see whether the translations and metadata are as clean as the paper suggests.\n\nRecommendation: send it to peer review. The core resource is new and likely useful, and the concerns are addressable through revision.","headline":"A genuinely useful open multilingual radiology text resource, but the 'largest' and representativeness claims are softer than the paper lets on.","tokens_in":12295,"tokens_out":2064,"would_cite":true,"duration_ms":19341,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PARROT assembles 2,658 fictional radiology reports in 13 languages and shows that even radiologists struggle to tell them from AI-generated reports.","keywords":["radiology reports","multilingual dataset","natural language processing","large language models","fictional medical data","human-AI differentiation","ICD-10 coding","open benchmark"],"falsifier":"Compare distributions of report structure, terminology, and length between PARROT reports and a matched sample of real reports from the same institutions and languages; if real and fictional reports differ systematically, or if an NLP model trained on PARROT fails to match a model trained on real reports on a held-out real-language report understanding task, the central utility claim is contradicted.","tokens_in":10810,"feed_emoji":"🩻","tokens_out":6131,"duration_ms":57582,"temperature":0.7,"pith_summary":"PARROT is an attempt to break the English-only bottleneck in radiology natural-language processing by building the largest openly available multilingual collection of radiology reports: 2,658 fictional but expert-written reports from 76 radiologists in 21 countries and 13 languages, released without patient-data restrictions. The paper argues that because the reports are written by working radiologists following their usual local formats and terminology, the dataset captures authentic cross-country variation in reporting style that synthetic English corpora and privacy-limited real records cannot. To show the reports are not trivially replaceable by AI output, the authors ran a discrimination study in which 154 participants—radiologists, other healthcare staff, and laypeople—could distinguish human PARROT reports from frontier-language-model-generated ones only 53.9% of the time overall, with radiologists at 56.9%. If the central claim holds, PARROT gives NLP researchers a legal, reusable benchmark for training and evaluating tools that must work across languages and healthcare systems.","feed_headline":"Fictional radiology reports in 13 languages are now open","feed_subtitle":"A 13-language, 21-country test bed for medical NLP—and humans classify human vs. AI reports at only 53.9% accuracy.","key_machinery":"The carrying object is the PARROT corpus itself: a version-controlled, JSONL-formatted collection of fictional radiology reports contributed by 76 radiologists, each writing in the format, terminology, and structure they would use in their own clinical practice, with metadata and ICD-10 codes attached. The human-versus-AI differentiation study acts as the proof-of-concept: using a frontier large language model to generate comparison reports and asking 154 participants to classify them tests whether PARROT's expert-written text is distinguishable from AI text, supporting the paper's positioning of the dataset as a middle ground between synthetic corpora and privacy-restricted real records.","core_discovery":"The paper's central claim is that fictional reports, when authored by practising radiologists following their own local reporting conventions, can serve as a proxy for real clinical text for natural-language-processing benchmarking. The dataset itself is the discovery: 2,658 reports with metadata on modality, anatomy, clinical context, ICD-10 codes, and English translations for non-English entries, spanning CT, MRI, radiography, and ultrasound. The validation experiment shows that humans, including radiologists, have difficulty telling these reports from frontier-LLM output; this is presented as evidence that the PARROT texts carry authentic stylistic and reasoning patterns that pure synthetic generation misses, while remaining free of privacy constraints.","pith_inferences":["If PARROT reports are representative, the differentiation result hints at a practical detector: radiologists outperform others by about seven points, so extracting the features they use could yield a machine classifier, although the paper does not attempt this.","A stronger test of the dataset's premise would be to collect real de-identified reports from the same contributing radiologists and measure the style gap; the present release does not include that comparison and could not detect it.","The dataset's open license and English translations make it a plausible starting point for cross-lingual training, but users should validate transfer to their own local clinical text before deployment."],"forward_implications":["Researchers can develop and compare radiology NLP tools on one legal, open dataset spanning 13 languages, without negotiating data-sharing agreements or de-identification protocols.","Workflows such as report summarization, structured extraction, and ICD-10 coding can be tested on expert-written input in languages that previously had no public benchmark.","Because French and Spanish each come from multiple countries, PARROT enables studying whether reporting-style variation within a language changes model performance.","The near-chance human-versus-AI result implies that quality control for AI-generated radiology text should not rely on human detection alone."],"supporting_citations":[{"why":"Documents the English-only, US-centric scope of a widely used public clinical text resource, defining the gap PARROT targets.","marker":"[11]"},{"why":"Provides a second large English-only radiology report corpus, reinforcing the need for multilingual coverage.","marker":"[12]"},{"why":"Systematic review of clinical text datasets showing the scarcity of openly available multilingual resources.","marker":"[13]"},{"why":"Shows radiology-report evaluation metrics and conventions fail to generalize across countries, motivating the cross-region design.","marker":"[14]"},{"why":"Describes the professional survey network used to recruit radiologist contributors from underrepresented regions.","marker":"[17]"},{"why":"System card of the frontier language model used to generate the AI reports in the differentiation study.","marker":"[19]"}],"fun_headline_variants":["13-language fictional radiology dataset tests medical AI","Fictional radiology reports fool humans nearly half the time","Open multilingual radiology reports: humans vs AI at 53.9%","2658 fictional radiology reports in 13 languages for AI testing","Radiology NLP benchmark: 13 languages, 21 countries, zero privacy risk"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset presupposes that fictional reports written on request by volunteer radiologists, who were told to follow their usual style and to include plausible findings, are representative enough of real radiology reporting in those countries and languages for NLP models benchmarked on them to transfer to clinical practice.","fun_headline_variants_meta":{"raw":{"variants":["13-language fictional radiology dataset tests medical AI","Fictional radiology reports fool humans nearly half the time","Open multilingual radiology reports: humans vs AI at 53.9%","2658 fictional radiology reports in 13 languages for AI testing","Radiology NLP benchmark: 13 languages, 21 countries, zero privacy risk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000841,"raw_usage":{"total_tokens":3698,"prompt_tokens":1012,"completion_tokens":2686,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":2595}},"tokens_in":628,"tokens_out":2686,"duration_ms":18252,"temperature":1.0,"reasoning_tokens":2595,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:01:37.917836+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare distributions of report structure, terminology, and length between PARROT reports and a matched sample of real reports from the same institutions and languages; if real and fictional reports differ systematically, or if an NLP model trained on PARROT fails to match a model trained on real reports on a held-out real-language report understanding task, the central utility claim is contradicted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the English-only, US-centric scope of a widely used public clinical text resource, defining the gap PARROT targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a second large English-only radiology report corpus, reinforcing the need for multilingual coverage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Systematic review of clinical text datasets showing the scarcity of openly available multilingual resources."},{"cited_title":"Banerjee, A","cited_arxiv_id":null,"evidence_quote":"Shows radiology-report evaluation metrics and conventions fail to generalize across countries, motivating the cross-region design."},{"cited_title":"Busch, L","cited_arxiv_id":null,"evidence_quote":"Describes the professional survey network used to recruit radiologist contributors from underrepresented regions."},{"cited_title":"Openai o1 system card","cited_arxiv_id":null,"evidence_quote":"System card of the frontier language model used to generate the AI reports in the differentiation study."}],"review_version":1}