{"id":"6e776f75-12db-4af4-9aca-d1b9b95afce8","arxiv_id":"2412.00230","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of German-language clinical and medical text corpora identifies 71 distinct resources, with only two real clinical corpora currently distributable under formal data-use agreements.","lead":"This paper maps German-language clinical text corpora and their stand-ins, finding 71 distinct collections but very few publicly shareable real ones. It gives researchers a reference guide and highlights how little is known about whether translated or synthetic substitutes really work.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The qualitative access divide is robust, but the headline ratio \"5 of 32\" rests on a document-uniqueness classification that is applied inconsistently to composite or subset corpora such as Llorca-23 and RadQA.","rationale":"The reader's conditional verdict rests mainly on truncated search coverage and subjective relevance screening. My stress-test identifies a related but distinct soft spot: the internal consistency of the document-unique classification that produces the exact 5/32 ratio. The paper's own definition of document-uniqueness is violated by at least two entries in Table 6, so the precise numeric claim is less secure than the qualitative divide. However, the central claim about a clear access divide does not depend on whether the denominator is 32 or 30, or on whether the numerator is 5 or 4; even under conservative reclassification, the overwhelming majority of real German clinical corpora remain inaccessible while nearly all substitutes are publicly available. The proposed check would settle whether the exact counts need revision. Since the reader already assigned a conditional verdict and my concern does not move the verdict in a different direction, UNCHANGED is the appropriate recommendation.","tokens_in":48822,"tokens_out":6271,"duration_ms":64992,"concrete_test":"Recompute the document-unique partition for the 46 real clinical corpus publications using the paper's own definitions and the document-provenance information in Table 1; specifically, test whether Llorca-23 [57] and RadQA [62] have zero document-set intersection with their stated source corpora (Bronco, Cardio:DE, GGPOnc 2.0, GraSCCo; Idrissi-Yaghir-24). Then independently verify the access mechanism for each of the five claimed accessible real corpora (Bronco, Cardio:DE, Ex4CDS, Böhringer-24, GeMTeX) by checking the cited DUA pages, repositories, and contact-based agreements. If the corrected ratio stays near 5/30 and all five access claims verify, the central claim stands; if the numerator or denominator moves by more than two, the headline figures should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical anchor of the claim is that only 5 of 32 document-unique real clinical corpora are externally accessible. The paper defines document-unique as zero intersection of document sets, or as genealogically aligned versions of the same corpus. Applying this definition to the supplementary tables reveals two likely inconsistencies. First, Llorca-23 [57] is explicitly a meta-dataset composed of selected documents from Bronco, Cardio:DE, GGPOnc 2.0, and GraSCCo 1.0; its document set therefore intersects four other real or proxy corpora, yet it is counted as a document-unique real corpus in Table 6. Second, RadQA [62] is a question-answering dataset built from 1,223 radiology reports that appear to be a subset of the Idrissi-Yaghir-24 clinical collection, again making it non-unique under the stated criterion. If these two are removed, the denominator changes from 32 to 30, and the accessibility ratio becomes 5/30 rather than 5/32. The numerator is also judgment-dependent: Böhringer-24 is counted as externally accessible only under the paper's explicitly optimistic assumption that informal private negotiations will result in access, while GeMTeX is described as a corpus in statu nascendi. The qualitative divide between locked real corpora and publicly available substitutes is likely to survive such corrections, but the exact headline percentages are not independently reproducible from the stated criteria without a formal document-set intersection check and an audit of each access claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys German-language clinical and medical text corpora, classifying them into real, translated, synthetic, and domain-proxy (close and distant) categories. It reports results from a PRISM-style literature search of four bibliographic sources, identifying 78 relevant documents that describe 92 corpus versions, which the author reduces to 71 document-unique and 69 annotation-unique corpora. The paper's central claim is that, while almost all authentic German clinical corpora are locked in hospital data silos, the accessible substitutes (translated, synthetic, or domain proxies) are largely publicly available, yet their validity as replacements remains an open empirical question. The paper also proposes a 'corpus card' template for standardizing future corpus documentation.","tokens_in":49163,"tokens_out":4112,"duration_ms":37645,"significance":"The survey addresses an important practical problem: the near-total inaccessibility of German-language clinical text data for NLP research. Its catalog, with per-corpus details on documents, tokens, genres, annotations, and availability, is a useful resource for the clinical NLP community, and the qualitative divide between locked authentic corpora and openly available substitutes is a timely and credible observation. The proposed corpus card template is a constructive contribution to documentation standards. These strengths are substantial even though the exact numerical headline (5 of 32 accessible) is, as discussed below, not fully reproducible from the stated criteria without additional document-set audits.","major_comments":[{"comment":"Llorca-23 [57] is counted as a document-unique real clinical corpus in Table 6, but Table 1 describes it as a meta-dataset composed of selected documents from Bronco, Cardio:DE, GGPOnc 2.0, and GraSCCo 1.0. Under the paper's own definition of document-uniqueness (zero intersection of document sets, or genealogical superset/subset alignment), Llorca-23 has non-empty intersection with four other corpora and is not genealogically aligned with any single one. It should therefore be excluded from the document-unique count, reducing the denominator from 32 to 31.","section":"Table 6 and Results (document-unique definition)"},{"comment":"RadQA [62] is also counted as document-unique, yet Table 1 indicates that its question-answer pairs derive from 1,223 radiology reports of brain CT scans, which appear to be a subset of the clinical collection described in the same publication as Idrissi-Yaghir-24 (25,023k documents). If those radiology reports are contained in the larger Idrissi-Yaghir-24 document set, RadQA violates the zero-intersection criterion. Excluding it would change the denominator to 30, and the headline ratio would become 5/30 rather than 5/32.","section":"Table 6 and Table 1 (RadQA/Idrissi-Yaghir-24)"},{"comment":"The statement that 5 of 32 document-unique real clinical corpora are externally accessible rests on an optimistic classification: Böhringer-24 [14] is available only upon informal private negotiation, and GeMTeX [59] is described as a corpus 'in statu nascendi' that is 'currently not ready for use.' Only Bronco, Cardio:DE, and Ex4CDS currently have concrete distribution channels. The paper should report the optimistic and the strictly verifiable counts separately (e.g., 3 of 30) and specify which corpora underlie each, so that the percentage is independently reproducible.","section":"Results and Discussion (accessibility numerator)"}],"minor_comments":[{"comment":"The search was truncated at 100 hits for ACL Anthology (5,510 hits) and Google Scholar (~443,000 hits). This limitation is acknowledged, but the abstract and objectives describe the survey as 'comprehensive'; the paper should state in the abstract or conclusions that the catalog may miss relevant corpora ranked below the first 100 hits, even though the qualitative divide is unlikely to change.","section":"Materials and Methods (search truncation)"},{"comment":"The abbreviation 'PRISM' is used throughout; the standard name of the reporting guideline is PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses).","section":"Abstract (PRISM vs. PRISMA)"},{"comment":"Several entries in the supplementary tables are hard to read or appear truncated, e.g., '1,245k' for DMP 'HerzMobil' and the fragmentary 'noi' entries; a consistent use of 'n/a' versus 'noi' and a clearer separation of multi-part availability symbols would improve usability.","section":"Supplementary Tables"},{"comment":"The competing-interests statement declares no competing interests, but the author is a co-author of several cataloged corpora (FraMed, 3000PA, JSynCC, GraSCCo, GGPOnc) and proposes the corpus card template. This should be disclosed as a potential bias in the description and assessment of those corpora.","section":"Competing Interests"}],"recommendation":"major_revision","confidential_remarks":"The central qualitative finding is robust and valuable, but the headline quantitative claim (5 of 32) needs correction and clearer documentation of both the document-uniqueness classification and the optimism level of the accessibility assessment. The conflict-of-interest statement should also be revisited given the author's direct involvement in many of the included corpora and in the proposed corpus card template."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper for one thing: it is the most systematic catalog of German clinical/medical corpora I have seen, and its central claim—real clinical corpora are locked away while translated, synthetic, and proxy substitutes are broadly accessible—is almost certainly correct in qualitative terms. The author deserves credit for a transparent PRISM-style search, clear taxonomy, and detailed supplementary tables that go beyond earlier surveys by Starlinger and by Zesch/Bewersdorff. The corpus card template is a modest but practical contribution.\n\nThat said, the exact numbers do not stand up to scrutiny, and the stress-test note is on target. The paper defines document-unique as zero intersection of document sets, but then counts Llorca-23, a meta-dataset deliberately assembled from Bronco, Cardio:DE, GGPOnc 2.0, and GraSCCo, as a document-unique real corpus. RadQA, built from a subset of the Idrissi-Yaghir-24 collection, gets the same treatment. Remove those two and the denominator drops from 32 to 30. The numerator is just as soft: Böhringer-24 is counted as accessible only under the author's explicitly optimistic assumption that informal negotiations will end in a DUA, and GeMTeX is still in statu nascendi. So the headline \"5 of 32\" and the derived 15%/6% figures are not independently reproducible from the stated criteria. The qualitative divide survives these corrections, but the precise ratios should not be quoted without an audit.\n\nThe truncated search—only the first 100 hits for ACL Anthology and Google Scholar—is a real completeness limitation, and the author's wording that the review covers \"hopefully all\" corpora is too confident. A revision should either estimate how many corpora were missed or hedge the completeness claim explicitly. The author is also a contributor to several cataloged corpora, which is not a flaw by itself, but it makes an independent check of the access classifications more important, not less.\n\nAll of this is fixable. The paper is a useful reference for anyone working in German clinical NLP, and it deserves peer review rather than desk rejection. I would send it out and ask the referee to focus on the document-uniqueness accounting and the access classifications.","headline":"A useful and generally careful catalog of German clinical corpora that gets the qualitative access divide right, but whose headline 5-of-32 number rests on inconsistent document-uniqueness calls and should be fixed before the counts are quoted.","tokens_in":49641,"tokens_out":1854,"would_cite":true,"duration_ms":19431,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of German-language clinical and medical text corpora establishes a sharp divide: authentic clinical documents stay locked in hospital data silos, and nearly every publicly accessible alternative is a translated…","keywords":["clinical text corpora","German-language NLP","corpus accessibility","data privacy","synthetic clinical data","domain proxies","systematic review","corpus documentation"],"falsifier":"A bibliographic sweep without the 100-hit truncation, plus a targeted solicitation to German hospital NLP groups, would test the census: if it surfaced even a handful of previously uncounted externally accessible authentic clinical corpora, the 5-of-32 accessibility figure would need revision. Separately, a head-to-head benchmark, training the same named-entity or coding model on real, translated, synthetic, close-proxy, and distant-proxy data and evaluating all on held-out real clinical reports, would settle whether the paper's central worry about substitute validity is justified: parity would dissolve it, while a gap would confirm it.","tokens_in":48617,"feed_emoji":"🔒","tokens_out":8953,"duration_ms":76853,"temperature":0.7,"pith_summary":"This paper surveys the full landscape of German-language clinical and medical text corpora and establishes a sharp, quantified divide: authentic clinical documents are almost entirely locked inside hospital data silos, while essentially everything that is publicly accessible is a substitute, such as translated English clinical data, synthetic documents about fictitious patients, or domain proxies like journal articles, guidelines, Wikipedia, and social media. The survey counts 71 document-unique corpora from 92 published versions and finds that only 5 of 32 real clinical corpora are accessible to outsiders at all, with only 2, namely Bronco and Cardio:DE, distributable through a formalized data-use agreement. The paper's central argument is that the data bottleneck only appears to be broken: substitutes differ from real clinical text in genre, style, terminology, and medical expertise in ways that are obvious qualitatively but have never been measured empirically. It therefore closes by defining the missing research agenda, a systematic, empirically grounded yardstick for comparing real corpora with their substitutes, and proposes a generic corpus-card template to standardize future corpus documentation.","feed_headline":"Only 2 of 32 German clinical corpora are shareable","feed_subtitle":"A systematic survey finds the data bottleneck is only apparently broken; substitutes differ sharply from real clinical text.","key_machinery":"The argument is carried by a five-category typology of corpora, namely real, translated, and synthetic clinical corpora and close versus distant domain proxies, applied through a PRISM-conformant systematic review of four bibliographic sources. Two counting distinctions do the quantitative work: document-unique corpora, meaning zero intersection of underlying document sets with superset and version relationships merged, versus annotation-unique corpora distinguished by their metadata, and the three-valued accessibility classification of inaccessible, DUA-based, and unrestricted. On top of this, the paper's qualitative analysis of the distance between categories, clinical reporting as performance under time pressure with local abbreviation dialects versus scholarly writing and lay medical discourse, is what motivates the claim that substitute validity is an open empirical question. The proposed corpus-card template documented in an appendix is the standardization mechanism intended to make future corpus descriptions comparable across the typology.","core_discovery":"On the paper's own terms, the discovery is a map and a census. From 362 bibliographic hits screened under a PRISM-conformant protocol, the review identifies 92 published corpus versions, of which 71 are document-unique, and organizes them into five types: real, translated, and synthetic clinical corpora, plus close and distant domain proxies. The census shows that of the 32 document-unique authentic clinical corpora, including heavily annotated resources like 3000PA 5.0 with about 2.1 million annotation units on 6,600 discharge summaries, only five are externally accessible under optimistic assumptions, and only two, Bronco and Cardio:DE, are ready for distribution under a standardized formal data-use agreement. All three translated, three synthetic, sixteen close-proxy, and seventeen distant-proxy corpora, by contrast, are publicly available, so the apparent data poverty of German clinical NLP is really a distribution blockade. The paper argues that the substitutes, whatever their surface usefulness, deviate systematically from clinical writing in syntactic well-formedness, jargon and abbreviation use, and the shift from individual-patient documentation to generalizable scholarly or lay discourse, and that no empirical measure yet says how much this deviation costs.","pith_inferences":["If the divide is real, benchmark results obtained on any substitute corpus probably overstate how a system will behave on genuine German clinical text; the community should expect a systematic performance gap until a substitution cost model exists.","The five-cell typology transfers directly to other privacy-restricted languages: the same real, translated, synthetic, and proxy structure and the same DUA-versus-substitute trade-off should reappear in, say, French, Japanese, or Italian clinical NLP, so the German survey can serve as a template for comparable national audits.","A concrete next experiment, implicit in the paper but not run there, would profile all five corpus categories with stylometric metrics of the kind the paper cites, then correlate the metric distances with downstream-task performance on held-out real clinical documents, turning the paper's open validity question into a measurable substitution-cost curve."],"forward_implications":["German clinical NLP systems that cannot obtain data-use agreement access will be trained and evaluated on domain-shifted data; the paper documents why the shift is systematic rather than incidental.","Bronco and Cardio:DE define the current distribution standard for privacy-sensitive German clinical corpora and are, for now, the only safe common evaluation grounds.","Releasing trained language models instead of raw text is a partial workaround, but published privacy attacks that can reconstruct sensitive content from model representations mean this route is not a clean solution.","The proposed corpus-card template, if adopted, would make future corpus descriptions comparable across all five categories and is a precondition for the missing substitution cost model.","Consent-based initiatives such as GeMTeX promise to enlarge the accessible pool of authentic corpora, but only through DUA-mediated access and only after a delay."],"supporting_citations":[{"why":"The Bronco oncological corpus is the pioneering DUA-accessible real German clinical corpus; the paper's accessibility count treats it as one of only two formally distributable corpora.","marker":"[11]"},{"why":"Cardio:DE is the first German clinical corpus distributed under a formal DUA with intact document structure; the paper treats it as the standard for privacy-compliant sharing.","marker":"[56]"},{"why":"Supplies the systematic-review protocol whose identification and screening flow produces the census of 78 relevant documents and 92 published corpus versions.","marker":"[20]"},{"why":"The 3000PA 5.0 final release, with about 2.1 million annotation units on 6,600 discharge summaries, is the leading example of a richly annotated authentic corpus that remains locked.","marker":"[60]"},{"why":"Analyzes the GDPR and German data-protection conditions, including informed consent and the unreasonable-effort clause, that ground the paper's explanation of why German hospitals block corpus distribution.","marker":"[17]"},{"why":"Reports preliminary evidence that models trained on synthetic clinical German data transfer poorly to real corpora such as Bronco and Cardio:DE, supporting the paper's validity concern.","marker":"[76]"},{"why":"Introduces the stylometric toolkit that the paper points to as the descriptive side of the missing empirically grounded comparison of corpus types.","marker":"[110]"}],"fun_headline_variants":["Only 2 of 32 authentic German clinical corpora are open","German clinical text: 30 of 32 corpora locked from researchers","Survey: Clinical corpus substitutes fail to match real German text","Data privacy stalls German clinical NLP: 30 of 32 corpora closed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The census behind the headline numbers rests on a literature screen that could only inspect the first 100 hits for two of its four databases, ACL Anthology with 5,510 hits and Google Scholar with about 443,000, so any relevant German clinical corpus ranked below those cutoffs would be missing from the counts, changing the exact figures though probably not the qualitative divide.","fun_headline_variants_meta":{"raw":{"variants":["Only 2 of 32 authentic German clinical corpora are open","German clinical text: 30 of 32 corpora locked from researchers","Survey: Clinical corpus substitutes fail to match real German text","Data privacy stalls German clinical NLP: 30 of 32 corpora closed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00065,"raw_usage":{"total_tokens":3069,"prompt_tokens":1120,"completion_tokens":1949,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":736,"completion_tokens_details":{"reasoning_tokens":1873}},"tokens_in":736,"tokens_out":1949,"duration_ms":13487,"temperature":1.0,"reasoning_tokens":1873,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:35:11.595084+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A bibliographic sweep without the 100-hit truncation, plus a targeted solicitation to German hospital NLP groups, would test the census: if it surfaced even a handful of previously uncounted externally accessible authentic clinical corpora, the 5-of-32 accessibility figure would need revision. Separately, a head-to-head benchmark, training the same named-entity or coding model on real, translated, synthetic, close-proxy, and distant-proxy data and evaluating all on held-out real clinical reports, would settle whether the paper's central worry about substitute validity is justified: parity would dissolve it, while a gap would confirm it.","supporting_citations":[],"review_version":1}