Pith. sign in

REVIEW 3 major objections 4 minor 200 references

Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A systematic review of German-language clinical and medical text corpora establishes a sharp divide: authentic clinical documents stay locked in hospital data silos, and nearly every publicly accessible alternative is a translated…

desk verdict A useful and generally careful catalog of German clinical corpora that gets the qualitative access divide right, but whose headline 5-of-32 number rests on inconsistent document-uniqueness calls and should be fixed before the counts are quoted. read the letter →

arxiv 2412.00230 v2 pith:N2CYWMTV submitted 2024-11-29 cs.CL

classification cs.CL
keywords clinicaltextcorporaGerman-languageNLPcorpusaccessibilitydataprivacysyntheticdomainproxiessystematicreviewdocumentation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper surveys the full landscape of German-language clinical and medical text corpora and establishes a sharp, quantified divide: authentic clinical documents are almost entirely locked inside hospital data silos, while essentially everything that is publicly accessible is a substitute, such as translated English clinical data, synthetic documents about fictitious patients, or domain proxies like journal articles, guidelines, Wikipedia, and social media. The survey counts 71 document-unique corpora from 92 published versions and finds that only 5 of 32 real clinical corpora are accessible to outsiders at all, with only 2, namely Bronco and Cardio:DE, distributable through a formalized data-use agreement. The paper's central argument is that the data bottleneck only appears to be broken: substitutes differ from real clinical text in genre, style, terminology, and medical expertise in ways that are obvious qualitatively but have never been measured empirically. It therefore closes by defining the missing research agenda, a systematic, empirically grounded yardstick for comparing real corpora with their substitutes, and proposes a generic corpus-card template to standardize future corpus documentation.

What carries the argument

The argument is carried by a five-category typology of corpora, namely real, translated, and synthetic clinical corpora and close versus distant domain proxies, applied through a PRISM-conformant systematic review of four bibliographic sources. Two counting distinctions do the quantitative work: document-unique corpora, meaning zero intersection of underlying document sets with superset and version relationships merged, versus annotation-unique corpora distinguished by their metadata, and the three-valued accessibility classification of inaccessible, DUA-based, and unrestricted. On top of this, the paper's qualitative analysis of the distance between categories, clinical reporting as performance under time pressure with local abbreviation dialects versus scholarly writing and lay medical discourse, is what motivates the claim that substitute validity is an open empirical question. The proposed corpus-card template documented in an appendix is the standardization mechanism intended to make future corpus descriptions comparable across the typology.

What would settle it

A bibliographic sweep without the 100-hit truncation, plus a targeted solicitation to German hospital NLP groups, would test the census: if it surfaced even a handful of previously uncounted externally accessible authentic clinical corpora, the 5-of-32 accessibility figure would need revision. Separately, a head-to-head benchmark, training the same named-entity or coding model on real, translated, synthetic, close-proxy, and distant-proxy data and evaluating all on held-out real clinical reports, would settle whether the paper's central worry about substitute validity is justified: parity would dissolve it, while a gap would confirm it.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a map and a census. From 362 bibliographic hits screened under a PRISM-conformant protocol, the review identifies 92 published corpus versions, of which 71 are document-unique, and organizes them into five types: real, translated, and synthetic clinical corpora, plus close and distant domain proxies. The census shows that of the 32 document-unique authentic clinical corpora, including heavily annotated resources like 3000PA 5.0 with about 2.1 million annotation units on 6,600 discharge summaries, only five are externally accessible under optimistic assumptions, and only two, Bronco and Cardio:DE, are ready for distribution under a standardized formal data-use agreement. All three translated, three synthetic, sixteen close-proxy, and seventeen distant-proxy corpora, by contrast, are publicly available, so the apparent data poverty of German clinical NLP is really a distribution blockade. The paper argues that the substitutes, whatever their surface usefulness, deviate systematically from clinical writing in syntactic well-formedness, jargon and abbreviation use, and the shift from individual-patient documentation to generalizable scholarly or lay discourse, and that no empirical measure yet says how much this deviation costs.

Load-bearing premise

The census behind the headline numbers rests on a literature screen that could only inspect the first 100 hits for two of its four databases, ACL Anthology with 5,510 hits and Google Scholar with about 443,000, so any relevant German clinical corpus ranked below those cutoffs would be missing from the counts, changing the exact figures though probably not the qualitative divide.

Editorial extensions

If this is right

  • German clinical NLP systems that cannot obtain data-use agreement access will be trained and evaluated on domain-shifted data; the paper documents why the shift is systematic rather than incidental.
  • Bronco and Cardio:DE define the current distribution standard for privacy-sensitive German clinical corpora and are, for now, the only safe common evaluation grounds.
  • Releasing trained language models instead of raw text is a partial workaround, but published privacy attacks that can reconstruct sensitive content from model representations mean this route is not a clean solution.
  • The proposed corpus-card template, if adopted, would make future corpus descriptions comparable across all five categories and is a precondition for the missing substitution cost model.
  • Consent-based initiatives such as GeMTeX promise to enlarge the accessible pool of authentic corpora, but only through DUA-mediated access and only after a delay.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the divide is real, benchmark results obtained on any substitute corpus probably overstate how a system will behave on genuine German clinical text; the community should expect a systematic performance gap until a substitution cost model exists.
  • The five-cell typology transfers directly to other privacy-restricted languages: the same real, translated, synthetic, and proxy structure and the same DUA-versus-substitute trade-off should reappear in, say, French, Japanese, or Italian clinical NLP, so the German survey can serve as a template for comparable national audits.
  • A concrete next experiment, implicit in the paper but not run there, would profile all five corpus categories with stylometric metrics of the kind the paper cites, then correlate the metric distances with downstream-task performance on held-out real clinical documents, turning the paper's open validity question into a measurable substitution-cost curve.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This manuscript surveys German-language clinical and medical text corpora, classifying them into real, translated, synthetic, and domain-proxy (close and distant) categories. It reports results from a PRISM-style literature search of four bibliographic sources, identifying 78 relevant documents that describe 92 corpus versions, which the author reduces to 71 document-unique and 69 annotation-unique corpora. The paper's central claim is that, while almost all authentic German clinical corpora are locked in hospital data silos, the accessible substitutes (translated, synthetic, or domain proxies) are largely publicly available, yet their validity as replacements remains an open empirical question. The paper also proposes a 'corpus card' template for standardizing future corpus documentation.

Significance. The survey addresses an important practical problem: the near-total inaccessibility of German-language clinical text data for NLP research. Its catalog, with per-corpus details on documents, tokens, genres, annotations, and availability, is a useful resource for the clinical NLP community, and the qualitative divide between locked authentic corpora and openly available substitutes is a timely and credible observation. The proposed corpus card template is a constructive contribution to documentation standards. These strengths are substantial even though the exact numerical headline (5 of 32 accessible) is, as discussed below, not fully reproducible from the stated criteria without additional document-set audits.

major comments (3)
  1. [Table 6 and Results (document-unique definition)] Llorca-23 [57] is counted as a document-unique real clinical corpus in Table 6, but Table 1 describes it as a meta-dataset composed of selected documents from Bronco, Cardio:DE, GGPOnc 2.0, and GraSCCo 1.0. Under the paper's own definition of document-uniqueness (zero intersection of document sets, or genealogical superset/subset alignment), Llorca-23 has non-empty intersection with four other corpora and is not genealogically aligned with any single one. It should therefore be excluded from the document-unique count, reducing the denominator from 32 to 31.
  2. [Table 6 and Table 1 (RadQA/Idrissi-Yaghir-24)] RadQA [62] is also counted as document-unique, yet Table 1 indicates that its question-answer pairs derive from 1,223 radiology reports of brain CT scans, which appear to be a subset of the clinical collection described in the same publication as Idrissi-Yaghir-24 (25,023k documents). If those radiology reports are contained in the larger Idrissi-Yaghir-24 document set, RadQA violates the zero-intersection criterion. Excluding it would change the denominator to 30, and the headline ratio would become 5/30 rather than 5/32.
  3. [Results and Discussion (accessibility numerator)] The statement that 5 of 32 document-unique real clinical corpora are externally accessible rests on an optimistic classification: Böhringer-24 [14] is available only upon informal private negotiation, and GeMTeX [59] is described as a corpus 'in statu nascendi' that is 'currently not ready for use.' Only Bronco, Cardio:DE, and Ex4CDS currently have concrete distribution channels. The paper should report the optimistic and the strictly verifiable counts separately (e.g., 3 of 30) and specify which corpora underlie each, so that the percentage is independently reproducible.
minor comments (4)
  1. [Materials and Methods (search truncation)] The search was truncated at 100 hits for ACL Anthology (5,510 hits) and Google Scholar (~443,000 hits). This limitation is acknowledged, but the abstract and objectives describe the survey as 'comprehensive'; the paper should state in the abstract or conclusions that the catalog may miss relevant corpora ranked below the first 100 hits, even though the qualitative divide is unlikely to change.
  2. [Abstract (PRISM vs. PRISMA)] The abbreviation 'PRISM' is used throughout; the standard name of the reporting guideline is PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses).
  3. [Supplementary Tables] Several entries in the supplementary tables are hard to read or appear truncated, e.g., '1,245k' for DMP 'HerzMobil' and the fragmentary 'noi' entries; a consistent use of 'n/a' versus 'noi' and a clearer separation of multi-part availability symbols would improve usability.
  4. [Competing Interests] The competing-interests statement declares no competing interests, but the author is a co-author of several cataloged corpora (FraMed, 3000PA, JSynCC, GraSCCo, GGPOnc) and proposes the corpus card template. This should be disclosed as a potential bias in the description and assessment of those corpora.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the review's access-divide claim is an empirical cataloging result, not a derivation from its own inputs.

full rationale

The paper is a systematic review, not a derivation. Its central claim—that real German clinical corpora are mostly inaccessible while translated, synthetic, and proxy substitutes are publicly available—is an empirical summary of independently published corpus descriptions. The 'derivation chain' consists of PRISM screening, eligibility criteria, and tabulation of document, token, and access information from the cited primary sources. No equation-level reduction occurs: accessibility labels and uniqueness counts are applied to external artifacts and can in principle be checked against the cited repositories and DOIs. Self-citations (FraMed, 3000PA, JSynCC, GraSCCo, GGPOnc, DoPA Meter) appear because the author is a primary contributor to several surveyed corpora, but they function as literature sources for those specific corpora, not as justification for the survey's typology or for the 5-of-32 access ratio. The proposed corpus card template is explicitly a recommendation for future documentation, not a result derived from or validated by the survey. Some individual classifications may be debatable—for example, counting the Llorca-23 meta-dataset or RadQA as document-unique is a potential internal-consistency issue—but that would be a correctness or completeness concern, not circularity. No load-bearing premise is defined in terms of the conclusion, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey's central claims rest on search and screening choices rather than fitted parameters. There are no free parameters or invented entities. The main burdens are the representativeness of truncated searches, the reliability of classifying corpora from publication descriptions, and the author's chosen inclusion thresholds.

assumptions (3)
  • domain assumption Truncated screening of ACL Anthology and Google Scholar results is representative of the full corpus literature.
    ACL yielded 5,510 hits and Google ~443,000, but only the first 100 from each were screened. The completeness of the 71-corpus catalog depends on this assumption.
  • domain assumption Corpus descriptions in papers allow reliable inference of document uniqueness and access status.
    The document-unique and annotation-unique distinctions are based on reported document sets and metadata. If descriptions are incomplete or inconsistent, the counts and categories could shift.
  • ad hoc to paper Inclusion thresholds such as at least 100 documents or 10,000 tokens, human medicine only, and written data only define the relevant corpus landscape.
    These cutoffs are chosen by the author and shape which corpora enter the counts. A different threshold would change the numbers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data." pith.science (2026). https://pith.science/paper/N2CYWMTV

@misc{pith2026241200230,
  author       = {Pith},
  title        = {Pith review of: Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N2CYWMTV}},
  note         = {Machine review of arXiv:2412.00230}
}
read the original abstract

We survey clinical document corpora, with focus on German textual data. Due to rigid data privacy legislation in Germany these resources, with only few exceptions, are stored in safe clinical data spaces and locked against clinic-external researchers. This situation stands in stark contrast with established workflows in the field of natural language processing where easy accessibility and reuse of data collections are common practice. Hence, alternative corpus designs have been examined to escape from this data poverty. Besides machine translation of English clinical datasets and the generation of synthetic corpora with fictitious clinical contents, several other types of domain proxies have come up as substitutes for clinical documents. Common instances of close proxies are medical journal publications, therapy guidelines, drug labels, etc., more distant proxies include online encyclopedic medical articles or medical contents from social media channels. After PRISM-conformant identification of 362 hits from 4 bibliographic systems, 78 relevant documents were finally selected for this review. They contained overall 92 different published versions of corpora from which 71 were truly unique in terms of their underlying document sets. Out of these, the majority were clinical corpora -- 46 real ones, 5 translated ones, and 6 synthetic ones. As to domain proxies, we identified 18 close and 17 distant ones. There is a clear divide between the large number of non-accessible authentic clinical German-language corpora and their publicly accessible substitutes: translated or synthetic, close or more distant proxies. So on first sight, the data bottleneck seems broken. Yet differences in genre-specific writing style, wording and medical domain expertise in this typological space are also obvious. This raises the question how valid alternative corpus designs really are.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

200 extracted references · 78 canonical work pages

  1. [57]

    Llorca, Ignacio, & Borchert, Florian, & Schapranow, Matthieu -P. (2023). A meta -dataset of German medical corpora: harmonization of annotations and cross -corpus NER evaluation. In: ClinicalNLP 2023 – Proceedings of the 5th Workshop on Clinical Natural Language Processing @ ACL 2023. Toronto, Ontario, Canada, July 14, 2023, pp. 171-181

  2. [62]

    Final report on the German clinical reference corpus 3000PA

    Hahn, Udo, & Modersohn, Luise, & Faller, Jakob, & Lohr, Christina (2024). Final report on the German clinical reference corpus 3000PA. In: MEDINFO 2023 – The Future Is Accessible. Proceedings of the 19th World Congress on Medical and Health Informatics. [Sydney, New South Wales, Australia, 8 -12 July 2023], pp. 599 -603 (Studies in Health Technology and I...

  3. [14]

    Böhringer, Daniel, & Angelova, P., & Fuhrmann, L., & Zimmermann, J., & Schargus, M., & Eter, N., & Reinhard, T. (2024). Automatic inference of ICD -10 codes from German ophthalmologic physicians' letters using natural language processing. Scientific Reports, 14:#9035 [6 pp.]

  4. [59]

    Announcement of the German Medical Text Corpus Project ( GeMTeX)

    Meineke, Frank, & Modersohn, Luise, & Loeffler, Markus, & Boeker, Martin (2023). Announcement of the German Medical Text Corpus Project ( GeMTeX). In: Caring is Sharing – Exploiting the Value in Data for Health and Innovation. Proceedings of [the 33rd Medical Informatics Europe Conference] MIE 2023. [Gothenburg, Sweden, 22-25 May 2023], pp. 835-836 (Studi...

  5. [1]

    arXiv:1904.01172 (v2)

    Storks, Shane, & Gao, Qiaozi , & Chai, Joyce Yue (2020): Recent advances in natural language inference: a survey of benchmarks, resources, and approaches. arXiv:1904.01172 (v2)

  6. [2]

    Data and its (dis)contents: a survey of dataset development and use in machine learning research

    Paullada, Amandalynne, & Raji, Inioluwa Deborah, & Bender, Emily M., & Denton, Emily, & Hanna, Alex (2021). Data and its (dis)contents: a survey of dataset development and use in machine learning research. Patterns, 2(11):#100336

  7. [3]

    Computational Methods for Corpus Annotation and Analysis

    Lu, Xiaofei (2014). Computational Methods for Corpus Annotation and Analysis. Springer

  8. [4]

    & Pustejovsky, James D., eds

    Ide, Nancy C. & Pustejovsky, James D., eds. (2017). Handbook of Linguistic Annotation. Springer

Show all 200 references
  1. [5]

    Campbell, David A., & Johnson, Stephen B. (2001). Comparing syntactic complexity in medical and non-medical corpora. In: AMIA 2001 – Proceedings of the 2001 Annual Symposium of the American Medical Informatics Association. A Medical Informatics Odyssey: Visions of the Future a...

  2. [6]

    Two biomedical sublanguages: a description based on the theories of Zellig Harris

    Friedman, Carol, & Kra, Pauline, & Rzhetsky, Andrey (2002). Two biomedical sublanguages: a description based on the theories of Zellig Harris. Journal of Biomedical Informatics, 35(4):222- 235

  3. [7]

    Zeng, Qing T., & Redd, Doug, & Divita, Guy, & SamahJarad, & Brandt, Cynthia A., & Nebeker, Jonathan R. (2011). Characterizing clinical text and sublanguage: a case study of the VA clinical notes. Journal of Health & Medical Informatics, 2011:S3

  4. [8]

    Document clustering of clinical narratives: a systematic study of clinical sublanguages

    Patterson, Olga V., & Hurdle, John Franklin (2011). Document clustering of clinical narratives: a systematic study of clinical sublanguages. In: AMIA 2011 – Proceedings of the 2011 Annual Symposium on Biomedical and Health Informatics of the American Medical Informatics Associ...

  5. [9]

    Stylistic features of case reports as a genre of medical discourse

    Lysanets, Yuliia, & Morokhovets, Halyna, & Bieliaieva, Olena (2017). Stylistic features of case reports as a genre of medical discourse. Journal of Medical Case Reports, 11:#83 (83:1–83:5)

  6. [10]

    Cross -domain German medical named entity recognition using a pre -trained language model and unified medical semantic types

    Liang, Siting, & Hartmann, Mareike, & Sonntag, Daniel (2023). Cross -domain German medical named entity recognition using a pre -trained language model and unified medical semantic types. In: ClinicalNLP 2023 – Proceedings of the 5th Workshop on Clinical Natural Language Proce...

  7. [11]

    Annotation and initial evaluation of a large annotated German oncological corpus

    Kittner, Madeleine, & Lamping, Mario, & Rieke, Damian T., & Götze, Julian, & Bajwa, Bariya, & Jelas, Ivan, & Rüter, Gina, & Hautow, Hanjo, & Sänger, Mario, & Habibi, Maryam, & Zettwitz, Marit, & de Bortoli, Till, & Ostermann, Leonie, & Ševa, Jurica, & Starlinger, Johannes, & K...

  8. [12]

    An annotated corpus of textual explanations for clinical decision support

    Roller, Roland, & Burchardt, Aljoscha, & Feldhus, Nils, & Seiffe, Laura, & Budde, Klemens, & Ronicke, Simon, & Osmanodja, Bilgin (2022). An annotated corpus of textual explanations for clinical decision support. In: LREC 2022 – Proceedings of the 13th International Conference ...

  9. [13]

    Building a German clinical named entity recognition system without in -domain training data

    Liang, Siting, & Profitlich, Hans-Jürgen, & Klass, Maximilian, & Möller -Grell, Niko, & Bergmann, Celine-Fabienne, & Heim, Simon, & Niklas, Christian, & Sonntag, Daniel (2024). Building a German clinical named entity recognition system without in -domain training data. In: Cli...

  10. [15]

    A short review of ethical challenges in clinical natural language processing

    Šuster, Simon, & Tulkens, Stéphan, & Daelemans, Walter (2017). A short review of ethical challenges in clinical natural language processing. In: Proceedings of the 1st ACL Workshop on Ethics in Natural Language Processing @ EACL 2017. Valencia, Spain, April 4, 2017, pp. 80-87

  11. [16]

    Semi -automated de- identification of German content sensitive reports for big data analytics

    Seuss, Hannes, & Dankerl, Peter, & Ihle, Matthias, & Grandjean, Andrea, & Hammon, Rebecca, & Kaestle, Nicola, & Fasching, Peter A., & Maier, Christian, & Christoph, Jan, & Sedlmayr, Martin, & Uder, Michael, & Cavallaro, Alexander, & Hammon, Matthias ( 2017). Semi -automated de...

  12. [17]

    De-identifying GraSCCo: a pilot study for the de -identification of the German Medical Text Project ( GeMTeX) corpus

    Lohr, Christina, & Matthies, Franz & Faller, Jakob & Modersohn, Luise & Riedel, Andrea & Hahn, Udo & Kiser, Rebekka & Boeker, Martin & Meineke, Frank (2024). De-identifying GraSCCo: a pilot study for the de -identification of the German Medical Text Project ( GeMTeX) corpus. I...

  13. [18]

    How to improve information extraction from German medical notes

    Starlinger, Johannes, & Kittner, Madeleine, & Blankenstein, Oliver, & Leser, Ulf (2016). How to improve information extraction from German medical notes. it – Information Technology , 58(10):1-8

  14. [19]

    German medical natural language processing: a data-centric survey

    Zesch, Torsten, & Bewersdorff, Jeanette (2022). German medical natural language processing: a data-centric survey. In: UR-AI 2022 – Proceedings of the 4th Upper -Rhine Artificial Intelligence Symposium: Artificial Intelligence Applications in Medicine and Manufacturing . Villi...

  15. [20]

    Preferred reporting items for systematic reviews and meta-analyses: the Prisma statement

    Moher, David, & Liberati, Alessandro, & Tetzlaff, Jennifer, & Altman, Douglas G., & The PRISMA Group (2009). Preferred reporting items for systematic reviews and meta-analyses: the Prisma statement. PLoS Medicine, 6(7):e1000097

  16. [21]

    An annotated German-language medical text corpus as language resource

    Wermter, Joachim, & Hahn, Udo (2004). An annotated German-language medical text corpus as language resource. In: LREC 2004 – Proceedings of the 4th International Conference on Language Resources and Evaluation. Lisbon, Portugal, 24-30 May 2004, pp. 473-476

  17. [22]

    Disclose models, hide the data: how to make use of confidential corpora without seeing sensitive raw data

    Faessler, Erik, & Hellrich, Johannes, & Hahn, Udo (2014). Disclose models, hide the data: how to make use of confidential corpora without seeing sensitive raw data. In: LREC 2014 – Proceedings of the 9th International Conference on Language Resources and Evaluation . Reykjavik...

  18. [23]

    Sharing models and tools for processing German clinical texts

    Hellrich, Johannes, & Matthies, Franz, & Faessler, Erik, & Hahn, Udo (2015). Sharing models and tools for processing German clinical texts. In: Digital Healthcare Empowering Europeans. MIE 2015 – Proceedings of the 26th Conference on Medical Informatics in Europe. Madrid, Spai...

  19. [24]

    Biomedical data mining in clinical routine: expanding the impact of hospital information systems

    Müller, Marcel, & Markó, Kornél G., & Daumke , Philipp, & Paetzold, Jan, & Roesner, Arnold, & Klar, Rüdiger (2007). Biomedical data mining in clinical routine: expanding the impact of hospital information systems. In: MedInfo 2007 – Proceedings of the 12th World Congress on He...

  20. [25]

    Enhanced information retrieval from narrative German-language clinical text documents using automated document classificati on

    Spat, Stephan, & Cadonna, Bruno, & Rakovac, Ivo, & Gütl, Christian, & Leitner, Hubert, & Stark, Günther, & Beck, Peter (2008). Enhanced information retrieval from narrative German-language clinical text documents using automated document classificati on. In: eHealth Beyond the...

  21. [26]

    Truecasing clinical narratives

    Kreuzthaler, Markus, & Schulz, Stefan (2011). Truecasing clinical narratives. In: User Centred Networked Health Care. MIE 2011 – Proceedings of the 23rd Conference of the European Federation of Medical Informatics . Oslo, Norway, August 28 -31, 2011, pp. 589 -593 (Studies in H...

  22. [27]

    Information extraction from unstructured electronic health records and integration into a data warehouse

    Fette, Georg, & Ertl, Maximilian, & Wörner, Anja, & Kluegl, Peter, & Störk, Stefan, & Puppe, Frank (2012). Information extraction from unstructured electronic health records and integration into a data warehouse. In: INFORMATIK 2012: Was bewegt uns in der/die Zukunft? Proceedings der

  23. [28]

    Identifying pathological findings in German radiology reports using a syntacto -semantic parsing approach

    Bretschneider, Claudia, & Zillner, Sonja, & Hammon, Matthias (2013). Identifying pathological findings in German radiology reports using a syntacto -semantic parsing approach. In: BioNLP 2013 – Proceedings of the 2013 Workshop on Biomedical Natural Language Processing @ ACL

  24. [29]

    Corpus-based translation of ontologies for improved multilingual semantic annotation

    Bretschneider, Claudia, & Oberkampf, Heiner, & Zillner, Sonja, & Bauer, Bernhard, & Hammon, Matthias (2014). Corpus-based translation of ontologies for improved multilingual semantic annotation. In: SWAIE 2014 – Proceedings of 3rd Workshop on Semantic Web and Information Extra...

  25. [30]

    Fine-grained information extraction from German transthoracic echocardiography reports

    Toepfer, Martin, & Corovic, Hamo, & Fette, Georg, & Kluegl, Peter, & Störk, Stefan, & Puppe, Frank (2015). Fine-grained information extraction from German transthoracic echocardiography reports. BMC Medical Informatics and Decision Making, 15:#91 (91:1–91:16)

  26. [31]

    A corpus of German clinical reports for ICD and OPS- based language modeling

    Lohr, Christina, & Herms, Robert (2016). A corpus of German clinical reports for ICD and OPS- based language modeling. In: CLAW 2016 – Proceedings of the 6th Workshop on Controlled Language Applications @ LREC 2016. Portorož, Slovenia, 28 May 2016, pp. 20-23

  27. [32]

    Automated classification of selected data elements from free -text diagnostic reports in clinical research

    Löpprich, Martin, & Krauss, Felix, & Ganzinger, Matthias, & Senghas, Karsten, & Riezler, Stefan, & Knaup, Petra (2016). Automated classification of selected data elements from free -text diagnostic reports in clinical research. Methods of Information in Medicine, 55(4):373-380

  28. [33]

    A fine -grained corpus annotation schema of German nephrology records

    Roller, Roland, & Uszkoreit, Hans, & Xu, Feiyu, & Seiffe, Laura, & Mikhailov, Michael, & Staeck, Oliver, & Budde, Klemens, & Halleck, Fabian, & Schmidt, Danilo (2016). A fine -grained corpus annotation schema of German nephrology records. In: ClinicalNLP 2016 – Proceedings of ...

  29. [34]

    Unsupervised abbreviation detection in clinical narratives

    Kreuzthaler, Markus, & Oleynik, Michel, & Avian, Alexander, & Schulz, Stefan (2016). Unsupervised abbreviation detection in clinical narratives. In: ClinicalNLP 2016 – Proceedings of the 1st Workshop on Clinical Natural Language Processing @ COLING 2016 . Osaka, Japan, Decembe...

  30. [35]

    Negation detection in clinical reports written in German

    Cotik, Viviana, & Roller, Roland, & Xu, Feiyu, & Uszkoreit, Hans, & Budde, Klemens, & Schmidt, Danilo (2016). Negation detection in clinical reports written in German. In: BioTxtM 2016 – Proceedings of the 5th Workshop on Building and Evaluating Resources for Biomedical Text M...

  31. [36]

    Unsupervised abbreviation expansion in clinical narratives

    Oleynik, Michel, & Kreuzthaler, Markus, & Schulz, Stefan (2017). Unsupervised abbreviation expansion in clinical narratives. In: MedInfo 2017 – Proceedings of the 16th World Congress on Medical and Health Informatics: Precision Healthcare through Informatics. Hangzhou, China, ...

  32. [37]

    Detecting named entities and relations in German clinical reports

    Roller, Roland, & Rethmeier, Nils, & Thomas, Philippe E., & Hübner, Marc, & Uszkoreit, Hans, & Staeck, Oliver, & Budde, Klemens, & Halleck, Fabian, & Schmidt, Danilo (2018). Detecting named entities and relations in German clinical reports. In: Language Technologies for the Ch...

  33. [38]

    Semi -automatic terminology generation for information extraction from German chest X-ray reports

    Krebs, Jonathan, & Corovic, Hamo, & Dietrich, Georg, & Ertl, Maximilian, & Fette, Georg, & Kaspar, Mathias, & Krug, Markus, & Störk, Stefan, & Puppe, Frank (2017). Semi -automatic terminology generation for information extraction from German chest X-ray reports. In: German Med...

  34. [39]

    3000PA: towards a national reference corpus of German clinical language

    Hahn, Udo, & Matthies, Franz, & Lohr, Christina, & Löffler, Markus (2018). 3000PA: towards a national reference corpus of German clinical language. In: MIE 2018 – Proceedings of the 29th Conference on Medical Informatics in Europe: Building Continents of Knowledge in Oceans of...

  35. [40]

    CDA- compliant section annotation of German -language discharge summaries: guideline development, annotation campaign, section classification

    Lohr, Christina, & Luther, Stephanie, & Matthies, Franz, & Modersohn, Luise, & Ammon, Danny, & Saleh, Kutaiba , & Henkel, Andreas, & Kiehntopf, Michael, & Hahn, Udo (2018). CDA- compliant section annotation of German -language discharge summaries: guideline development, annota...

  36. [41]

    Natural language processing of German clinical colorectal cancer notes for guideline - based treatment evaluation

    Becker, Matthias, & Kasper, Stefan, & Böckmann, Britta, & Jöckel, Karl-Heinz, & Virchow, Isabel (2019). Natural language processing of German clinical colorectal cancer notes for guideline - based treatment evaluation. International Journal of Medical Informatics, 127:141-146

  37. [42]

    Jahrestagung der Gesellschaft für Informatik e.V. (GI). Braunschweig, Deutschland, 16.-21. September 2012, pp. 1237-1251 (GI-Edition - Lecture Notes in Informatics, P-208)

  38. [43]

    Deep learning approaches outperform conventional strategies in de-identification of German medical reports

    Richter-Pechanski, Phillip, & Amr, Ali, & Katus, Hugo A., & Dieterich, Christoph (2019). Deep learning approaches outperform conventional strategies in de-identification of German medical reports. In: German Medical Data Sciences: Shaping Change – Creative Solutions for Innova...

  39. [44]

    Annotating German clinical documents for de - identification

    Kolditz, Tobias, & Lohr, Christina, & Hellrich, Johannes, & Modersohn, Luise, & Betz, Boris, & Kiehntopf, Michael, & Hahn, Udo (2019). Annotating German clinical documents for de - identification. In: MEDINFO 2019 – Proceedings of the 17th World Congress on Medical and Health ...

  40. [45]

    An evolutionary approach to the annotation of discharge summaries

    Lohr, Christina, & Modersohn, Luise, & Hellrich, Johannes, & Kolditz, Tobias, & Hahn, Udo (2020). An evolutionary approach to the annotation of discharge summaries. In: Digital Personalized Health and Medicine. MIE 2020 – Proceedings of the 30th Conference on Medical Informati...

  41. [46]

    Knowledge-based best of breed approach for automated detection of clinical events based on German free text digital hospital discharge letters

    König, Maximilian, & Sander, André, & Demuth, Ilja, & Diekmann, Daniel, & Steinhagen - Thiessen, Elisabeth (2019). Knowledge-based best of breed approach for automated detection of clinical events based on German free text digital hospital discharge letters. PLoS ONE, 14: #e0224916

  42. [47]

    Information extraction models for German clinical text

    Roller, Roland, & Seiffe, Laura, & Ayach, Ammer, & Möller, Sebastian, & Marten, Oliver, & Mikhailov, Michael, & Alt, Christoph, & Schmidt, Danilo, & Halleck, Fabian, & Naik, Marcel, & Duettmann, Wiebke, & Budde, Klemens (2020). Information extraction models for German clinical...

  43. [48]

    Bressem, Keno K., & Adams, Lisa C., & Gaudin, Robert A., & Tröltzsch, Daniel, & Hamm, Bernd, & Makowski, Marcus R., & Schüle, Chan-Yong, & Vahldiek, Janis L., & Niehues, Stefan M. (2020). Highly accurate classification of chest radiographic reports u sing a deep learning natur...

  44. [49]

    Automatic extraction of 12 cardiovascular concepts from German discharge letters using pre -trained language models

    Richter-Pechanski, Phillip, & Geis, Nicolas A., & Kiriakou, Christina, & Schwab, Dominic M., & Dieterich, Christoph (2021). Automatic extraction of 12 cardiovascular concepts from German discharge letters using pre -trained language models. Digital Health , 7:#10.1177/20552076...

  45. [50]

    Merkmalsextraktion aus klinischen Routinedaten mittels Text-Mining

    Grundel, Bastian, & Bernardeau, Marc -Antoine, & Langner, Holger, & Schmidt, Christoph, & Böhringer, Daniel, & Ritter, Marc, & Rosenthal, Paul, & Grandjean, Andrea, & Schulz, Stefan, & Daumke, Philipp, & Stahl, Andreas (2021). Merkmalsextraktion aus klinischen Routinedaten mit...

  46. [51]

    Using a corpus-assisted discourse studies approach to analyse gender: a case study of German radiology reports

    Irschara, Karoline (2022). Using a corpus-assisted discourse studies approach to analyse gender: a case study of German radiology reports. Gender a Výzkum, 23(2):114-139

  47. [52]

    Irschara, Karoline, & Posch, Claudia, & Waldner, Birgit, & Huber, Anna-Lena, & Glodny, Bernhard, & Gruber, Leonhard, & Mangesius, Stephanie (2022). Building the MedCorpInn corpus: issues and goals, In: Posch, Claudia & Irschara, Karoline & Rampl , Gerhard (eds.), Wort – Satz –...

  48. [53]

    arXiv preprint arXiv:2207.03885

    Roller, Roland, & Seiffe, Laura, & Ayach, Ammer, & Möller, Sebastian, & Marten, Oliver, & Mikhailov, Michael, & Alt, Christoph, & Schmidt, Danilo, & Halleck, Fabian, & Naik, Marcel G., & Duettmann, Wiebke, & Budde, Klemens (2022): A medical information extraction workbench to ...

  49. [54]

    Deep learning-based detection of psychiatric attributes from German mental health records

    Madan, Sumit, & Zimmer, Fabian Julius, & Balabin, Helena, & Schaaf, Sebastian, & Fröhlich, Holger, & Fluck, Juliane, & Neuner, Irene, & Mathiak, Klaus, & Hofmann -Apitius, Martin, & Sarkheil, Pegah (2022). Deep learning-based detection of psychiatric attributes from German men...

  50. [55]

    Patient - friendly clinical notes: towards a new text simplification dataset

    Trienes, Jan, & Schlötterer, Jörg, & Schildhaus, Hans -Ulrich, & Seifert, Christin (2022). Patient - friendly clinical notes: towards a new text simplification dataset. In: TSAR 2022 – Proceedings of the [1st] Workshop on Text Simplification, Accessibility, and Readability @ E...

  51. [56]

    A domain -adapted dependency parser for German clinical text

    Kara, Elif, & Zeen, Tatjana, & Gabryszak, Aleksandra, & Budde, Klemens, & Schmidt, Danilo, & Roller, Roland (2018). A domain -adapted dependency parser for German clinical text. In: KONVENS 2018 – Proceedings of the 14th Conference on Natural Language Processing. Vienna, Austr...

  52. [58]

    Richter-Pechanski, Phillip, & Wiesenbach, Philipp, & Schwab, Dominic M., & Kiriakou, Christina, & He, Mingyang, & Allers, Michael M., & Tiefenbacher, Anna S., & Kunz, Nicola, & Martynova, Anna, & Spiller, Noemie, & Mierisch, Julian, & Borchert, Flori an, & Schwind, Charlotte, ...

  53. [60]

    Fries, Jason Alan, & Weber, Leon, & Seelam, Natasha, & Altay, Gabriel, & Datta, Debajyoti, & Garda, Samuele, & Kang, Sunny M. S., & Su, Ruisi, & Kusa, Wojciech, & Cahyawijaya, Samuel, & Barth, Fabio, & Ott, Simon, & Samwald, Matthias, & Bach, Stephen H., & Biderman, Stella, & ...

  54. [61]

    Bressem, Keno K., & Papaioannou, Jens -Michalis, & Grundmann, Paul, & Borchert, Florian, & Adams, Lisa C., & Liu, Leonhard, & Busch, Felix, & Xu, Lina, & Loyen, Jan P., & Niehues, Stefan M., & Augustin, Moritz, & Grosser, Lennart, & Makowski, Marcus R ., & Aerts, Hugo J. W. L....

  55. [63]

    Masketeer: an ensemble-based pseudonymization tool with entity recognition for German unstructured medical free text

    Baumgartner, Martin, & Kreiner, Karl, & Wiesmüller, Fabian, & Hayn, Dieter, & Puelacher, Christian, & Schreier, Günter (2024). Masketeer: an ensemble-based pseudonymization tool with entity recognition for German unstructured medical free text. Future Internet, 16:#281

  56. [64]

    Idrissi-Yaghir, Ahmad, & Dada, Amin, & Schäfer, Henning, & Arzideh, Kamyar, & Baldini, Giulia, & Trienes, Jan, & Hasin, Max, & Bewersdorff, Jeanette, & Schmidt, Cynthia S., & Bauer, Marie, & Smith, Kaleb E., & Bian, Jiang, & Wu, Yonghui, & Schlöttere r, Jörg, & Zesch, Torsten,...

  57. [65]

    Extraction of UMLS® concepts using Apache cTakes™ for German language

    Becker, Matthias, & Böckmann, Britta (2016). Extraction of UMLS® concepts using Apache cTakes™ for German language. In: Health Informatics Meets eHealth. Predictive Modeling in Healthcare – From Prediction to Prevention. Proceedings of the 10th eHealth2016 Conference . Vienna,...

  58. [66]

    Zero -shot LLMs for named entity recognition: targeting cardiac 28 function indicators in German clinical texts

    Plagwitz, Lucas, & Neuhaus, Philipp, & Yildirim, Kemal, & Losch, Noah, & Varghese, Julian, & Büscher, Antonius (2024). Zero -shot LLMs for named entity recognition: targeting cardiac 28 function indicators in German clinical texts. In: German Medical Data Sciences 2024. Health...

  59. [67]

    Software Impacts, 11:#100212 [4 pp.]

    Frei, Johann, & Kramer, Frank (2022): GerNERMed: an open German medical NER model. Software Impacts, 11:#100212 [4 pp.]

  60. [68]

    F., & Leveling, Johannes, & Kelly, Liadh, & Goeuriot, Lorraine, & Martínez, David, & Zuccon, Guido (2013)

    Suominen, Hanna, & Salanterä, Sanna, & Velupillai, Sumithra, & Chapman, Wendy W., & Savova, Guergana K., & Elhadad, Noémie, & Pradhan, Sameer S., & South, Brett R., & Mowery, Danielle L., & Jones, Gareth J. F., & Leveling, Johannes, & Kelly, Liadh, & Goeuriot, Lorraine, & Mart...

  61. [69]

    2018 n2c2 Shared Task o n Adverse Drug Events and Medication Extraction in Electronic Health Records

    Henry, Samuel, & Buchan, Kevin, & Filannino, Michele, & Stubbs, Amber, & Uzuner, Özlem (2020). 2018 n2c2 Shared Task o n Adverse Drug Events and Medication Extraction in Electronic Health Records. Journal of the American Medical Informatics Association, 27(1):3-12

  62. [70]

    German medical named entity recognition model and data set creation using machine translation and word alignment: algorithm development and validation

    Frei, Johann, & Kramer, Frank (2023). German medical named entity recognition model and data set creation using machine translation and word alignment: algorithm development and validation. JMIR Formative Research, 7:e39077 [13 pp.]

  63. [71]

    Harnessing the power of LLMs in practice: a survey on ChatGPT and beyond

    Yang, Jingfeng, & Jin, Hongye, & Tang, Ruixiang, & Han, Xiaotian, & Feng, Qizhang, & Jiang, Haoming, & Zhong, Shaochen , & Yin, Bing, & Hu, Xia (2024). Harnessing the power of LLMs in practice: a survey on ChatGPT and beyond. ACM Transactions on Knowledge Discovery from Data, ...

  64. [72]

    GerNERMed++ : semantic annotation in German medical NLP through transfer -learning, translation and word alignment

    Frei, Johann, & Frei -Stuber, Ludwig, & Kramer, Frank (2023). GerNERMed++ : semantic annotation in German medical NLP through transfer -learning, translation and word alignment. Journal of Biomedical Informatics, 147:#104513 [8 pp.]

  65. [73]

    Sharing copies of synthetic clinical corpora without physical distribution: a case study to get around IPRs and privacy constraints featuring the German JSynCC corpus

    Lohr, Christina, & Buechel, Sven, & Hahn, Udo (2018). Sharing copies of synthetic clinical corpora without physical distribution: a case study to get around IPRs and privacy constraints featuring the German JSynCC corpus. In: LREC 2018 – Proceedings of the 11th International C...

  66. [74]

    A comprehensive survey on pretrained foundation models: a history from Bert to ChatGPT

    Zhou, Ce, & Li, Qian, & Li, Chen, & Yu, Jun, & Liu, Yixin, & Wang, Guangjing, & Zhang, Kai, & Ji, Cheng, & Yan, Qiben, & He, Lifang, & Peng, Hao, & Li, Jianxin, & Wu, Jia, & Liu, Ziwei, & Xie, Pengtao, & Xiong, Caiming, & Pei, Jian, & Yu, Philip S., & Sun, Lichao (2024). A com...

  67. [75]

    Annotated dataset creation through large language models for non-English medical NLP

    Frei, Johann, & Kramer, Frank (2023). Annotated dataset creation through large language models for non-English medical NLP. Journal of Biomedical Informatics, 145:#104478 [9 pp.]

  68. [76]

    GraSCCo : the first publicly shareable, multiply-alienated German clinical text corpus

    Modersohn, Luise, & Schulz, Stefan, & Lohr, Christina, & Hahn, Udo (2022). GraSCCo : the first publicly shareable, multiply-alienated German clinical text corpus. In: German Medical Data Sciences 2022 – Future Medicine: More Precise, More Integrative, More Sustainable! Proceed...

  69. [77]

    Privacy risks of general -purpose language models

    Pan, Xudong, & Zhang, Mi, & Ji, Shouling, & Yang, Min (2020). Privacy risks of general -purpose language models. In: SP 2020 – Proceedings of the 2020 IEEE Symposium on Security and Privacy. San Francisco, California, USA, 18-21 May 2020, pp. 1314-1331

  70. [78]

    Lernen, Wissen, Daten, Analysen

    Şerbetçi, Oğuz, & Leser, Ulf (2023). Applicability of models trained on generated clinical German datasets on out -domain data. In: LWDA 2023 – Proceedings of the Conference on “Lernen, Wissen, Daten, Analysen.” Marburg, Germany, October 9-11, 2023, pp. 521-525

  71. [79]

    Clinical text anonymization, its influence on downstream NLP tasks and the risk of re -identification

    Larbi, Iyadh Ben Cheikh, & Burchardt, Aljoscha, & Roller, Roland (2023). Clinical text anonymization, its influence on downstream NLP tasks and the risk of re -identification. In: Proceedings of the Student Research Workshop @ EACL 2023. [Dubrovnik, Croatia,] May 2 -4, 2023 (H...

  72. [80]

    Extracting training data f rom large language models

    Carlini, Nicholas, & Tramèr, Florian, & Wallace, Eric, & Jagielski, Matthew, & Herbert-Voss, Ariel, & Lee, Katherine, & Roberts, Adam, & Brown, Tom, & Song, Dawn, & Erlingsson, Úlfar, & Oprea, Alina, & Raffel, Colin (2021). Extracting training data f rom large language models....

  73. [81]

    Semantic annotation for concept -based cross -language medical information retrieval

    Volk, Martin, & Ripplinger, Bärbel, & Vintar, Špela, & Buitelaar, Paul, & Raileanu, Diana, & Sacaleanu, Bogdan (2002). Semantic annotation for concept -based cross -language medical information retrieval. International Journal of Medical Informatics, 67(1-3):79-112

  74. [82]

    Brown, Ralf D. (2002). Corpus -driven splitting of compound words. In: Proceedings of the 9th Conference on Theoretical and Methodological Issues in Machine Translation of Natural Languages: Papers. Keihanna, Japan, March 13-17, 2002, #3 (3:1–3:10)

  75. [83]

    Customizing parallel corpora at the document level

    Rogati, Monica, & Yang, Yiming (2004). Customizing parallel corpora at the document level. In: ACL '04 – Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics: Interactive Poster and Demonstration Sessions. Barcelona, Spain, 21 –26 July 2004, ...

  76. [84]

    Cross -language MeSH indexing using morpho-semantic normalization

    Markó, Kornél G., & Daumke, Philipp, & Schulz, Stefan, & Hahn, Udo (2003). Cross -language MeSH indexing using morpho-semantic normalization. In: AMIA 2003 – Proceedings of the 2003 Annual Symposium of the American Medical Informatics Association. Biomedical and Health Informa...

  77. [85]

    Collaboratively annotating multilingual parallel corpora in the biomedical domain: some Mantras

    Hellrich, Johannes, & Clematide, Simon, & Hahn, Udo, & Rebholz -Schuhmann, Dietrich (2014). Collaboratively annotating multilingual parallel corpora in the biomedical domain: some Mantras. In: LREC 2014 – Proceedings of the 9th International Conference on Language Resources an...

  78. [86]

    Revising the compositional method for terminology acquisition from comparable corpora

    Morin, Émmanuel, & Daille, Béatrice (2012). Revising the compositional method for terminology acquisition from comparable corpora. In: COLING 2012 – Proceedings of the 24th International Conference on Computational Linguistics. Mumbai, India, 8-15 December 2012, pp. 1797-1810

  79. [87]

    Version 1.0

    Bojar, Ondřej, & Haddow, Barry, & Mareček, David, & Sudarikov, Roman, & Tamchyna, Aleš, & Variš, Dušan (2017): HimL D1.1 : Report on Building Translation Systems for Public Health 30 Domain. Version 1.0. (European Union’s Horizon 2020 Research and Innovation Programme under gr...

  80. [88]

    A multilingual gold -standard corpus for biomedical concept recognition: the Mantra GSC

    Kors, Jan A., & Clematide, Simon, & Akhondi, Saber A., & van Mulligen, Erik M., & Rebholz - Schuhmann, Dietrich (2015). A multilingual gold -standard corpus for biomedical concept recognition: the Mantra GSC . Journal of the American Medical Informatics Association , 22(5):948-956

  81. [89]

    On the construction of multilingual corpora for clinical text mining

    Villena, Fabián, & Eisenmann, Urs, & Knaup, Petra, & Dunstan, Jocelyn, & Ganzinger, Matthias (2020). On the construction of multilingual corpora for clinical text mining. In: Digital Personalized Health and Medicine. MIE 2020 – Proceedings of the 30th Conference on Medical Inf...

  82. [90]

    Leveraging Wikipedia knowledge to classify multilingual biomedical documents

    Mouriño García, Marcos Antonio, & Pérez Rodríguez, Roberto, & Rifón, Luis Anido (2018). Leveraging Wikipedia knowledge to classify multilingual biomedical documents. Artificial Intelligence in Medicine, 88:37-57

  83. [91]

    Borchert, Florian, & Lohr, Christina, & Modersohn, Luise, & Witt, Jonas, & Langer, Thomas, & Follmann, Markus, & Gietzelt, Matthias, & Arnrich, Bert, & Hahn, Udo, & Schapranow, Matthieu- P. (2022). GGPOnc 2.0 —the German Clinical Guideline Corpus for Oncology: curation workflo...

  84. [92]

    Borchert, Florian, & Lohr, Christina, & Modersohn, Luise, & Langer, Thomas, & Follmann, Markus, & Sachs, Jan Philipp, & Hahn, Udo, & Schapranow, Matthieu -P. (2020). GGPOnc: a corpus of German medical text with rich metadata based on clinical practice guidelines. In: LOUHI 202...

  85. [93]

    Sector: a neural model for coherent topic segmentation and classification

    Arnold, Sebastian, & Schneider, Rudolf, & Cudré -Mauroux, Philippe, & Gers, Felix A., & Löser, Alexander (2019). Sector: a neural model for coherent topic segmentation and classification. Transactions of the Association for Computational Linguistics, 7:169-184

  86. [94]

    Crit ical assessment of transformer-based AI models for German clinical notes

    Lentzen, Manuel, & Madan, Sumit, & Lage-Rupprecht, Vanessa, & Kühnel, Lisa, & Fluck, Juliane, & Jacobs, Marc, & Mittermaier, Mirja, & Witzenrath, Martin, & Brunecker, Peter, & Hofmann - Apitius, Martin, & Weber, Joachim, & Fröhlich, Holger (2022). Crit ical assessment of trans...

  87. [95]

    Tracking and analyzing recent developments in German-language online press in the face of the coronavirus crisis: cOWIDplus Analysis and cOWIDplus Viewer

    Wolfer, Sascha, & Koplenig, Alexander, & Michaelis, Frank, & Müller -Spitzer, Carolin (2020). Tracking and analyzing recent developments in German-language online press in the face of the coronavirus crisis: cOWIDplus Analysis and cOWIDplus Viewer. International Journal of Cor...

  88. [96]

    From witch’s shot to music making bones: resources for medical laymen to technical language and vice versa

    Seiffe, Laura, & Marten, Oliver, & Mikhailov, Michael, & Schmeier, Sven, & Möller, Sebastian, & Roller, Roland (2020). From witch’s shot to music making bones: resources for medical laymen to technical language and vice versa. In: LREC 2020 – Proceedings of the 12th Internatio...

  89. [97]

    Fang-Covid: a new large -scale benchmark dataset for fake news detection in German

    Mattern, Justus, & Qiao, Yu, & Kerz, Elma, & Wiechmann, Daniel, & Strohmaier, Markus (2021). Fang-Covid: a new large -scale benchmark dataset for fake news detection in German. In: FEVER 2021 – Proceedings of the 4th Workshop on Fact Extraction and VERification @ EMNLP

  90. [98]

    Investigating label suggestions for opinion mining in German Covid-19 social media

    Beck, Tilman, & Lee, Ji -Ung, & Viehmann, Christina, & Maurer, Marcus, & Quiring, Oliver, & Gurevych, Iryna (2021). Investigating label suggestions for opinion mining in German Covid-19 social media. In: ACL-IJCNLP 2021 – Proceedings of the 59th Annual Meeting of the Associati...

  91. [99]

    A dataset for pharmacovigilance in German, French, and Japanese: annotating adverse drug reactions across languages

    Raithel, Lisa, & Yeh, Hui -Syuan, & Yada, Shuntaro, & Grouin, Cyril, & Lavergne, Thomas, & Névéol, Aurélie, & Paroubek, Patrick, & Thomas, Philippe E., & Nishiyama, Tomohiro, & Möller, Sebastian, & Aramaki, Eiji, & Matsumoto, Yuji, & Roller, Roland, & Zweigenbaum, Pierre (2024...

  92. [100]

    Automatic identification of COVID -19-related narratives in German Telegram channels and chats

    Heinrich, Philipp, & Blombach, Andreas, & Dang, Bao Minh Doan, & Zilio, Leonardo, & Havenstein, Linda, & Dykes, Nathan, & Evert, Stephanie, & Schäfer, Fabian (2024). Automatic identification of COVID -19-related narratives in German Telegram channels and chats. In: LREC-COLING...

  93. [101]

    Cross -lingual approaches for the detection of adverse drug reactions in German from a patient’s perspective

    Raithel, Lisa, & Thomas, Philippe E., & Roller, Roland, & Sapina, Oliver, & Möller, Sebastian, & Zweigenbaum, Pierre (2022). Cross -lingual approaches for the detection of adverse drug reactions in German from a patient’s perspective. In: LREC 2022 – Proceedings of the 13th In...

  94. [102]

    Between Plain Language and Einfache Sprache: a Corpus Analysis of Layperson Summaries of Clinical Trials in English, German, and Italian

    Pedrini, Giulia (2024). Between Plain Language and Einfache Sprache: a Corpus Analysis of Layperson Summaries of Clinical Trials in English, German, and Italian. Frank & Timme Verlag

  95. [103]

    Creating ontology-annotated corpora from Wikipedia for medical named-entity recognition

    Frei, Johann, & Kramer, Frank (2024). Creating ontology-annotated corpora from Wikipedia for medical named-entity recognition. In: BioNLP 2024 – Proceedings of the 23rd Meeting of the ACL Special Interest Group on Biomedical Natural Language Processing: Workshop and Shared Tas...

  96. [104]

    HealthFC: verifying health claims with evidence-based medical fact-checking

    Vladika, Juraj, & Schneider, Phillip, & Matthes, Florian (2024). HealthFC: verifying health claims with evidence-based medical fact-checking. In: LREC-COLING 2024 – Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Eval...

  97. [105]

    Privacy aware question -answering system for online mental health risk assessment

    Chhikara, Prateek, & Pasupulety, Ujjwal, & Marshall, John, & Chaurasia, Dhiraj, & Kumari, Shweta (2023). Privacy aware question -answering system for online mental health risk assessment. In: BioNLP 2023 – Proceedings of the 22nd Workshop on Biomedical Language Processing (Bio...

  98. [106]

    Adams, Lisa C., & Truhn, Daniel, & Busch, Felix, & Kader, Avan, & Niehues, Stefan M., & Makowski, Marcus R., & Bressem, Keno K. (2023). Leveraging GPT-4 for post hoc transformation of free -text radiology reports into structured reporting: a multilingual feasibility study. Rad...

  99. [107]

    Don’t quote me: reverse identification of research participants in social media studies

    Ayers, John W., & Caputi, Theodore L., & Nebeker, Camille, & Dredze, Mark (2018). Don’t quote me: reverse identification of research participants in social media studies. npj Digital Medicine, 1:#30 [2 pp.]

  100. [108]

    On the impact of cross - domain data on German language models

    Dada, Amin, & Chen, Aokun, & Peng, Cheng, & Smith, Kaleb E., & Idrissi -Yaghir, Ahmad, & Seibold, Constantin, & Li, Jianning, & Heiliger, Lars, & Friedrich, Christoph M., & Truhn, Daniel, & Egger, Jan, & Bian, Jiang, & Kleesiek, Jens, & Wu, Yonghui (2023). On the impact of cro...

  101. [109]

    Viability of open large language models for clinical documentation in German health care: real -world model evaluation stud y

    Heilmeyer, Felix, & Böhringer, Daniel, & Reinhard, Thomas, & Arens, Sebastian, & Lyssenko, Lisa, & Haverkamp, Christian (2024). Viability of open large language models for clinical documentation in German health care: real -world model evaluation stud y. JMIR Medical Informati...

  102. [110]

    Few -shot and prompt training for text classification in German doctor' s letters

    Richter-Pechanski, Phillip, & Wiesenbach, Philipp, & Schwab, Dominic M., & Kiriakou, Christina, & He, Mingyang, & Geis, Nicolas A., & Frank, Anette, & Dieterich, Christoph (2023). Few -shot and prompt training for text classification in German doctor' s letters. In: Caring is ...

  103. [111]

    Reusable templates and guides for documenting datasets and models for natural language processing and generation: a case study of the HuggingFace and GEM data and model cards

    McMillan-Major, Angelina, & Osei, Salomey, & Rodriguez, Juan Diego, & Ammanamanchi, Pawan Sasanka, & Gehrmann, Sebastian, & Jernite, Yacine (2021). Reusable templates and guides for documenting datasets and models for natural language processing and generation: a case study of...

  104. [112]

    noi” indicates that “no information

    Gebru, Timnit, & Morgenstern, Jamie, & Vecchione, Briana, & Wortman Vaughan, Jennifer, & Wallach, Hanna M., & Daumé III, Hal, & Crawford, Kate (2021). Datasheets for datasets. In: Communications of the ACM, 64(12):86-92. 33 Supplementary Material A. Tables of German-Language C...

  105. [113]

    DoPA Meter : a tool suite for metrical document profiling and aggregation

    Lohr, Christina, & Hahn, Udo (2023). DoPA Meter : a tool suite for metrical document profiling and aggregation. In: EMNLP 2023 – Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Singapore, Singapore, December 6-10, ...

  106. [116]

    – 2004 noi (~6,500 sentences) 100k Various clinical report types (discharge, pathology, histol- ogy, and surgery reports), a medical textbook, and Web documents taken from a consumer health care portal (netdoktor) Annotation Types Sentence & token splits, parts of speech (PoS)...

  107. [117]

    – 2007 ~ 30,000 noi Mainly discharge letters, but also surgical reports, immunodermatological findings and other narrative reports of clinical results (dermatology) none ⚫ Spat-08

  108. [118]

    – 2008 1,500 subset from 18k noi 26 clinical document types from 8 medical fields (vascular & casualty surgery, internal medicine, neurology, anaesthesia, intensive care, radiology, physiotherapy) Annotation Types Classification into document types and medical fields Entity No...

  109. [119]

    – 2011 3,542 84k Pathology reports Annotation (Automatic) rewriting of fully capitalized texts as mixed capitalized and lower-cased texts (following German orthography rules) Entity Normalization: n/a Annotation Guideline: n/a IAA Measurement: n/a ⚫ Fette-12

  110. [120]

    – 2012 544 subset from 193k noi Clinical reports from 5 clinical domains (echocardiography, ECG, lung function, X-ray thorax, bicycle stress test) Annotation Types Automatic extraction of attribute-value pairs from the 5 clinical domains Entity Normalization: Y (local terminol...

  111. [121]

    – 2013 174 subset from 2,7k 28k Radiology reports (lymphoma) Annotation classification into “pathological“ or “non- pathological“ sentences Entity Normalization: N Annotation Guideline: N IAA Measurement: N ⚫ Bretschnei- der-14

  112. [122]

    – 2014 2,713 347k Radiology reports (lymphoma) Annotation Types (Annotated items) Automatic concept annotation with (ma- chine-translated) German RadLex terms (: 148k tokens (= 42.6 %)) Entity Normalization: Y (RadLex) Annotation Guideline: n/a IAA Measurement: n/a ⚫ 10 https...

  113. [123]

    – 2015 140 subset from 69k/70k noi (Transthoracic) echocardio- graphy reports Annotation Types (Annotated items) Automatic extraction of 440 attribute-value pairs from the echocardiography domain (e.g., attributes: Aortic Valve, Mitral Valve, Tricuspid Valve, regurgitation (ao...

  114. [124]

    patient”, “surgery

    – 2016 450 subset from 22,4k 5,8m 266k 125,9m Operative reports (digestive tract) (Fragments of) newspaper articles with medical content extracted from DWDS (Digitales Wörterbuch der Deutschen Sprache) Annotation Types Diagnoses, Procedures Entity Normalization: Y (ICD for dia...

  115. [125]

    – 2016 737 (para- graphs only) noi Main diagnosis paragraphs split from discharge sum- maries (oncology: multiple myeloma) Annotation Types (Annotated items) Diagnosis (0,9k), State of Disease (specific data elements characteristic for multiple myeloma; 7,7k) (: 8,6k) Entity ...

  116. [126]

    However, upon inspection of the supplement, the corpus was not listed

    – 2016 118 + 1,607 = 1,725 90k + 68k = 158k Discharge summaries & clinical notes (nephrology) Annotation Types (Annotated items) 23 entity types, grouped into 7 major categories: [Time: Date, Temporal Course; Person/Body: Person, Body Part, Tissue, Body Fluid, Localization; Pr...

  117. [127]

    – 2016 1,696 noi Discharge letters (dermatology) Annotation (Annotated items) abbreviated word forms (: 2,3 k) Entity Normalization: N Annotation Guideline: N IAA Measurement: Y ⚫ Cotik-16

  118. [128]

    – 2016 8 + 175 = 183 6,2k + 6,7k = 12,9k Discharge summaries & clinical notes (nephrology) Annotation Types (Annotated items) Negation (0,4k) & Factuality: affirmed (0,6k), speculated (<0,1k) of Findings (: 1,1k) Entity Normalization: Y (UMLS) Annotation Guideline: N (schema ...

  119. [129]

    – 2017 1,400 [subset from 4,671 + 2,804 + 1,008 + 6,223 = 14,706 ~5,000k ~50,000k pathology reports medical reports (Gynecology) operative reports (Gynecology) radiology reports Annotation Types (Annotated items) 9 Personally Identifiable Information (PII) types [Name, Age, Co...

  120. [130]

    – 2017 30,000 noi discharge summaries (cardiology) none (200 abbreviations) ⚫ Roller-18

  121. [131]

    – 2018 626 (subset from [33]) 26,5k* (*estimated from averages) Clinical notes & discharge summaries (nephrology) Annotation Types (Annotated items) 8 named entity types: [Medical Condition: Symptom, Finding, Diagnosis (2,5k), Treatment (1,7k), State of Health (1,5k), Medi- ca...

  122. [132]

    – 2017 Radiology reports (chest)

  123. [133]

    Semi-automatic acquisition of a local clinical terminology composed of 258 attributes for processing radiology reports

  124. [134]

    ⚫ 38 100 subset from 3,000 noi 3

    Value categories for attributes: negation, laterality (right, left, both sides), location, degree of severity, condition-after, & progression note. ⚫ 38 100 subset from 3,000 noi 3. Automatic extraction of 735 attribute- value pairs. Entity Normalization: Y (local terminology)...

  125. [135]

    – 2018 2,360 (from 3 different clinical sites) 3,997k (mostly) Discharge sum- maries, few transfer letters Annotation Types 1 Medication entity + 5 Medication relation types [Medication/Drug: Dosage, Mode, Frequency, Duration, Medical Reason] Entity Normalization: N Annotation...

  126. [136]

    – 2018 1,106 subset from 3000PA 1,500k (mostly) Discharge sum- maries, few transfer letters Annotation Types (Annotated items) 18 Section Heading types [Salutation (12,9k), Anamnesis (0,6k): Patient history (6,0k) & Family history (<0,1k), Diagnosis (4,0k): Admission diagnosis...

  127. [137]

    – 2019 820 + 817 + 107 + 326 + 20 + 423 = 2,513 subset from 5,506 noi (Mixed) clinical reports: medical reports, radiology reports, microbiology reports, pathology reports, virology reports, and tumor board protocols Annotation Types (Annotated items) 11 named entity types, at...

  128. [138]

    – 2019 1,106 subset from 3000PA 1,400k (mostly) Discharge sum- maries, few transfer letters Annotation Types (Annotated items) 13 PII types [Age (0,5k), Contact (phone, email, URL; 0,6k), Date (20,6k), Birthdate (1,1k), ID (patient, e.g., EPR number; 0,4k; Typist; 0,7k), Locat...

  129. [139]

    – 2019 113 107k Medical reports (cardiology) Annotation Types (Annotated items) 8 PII types [person, location, date, phone, organization, title, salutation, zip code]: (: 5,2k) Entity Normalization: N Annotation Guideline: N IAA Measurement: N ⚫ König-19

  130. [140]

    proton-pump inhibitor use – osteoporosis

    – 2019 1,982 2,001k Discharge summaries (osteoporosis) Annotation Types (Annotated items) 1 Drug-Disease relation [“proton-pump inhibitor use – osteoporosis”] (2,0k), including concept recognition for PPI and osteoporosis [extracted from the hospital-internal study database as...

  131. [141]

    – 2020 1,106 subset from 3000PA 1,500k (mostly) Discharge sum- maries, few transfer letters Annotation Types (Annotated items) 3 named entity types [Diagnosis (55k), Findings (155k), Symptoms (8k)] & 3 attributes of NE types [Time (previous, recurrent, uncertain), Modality (su...

  132. [142]

    – 2020 5,783 subset from 3,8m radiology reports used for model pre- training 399k* (*estimated from averages) 416m Radiology reports (chest radiographs, chest CT scans) Annotation Types (Annotated items) 9 Finding types, incl. Medical Devices [Congestion (1,5k), Opacity (e.g.,...

  133. [143]

    – 2020 118 + 1,607 = 1,725 (data taken from [33]) 90k + 68k = 158k (data taken from [33]) Discharge summaries & clinical notes (nephrology) Annotation Types (Annotated items) 17 Named entity types [Medical condition (11,6k), Measurement (5,9k), Body part (5,4k), Treatment (5,3...

  134. [144]

    – 2021 40,485 noi Discharge summaries (ophthalmology) Annotation Types (Annotated items) Extraction of Visus (visual acuity; 47,6k), Tensio (intraocular pressure; 40,4k) and Diagnoses for macular diseases (3,2k) Entity Normalization: Y (SNOMED-CT) Annotation Guideline: n/a (ex...

  135. [145]

    – 2021 200 set of 11,4k shuffled sentences 90k Discharge summaries (oncology: hepatocellular carcinoma or melanoma) from two national hospitals (Berlin, Tübingen) Annotation Types (Annotated items) Section Headings Entity Normalization: N Annotation Guideline: Y (see Supplemen...

  136. [146]

    – 2021 204 subset from ~200,000 382k subset from ~218m Discharge summaries (cardiology) Annotation Types (Annotated items) 12 cardiovascular concepts [angina pectoris (0,2k), dyspnea (0,2k), nycturia (0,1k), edema (0,1k), palpitation (0,1k), vertigo (0,1k), syncope (0,2k), art...

  137. [147]

    – 2022 150 510 subset from 30k noi noi Discharge summaries (Psychiatry: Mental Status Examination (MSE) reports) Annotation Types (Annotated items) psychiatric attributes (3,4k), normal (1,7k) and pathological assessments (1,3k), and grounding of pathological assessments in th...

  138. [148]

    – 2022 720 13,4k* (*estimated from averages) Physicians’ justifications supporting their estimated likelihood of future possible negative patient outcomes after transplantation (kidney disease endpoints: rejection, death-censored graft loss, and infection within the next 90 da...

  139. [149]

    – 2022 (updated version of [47]) 61 + 1,300 = 1,361 57,2k + 54,2k = 111,4k Discharge summaries & clinical notes (nephrology: kidney transplantations) Annotation Types (Annotated items) 17 Named entity types [Medical condition (9,0k), Measurement (5,4k), Body part (3,4k), Treat...

  140. [150]

    – 2022 851 327k (expert) 463k (simplified) : 790k pathology reports of sarcoma patients Parallel corpus of expert-level and layman- directed, patient-friendly parallel versions of pathology reports ⚫ (efforts for data sharing under way) Cardio:DE

  141. [151]

    – 2023 500 993k clinical notes and reports (cardiology: 311 in-patient & 172 out-patient letters, and 17 letters of the cardiac emergency room) Annotation Types (Annotated items) 14 named entity types for section headings [salutation (0.5k), anamnesis (1,5k), diagnosis (admiss...

  142. [152]

    – 2023 150 < 500 30 71k 800k 1,877k Discharge summaries (oncology) from Bronco Discharge summaries (cardiology) from Cardio:DE Clinical guidelines (oncology) from GGPOnc 2.0 Harmonizing approach for four German medical corpora (Bronco, Cardio:DE, GGPOnc 2.0, GraSCCo 1.0) using...

  143. [153]

    – 2023 > 150k Clinical reports covering 4 medical areas (cardiology, pathology, pharmacy, and neurology) from 6 different clinical sites (e.g., discharge summaries, findings reports) Annotation Types Multiple annotation layers Entity Normalization: Y (SNOMED CT, ICD- 10; plann...

  144. [154]

    – 2024 J: 1,106 A: 1,715 L: 3,823 =6,644 J: 1,8m A: 1,7m L: 3,8m = 7,3m Clinical reports from 3 different clinical sites (Jena, Aachen, and Leipzig) – (mainly discharge summaries and transfer reports) Automatic tagging with token and sentence boundaries (silver standard) Ann...

  145. [155]

    – 2024 2,000 + 2,000 + 2,000 = 6,000 subset from 3,7m radiology reports 4,369 62 63,884 11,322 12,139 257,999 330,994 373,421 7,486 3,639  4,723,010 854k* (*estimated) 520,718k + 1,194k + 44k + 12.299k + 9,324k + 1,984k + 259,285k + 186,201k + 69,639k + 90,381k + 2,800k  1,1...

  146. [156]

    – 2024 100 + 100 + 100 = 300 noi ophthalmologic physicians’ letters from three different German hospitals Annotation Types 771 + 1226 + 809 = 2,806 diagnoses (manually curated silver standard composed of ICD-10 codes) Entity Normalization: Y (ICD-10) Annotation Guideline: N IA...

  147. [157]

    HerzMobil

    – 2024 25,023k 29,273 (question- answer pairs) 3,060,845 k noi different types of clinical reports, clinical notes, and doctor’s letters question-answer pairs created from 1,223 radiology reports of brain CT scans none one custom question for every third report (covering ~400 ...

  148. [158]

    – 2024 35,579 1.245k* (estimated from mean length) Clinical notes Annotation Types (Annotated items) 9 PII types: First and last Names of Health- care professional (21,9k), Patient (16,1k), other Person (7,0k), Medical site (3,2k), Website URL (< 0,0k), Email address (~0,0k), ...

  149. [159]

    – 2024 498 noi Cardiac magnetic resonance imaging (MRI) reports Annotation Types (Annotated items) Attribute-value pairs for 14 cardiac function indicators, such as ejection fraction or volumes for the left and right ventricle Entity Normalization: N Annotation Guideline: N IA...

  150. [160]

    – 2016 61 + 54 + 42 + 42 = 199 noi (Mixed) clinical reports: discharge summaries, ECG reports, echo reports, and radiology reports (taken from the ShARe/CLEF eHealth 2013 Shared Task 1 (MIMIC II) [66] → automatic translation from English to German using Google Translate) Annot...

  151. [161]

    – 2023 404 367k discharge summaries [taken from the n2c2 2018 Shared Task Track 2 (MIMIC III) [69] → automatic translation from English to German using a pretrained neural machine translation model from fairseq & alignments from Awesome- Align Annotation Types (Annotated items...

  152. [162]

    – 2024 noi 6,000k (abstracts) 695,000k 1,700,000 k MIMIC III clinical notes & PubMed articles automatic translation from English to German using a pretrained neural machine translation model from fairseq none none ◆ Translation -based model25 Table 2: Translated Real Clinical ...

  153. [163]

    – 2018 399 + 468 = 867 193k + 119k = 313k Operative reports (orthopedics, trauma & general surgery) Case reports/descriptions (emergency and internal medicine, general surgery, anesthetics, ophthalmology) [taken from e-book versions of medical textbooks, manually generated] An...

  154. [164]

    – 2022 63 44k Discharge summaries [manually generated from real mixed- domain clinical (hospitals in Germany and Austria) and published Web resources] Case reports [from Open Access journals] none ✓ (public)27 Frei-23

  155. [165]

    – 2023 (9,845 sen- tences) 121k (sentences automatically generated via few-shot prompts (12 manually created sentences) from a large language model: GPT NeoX from EleutherAI) Annotation Types (Annotated items) 3 named entity types (automatically generated silver standard) [Med...

  156. [166]

    – 2024 399 200k Operative reports (orthopedics, trauma & general surgery) Case reports/descriptions (emergency and internal medicine, general surgery, anesthetics, ophthalmology) [taken from e-book versions of medical textbooks, manually generated] Annotation Types (Annotated ...

  157. [167]

    – 2024 63 44k Discharge summaries [manually generated from real mixed- domain clinical (hospitals in Germany and Austria) and published Web resources] Case reports [from Open Access journals] Annotation Types (Annotated items) Named Entities and Semantic Relations, Temporal Re...

  158. [168]

    – 2024 63 44k Discharge summaries [manually generated from real mixed- domain clinical (hospitals in Germany and Austria) and published Web resources] Case reports Annotation Types (Annotated items) 19 PII types [Name – Patient (0,2k), Doctor (0,2k), Title (0,1k), etc., ✓ (pub...

  159. [169]

    – 2002 531,690 (journal article titles) ~ [4,000- 5,000]k Parallel corpus (English–German) of paired journal article titles retrieved from PubMed none ✓ MuchMore

  160. [170]

    – 2002 ~ 9,000 (abstracts for each language) ~ 1,000k Parallel corpus (English–German) of abstracts from 41 medical journals hosted at the Springer Web site covering various medical sub- domains (e.g. neurology, radiology) Annotation Types Sentences, tokens, parts of speech (P...

  161. [171]

    – 2003 5,271 ~ 910k Abstracts of German medical jour- nal publications, available from an online library for medicine (SpringerLink) Annotation Types Automatically derived index terms Entity Normalization: Y (local dictionary linked with MeSH terms) Annotation Guideline: n/a I...

  162. [172]

    – 2004 9,640 ~ 450k [30k sentences] ~ 5,500k [549k sentences] titles plus abstracts of medical journal articles from Springer, each in German (& in English); paired titles of medical journal articles (from PubMed) none ✓ FraMed

  163. [173]

    – 2004 noi (~6,500 sentences) 100k Various clinical report types (discharge, pathology, histology, and surgery reports), a medical textbook, and Web documents taken from a consumer health care portal (netdoktor) Annotation Types Sentence & token splits, parts of speech (PoS) E...

  164. [174]

    breast cancer

    – 2012 103 118 (English) 130 (French) 220k 265k (English) 265k (French) Multilingual comparable corpus (English, French, German) from scientific paper websites, with hits for “breast cancer” (‘cancer du sein’ in French and ‘Brustkrebs’ in German) in titles & keyword sections o...

  165. [175]

    – 2014 EFGSD: 4,255k 719k + 141k + 121k = 981k EFGSD: 60,424k 5,997k + 2,100k + 5,194k =13,291k Multilingual parallel corpus: English, French, German, Spanish, Dutch), including Medline titles (PubMed) Drug labels (EMEA) Patent claims (EPO) Annotation Types (Annotated items)...

  166. [176]

    – 2015 EFGSD: 1,450 + 100 + 100 + 50 = 250 (Subset of [85]) EFGSD: 29,329 947 + 1,956 + 3,117 = 6,020 (Subset of [85]) Multilingual parallel corpus: English, French, German, Spanish, Dutch), including Medline titles (PubMed) Drug labels (EMEA) Patent claims (EPO) Annotation ...

  167. [177]

    – 2017 781k + 33k + 1,848k = 2,662k > 60,000k (estimated) Multilingual parallel corpus (English, German), including EMEA (European Medicines Agency) documents MuchMore segments Marec patent documents none ✓ (upon request) EFSG- UVigoMED ML– UVigoMED

  168. [178]

    – 2018 2,130 all: 19,210 3,147 all: 23,647 ~ 500k Multi-lingual corpus: Medline/PubMed abstracts (English, French, Spanish, German) about 26 types of Diseases Wikipedia articles (German, English, French, Spanish, Italian, Galician, Romanian, Slovene, and Icelandic) about Hum...

  169. [179]

    – 2020 59,539 all: 93,969 20,438k all: 83.869k (Web-scraped) Multilingual corpus (German, English, Spanish), with a 63% share of German-language medical full-text articles/abstracts none ✓ Zenodo30 GGPOnc 1.0

  170. [180]

    ✓ (DUA)31 GGPOnc 2.0

    – 2020 25 (4.2k annotated text segments) Subset of 8.4k text segments 664k Subset of 1,340k (all) Clinical Practice Guidelines of the German Cancer Society (oncology) Annotation Types (Annotated items) 7 named entity types [UMLS Semantic Groups: Anatomical Structure, Chemicals...

  171. [181]

    – 2022 30 (5k annotated text segments) (Subset of 10.2k text segments) (Superset of [90]) 830k * (estimated) Subset of 1,877k (all) Clinical Practice Guidelines of the German Cancer Society (oncology) Annotation Types (Annotated items) 3 named entity types [SNOMED-CT top-level...

  172. [183]

    – 2022 50 32k Discharge summaries (neurology) [because of the small number of tokens & documents the clinical portion of this corpus is excluded from deeper consideration] Annotation Types Section Headings (8 categories) [Header and Footer, Personal Data, Diagnoses, Anamneses,...

  173. [184]

    – 2024 2,000 + 2,000 + 2,000 = 6,000 subset from 3,7m radiology reports 4,369 62 63,884 11,322 12,139 257,999 330,994 373,421 854k* (*estimated) 520,718k + 1,194k + 44k + 12.299k + 9,324k + 1,984k + 259,285k + 186,201k + 69,639k Radiology reports (chest radiographs, chest CT s...

  174. [185]

    – 2024 ~ 6,000k 6,000k (abstracts) 695,000k 1,700,000 k MIMIC III clinical notes & PubMed articles automatic translation from English to German using a pretrained neural machine translation model from fairseq none none ◆ Translation -based model34 Table 4: Close Domain Proxies...

  175. [186]

    – 2004 noi (~6,500 sentences) 100k Various clinical report types (discharge, pathology, histology, and surgery reports), a medical textbook, and Web documents taken from a consumer health care portal (netdoktor) Annotation Types Sentence & token splits, parts of speech (PoS) E...

  176. [187]

    patient”, “surgery

    – 2016 450 subset from 22,4k 5,8m 266k 125,9m Operative reports (digestive tract) (Fragments of) newspaper arti- cles with medical content extract- ed from DWDS (Digitales Wörter- buch der Deutschen Sprache) Annotation Types Diagnoses, Procedures Entity Normalization: Y (ICD f...

  177. [188]

    – 2018 2,130 all: 19,210 3,147 all: 23,647 ~ 500k Multi-lingual corpus: Medline/PubMed abstracts (German, English, French, Spanish) about 26 types of Diseases Wikipedia articles (German, English, French, Spanish, Italian, Galician, Romanian, Slovene, and Icelandic) about Hum...

  178. [189]

    – 2019 2,3k (Diseases, German only) subset from 38k ~2,000k* (*estimate) (45.7 sentences/ article) Wikipedia articles (German, En- glish) about Diseases (and Cities) Annotation Types (Annotated items) 25 topic classes [Diagnosis, Treatment, Symptoms, Mechanism, Medication, Cla...

  179. [190]

    – 2020 2k (kidney diseases) 2k (stomach and intestines) (: 4k) 204k (kidney diseases) 235k (stomach and intestines) (: 439k) Threads from the German-lan- guage patient forum Med1 Annotation Types (Annotated items) Paraphrase equivalence links between medical expert (: 1,7k)...

  180. [191]

    – 2020 noi 13,649k RSS feeds about the corona-virus pandemic from 13 German news- papers and 3 non-print outlets: print: Focus Online, Frankfurter Allgemei- ne Zeitung, Frankfurter Rundschau, Süd- deutsche Zeitung, Neue Zürcher Zeitung, SpiegelOnline, Standard, tageszeitung (T...

  181. [192]

    – 2021 3k subset from 238k (~555k*) (*estimate) Tweets (selected by search terms, such as Corona, Pandemic, Covid 19, Social distance, etc.) Annotation Types (Annotated items) 4-category label system indicating the tweet’s stance towards governmental measures taken against the...

  182. [193]

    – 2021 28,1k + 13,2k = 41,3k (complete news arti- cle / tweet) 22,000k* + 10,600k* = 32,600k* (*estimate) Real news articles and tweets Fake news articles and tweets (selected by the query terms: Corona, Covid, Infektion, Lockdown, Impfen, Impfung, Impfstoff) Automatically gen...

  183. [194]

    – 2022 101 (complete forum post) subset from 4,169 11,k 463k* (*estimate) Threads about Adverse Drug Reactions (ADRs) from the German-language patient forum Lifeline Annotation Types (Annotated items) (Binary) categorization of documents into those reporting ADRs (101 posts) a...

  184. [195]

    – 2022 noi (~7,7GB) noi Web documents taken from a consumer health care portal & Medical newspapers & (German) PubMed abstracts & Clinical case studies & Medical textbooks none ✓ ChaDL

  185. [196]

    – 2022 50 32k 7,069k + 38,374k + 20,637k = 66,080k Discharge summaries (neurology) [because of the small number of tokens & documents the clinical portion of this cor-pus is excluded from deeper consideration] Drug labels Bio-medical abstracts (LIVIVO) Medical Wikipedia articl...

  186. [197]

    – 2024 2,000 + 2,000 + 2,000 = 6,000 subset from 3,7m radiology reports 4,369 62 63,884 11,322 854k* (*estimated) 520,718k + 1,194k + 44k + 12.299k + 9,324k Radiology reports (chest radiographs, chest CT scans, CT/radiograph examinations of the wrist covering a wide range of b...

  187. [198]

    – 2024 1099 (posts) Subset from > 13 million posts collected from over 200 different Telegram channels ~198k Subset from ~ 400 million tokens Posts from Telegram on conspir- acy narratives surrounding the COVID-19 pandemic Annotation Types (Annotated items) 14 labels for consp...

  188. [199]

    – 2024 750 (health- related claims & evidence informa- tion) ~ 675k Bilingual corpus (English – Ger- man) for medical fact checking selected from the Web portal Medizin Transparent Annotation Types (Annotated items) Claim – evidence – verdict text triples: (Public health) clai...

  189. [200]

    – 2024 60 (CT summaries re-phrased in layperson language) 145k Parallel corpus (English, German, Italian, so altogether 180 CT summaries) of layperson sum- maries of clinical trials (CT) none ✓ 40 https://corpora.linguistik.uni-erlangen.de/cqpweb/schwurpus_v2/ and https://gith...

  190. [201]

    words” # Types Absolute number of distinct single “words

    – 2024 84,478 (text fragments) 2,023k Wikipedia text fragments, labelled with an Anatomical Therapeutic Chemical (ATC) code Annotation Types (Annotated items) ATC code tags (automatically extracted from WikiData) (: 105,2k codes) Entity Normalization: Y (WikiData QID numbers ...

  191. [2013]

    Sofia, Bulgaria, August 8, 2013, pp. 27-35

  192. [2021]

    November 10, 2021 (Virtual Event), pp. 78-91. 31

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.