Pith. sign in

REVIEW 3 cited by

Low-resource Information Extraction with the European Clinical Case Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.20568 v1 pith:7OTEIBFH submitted 2025-03-26 cs.CL

Low-resource Information Extraction with the European Clinical Case Corpus

classification cs.CL
keywords datadatasetlanguagesclinicale3c-3englishfiveitalian
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present E3C-3.0, a multilingual dataset in the medical domain, comprising clinical cases annotated with diseases and test-result relations. The dataset includes both native texts in five languages (English, French, Italian, Spanish and Basque) and texts translated and projected from the English source into five target languages (Greek, Italian, Polish, Slovak, and Slovenian). A semi-automatic approach has been implemented, including automatic annotation projection based on Large Language Models (LLMs) and human revision. We present several experiments showing that current state-of-the-art LLMs can benefit from being fine-tuned on the E3C-3.0 dataset. We also show that transfer learning in different languages is very effective, mitigating the scarcity of data. Finally, we compare performance both on native data and on projected data. We release the data at https://huggingface.co/collections/NLP-FBK/e3c-projected-676a7d6221608d60e4e9fd89 .

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

    cs.CL 2026-06 unverdicted novelty 8.0

    EDEN releases the largest freely available Italian clinical notes corpus (4M notes, 6k annotated) and proposes CRF-filling as a structured extraction benchmark with zero-shot baselines from Gemma models.

  2. eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

    cs.CL 2026-06 unverdicted novelty 8.0

    Introduces the largest freely available Italian clinical notes corpus with 4M notes and expert-annotated subset for a new CRF-filling benchmark.

  3. eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

    cs.CL 2026-06 conditional novelty 6.0

    eCREAM-MedCorpus releases ~4M anonymized Italian ED clinical notes and a 6k-note 132-item CRF annotation set, with zero-shot Gemma/MedGemma CRF-filling baselines.