Pith. sign in

REVIEW 4 cited by

EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.13448 v1 pith:TZBODP5D submitted 2022-10-24 cs.CL

classification cs.CL
keywords cross-lingualsummarizationdatadataseteur-lex-sumlegalworkaccess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing summarization datasets come with two main drawbacks: (1) They tend to focus on overly exposed domains, such as news articles or wiki-like texts, and (2) are primarily monolingual, with few multilingual datasets. In this work, we propose a novel dataset, called EUR-Lex-Sum, based on manually curated document summaries of legal acts from the European Union law platform (EUR-Lex). Documents and their respective summaries exist as cross-lingual paragraph-aligned data in several of the 24 official European languages, enabling access to various cross-lingual and lower-resourced summarization setups. We obtain up to 1,500 document/summary pairs per language, including a subset of 375 cross-lingually aligned legal acts with texts available in all 24 languages. In this work, the data acquisition process is detailed and key characteristics of the resource are compared to existing summarization resources. In particular, we illustrate challenging sub-problems and open questions on the dataset that could help the facilitation of future research in the direction of domain-specific cross-lingual summarization. Limited by the extreme length and language diversity of samples, we further conduct experiments with suitable extractive monolingual and cross-lingual baselines for future work. Code for the extraction as well as access to our data and baselines is available online at: https://github.com/achouhan93/eur-lex-sum.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SteuerLLM: Local specialized large language model for German tax law analysis

    cs.CL 2026-02 reject novelty 6.0 of 10

    A tax-specialized 28B model beats larger general-purpose LLMs on a new authentic German tax-law exam benchmark, but its edge may be inflated by overlap between training and evaluation exams.

  2. Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization

    cs.CL 2025-09 reject novelty 5.0 of 10

    PEFT adapters trained on high-resource summarization domains can improve Llama-3-8B's summaries on unseen domains, but the reported gains are weakened by test-set selection and missing significance tests.

  3. Population change, age structure, and socio-economic performance

    econ.GN 2025-08 unverdicted novelty 4.0 of 10

    Across nine socio-economic indices, countries with low or negative population growth perform better on average, with no evidence that ageing populations worsen outcomes.

  4. A Comprehensive Framework for Reliable Legal AI: Combining Specialized Expert Systems and Adaptive Refinement

    cs.AI 2024-12 reject novelty 4.0 of 10

    The paper proposes a hybrid legal AI architecture and claims large accuracy gains, but it presents no numerical evidence, code, or data to support the claim.

Pith tools