Pith. sign in

REVIEW 1 cited by

CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.06154 v2 pith:56K3JH54 submitted 2019-11-10 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords pairsdocumentscross-lingualdatasetidentifytranslationsacrosscommon
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs. We mine sixty-eight snapshots of the Common Crawl corpus and identify web document pairs that are translations of each other. We release a new web dataset consisting of over 392 million URL pairs from Common Crawl covering documents in 8144 language pairs of which 137 pairs include English. In addition to curating this massive dataset, we introduce baseline methods that leverage cross-lingual representations to identify aligned documents based on their textual content. Finally, we demonstrate the value of this parallel documents dataset through a downstream task of mining parallel sentences and measuring the quality of machine translations from models trained on this mined data. Our objective in releasing this dataset is to foster new research in cross-lingual NLP across a variety of low, medium, and high-resource languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In-Domain African Languages Translation Using LLMs and Multi-armed Bandits

    cs.CL 2025-05 reject novelty 4.0 of 10

    Bandit-based model selection matches or slightly improves on the best single NMT system for in-domain English-to-African translation, but the claimed high-confidence statistical support is absent.

Pith tools