Pith. sign in

REVIEW 2 cited by

Towards a Cleaner Document-Oriented Multilingual Crawled Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.06642 v1 pith:6MVUZ5BL submitted 2022-01-17 cs.CL

classification cs.CL
keywords languagedatalargeautomaticcorpusdocument-orientedlearningmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The need for raw large raw corpora has dramatically increased in recent years with the introduction of transfer learning and semi-supervised learning methods to Natural Language Processing. And while there have been some recent attempts to manually curate the amount of data necessary to train large language models, the main way to obtain this data is still through automatic web crawling. In this paper we take the existing multilingual web corpus OSCAR and its pipeline Ungoliant that extracts and classifies data from Common Crawl at the line level, and propose a set of improvements and automatic annotations in order to produce a new document-oriented version of OSCAR that could prove more suitable to pre-train large generative language models as well as hopefully other applications in Natural Language Processing and Digital Humanities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HalleluBERT: Let Every Token That Has Meaning Bear Its Weight

    cs.CL 2025-10 conditional novelty 5.0 of 10

    HalleluBERT, a Hebrew-only RoBERTa encoder family trained from scratch at scale, reports the highest unweighted mean scores on BMC, NEMO, and SMCD benchmarks, but without statistical significance testing.

  2. Salamandra Technical Report

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.

Pith tools