REVIEW 1 cited by
FuLG: 150B Romanian Corpus for Language Model Pretraining
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Research in the field of language models is rapidly evolving, with many open models being released to the public. Openly available pretraining corpora usually focus on only a handful of languages, with many others either missing completely or extremely underrepresented. In this report, we introduce FuLG, a hundred-fifty-billion-token Romanian corpus extracted from CommonCrawl. We present our methodology for filtering FuLG and compare it via ablation studies against existing Romanian corpora.
Forward citations
Cited by 1 Pith paper
-
LLMic: Romanian Foundation Language Model
LLMic, a 3B Romanian-English model with a diacritic-free tokenizer, reports higher WMT English-to-Romanian translation scores than larger open models, though the evaluation setup is biased.
Discussion (0). Continue with ORCID to comment.