Pith. sign in

REVIEW 3 cited by

CamemBERT: a Tasty French Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03894 v3 pith:T63CN2HG submitted 2019-11-10 cs.CL

classification cs.CL
keywords languagemodelsdatalanguagescamembertcrawledfrenchmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pretrained language models are now ubiquitous in Natural Language Processing. Despite their success, most available models have either been trained on English data or on the concatenation of data in multiple languages. This makes practical use of such models --in all languages except English-- very limited. In this paper, we investigate the feasibility of training monolingual Transformer-based language models for other languages, taking French as an example and evaluating our language models on part-of-speech tagging, dependency parsing, named entity recognition and natural language inference tasks. We show that the use of web crawled data is preferable to the use of Wikipedia data. More surprisingly, we show that a relatively small web crawled dataset (4GB) leads to results that are as good as those obtained using larger datasets (130+GB). Our best performing model CamemBERT reaches or improves the state of the art in all four downstream tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Automated Journalistic Questions: A New Method for Extracting 5W1H in French

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The first French 5W1H extraction pipeline matches GPT-4o in agreement with human annotators on a new 250-article Quebec news corpus.

  2. Assessing the Role of Data Quality in Training Bilingual Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A quality filter trained only on English labels can select better French, German, and Chinese pretraining data, improving bilingual model performance and cutting the monolingual-bilingual gap to about 1%.

  3. Extracting Structured Requirements from Unstructured Building Technical Specifications for Building Information Modeling

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    A study showing that CamemBERT and Fr_core_news_lg achieve over 90% F1 for named entity recognition and Random Forest achieves over 80% F1 for relation extraction on French building technical specifications.

Pith tools