Pith. sign in

REVIEW 7 cited by

CroissantLLM: A Truly Bilingual French-English Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.00786 v5 pith:JJ5IG3H3 submitted 2024-02-01 cs.CL cs.LG

classification cs.CLcs.LG
keywords modellanguagebilingualtrainingdatafrenchmodelscroissantllm
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce CroissantLLM, a 1.3B language model pretrained on a set of 3T English and French tokens, to bring to the research and industrial community a high-performance, fully open-sourced bilingual model that runs swiftly on consumer-grade local hardware. To that end, we pioneer the approach of training an intrinsically bilingual model with a 1:1 English-to-French pretraining data ratio, a custom tokenizer, and bilingual finetuning datasets. We release the training dataset, notably containing a French split with manually curated, high-quality, and varied data sources. To assess performance outside of English, we craft a novel benchmark, FrenchBench, consisting of an array of classification and generation tasks, covering various orthogonal aspects of model performance in the French Language. Additionally, rooted in transparency and to foster further Large Language Model research, we release codebases, and dozens of checkpoints across various model sizes, training data distributions, and training steps, as well as fine-tuned Chat models, and strong translation models. We evaluate our model through the FMTI framework, and validate 81 % of the transparency criteria, far beyond the scores of even most open initiatives. This work enriches the NLP landscape, breaking away from previous English-centric work in order to strengthen our understanding of multilinguality in language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 8 citations worldwide. Full citation record

  1. Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation

    cs.CL 2025-09 conditional novelty 6.0 of 10

    At matched training or inference FLOPs, plain instruction tuning is usually as good as or better than reasoning distillation, and reasoning only wins on open-ended tasks at 7B scale and up.

  2. Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness

    cs.LG 2025-06 conditional novelty 6.0 of 10

    IAM interpolates between an original model and a shadow model to score each sample's unlearning completeness, achieving top AUC for exact unlearning and top correlation for approximate unlearning, and exposing under- ...

  3. CMLFormer: A Dual Decoder Transformer with Switching Point Learning for Code-Mixed Language Modeling

    cs.CL 2025-05 conditional novelty 6.0 of 10

    CMLFormer, a dual-decoder Transformer with synchronized cross-attention and switching point prediction objectives, improves F1 on HASOC-2021 Hinglish hate speech detection by up to 0.18 over a same-data BERTbase.

  4. Training Bilingual LMs with Data Constraints in the Targeted Language

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Higher-quality auxiliary English pretraining data improves target-language performance for languages close to English (about 2% on translated QA tasks), but not for distant languages, when target-language data is limi...

  5. Assessing the Role of Data Quality in Training Bilingual Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A quality filter trained only on English labels can select better French, German, and Chinese pretraining data, improving bilingual model performance and cutting the monolingual-bilingual gap to about 1%.

  6. ChocoLlama: Lessons Learned From Teaching Llamas Dutch

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Continued pretraining with LoRA and a Dutch-specific tokenizer improves Llama-2's Dutch, but gives limited gains for already-multilingual Llama-3.

  7. The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A code LLM learning a new language first translates through its dominant language's internal system, then builds a separate system; this pattern can guide optimal training data mixing.

Pith tools