REVIEW 7 cited by
CroissantLLM: A Truly Bilingual French-English Language Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce CroissantLLM, a 1.3B language model pretrained on a set of 3T English and French tokens, to bring to the research and industrial community a high-performance, fully open-sourced bilingual model that runs swiftly on consumer-grade local hardware. To that end, we pioneer the approach of training an intrinsically bilingual model with a 1:1 English-to-French pretraining data ratio, a custom tokenizer, and bilingual finetuning datasets. We release the training dataset, notably containing a French split with manually curated, high-quality, and varied data sources. To assess performance outside of English, we craft a novel benchmark, FrenchBench, consisting of an array of classification and generation tasks, covering various orthogonal aspects of model performance in the French Language. Additionally, rooted in transparency and to foster further Large Language Model research, we release codebases, and dozens of checkpoints across various model sizes, training data distributions, and training steps, as well as fine-tuned Chat models, and strong translation models. We evaluate our model through the FMTI framework, and validate 81 % of the transparency criteria, far beyond the scores of even most open initiatives. This work enriches the NLP landscape, breaking away from previous English-centric work in order to strengthen our understanding of multilinguality in language models.
Forward citations
Cited by 7 Pith papers
-
Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation
At matched training or inference FLOPs, plain instruction tuning is usually as good as or better than reasoning distillation, and reasoning only wins on open-ended tasks at 7B scale and up.
-
Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness
IAM interpolates between an original model and a shadow model to score each sample's unlearning completeness, achieving top AUC for exact unlearning and top correlation for approximate unlearning, and exposing under- ...
-
CMLFormer: A Dual Decoder Transformer with Switching Point Learning for Code-Mixed Language Modeling
CMLFormer, a dual-decoder Transformer with synchronized cross-attention and switching point prediction objectives, improves F1 on HASOC-2021 Hinglish hate speech detection by up to 0.18 over a same-data BERTbase.
-
Training Bilingual LMs with Data Constraints in the Targeted Language
Higher-quality auxiliary English pretraining data improves target-language performance for languages close to English (about 2% on translated QA tasks), but not for distant languages, when target-language data is limi...
-
Assessing the Role of Data Quality in Training Bilingual Language Models
A quality filter trained only on English labels can select better French, German, and Chinese pretraining data, improving bilingual model performance and cutting the monolingual-bilingual gap to about 1%.
-
ChocoLlama: Lessons Learned From Teaching Llamas Dutch
Continued pretraining with LoRA and a Dutch-specific tokenizer improves Llama-2's Dutch, but gives limited gains for already-multilingual Llama-3.
-
The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model
A code LLM learning a new language first translates through its dominant language's internal system, then builds a separate system; this pattern can guide optimal training data mixing.
Discussion (0). Continue with ORCID to comment.