Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Marco-LLM claims that a two-stage multilingual continual-pretraining and post-training recipe lifts Qwen2's average score across 29 languages from 69.1 to 75.5 at 7B scale and from 85.2 to 87.9 at 72B scale while also improving direct…

desk verdict A credible industrial recipe with large claimed multilingual gains, but the evaluation protocol is under-specified and the results are not yet independently verifiable. read the letter →

arxiv 2412.04003 v1 pith:B3RNZNUE submitted 2024-12-05 cs.CL

classification cs.CL
keywords multilingualLLMcontinualpretraininglow-resourcelanguagescross-lingualtransfermachinetranslationsupervisedfine-tuningpreferencealignmentdatacuration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to show that a strong multilingual base model can be extended to underrepresented languages without retraining from scratch. The recipe is continual pretraining on a curated 300B-token corpus covering 29 languages in two stages, then multilingual supervised fine-tuning and preference alignment. The authors report that this lifts the average score across 29 languages from 69.1 to 75.5 at 7B scale and from 85.2 to 87.9 at 72B scale, with the largest gains in low-resource languages such as Nepali and Kazakh, while keeping English and Chinese performance intact. If the reported margins hold, the implication is that data curation plus targeted continual training can narrow the gap between high- and low-resource language capabilities.

What carries the argument

The load-bearing mechanism is a two-stage continual pretraining schedule on a curated 300B-token multilingual corpus, followed by multilingual supervised fine-tuning and direct preference optimization. Stage-I uses 160B tokens at a peak learning rate of 1e-5 with a mixture that keeps 32% English and 17% Chinese to limit catastrophic forgetting; Stage-II uses 140B tokens at 6e-6 and raises the low-resource share from 9% to 15% to push multilingual capability. The corpus work that makes this work is heavy filtering, MinHash deduplication, and the inclusion of parallel data wrapped in diverse translation templates to create cross-lingual alignment. The authors attribute part of the efficiency to Qwen2's 150k-token vocabulary, which keeps low-resource text highly compressible.

What would settle it

Score Marco-7B, Marco-72B, and all baselines on a freshly produced parallel test set in the 19 low-resource languages using one fixed prompt, one decoder, and one tokenizer, and also search the training corpora for near-duplicates of Flores, Belebele, and MMMLU items; the central claim fails if the margins vanish under that protocol or if contamination turns up.

Watch

Extended reading notes

Core claim

The paper's central claim is that its two-stage continual pretraining and post-training recipe turns Qwen2 into a model whose low-resource language performance is substantially better than the base and than comparable open models, without sacrificing high-resource performance. The supporting evidence is an average of 75.5 for Marco-7B across 29 languages versus 69.1 for Qwen2.5-7B, and 87.9 for Marco-72B versus 85.2 for Qwen2.5-72B, plus large gains on non-English-pivot Flores translation (19.7 versus 14.6 BLEU at 7B scale). It also reports that Marco-72B beats GPT-4 on many MMMLU languages and beats Google Translate on several Flores directions.

Load-bearing premise

Every reported margin depends on the baselines being evaluated under the identical prompt, decoding, and scoring protocol, and on none of the training corpora containing test-set sentences.

Editorial extensions

If this is right

  • A 7B-parameter model built with this recipe can outperform much larger general-purpose models on low-resource language benchmarks.
  • English, Chinese, and other high-resource languages do not have to be traded away when extending a model to low-resource languages.
  • Direct translation between non-English pairs becomes practical without an English pivot, which is relevant for language pairs that commercial systems serve poorly.
  • The parallel-data ablation implies that data filtering is load-bearing at 72B scale, so small-scale pilots may not predict what large models need.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural test of generalizability would be to apply the same two-stage recipe to a different base model with a smaller vocabulary; if the gains shrink, the 150k vocabulary is doing much of the work.
  • The 5.9% overall data utilization rate suggests the bottleneck is selection rather than crawl size; an explicit tokens-per-language saturation curve would show where extra low-resource data stops paying.
  • The authors' observed gap between high- and low-resource languages after SFT suggests that more pretraining tokens for low-resource languages, not longer SFT, is the lever for closing the remaining gap; this is an inference from their Figure 7, not a claim they test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. Marco-LLM is a recipe for extending an existing multilingual base model (Qwen2) to 29 languages by (i) continual pretraining on a curated 300B-token mixture with a two-stage curriculum and a lowered learning rate, and (ii) multilingual SFT and DPO. The paper reports large average gains: Marco-7B reaches 75.5 versus 69.1 for Qwen2.5-7B across 29 languages, and Marco-72B reaches 87.9 versus 85.2 for Qwen2.5-72B; the largest margins are in low-resource languages such as Nepali and Kazakh. It also reports improvements on MMMLU, Belebele, TyDiQA, and Flores, including English-pivot and any-to-any translation.

Significance. If the numbers are trustworthy, this is a practically valuable demonstration that a comparatively small amount (300B tokens) of well-curated multilingual continual pretraining plus multilingual post-training can substantially close the low-resource gap of a strong open model. The two-stage continual-pretraining design and the parallel-data filtering ablation (Section 3.5) are useful contributions, and the data collection pipeline is described in rare detail. However, the evidence is entirely benchmark-based, and the manuscript currently does not supply enough protocol detail or contamination checks to verify the headline margins. No code, model weights, or evaluation harness is promised, so independent verification is not possible from the paper alone.

major comments (4)
  1. [Section 3.3.1, Tables 6 and 11] The evaluation protocol is under-specified. The paper lists datasets, splits, shots, and metrics, but not the exact prompts, answer extraction rules, decoding hyperparameters, or BLEU tokenization/normalization used for any model. Because the central comparison is across different base models with possibly different tokenizers, small protocol differences can move scores by several points. Two values in the tables suggest protocol issues: Llama3-70B obtains 0.1 BLEU on En to Ko in Table 11, and Qwen2.5-7B drops from 80.2 to 71.6 on Dutch in Table 6. Provide the exact harness or a public reference, and report at least one per-language input/output example for each benchmark.
  2. [Section 4.1.1 and Section 3.1.3] No contamination audit is given for the SFT and DPO corpora. The only exclusion claim concerns high-quality knowledge data (Section 3.1.3); no such guarantee is made for the SFT mixture, which explicitly includes Aya collection, MetaMathQA, MathInstruct, Belle, Orca, WMT dev sets, WikiMatrix, translated preference data, and synthetic data. At least one of these sources has been reported to contain benchmark items, and parallel dev sets can overlap with test sets used for WMT16. Run an n-gram or embedding-based overlap analysis between every training component and every evaluation benchmark (MMMLU, AGIEval, CEval, Belebele, TyDiQA, Flores-200 devtest, XCOPA, XStoryCloze, XWinograd) and report the maximum overlap per benchmark.
  3. [Section 4.1.5, Table 12] The any-to-any translation section is internally inconsistent. The text refers to Table??, quotes averages of 19.5 and 14.4, while Table 12 reports averages of 19.7 and 14.6. Moreover, Table 12 contains only 7B models, so the abstract's claim of substantial enhancements in any-to-any machine translation tasks is not supported for the 72B model. Fix the reference, correct the numbers, and add the 72B any-to-any results or explicitly limit the claim to the 7B model.
  4. [Tables 6-11] All reported scores are single-run values without error bars or significance tests. While most claimed gains are large, several comparisons are close, for example Marco-72B versus Qwen2.5-72B on Polish in Table 7 (88.2 versus 88.8) and Marco-7B versus Qwen2.5-7B on Thai in Table 6 (72.9 versus 73.7). For such entries the textual claim of consistent outperformance is not supported without repeated evaluations or a paired test over the per-language subtasks.
minor comments (5)
  1. [Section 4.1.5 and Appendix A.2] Model names are used inconsistently: Marco, Marco-7B, Marco-Chat-7B, Marco-72B, and Marco-Chat appear for what seem to be the same models. Please fix the terminology once and for all.
  2. [Throughout] There are several typos and formatting errors: re-warned should be re-warmed in Section 3.2; Averge in Figure 7; Kazakh(he) in Section 4.2.3; Macro model in Section 4; truction (existing preference dataset) in Section 4.1.5; and the ratio of digits„ in Section 3.1.2.
  3. [Section 3.3.1 and Section 4.1.3] The relationship between X-MMLU (13 languages, Section 3.3.1) and MMMLU (14 languages, Section 4.1.3) is unclear. Clarify which dataset is used in which table, since both appear in the evaluation suite.
  4. [Section 4.2.3] The multilingual MT-bench comparison reports win/loss/tie rates from GPT-4o-mini but does not state the number of prompts per language, the judge prompt, the decoding temperature, or the tie-breaking rule. Add these details so the pairwise comparison is reproducible.
  5. [Section 3.5, Figure 5] The ablation on parallel data filtering is described as showing significant improvements, but Figure 5 has no numerical values and no statistical test. Report the underlying numbers and the number of evaluation examples used.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical training-and-evaluation report whose headline numbers are measured after held-out evaluation, not derived from fitted quantities or self-citations.

full rationale

Marco-LLM's central claims rest on multi-benchmark evaluations after continual pretraining, SFT, and DPO. No equation in the paper defines a predicted benchmark score from a fitted parameter, and no claimed result is equivalent to a training input by construction. The data mixture and learning rate are tuned on Marco-1.5B using evaluation families such as XStoryCloze, Belebele, and Flores, but this is hyperparameter selection, not a statistical reduction: the final 7B/72B scores are measured outcomes, not quantities forced by the tuning procedure. The paper does not invoke a uniqueness theorem, does not smuggle in an ansatz via citation, and contains no load-bearing self-citation; its baselines are external open models evaluated on public benchmarks. Concerns about possible benchmark contamination or inconsistent evaluation protocols are validity risks, not circularity, and the paper's Section 3.1.3 claim that common benchmark training sets are excluded from high-quality knowledge data is a data-handling statement rather than a derivation. Therefore the derivation chain is self-contained as an empirical system report, and no circular step can be quoted or exhibited.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The headline gains depend on unreleased data, scale-transfer of hyperparameters, and translated evaluation sets. The paper contributes an industrial recipe rather than a first-principles result, and the number of tuned quantities is significant.

free parameters (6)
  • Data mixture proportions = Stage-I: en 32%, zh 17%, other HR 30%, LR 9%, parallel 6%, HQ 5%, synthetic 1%; Stage-II: 28/15/26/15/8/6/2
    Determined by experiments on Marco-1.5B (Section 3.2) to balance forgetting and multilingual gains; the final models use these tuned proportions.
  • Peak learning rates = 1e-5 (Stage-I), 6e-6 (Stage-II)
    Selected from {1e-5, 2e-5, 3e-5} on Marco-1.5B based on English/Chinese forgetting versus multilingual accuracy (Section 3.6).
  • Per-language token caps = 10.6B per high-resource language, 1.9B per low-resource language
    Caps chosen experimentally from available corpora; Table 1 and Section 3.1.6 describe the utilization rates.
  • Filtering thresholds = Not specified numerically
    KenLM perplexity 'largely above average', LASER embedding similarity threshold, embedding similarity 0.7, QA similarity threshold, and IFD score threshold are referenced without exact values (Sections 3.1.1, 3.1.2, 4.1.1).
  • Stage token budgets = 160B (Stage-I), 140B (Stage-II)
    Total 300B is the chosen training budget; no ablation of budget size is reported (Section 3.2, Table 3).
  • SFT learning rate range = 6e-6 maximum, 6e-7 minimum
    Adjusted from the pre-training minimum learning rate and batch size (Section 4.1.2).
assumptions (4)
  • domain assumption Qwen2's tokenizer and pretrained representations transfer to low-resource languages with continued training alone.
    Section 3.2 says Qwen2 is multilingual-friendly with a 150k vocabulary; no experiment compares against training from scratch or adding new language-specific embeddings.
  • domain assumption Translated versions of benchmarks and preference data are valid for measuring multilingual ability.
    Sections 4.2.1 and 4.2.2 translate LMSYS and UltraFeedback into 28 languages and translate MT-Bench; translationese may not reflect native usage.
  • ad hoc to paper Small-scale (1.5B) choices of learning rate and data mixture transfer to 7B and 72B.
    Section 3.2 states that experiments in Marco-1.5B are used to optimize the data mix and learning rate, then scaled to 7B and 72B. No evidence is given that optimum hyperparameters are scale-invariant, and Section 3.5 itself shows scale-dependent effects of parallel data filtering.
  • domain assumption Machine translation BLEU on Flores is an appropriate proxy for cross-lingual ability.
    Section 3.3.1 lists Flores as an MT benchmark; BLEU scores are reported without tokenizer, detokenization, or statistical significance details.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement." pith.science (2026). https://pith.science/paper/B3RNZNUE

@misc{pith2026241204003,
  author       = {Pith},
  title        = {Pith review of: Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B3RNZNUE}},
  note         = {Machine review of arXiv:2412.04003}
}
read the original abstract

Large Language Models (LLMs) have achieved remarkable progress in recent years; however, their excellent performance is still largely limited to major world languages, primarily English. Many LLMs continue to face challenges with multilingual tasks, especially when it comes to low-resource languages. To address this issue, we introduced Marco-LLM: Massive multilingual training for cross-lingual enhancement LLM. We have collected a substantial amount of multilingual data for several low-resource languages and conducted extensive continual pre-training using the Qwen2 models. This effort has resulted in a multilingual LLM named Marco-LLM. Through comprehensive evaluations on various multilingual benchmarks, including MMMLU, AGIEval, Belebele, Flores-200, XCOPA and many others, Marco-LLM has demonstrated substantial improvements over state-of-the-art LLMs. Furthermore, Marco-LLM achieved substantial enhancements in any-to-any machine translation tasks, showing the effectiveness of our multilingual LLM. Marco-LLM is a pioneering multilingual LLM designed to not only perform exceptionally well in multilingual tasks, including low-resource languages, but also maintain strong performance in English and other major languages, closing the performance gap between high- and low-resource language capabilities. By bridging languages, this effort demonstrates our dedication to ensuring LLMs work accurately across various languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A multilingual LLM training method that groups similar languages, converts high-deviation layers into mixture-of-experts layers, and assigns one expert per language group improves perplexity across 18 to 128 languages.

  2. Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation

    cs.CL 2024-12 conditional novelty 4.0 of 10

    The WMT 2024 literary translation shared task finds that domain-enhanced systems lead in d-BLEU for Chinese-English, but human evaluators rank NLP2CT-UM and SJTU-LoveFiction at the top.

Reference graph

Works this paper leans on

76 extracted references · 28 canonical work pages · cited by 2 Pith papers

  1. [1]

    R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier - Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. \' A brego, J. Ahn, J. Austin, P. Barham, J. A. Botha, J. Bradbury, S. Brahma, K. ...

  2. [2]

    Artetxe, G

    M. Artetxe, G. Labaka, E. Agirre, and K. Cho. Unsupervised neural machine translation. In International Conference on Learning Representations, 2018

  3. [3]

    Aryabumi, J

    V. Aryabumi, J. Dang, D. Talupuru, S. Dash, D. Cairuz, H. Lin, B. Venkitesh, M. Smith, J. A. Campos, Y. C. Tan, K. Marchisio, M. Bartolo, S. Ruder, A. Locatelli, J. Kreutzer, N. Frosst, A. Gomez, P. Blunsom, M. Fadaee, A. Üstün, and S. Hooker. Aya 23: Open weight releases to further multilingual progress, 2024. URL https://arxiv.org/abs/2405.15032

  4. [4]

    J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023

  5. [5]

    Bandarkar, D

    L. Bandarkar, D. Liang, B. Muller, M. Artetxe, S. N. Shukla, D. Husa, N. Goyal, A. Krishnan, L. Zettlemoyer, and M. Khabsa. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...

  6. [6]

    Y. Bang, S. Cahyawijaya, N. Lee, W. Dai, D. Su, B. Wilie, H. Lovenia, Z. Ji, T. Yu, W. Chung, et al. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023, 2023

  7. [7]

    Barrault, O

    L. Barrault, O. Bojar, F. Bougares, R. Chatterjee, M. R. Costa-juss \`a , C. Federmann, M. Fishel, A. Fraser, Y. Graham, P. Guzman, B. Haddow, M. Huck, A. J. Yepes, P. Koehn, A. Martins, M. Morishita, C. Monz, M. Nagata, T. Nakazawa, and M. Negri, editors. Proceedings of the Fifth Conference on Machine Translation, Online, Nov. 2020. Association for Compu...

  8. [8]

    Bojar, C

    O. Bojar, C. Buck, R. Chatterjee, C. Federmann, L. Guillou, B. Haddow, M. Huck, A. J. Yepes, A. N \'e v \'e ol, M. Neves, P. Pecina, M. Popel, P. Koehn, C. Monz, M. Negri, M. Post, L. Specia, K. Verspoor, J. Tiedemann, and M. Turchi, editors. Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, Berlin, Germany, Aug. 20...

Show all 76 references
  1. [9]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Che...

  2. [10]

    Chaudhary, Y

    V. Chaudhary, Y. Tang, F. Guzmán, H. Schwenk, and P. Koehn. Low-resource corpus filtering using multilingual sentence embeddings. In Proceedings of the Fourth Conference on Machine Translation (Volume 3: Shared Task Papers, Day 2), pages 263--268, Florence, Italy, August 2019...

  3. [11]

    Chiang, L

    W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, and I. Stoica. Chatbot arena: An open platform for evaluating llms by human preference, 2024

  4. [12]

    Chowdhery, S

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Au...

  5. [13]

    J. H. Clark, E. Choi, M. Collins, D. Garrette, T. Kwiatkowski, V. Nikolaev, and J. Palomaki. T y D i QA : A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8: 0 454--470, 20...

  6. [14]

    Clark, I

    P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018. URL https://arxiv.org/abs/1803.05457

  7. [15]

    c4ai-command-r-plus-08-2024, 2024

    Cohere For AI . c4ai-command-r-plus-08-2024, 2024. URL https://huggingface.co/CohereForAI/c4ai-command-r-plus-08-2024

  8. [16]

    Conneau and G

    A. Conneau and G. Lample. Cross-lingual language model pretraining. Advances in Neural Information Processing Systems, 32: 0 7059--7069, 2019

  9. [17]

    Conneau, R

    A. Conneau, R. Rinott, G. Lample, A. Williams, S. Bowman, H. Schwenk, and V. Stoyanov. XNLI : Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475--2485, Brussels, Belgium, Oct....

  10. [18]

    G. Cui, L. Yuan, N. Ding, G. Yao, W. Zhu, Y. Ni, G. Xie, Z. Liu, and M. Sun. Ultrafeedback: Boosting language models with high-quality feedback, 2023

  11. [19]

    Dac Lai, C

    V. Dac Lai, C. Van Nguyen, N. T. Ngo, T. Nguyen, F. Dernoncourt, R. A. Rossi, and T. H. Nguyen. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback. arXiv e-prints, pages arXiv--2307, 2023

  12. [20]

    DeepSeek - AI, A. Liu, B. Feng, B. Wang, B. Wang, B. Liu, C. Zhao, C. Deng, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Yang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, ...

  13. [21]

    Dubey and et al

    A. Dubey and et al. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783

  14. [22]

    El-Kishky, V

    A. El-Kishky, V. Chaudhary, F. Guzm \'a n, and P. Koehn. CCAligned : A massive collection of cross-lingual web-document pairs. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), pages 5960--5969, Online, November 2020. Assoc...

  15. [23]

    A. Fan, S. Bhosale, H. Schwenk, Z. Ma, A. El-Kishky, S. Goyal, M. Baines, O. Celebi, G. Wenzek, V. Chaudhary, N. Goyal, T. Birch, V. Liptchinsky, S. Edunov, E. Grave, M. Auli, and A. Joulin. Beyond english-centric multilingual machine translation, 2020. URL https://arxiv.org/a...

  16. [24]

    Goyal, C

    N. Goyal, C. Gao, V. Chaudhary, P.-J. Chen, G. Wenzek, D. Ju, S. Krishnan, M. Ranzato, F. Guzm\' a n, and A. Fan. The flores-101 evaluation benchmark for low-resource and multilingual machine translation. 2021

  17. [25]

    Gunasekar, Y

    S. Gunasekar, Y. Zhang, J. Aneja, C. C. T. Mendes, A. D. Giorno, S. Gopi, M. Javaheripi, P. Kauffmann, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, H. S. Behl, X. Wang, S. Bubeck, R. Eldan, A. T. Kalai, Y. T. Lee, and Y. Li. Textbooks are all you need. CoRR, abs/2306.11644, 2023

  18. [26]

    D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang. Deepseek-coder: When the large language model meets programming -- the rise of code intelligence, 2024. URL https://arxiv.org/abs/2401.14196

  19. [27]

    Gurnee, N

    W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas. Finding neurons in a haystack: Case studies with sparse probing. Trans. Mach. Learn. Res., 2023, 2023

  20. [28]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. In ICLR . OpenReview.net, 2021

  21. [29]

    S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, X. Zhang, Z. L. Thai, K. Zhang, C. Wang, Y. Yao, C. Zhao, J. Zhou, J. Cai, Z. Zhai, N. Ding, C. Jia, G. Zeng, D. Li, Z. Liu, and M. Sun. Minicpm: Unveiling the potential of small language model...

  22. [30]

    Huang, Y

    Y. Huang, Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, J. Lei, Y. Fu, M. Sun, and J. He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. In Advances in Neural Information Processing Systems, 2023

  23. [31]

    B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Dang, A. Yang, R. Men, F. Huang, X. Ren, X. Ren, J. Zhou, and J. Lin. Qwen2.5-coder technical report, 2024. URL https://arxiv.org/abs/2409.12186

  24. [32]

    Hurst, A

    A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024

  25. [33]

    Ibrahim, B

    A. Ibrahim, B. Th \' e rien, K. Gupta, M. L. Richter, Q. G. Anthony, E. Belilovsky, T. Lesort, and I. Rish. Simple and scalable strategies to continually pre-train large language models. Trans. Mach. Learn. Res., 2024, 2024

  26. [34]

    W. Jiao, W. Wang, J.-t. Huang, X. Wang, and Z. Tu. Is chatgpt a good translator? a preliminary study. arXiv preprint arXiv:2301.08745, 1 0 (10), 2023

  27. [35]

    Joulin, E

    A. Joulin, E. Grave, P. Bojanowski, M. Douze, H. Jégou, and T. Mikolov. Fasttext.zip: Compressing text classification models. arXiv: Computation and Language,arXiv: Computation and Language, Nov 2016

  28. [36]

    Z. Ke, Y. Shao, H. Lin, T. Konishi, G. Kim, and B. Liu. Continual pre-training of language models, 2023. URL https://arxiv.org/abs/2302.03241

  29. [37]

    V. D. Lai, N. T. Ngo, A. P. B. Veyseh, H. Man, F. Dernoncourt, T. Bui, and T. H. Nguyen. Chatgpt beyond english: Towards a comprehensive evaluation of large language models in multilingual learning. arXiv preprint arXiv:2304.05613, 2023

  30. [38]

    Y. Li, S. Bubeck, R. Eldan, A. D. Giorno, S. Gunasekar, and Y. T. Lee. Textbooks are all you need II: phi-1.5 technical report. CoRR, abs/2309.05463, 2023

  31. [39]

    X. V. Lin, T. Mihaylov, M. Artetxe, T. Wang, S. Chen, D. Simig, M. Ott, N. Goyal, S. Bhosale, J. Du, R. Pasunuru, S. Shleifer, P. S. Koura, V. Chaudhary, B. O'Horo, J. Wang, L. Zettlemoyer, Z. Kozareva, M. T. Diab, V. Stoyanov, and X. Li. Few-shot learning with multilingual la...

  32. [40]

    Lovenia, R

    H. Lovenia, R. Mahendra, S. M. Akbar, L. J. V. Miranda, J. Santoso, E. Aco, A. Fadhilah, J. Mansurov, J. M. Imperial, O. P. Kampman, et al. Seacrowd: A multilingual multimodal data hub and benchmark suite for southeast asian languages. arXiv preprint arXiv:2406.10118, 2024

  33. [41]

    H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, and D. Zhang. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. CoRR, abs/2308.09583, 2023

  34. [42]

    Z. Luo, C. Xu, P. Zhao, Q. Sun, X. Geng, W. Hu, C. Tao, J. Ma, Q. Lin, and D. Jiang. Wizardcoder: Empowering code large language models with evol-instruct. In ICLR . OpenReview.net, 2024

  35. [43]

    Nguyen, C

    T. Nguyen, C. V. Nguyen, V. D. Lai, H. Man, N. T. Ngo, F. Dernoncourt, R. A. Rossi, and T. H. Nguyen. Culturax: A cleaned, enormous, and multilingual dataset for large language models in 167 languages. In LREC/COLING , pages 4226--4237. ELRA and ICCL , 2024

  36. [44]

    GPT-4 Technical Report

    OpenAI. GPT-4 Technical Report . arXiv preprint arXiv:2303.08774, 2023

  37. [45]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Gray, et al. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, 2022

  38. [46]

    Penedo, Q

    G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, H. Alobeidli, A. Cappelli, B. Pannier, E. Almazrouei, and J. Launay. The refinedweb dataset for falcon LLM: outperforming curated corpora with web data only. In NeurIPS, 2023

  39. [47]

    Pires, E

    T. Pires, E. Schlinger, and D. Garrette. How multilingual is multilingual bert? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996--5001, 2019

  40. [48]

    E. M. Ponti, G. Glava s , O. Majewska, Q. Liu, I. Vuli \'c , and A. Korhonen. XCOPA : A multilingual dataset for causal commonsense reasoning. In B. Webber, T. Cohn, Y. He, and Y. Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Process...

  41. [49]

    Rafailov, A

    R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=HPuSIXJaa9

  42. [50]

    Schwenk, V

    H. Schwenk, V. Chaudhary, S. Sun, H. Gong, and F. Guzm \'a n. W iki M atrix: Mining 135 M parallel sentences in 1620 language pairs from W ikipedia. In P. Merlo, J. Tiedemann, and R. Tsarfaty, editors, Proceedings of the 16th Conference of the European Chapter of the Associati...

  43. [51]

    S. She, W. Zou, S. Huang, W. Zhu, X. Liu, X. Geng, and J. Chen. MAPO : Advancing multilingual reasoning through multilingual-alignment-as-preference optimization. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for C...

  44. [52]

    F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, D. Das, and J. Wei. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations, 2023. URL https://ope...

  45. [53]

    Singh, N

    H. Singh, N. Gupta, S. Bharadwaj, D. Tewari, and P. Talukdar. Indicgenbench: A multilingual benchmark to evaluate generation capabilities of llms on indic languages, 2024 a . URL https://arxiv.org/abs/2404.16816

  46. [54]

    Singh, F

    S. Singh, F. Vargus, D. Dsouza, B. F. Karlsson, A. Mahendiran, W.-Y. Ko, H. Shandilya, J. Patel, D. Mataciunas, L. OMahony, M. Zhang, R. Hettiarachchi, J. Wilson, M. Machado, L. S. Moura, D. Krzemiński, H. Fadaei, I. Ergün, I. Okoh, A. Alaagib, O. Mudannayake, Z. Alyafeai, V. ...

  47. [55]

    N. Team, M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. M. Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe, ...

  48. [56]

    Tiedemann

    J. Tiedemann. Parallel data, tools and interfaces in OPUS . In N. Calzolari, K. Choukri, T. Declerck, M. U. Do g an, B. Maegaard, J. Mariani, A. Moreno, J. Odijk, and S. Piperidis, editors, Proceedings of the Eighth International Conference on Language Resources and Evaluation...

  49. [58]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample. Llama: Open and efficient foundation language models, 2023 b . URL https://arxiv.org/abs/2302.13971

  50. [59]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hoss...

  51. [60]

    \" U st \" u n, V

    A. \" U st \" u n, V. Aryabumi, Z. X. Yong, W. Ko, D. D'souza, G. Onilude, N. Bhandari, S. Singh, H. Ooi, A. Kayid, F. Vargus, P. Blunsom, S. Longpre, N. Muennighoff, M. Fadaee, J. Kreutzer, and S. Hooker. Aya model: An instruction finetuned open-access multilingual language m...

  52. [61]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. CoRR, abs/1706.03762, 2017. URL http://arxiv.org/abs/1706.03762

  53. [62]

    L. Wang, C. Lyu, T. Ji, Z. Zhang, D. Yu, S. Shi, and Z. Tu. Document-level machine translation with large language models. arXiv preprint arXiv:2304.02210, 2023

  54. [63]

    J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le. Finetuned language models are zero-shot learners. In International Conference on Learning Representations, 2022

  55. [64]

    T. Wei, L. Zhao, L. Zhang, B. Zhu, L. Wang, H. Yang, B. Li, C. Cheng, W. L \" u , R. Hu, C. Li, L. Yang, X. Luo, X. Wu, L. Liu, W. Cheng, P. Cheng, J. Zhang, X. Zhang, L. Lin, X. Wang, Y. Ma, C. Dong, Y. Sun, Y. Chen, Y. Peng, X. Liang, S. Yan, H. Fang, and Y. Zhou. Skywork: A...

  56. [65]

    X. Wei, H. Wei, H. Lin, T. Li, P. Zhang, X. Ren, M. Li, Y. Wan, Z. Cao, B. Xie, T. Hu, S. Li, B. Hui, B. Yu, D. Liu, B. Yang, F. Huang, and J. Xie. Polylm: An open source polyglot large language model. CoRR, abs/2307.06018, 2023 b

  57. [66]

    Wenzek, M

    G. Wenzek, M. Lachaux, A. Conneau, V. Chaudhary, F. Guzm \' a n, A. Joulin, and E. Grave. Ccnet: Extracting high quality monolingual datasets from web crawl data. In LREC , pages 4003--4012. European Language Resources Association, 2020

  58. [67]

    Whitehouse, M

    C. Whitehouse, M. Choudhury, and A. F. Aji. Llm-powered data augmentation for enhanced cross-lingual performance. In EMNLP , pages 671--686. Association for Computational Linguistics, 2023

  59. [68]

    C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, Q. Lin, and D. Jiang. Wizardlm: Empowering large pre-trained language models to follow complex instructions. In ICLR . OpenReview.net, 2024

  60. [69]

    L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel. m T 5: A massively multilingual pre-trained text-to-text transformer. In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty...

  61. [71]

    A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Yang, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. ...

  62. [72]

    S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36, 2024

  63. [73]

    Y. Yu, Y. Zhuang, J. Zhang, Y. Meng, A. J. Ratner, R. Krishna, J. Shen, and C. Zhang. Large language model as attributed training data generator: A tale of diversity and bias. In NeurIPS, 2023

  64. [74]

    Zellers, A

    R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

  65. [75]

    Zhang, P

    B. Zhang, P. Williams, I. Titov, and R. Sennrich. Improving massively multilingual neural machine translation and zero-shot translation. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational...

  66. [76]

    Zhong, R

    W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan. Agieval: A human-centric benchmark for evaluating foundation models, 2023. URL https://arxiv.org/abs/2304.06364

  67. [77]

    Çağatay Yıldız, N. K. Ravichandran, P. Punia, M. Bethge, and B. Ermis. Investigating continual pretraining in large language models: Insights and implications, 2024. URL https://arxiv.org/abs/2402.17400

  68. [78]

    Üstün, V

    A. Üstün, V. Aryabumi, Z.-X. Yong, W.-Y. Ko, D. D'souza, G. Onilude, N. Bhandari, S. Singh, H.-L. Ooi, A. Kayid, F. Vargus, P. Blunsom, S. Longpre, N. Muennighoff, M. Fadaee, J. Kreutzer, and S. Hooker. Aya model: An instruction finetuned open-access multilingual language mode...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.