Pith. sign in

REVIEW 4 major objections 5 minor 10 cited by

Salamandra Technical Report

T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This report claims Salamandra, a from-scratch family of 2B, 7B, and 40B open models trained on 35 European languages plus code, reaches competitive performance against similarly sized open-source models.

desk verdict A genuinely open multilingual model family with a strong engineering report, but the headline competitive claim for instructed models hangs on an unvalidated LLM judge and needs referee attention. read the letter →

arxiv 2502.08489 v2 pith:PXH5SZVV submitted 2025-02-12 cs.CL

classification cs.CL
keywords open-sourcelanguagemodelsmultilingualEuropeanlanguagesSpanishCatalanBasqueGalicianinstructiontuningbiasandsafetyevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces Salamandra, a suite of decoder-only language models in three sizes (2B, 7B, and 40B), trained from scratch on open-access text from 35 European languages plus code. The central claim is that these models achieve competitive performance against similarly sized open-source models on multilingual benchmarks, especially in Spanish, Catalan, Basque, and Galician. If the claim holds, researchers and companies get a fully open, Apache-2.0 alternative for European languages where strong open models are scarce. The paper additionally makes its training recipes, data curation methods, and evaluation scripts publicly available, so the suite is reproducible in spirit.

What carries the argument

The load-bearing machinery is the training and evaluation stack: a 256,000-token byte-pair-encoding tokenizer trained on a per-language balanced sample; multilayered data curation with language detection, deduplication, and quality scoring; factor and epoch sampling that emphasizes Spanish, Catalan, Galician, and Basque while undersampling English and code; a two-stage pre-training schedule with a continued high-quality-data phase; supervised instruction tuning on multilingual chat data with added embedding noise; and an evaluation pipeline built on an Iberian-language benchmark plus a separate judge model for open-ended tasks. The evaluation pipeline carries the competitive-performance claim, while the data pipelines are what make the models genuinely multilingual.

What would settle it

Give independent human annotators the same prompts used in the judge-based evaluation (story completion, math, paraphrase, reading comprehension, summarization, and translation) for Salamandra and the comparison models in Catalan, Basque, Galician, and Spanish, with model identity hidden, and see whether human preference rankings match the reported score ordering; a clear reversal would falsify the claim.

Watch

Extended reading notes

Core claim

The authors claim that Salamandra models, trained entirely on open-access data, reach performance competitive with comparable open baselines across a broad multilingual evaluation. On the reported benchmarks, the 7B and 40B variants often lead or tie in Catalan, Basque, and translation tasks, while mathematical reasoning and natural language inference remain relatively weak. The instruction-tuned checkpoints improve instruction following and several downstream categories, but they hurt translation and generative truthfulness. The 40B model is described as an intermediate checkpoint whose training has not finished and has not undergone the annealing phase. No released checkpoint has received preference-based alignment, and the paper states such alignment is future work.

Load-bearing premise

The conclusion stands only if the automatic benchmarks and the judge model actually measure real capability in Catalan, Basque, Galician, and Spanish; if those instruments are biased, the competitive result is not established.

Editorial extensions

If this is right

  • If the reported capabilities hold, the Apache-2.0 release gives researchers and companies deployable open models for Spanish, Catalan, Basque, Galician, and many other European languages without licensing barriers.
  • The published training recipes, data-selection logic, and evaluation scripts make the suite reproducible in spirit, so others can audit or extend the approach.
  • The instructed checkpoints should be used knowing they improve instruction following but lose ground on translation and truthfulness; task choice matters.
  • The 40B checkpoint, though unfinished, already leads the family and several comparisons, so completing its training and adding an annealing phase should improve it further.
  • The safety analysis shows that larger and instructed models lean more on social stereotypes and that red-teaming resistance varies by language, so downstream deployments need their own alignment and safety testing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because several evaluation instruments were built or translated by the same team and the judge-based scores were not checked against human raters, an independent multilingual human evaluation is the natural next test before betting on the competitive claim.
  • If the balanced tokenizer training transfers as the fertility numbers suggest, similar uniform-language tokenizer recipes could help other low-resource languages, though the report does not directly measure downstream transfer.
  • The vision experiments are explicitly a proof of concept; treating them as production multimodal capability would go beyond what the paper claims.
  • Given the absence of preference alignment, safety-sensitive applications should treat the released checkpoints as base material for further tuning, not as final products.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This technical report introduces Salamandra, a family of from-scratch decoder-only large language models of 2B, 7B, and 40B parameters, trained on a corpus covering 35 European languages and code, with instruction-tuned checkpoints, a vision proof-of-concept, and open releases under the Apache 2.0 license. The paper documents architecture, tokenizer design, data curation and processing pipelines, pre-training and continued-training recipes, post-training choices, and a large evaluation campaign: base and instructed model comparisons in Tables 9–18 using the LM Evaluation Harness, IberoBench and other benchmarks, an LLM-as-a-Judge setup with Prometheus-2 8x7B, and bias, cognitive-bias, and red-teaming analyses. The central claim is that Salamandra achieves competitive performance against similarly sized open-source models, especially in Iberian languages. The report is transparent about limitations, including an unfinished 40B checkpoint, the absence of preference alignment, and the lack of a comprehensive human evaluation setup.

Significance. If the evaluation results hold, this is a useful contribution to the European-language open-model landscape: the suite provides open weights at three sizes, public training and evaluation scripts, detailed data curation documentation, a model card, a datasheet, and an unusually honest discussion of weaknesses. The tokenizer fertility analysis and the open red-teaming pipeline are also valuable artifacts. However, the headline competitiveness claim is not yet fully established. The instructed-model comparisons in Tables 11–18 rely on an LLM judge that is not validated against human judgments for the non-English languages, and several headline Iberian-language benchmarks are developed or extended by the same team. With added human validation, uncertainty quantification, and clearer separation of preliminary 40B results, this could become a solid reference for multilingual open-weight models.

major comments (4)
  1. [§5.1, §5.2.2, Tables 11–18] The instructed-model evaluation uses Prometheus-2 8x7B as judge with an English system prompt and English rubrics while the queries and responses are in Spanish, Catalan, Basque, Galician, French, German, and Italian. The paper reports in §5.1 that Prometheus-2 'reacted' to non-English languages and required prompt iteration, and that a comprehensive human evaluation setup is 'on our current roadmap.' No human agreement or calibration check for the judge is reported. Because the instructed-model competitiveness claim rests on these scores, please add a human validation sample (e.g., 100–200 instances per language and task) with judge–human agreement metrics, or remove or substantially qualify the comparative instructed-model claims.
  2. [§5.1, Table 9] The base-model claim of 'strong capabilities' and 'competitive performance' is based on single point estimates without confidence intervals, standard errors, or significance tests. Several comparisons are within a few points (e.g., xstorycloze en: Salamandra 7B 79.09 vs. EuroLLM 9B 80.41), and §5.2.1 notes that library versions, tensor parallelism, and the use of vLLM can shift scores by 1–2% and in some cases more. The authors should either provide bootstrap intervals or other uncertainty estimates, state a minimum meaningful difference, and refrain from ranking models on gaps smaller than the reported variability.
  3. [§5.2.1, §6.2.1] IberoBench and EsBBQ are developed or extended by the same research group and are used for several headline Catalan, Basque, and Galician results. This is not invalid by itself, but it creates a correctness risk for the central comparative claim because these instruments have not been independently audited. Please (i) identify which headline conclusions depend exclusively on these in-house benchmarks, (ii) report contamination checks between these benchmarks and the pre-training or instruction-tuning data (the §3.6.1 Aya filtering indicates that evaluation-benchmark overlap is already a considered concern), and (iii) add at least one external or third-party benchmark per headline language.
  4. [§3.4, §5.2.3, Table 9] The 40B results in Table 9 come from an intermediate checkpoint whose training is still ongoing and that has not undergone the annealing phase, as stated in §3.4. Presenting these results among the final released checkpoints, and referring in the abstract to 'three different sizes,' can mislead readers about what is actually shipped. Please move the 40B results to a clearly separated preliminary section or table, and state in the abstract that the 40B release is an incomplete, non-final checkpoint.
minor comments (5)
  1. [Table 2] The row labelled 'Hungarian hr Balto-Slavic' should be labelled 'Croatian hr Balto-Slavic'; Hungarian is 'hu' and is listed separately in the same table and in Figure 4.
  2. [§2.3.2] The sentence 'A script for re-setting reserved tokens is provided is provided' contains a duplicated phrase and should be corrected.
  3. [§5.2.2] The text 'LM Evalaution Harness' should read 'LM Evaluation Harness.'
  4. [Table 9] There are unexplained 'nan' entries in the table (e.g., xquad_en for Occiglot-eu5 7B and Teuken 7B); please state how missing values were handled and whether they affect the comparisons.
  5. [§6.4] There are small typos: 'Aya 23 8B us generally more resistant' should read 'is generally more resistant,' and 'One significant issues' should be 'One significant issue.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: benchmark and judge limitations are validity risks, not circular reductions.

full rationale

The report makes no formal derivation claim; its central assertion is empirical: Salamandra 'achiev[es] competitive performance when compared to similarly sized open-source models' (Abstract) on the basis of Tables 9-18. These tables are produced by running released checkpoints on external benchmarks (Belebele, MGSM, FLORES-200, XStoryCloze, BBQ, PAWS, XLSum) plus IberoBench, a same-team benchmark that is itself built from human-annotated or human-translated datasets (Section 5.2.1, reference [19]). Using a benchmark authored partly by the same group is a validity and independence concern, not a circular reduction: no parameter is fitted to the benchmark, and the benchmark is not defined in terms of Salamandra outputs. The LLM-as-a-judge results (Section 5.2.2) rely on Prometheus-2, an external model, with English prompts and rubrics and no human validation for Catalan, Basque, or Galician; the paper explicitly notes that its human-evaluation setup is 'on our current roadmap' (Section 5.1) and that Prometheus-2 'reacted' to non-English languages. This is an openly acknowledged validation gap, not circularity. Likewise, the red-teaming section admits the 'absence of human annotation and evaluation' (Section 6.4.2). No equation is reused as its own input, no fitted prediction is renamed as a result, and no uniqueness claim is imported from the authors' prior work. The released weights and scripts make the headline results externally checkable, so the derivation chain is self-contained.

Assumptions & free parameters 7 free parameters · 6 assumptions · 0 invented entities

The central claim is conditioned on many hand-chosen engineering decisions: tokenizer size, data mixture ratios, learning rates, and evaluation protocols. None of these are derived from first principles; they are standard choices backed by preliminary experiments or prior literature. The architecture uses standard Transformer components from cited works. No new physical or theoretical entities are introduced. The paper's contribution is the combination and disclosure of these choices, not a mathematical derivation.

free parameters (7)
  • Tokenizer vocabulary size = 256,000 tokens
    Chosen after preliminary experiments (Section 2.3.1); larger vocabularies increase embedding size and affect model capacity and token fertility.
  • GQA groups = 8
    Set to 8 for 7B and 40B because it seemed to be a good trade-off in preliminary experiments (Section 2.1).
  • Peak learning rates = 2B: 2e-4; 7B: 3e-4; 40B: 5e-5
    Hand-selected per model size (Section 3.5, Table 4); standard practice rather than derived from theory.
  • Continued training data mixture = 55.51% FineWeb-Edu, 25.32% Colossal Oscar, 8.38% Wikipedia, 7.17% Aya, 3.63% StarCoder
    Chosen based on prior literature (Llama 3, Nemotron 4) rather than a quantitative search (Section 3.6.1).
  • Language sampling factors = English/code 0.5x; Spanish, Catalan, Galician, Basque oversampled; epochs 0.5 to 2
    Manual balancing to counteract English dominance and support the official Spanish languages (Section 3.1.3, Table 2).
  • NEFTune noise alpha = 5
    Chosen for instruction tuning (Section 4.1.2); no ablation is shown in the report.
  • Red-teaming inference parameters = temperature 0.8, top-p 0.95, repetition penalty 1.2
    Sampling configuration for safety evaluation (Section 6.4.1); attack success rates depend on these settings.
assumptions (6)
  • domain assumption Next-token prediction on curated web and text data transfers to downstream benchmark performance.
    The causal language modeling objective (Section 3.5) is assumed to produce models whose benchmark scores reflect useful capabilities; this is standard in LLM literature but not formally guaranteed.
  • domain assumption The evaluation benchmarks measure the capabilities claimed.
    Section 5 relies on IberoBench, MGSM, Belebele, FLORES, and other datasets as ground truth for strong and competitive performance; benchmark validity is assumed.
  • domain assumption Prometheus-2 8x7B scores correlate with human judgments in all evaluated languages.
    The LLM-as-a-judge setup (Section 5.2.2) uses Prometheus-2 without human validation; the report itself defers human evaluation to future work.
  • domain assumption Llama Guard 3 correctly identifies unsafe content in English, Spanish, and Catalan.
    Red-teaming attack success rates (Section 6.4) are computed from Llama Guard 3 judgments; the paper notes it was not trained for Catalan and has blind spots.
  • domain assumption Machine-translated red-teaming prompts preserve harmfulness across languages.
    Catalan prompts are NLLB translations (Section 6.4.1); authors acknowledge that translation may create a false impression of model quality.
  • standard math Base Transformer and scaling assumptions from prior work hold.
    Architecture choices such as RoPE, SwiGLU, RMSNorm, and GQA are adopted from cited prior work (Section 2.1), not derived in this report.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Salamandra Technical Report." pith.science (2026). https://pith.science/paper/PXH5SZVV

@misc{pith2026250208489,
  author       = {Pith},
  title        = {Pith review of: Salamandra Technical Report},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PXH5SZVV}},
  note         = {Machine review of arXiv:2502.08489}
}
read the original abstract

This work introduces Salamandra, a suite of open-source decoder-only large language models available in three different sizes: 2, 7, and 40 billion parameters. The models were trained from scratch on highly multilingual data that comprises text in 35 European languages and code. Our carefully curated corpus is made exclusively from open-access data compiled from a wide variety of sources. Along with the base models, supplementary checkpoints that were fine-tuned on public-domain instruction data are also released for chat applications. Additionally, we also share our preliminary experiments on multimodality, which serve as proof-of-concept to showcase potential applications for the Salamandra family. Our extensive evaluations on multilingual benchmarks reveal that Salamandra has strong capabilities, achieving competitive performance when compared to similarly sized open-source models. We provide comprehensive evaluation results both on standard downstream tasks as well as key aspects related to bias and safety.With this technical report, we intend to promote open science by sharing all the details behind our design choices, data curation strategy and evaluation methodology. In addition to that, we deviate from the usual practice by making our training and evaluation scripts publicly accessible. We release all models under a permissive Apache 2.0 license in order to foster future research and facilitate commercial use, thereby contributing to the open-source ecosystem of large language models.

Figures

Figures reproduced from arXiv: 2502.08489 by the authors.

Figure 1
Figure 1. Comparison of tokenizer fertility (i.e. tokens-per-word) across multiple languages: Catalan, [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Fertility score of a tokenizer trained on a balanced dataset where each language is repre [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Distribution of sources in the Salamandra pre-training dataset. Each data point represents [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (19 more)
Figure 4
Figure 4. Figure 4: Distribution of tokens in the pre-training and continued training phase corpus after applying [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Overview of data distribution in visual instruction tuning phases. In total, the dataset [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Visual summary of the setup used to evaluate the capabilities of the Salamandra family of [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]
Figure 7
Figure 7. Figure 7: Accuracy and difference scores in ambiguous and disambiguating contexts for each category [PITH_FULL_IMAGE:figures/full_fig_p036_7.png]
Figure 8
Figure 8. Figure 8: Accuracy and difference scores in ambiguous and disambiguating contexts for each category [PITH_FULL_IMAGE:figures/full_fig_p036_8.png]
Figure 9
Figure 9. Figure 9: Frequency distributions of predicted answers on ARC Easy and Challenge subsets depending [PITH_FULL_IMAGE:figures/full_fig_p038_9.png]
Figure 10
Figure 10. Figure 10: Frequency distributions of class 0 predictions on SST-2 dataset depending on the class distribution in few-shot. 0 denotes the negative class, while 1 denotes the positive class. English. While this approach has yielded valuable insights, it is somewhat limited by the…
Figure 11
Figure 11. Figure 11: M-AdvBench Dataset - Prompts per Hazard Category. Instances translated from EN are [PITH_FULL_IMAGE:figures/full_fig_p040_11.png]
Figure 12
Figure 12. Figure 12: HH-RLHF RT Dataset - Prompts per Hazard Category. Instances translated from EN are [PITH_FULL_IMAGE:figures/full_fig_p040_12.png]
Figure 13
Figure 13. Figure 13: Aya RT Dataset - Prompts per Hazard Category. Instances translated from EN are [PITH_FULL_IMAGE:figures/full_fig_p041_13.png]
Figure 14
Figure 14. Figure 14: Fertility scores for Germanic languages, namely Danish, German, English, Dutch, [PITH_FULL_IMAGE:figures/full_fig_p072_14.png]
Figure 15
Figure 15. Figure 15: Fertility scores for Romance languages, namely Catalan, Spanish, French, Galician, Italian, [PITH_FULL_IMAGE:figures/full_fig_p072_15.png]
Figure 16
Figure 16. Figure 16: Fertility scores for Balto-Slavic languages, namely Bulgarian, Czech, [PITH_FULL_IMAGE:figures/full_fig_p073_16.png]
Figure 17
Figure 17. Figure 17: Fertility scores for code and languages that belong to smaller families, namely Welsh, [PITH_FULL_IMAGE:figures/full_fig_p073_17.png]
Figure 18
Figure 18. Figure 18: Optical Character Recognition examples in English (up) and Spanish (down). [PITH_FULL_IMAGE:figures/full_fig_p091_18.png]
Figure 19
Figure 19. Figure 19: Captioning example in Spanish. Salamandra is asked to provide a detailed description [PITH_FULL_IMAGE:figures/full_fig_p091_19.png]
Figure 20
Figure 20. Figure 20: Multi-image example in Catalan. Salamandra responds correctly when asked which dish [PITH_FULL_IMAGE:figures/full_fig_p092_20.png]
Figure 21
Figure 21. Figure 21: Grounding example in Spanish. Salamandra provides a numbered list to describe the colors [PITH_FULL_IMAGE:figures/full_fig_p092_21.png]
Figure 22
Figure 22. Figure 22: Video Analysis example in English. Given a low-quality video of a meeting room, [PITH_FULL_IMAGE:figures/full_fig_p092_22.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering

    cs.CL 2025-07 conditional novelty 7.0 of 10

    EsBBQ and CaBBQ are new Spanish and Catalan bias benchmarks for multiple-choice QA, built with survey-validated stereotypes from Spain and evaluated on 17 language models.

  2. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Unsigned differential activations locate a few GLU-MLP neurons whose zeroing surgically destabilizes demographic bias while retaining ~99.5% of measured capabilities.

  3. Do LLMs exhibit the same commonsense capabilities across languages?

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LLMs produce more commonsensical sentences in English than in Spanish, Dutch, or Valencian, across automatic, LLM-judge, and human evaluations on the new MULTICOM benchmark.

  4. MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 11-task Maltese benchmark shows that 55 large language models lag behind small fine-tuned models, with prior Maltese exposure the strongest predictor.

  5. IberBench: LLM Evaluation on Iberian Languages

    cs.CL 2025-04 conditional novelty 6.0 of 10

    A community-run benchmark for Iberian languages shows that LLMs underperform on industry-relevant NLP tasks and on Basque and Galician relative to fundamental tasks and other Iberian languages.

  6. Emo Pillars: Knowledge Distillation to Support Fine-Grained Context-Aware and Context-Less Emotion Classification

    cs.CL 2025-04 conditional novelty 6.0 of 10

    Emo Pillars is a synthetic 400K-utterance, 28-class emotion dataset generated by Mistral-7B from narrative plots; fine-tuned RoBERTa/BERT models achieve SOTA or competitive F1 on GoEmotions, ISEAR, IEMOCAP, and EmoContext.

  7. A Sovereign, Open-Source Foundation Model for German and English

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Soofi S 30B-A3B, a hybrid Mamba-MoE model pretrained on ~27T tokens with deliberately up-weighted German, reports the highest English and German aggregate scores among fully open base models in its comparison while ma...

  8. S-DiverSe: Spanish Diverse Speech

    cs.CL 2026-07 conditional novelty 5.5 of 10

    S-DiverSe is a 3.2-hour multi-pathology Spanish pathological-speech benchmark on which heuristic post-processing outperforms fine-tuning for out-of-domain ASR.

  9. Analysis of Numerical Localisation in LLM Translations

    cs.CL 2026-08 conditional novelty 5.0 of 10

    For English-German numerical localisation with five small LLMs, in-context learning with explicit formatting rules beats direct translation, chain-of-thought, and post-editing in mean accuracy.

  10. Bielik 11B v2 Technical Report

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Bielik 11B v2, a depth-upscaled Mistral model continued-pretrained on Polish data, scores at or near the top of several Polish benchmarks despite having far fewer parameters than leading rivals.

Reference graph

Works this paper leans on

241 extracted references · 3 canonical work pages · cited by 10 Pith papers

  1. [1]

    The multilingual alignment prism: Aligning global and local preferences to reduce harm, 2024

    Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker. The multilingual alignment prism: Aligning global and local preferences to reduce harm, 2024. URL https://arxiv.org/abs/2406.18682

  2. [2]

    Ungoliant: An optimized pipeline for the generation of a very large-scale multilingual web corpus

    Julien Abadji, Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. Ungoliant: An optimized pipeline for the generation of a very large-scale multilingual web corpus. In Harald Lüngen, Marc Kupietz, Piotr Ba´nski, Adrien Barbaresi, Simon Clematide, and Ines Pisetta, editors, Proceedings of the Workshop on Challenges in the Management of Large Corp...

  3. [3]

    Towards a cleaner document-oriented multilingual crawled corpus, 2022

    Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, and Benoît Sagot. Towards a cleaner document-oriented multilingual crawled corpus, 2022. URL https://arxiv.org/abs/ 2201.06642

  4. [4]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  5. [5]

    Nemotron-4 340b technical report

    Bo Adler, Niket Agarwal, Ashwath Aithal, Dong H Anh, Pallab Bhattacharya, Annika Brundyn, Jared Casper, Bryan Catanzaro, Sharon Clay, Jonathan Cohen, et al. Nemotron-4 340b technical report. arXiv preprint arXiv:2406.11704, 2024

  6. [6]

    Gqa: Training generalized multi-query transformer models from multi-head checkpoints

    Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. arXiv preprint arXiv:2305.13245, 2023

  7. [7]

    Teuken-7b-base & teuken-7b-instruct: Towards european LLMs,

    Mehdi Ali, Michael Fromm, Klaudia Thellmann, Jan Ebert, Alexander Arno Weber, Richard Rutmann, Charvi Jain, Max Lübbering, Daniel Steinigen, Johannes Leveling, Katrin Klug, Jasper Schulze Buschhoff, Lena Jurkschat, Hammam Abdelwahab, Benny Jörg Stein, Karl- Heinz Sylla, Pavel Denisov, Nicolo’ Brandizzi, Qasid Saleem, Anirban Bhowmick, Lennard Helmer, Chel...

  8. [8]

    The falcon series of open language models

    Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, et al. The falcon series of open language models. arXiv preprint arXiv:2311.16867, 2023

Show all 241 references
  1. [9]

    Alves, José Pombal, Nuno M

    Duarte M. Alves, José Pombal, Nuno M. Guerreiro, Pedro H. Martins, João Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, José G. C. de Souza, and André F. T. Martins. Tower: An open multilingual large language model for translati...

  2. [10]

    The claude 3 model family: Opus, sonnet, haiku

    Anthropic. The claude 3 model family: Opus, sonnet, haiku. https://www-cdn.anthropic. com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf,

  3. [11]

    Does corpus quality really matter for low-resource languages?, 2022

    Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de Viñaspre, and Aitor Soroa. Does corpus quality really matter for low-resource languages?, 2022

  4. [12]

    Accessed: 2024-12-01

  5. [13]

    Occiglot at WMT24: European open-source large language models evaluated on translation

    Eleftherios Avramidis, Annika Grützner-Zahn, Manuel Brack, Patrick Schramowski, Pedro Ortiz Suarez, Malte Ostendorff, Fabio Barth, Shushen Manakhimova, Vivien Macketanz, Georg Rehm, and Kristian Kersting. Occiglot at WMT24: European open-source large language models evaluated ...

  6. [14]

    Aya 23: Open weight releases to further multilingual progress

    Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, et al. Aya 23: Open weight releases to further multilingual progress. arXiv preprint arXiv:2405.15032, 2024. 44

  7. [15]

    A critical analysis of the largest source for generative ai training data: Com- mon crawl

    Stefan Baack. A critical analysis of the largest source for generative ai training data: Com- mon crawl. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’24, page 2199–2208, New York, NY , USA, 2024. Association for Computing Mach...

  8. [16]

    Layer normalization

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016

  9. [17]

    The belebele benchmark: a parallel reading comprehension dataset in 122 language variants

    Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. In Lun-Wei Ku, And...

  10. [18]

    Qwen technical report

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023

  11. [19]

    IberoBench: A benchmark for LLM evaluation in Iberian languages

    Irene Baucells, Javier Aula-Blasco, Iria de-Dios-Flores, Silvia Paniagua Suárez, Naiara Perez, Anna Salles, Susana Sotelo Docio, Júlia Falcão, José Javier Saiz, Robiert Sepulveda Torres, Jeremy Barnes, Pablo Gamallo, Aitor Gonzalez-Agirre, German Rigau, and Marta Villegas. Ibe...

  12. [20]

    A unified taxonomy of harmful content

    Michele Banko, Brendon MacKeen, and Laurie Ray. A unified taxonomy of harmful content. In Workshop on Abusive Language Online, 2020. URL https://api.semanticscholar. org/CorpusID:226283543

  13. [21]

    Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell

    Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623, New York, ...

  14. [22]

    Marta Bañón, Miquel Esplà-Gomis, Mikel L. Forcada, Cristian García-Romero, Taja Kuzman, Nikola Ljubeši ´c, Rik van Noord, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Peter Rupnik, Vít Suchomel, Antonio Toral, Tobias van der Werff, and Jaume Zaragoza. MaCoCu: Massive collection...

  15. [23]

    Gpt-neox-20b: An open-source autoregressive language model

    Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint arXiv:2204.06745, 2022. 45

  16. [24]

    Pythia: A suite for analyzing large language models across training and scaling

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In Intern...

  17. [25]

    Community OSCAR: A community effort for multilingual web data,

    Manuel Brack, Malte Ostendorff, Pedro Ortiz Suarez, José Javier Saiz, Iñaki Lacunza Castilla, Jorge Palomar-Giner, Aleksandr Shvets, Patrick Schramowski, Georg Rehm, Marta Villegas, and Kristian Kersting. Community OSCAR: A community effort for multilingual web data,

  18. [26]

    Smith, and Luke Zettlemoyer

    Terra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, and Luke Zettlemoyer. Breaking the curse of multilinguality with cross-lingual expert language models, 2024. URL http://arxiv.org/abs/2401.10440

  19. [27]

    Pretrained biomedical language models for clinical NLP in Spanish

    Casimiro Pio Carrino, Joan Llop, Marc Pàmies, Asier Gutiérrez-Fandiño, Jordi Armengol- Estapé, Joaquín Silveira-Ocampo, Alfonso Valencia, Aitor Gonzalez-Agirre, and Marta Ville- gas. Pretrained biomedical language models for clinical NLP in Spanish. In Proceedings of the 21st ...

  20. [28]

    URL https://occiglot.eu/papers/Community_Oscar.pdf

  21. [29]

    Language models are few-shot learners

    Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020

  22. [30]

    Cross-lingual ability of multilingual masked language models: A study of language structure, 2022

    Yuan Chai, Yaobo Liang, and Nan Duan. Cross-lingual ability of multilingual masked language models: A study of language structure, 2022. URL https://arxiv.org/abs/2203.08430. Version Number: 1

  23. [31]

    Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K

    Tyler A. Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K. Bergen. When is multilin- guality a curse? language modeling for 250 high- and low-resource languages, 2023. URL https://arxiv.org/abs/2311.09205. Version Number: 1

  24. [32]

    Isaac Caswell, Julia Kreutzer, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Auguste Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios Gonzales, I...

  25. [33]

    Ernie-code: Beyond english-centric cross-lingual pretraining for programming languages

    Yekun Chai, Shuohuan Wang, Chao Pang, Yu Sun, Hao Tian, and Hua Wu. Ernie-code: Beyond english-centric cross-lingual pretraining for programming languages. arXiv preprint arXiv:2212.06742, 2022

  26. [34]

    Is it good data for multilingual instruction tuning or just bad multilingual evaluation for large language models?arXiv preprint arXiv:2406.12822, 2024

    Pinzhen Chen, Simon Yu, Zhicheng Guo, and Barry Haddow. Is it good data for multilingual instruction tuning or just bad multilingual evaluation for large language models?arXiv preprint arXiv:2406.12822, 2024

  27. [35]

    Palm: Scaling language modeling with pathways

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240): 1–113, 2023. 46

  28. [37]

    Evaluating large language models trained on code

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021

  29. [38]

    Think you have solved question answering? try arc, the ai2 reasoning challenge

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1, 2018

  30. [39]

    COMMISSION DECISION C/2024/1459 Establishing the European Artificial Intelligence Office

    European Commission. COMMISSION DECISION C/2024/1459 Establishing the European Artificial Intelligence Office. Official Journal of the European Union , 2024-01-24. URL http://data.europa.eu/eli/C/2024/1459/oj

  31. [40]

    Breaking down the defenses: A comparative survey of attacks on large language models

    Arijit Ghosh Chowdhury, Md Mofijul Islam, Vaibhav Kumar, Faysal Hossain Shezan, Vinija Jain, and Aman Chadha. Breaking down the defenses: A comparative survey of attacks on large language models. arXiv preprint arXiv:2403.04786, 2024

  32. [41]

    Scaling instruction-finetuned language models

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1–53, 2024

  33. [42]

    Free dolly: Introducing the world’s first truly open instruction-tuned llm

    Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. Free dolly: Introducing the world’s first truly open instruction-tuned llm. Databricks, 2023

  34. [43]

    Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

    Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023. URL https://www.databricks.com/blog/2023/ 04/12/dolly-first-op...

  35. [44]

    RedPajama: An open source recipe to reproduce LLaMA training dataset,

    Together Computer. RedPajama: An open source recipe to reproduce LLaMA training dataset,

  36. [45]

    Flor: On the effectiveness of language adaptation

    Severino Da Dalt, Joan Llop, Irene Baucells, Marc Pàmies, Yishi Xu, Aitor Gonzalez-Agirre, and Marta Villegas. Flor: On the effectiveness of language adaptation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Eval...

  37. [46]

    Unsupervised cross-lingual representation learning at scale, 2020

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale, 2020. URL http://arxiv.org/ abs/1911.02116

  38. [47]

    Flashattention-2: Faster attention with better parallelism and work partitioning (2023)

    Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning (2023). arXiv preprint arXiv:2307.08691, 2023

  39. [48]

    Flashattention: Fast and memory-efficient exact attention with io-awareness

    Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 35:16344–16359, 2022

  40. [49]

    Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansan...

  41. [50]

    CorpusNÓS: A massive Galician corpus for training large language models

    Iria de Dios-Flores, Silvia Paniagua Suárez, Cristina Carbajal Pérez, Daniel Bardanca Out- eiriño, Marcos Garcia, and Pablo Gamallo. CorpusNÓS: A massive Galician corpus for training large language models. In Pablo Gamallo, Daniela Claro, António Teixeira, Livy Real, Marcos Ga...

  42. [51]

    Rlhf can speak many languages: Unlocking multilingual preference optimization for llms

    John Dang, Arash Ahmadian, Kelly Marchisio, Julia Kreutzer, Ahmet Üstün, and Sara Hooker. Rlhf can speak many languages: Unlocking multilingual preference optimization for llms. arXiv preprint arXiv:2407.02552, 2024

  43. [52]

    Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021

    Jesse Dodge, Maarten Sap, Ana Marasovi´c, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021. URL https://arxiv.org/abs/2104.08758

  44. [53]

    The llama 3 herd of models

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024

  45. [54]

    Language modeling with gated convolutional networks

    Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. Language modeling with gated convolutional networks. In International conference on machine learning, pages 933–941. PMLR, 2017. 47

  46. [55]

    Tomaž Erjavec, Maciej Ogrodniczuk, Petya Osenova, Nikola Ljubeši ´c, Kiril Simov, Vladislava Grigorova, Michał Rudolf, Andrej Pan ˇcur, Matyáš Kopp, Starkaður Barkarson, Stein{\textbackslash}t hór Steingrímsson, Henk van der Pol, Griet Depoorter, Jesse de Does, Bart Jongejan, ...

  47. [56]

    A new massive multilingual dataset for high-performance language technologies, 2024

    Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, and Jörg Tiedemann. A new massive multilingual dataset for high-performance ...

  48. [57]

    Towards ana- lyzing and understanding the limitations of dpo: A theoretical perspective

    Duanyu Feng, Bowen Qin, Chen Huang, Zheng Zhang, and Wenqiang Lei. Towards ana- lyzing and understanding the limitations of dpo: A theoretical perspective. arXiv preprint arXiv:2404.04626, 2024

  49. [58]

    BLEU might be guilty but references are not innocent

    Markus Freitag, David Grangier, and Isaac Caswell. BLEU might be guilty but references are not innocent. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors,Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 61–71, ...

  50. [59]

    Ljubeši´c, and N

    Tomaž Erjavec, N. Ljubeši´c, and N. Logar. The slWaC corpus of the slovene web. Informatica (Slovenia), 39:35–42, 2015

  51. [61]

    DoGE: Domain reweighting with general- ization estimation, 2023

    Simin Fan, Matteo Pagliardini, and Martin Jaggi. DoGE: Domain reweighting with general- ization estimation, 2023. URL https://arxiv.org/abs/2310.15393. Version Number: 2

  52. [62]

    A framework for few-shot language model evaluation, 12 2023

    Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...

  53. [63]

    Datasheets for datasets

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. Datasheets for datasets. Communications of the ACM, 64(12):86–92, 2021

  54. [64]

    Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned

    Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858, 2022

  55. [65]

    Building a data infrastructure for a mid-resource language: The case of Catalan

    Aitor Gonzalez-Agirre, Montserrat Marimon, Carlos Rodriguez-Penagos, Javier Aula- Blasco, Irene Baucells, Carme Armentano-Oller, Jorge Palomar-Giner, Baybars Kulebi, and Marta Villegas. Building a data infrastructure for a mid-resource language: The case of Catalan. In Nicolet...

  56. [66]

    The pile: An 800gb dataset of diverse text for language modeling

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling. ArXiv, abs/2101.00027, 2021. URL https://...

  57. [67]

    Accelerate: Training and inference at scale made simple, efficient and adaptable

    Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan. Accelerate: Training and inference at scale made simple, efficient and adaptable. https://github.com/huggingface/accelerate, 2022

  58. [68]

    Spanish legalese language model and corpora, 2021

    Asier Gutiérrez-Fandiño, Jordi Armengol-Estapé, Aitor Gonzalez-Agirre, and Marta Villegas. Spanish legalese language model and corpora, 2021

  59. [69]

    Building a data infrastructure for a mid-resource language: The case of Catalan

    Aitor Gonzalez-Agirre, Montserrat Marimon, Carlos Rodriguez-Penagos, Javier Aula- Blasco, Irene Baucells, Carme Armentano-Oller, Jorge Palomar-Giner, Baybars Kulebi, and Marta Villegas. Building a data infrastructure for a mid-resource language: The case of Catalan. In Nicolet...

  60. [70]

    Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong- Bin Kang, M

    Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong- Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. XL-sum: Large-scale multilingual ab- stractive summarization for 44 languages. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, e...

  61. [71]

    Olmo: Accelerating the science of language models

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, A Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. Olmo: Accelerating the science of language models. arxiv preprint, 2024. URL https://api. semanticscholar. org/CorpusID, 267365485, 2024

  62. [72]

    Pile of law: learning responsible data filtering from the law and a 256gb open-source legal dataset

    Peter Henderson, Mark S Krass, Lucia Zheng, Neel Guha, Christopher D Manning, Dan Jurafsky, and Daniel E Ho. Pile of law: learning responsible data filtering from the law and a 256gb open-source legal dataset. In Proceedings of the 36th International Conference on Neural Infor...

  63. [73]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020

  64. [74]

    The danish parliament corpus 2009 - 2017, v1, 2018

    Dorte Haltrup Hansen. The danish parliament corpus 2009 - 2017, v1, 2018. URL http: //hdl.handle.net/20.500.12115/8

  65. [75]

    Training compute-optimal large language models

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022

  66. [76]

    Model performance scaling with multiple data sources

    Tatsunori Hashimoto. Model performance scaling with multiple data sources. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139, pages 4107–4116. PMLR, 2021. URL https://proceedings.mlr. press/v139/hashimoto2...

  67. [77]

    Perplexed by quality: A perplexity-based method for adult and harmful content detection in multilingual heterogeneous web data, 2022

    Tim Jansen, Yangling Tong, Victoria Zevallos, and Pedro Ortiz Suarez. Perplexed by quality: A perplexity-based method for adult and harmful content detection in multilingual heterogeneous web data, 2022. URL https://arxiv.org/abs/2212.10440

  68. [78]

    Mistral 7b

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023

  69. [79]

    Measuring mathematical problem solving with the MATH dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. ArXiv, 2021

  70. [80]

    Kobbq: Korean bias benchmark for question answering

    Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Alice Oh, and Hwaran Lee. Kobbq: Korean bias benchmark for question answering. Transactions of the Association for Computational Linguistics, 12:507–524, 2024

  71. [81]

    Neftune: Noisy embeddings improve instruction finetuning

    Neel Jain, Ping-yeh Chiang, Yuxin Wen, John Kirchenbauer, Hong-Min Chu, Gowthami Somepalli, Brian R Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Aniruddha Saha, et al. Neftune: Noisy embeddings improve instruction finetuning. arXiv preprint arXiv:2310.05914, 2023

  72. [82]

    The state and fate of linguistic diversity and inclusion in the NLP world, 2021

    Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. The state and fate of linguistic diversity and inclusion in the NLP world, 2021. URL http: //arxiv.org/abs/2004.09095

  73. [83]

    Fasttext.zip: Compressing text classification models, 2016

    Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. Fasttext.zip: Compressing text classification models, 2016. URL https://arxiv. org/abs/1612.03651

  74. [84]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...

  75. [85]

    Scaling laws for neural language models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020

  76. [86]

    Natural language processing for dialects of a language: A survey

    Aditya Joshi, Raj Dabre, Diptesh Kanojia, Zhuang Li, Haolan Zhan, Gholamreza Haffari, and Doris Dippold. Natural language processing for dialects of a language: A survey. arXiv preprint arXiv:2401.05632, 2024

  77. [87]

    The biggen bench: A principled benchmark for fine-grained evaluation of language models with language models, 2024

    Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko...

  78. [88]

    Prometheus 2: An open source language model specialized in evaluating other language models, 2024

    Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. Prometheus 2: An open source language model specialized in evaluating other language models, 2024. URL https: //arxiv.org/abs/2405.01535

  79. [89]

    Bag of tricks for efficient text classification, 2016

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of tricks for efficient text classification, 2016. URL https://arxiv.org/abs/1607.01759

  80. [90]

    Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36, 2024

    Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, et al. Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Proce...

  81. [91]

    A diagram is worth a dozen images, 2016

    Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images, 2016. URL https://arxiv.org/abs/1603. 07396. 50

  82. [92]

    Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii- Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Is- abel Papadimi...

  83. [93]

    Nemo: a toolkit for building ai applications using neural modules

    Oleksii Kuchaiev, Jason Li, Huyen Nguyen, Oleksii Hrinchuk, Ryan Leary, Boris Ginsburg, Samuel Kriman, Stanislav Beliaev, Vitaly Lavrukhin, Jack Cook, et al. Nemo: a toolkit for building ai applications using neural modules. arXiv preprint arXiv:1909.09577, 2019

  84. [94]

    Adam: A method for stochastic optimization

    Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014

  85. [95]

    Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing

    Taku Kudo and John Richardson. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226, 2018

  86. [96]

    Eesti keele ühendkorpuste sari 2013–2021: mahukaim eestikeelsete digitekstide kogu

    Kristina Koppel and Jelena Kallas. Eesti keele ühendkorpuste sari 2013–2021: mahukaim eestikeelsete digitekstide kogu. Eesti Rakenduslingvistika Ühingu aastaraamat Estonian Papers in Applied Linguistics, 18:207–228, 2022. ISSN 1736-2563, 2228-0677. doi: 10.5128/ erya18.12. URL...

  87. [97]

    Towards a comprehensive taxonomy and large-scale annotated corpus for online slur usage

    Jana Kurrek, Haji Mohammad Saleem, and Derek Ruths. Towards a comprehensive taxonomy and large-scale annotated corpus for online slur usage. In Workshop on Abusive Language Online, 2020. URL https://api.semanticscholar.org/CorpusID:226283752. 51

  88. [98]

    URL https://aclanthology.org/2022.tacl-1.4

    doi: 10.1162/tacl_a_00447. URL https://aclanthology.org/2022.tacl-1.4

  89. [99]

    Openassistant conversations – democratizing large language model alignment, 2023

    Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and A...

  90. [100]

    Subword regularization: Improving neural network translation models with multiple subword candidates

    Taku Kudo. Subword regularization: Improving neural network translation models with multiple subword candidates. arXiv preprint arXiv:1804.10959, 2018

  91. [101]

    T \" ulu 3: Pushing frontiers in open language model post-training

    Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, et al. T \" ulu 3: Pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124, 2024

  92. [102]

    Operational- izing a national digital library: The case for a norwegian transformer model

    Per E Kummervold, Javier De la Rosa, Freddy Wetjen, and Svein Arne Brygfjeld. Operational- izing a national digital library: The case for a norwegian transformer model. In Simon Dobnik and Lilja Øvrelid, editors, Proceedings of the 23rd Nordic Conference on Computational Lingu...

  93. [103]

    The national corpus of polish (NKJP)

    Barbara Lewandowska-Tomaszczyk, Rafał Górski, Marek Łazi´nski, and Adam Przepiórkowski. The national corpus of polish (NKJP). language use and data analysis, 2013

  94. [104]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems...

  95. [105]

    StarCoder: may the source be with you! ArXiv, 2023

    Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, ...

  96. [106]

    SYN v9: large corpus of written czech,

    Michal K ˇren, Václav Cvr ˇcek, Jan Henyš, Milena Hnátková, Tomáš Jelínek, Jan Kocek, Dominika Kováˇríková, Jan Kˇrivan, Jiˇrí Miliˇcka, Vladimír Petkeviˇc, Pavel Procházka, Hana Skoumalová, Jana Šindlerová, and Michal Škrabal. SYN v9: large corpus of written czech,

  97. [107]

    Openorca: An open dataset of gpt augmented flan reasoning traces

    Wing Lian, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet V ong, and "Teknium". Openorca: An open dataset of gpt augmented flan reasoning traces. https://https:// huggingface.co/Open-Orca/OpenOrca, 2023

  98. [108]

    Openorca: An open dataset of gpt augmented flan reasoning traces

    Wing Lian, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet V ong, and "Teknium". Openorca: An open dataset of gpt augmented flan reasoning traces. https://https:// huggingface.co/Open-Orca/OpenOrca, 2023. 52

  99. [109]

    Bloom: A 176b-parameter open-access multilingual language model

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili´c, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. Bloom: A 176b-parameter open-access multilingual language model. ArXiv, 2022

  100. [110]

    ROUGE: A package for automatic evaluation of summaries

    Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. InText Summariza- tion Branches Out, pages 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. URL https://aclanthology.org/W04-1013/

  101. [111]

    Llava-onevision: Easy visual task transfer, 2024

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer, 2024. URL https://arxiv.org/abs/2408.03326

  102. [112]

    OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles

    Pierre Lison and Jörg Tiedemann. OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, an...

  103. [113]

    Leveraging large language models for NLG evaluation: Advances and challenges

    Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, Yuxuan Lai, Chongyang Tao, and Shuai Ma. Leveraging large language models for NLG evaluation: Advances and challenges. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empir...

  104. [114]

    Best practices and lessons learned on synthetic data for language models

    Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zhang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, et al. Best practices and lessons learned on synthetic data for language models. arXiv preprint arXiv:2404.07503, 2024

  105. [115]

    bs,hr,srWaC - web corpora of bosnian, croatian and serbian

    Nikola Ljubeši´c and Filip Klubi ˇcka. bs,hr,srWaC - web corpora of bosnian, croatian and serbian. In Felix Bildhauer and Roland Schäfer, editors, Proceedings of the 9th Web as Corpus Workshop (WaC-9), pages 29–35. Association for Computational Linguistics, 2014. doi: 10.3115/...

  106. [116]

    Manning, Christopher Ré, Diana Acosta-Navas, Drew A

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas...

  107. [117]

    Decoupled weight decay regularization

    I Loshchilov. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017

  108. [118]

    Few-shot learning with multilingual generative language models

    Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mon...

  109. [120]

    Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

    Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, January 2024. URL https://llava-vl.github.io/blog/2024-01-30-llava-next/

  110. [121]

    Listening to affected communities to define extreme speech: Dataset and experiments

    Antonis Maronikolakis, Axel Wisiorek, Leah Nann, Haris Jabbar, Sahana Udupa, and Hinrich Schütze. Listening to affected communities to define extreme speech: Dataset and experiments. In Findings, 2022. URL https://api.semanticscholar.org/CorpusID:247596954

  111. [123]

    On llms-driven synthetic data generation, curation, and evaluation: A survey, 2024

    Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, and Haobo Wang. On llms-driven synthetic data generation, curation, and evaluation: A survey, 2024. URL https://arxiv. org/abs/2406.15126, 2024

  112. [124]

    Pre-training data quality and quantity for a low-resource language: New corpus and BERT models for maltese

    Kurt Micallef, Albert Gatt, Marc Tanti, Lonneke van der Plas, and Claudia Borg. Pre-training data quality and quantity for a low-resource language: New corpus and BERT models for maltese. In Proceedings of the Third Workshop on Deep Learning for Low-Resource Natural Language P...

  113. [125]

    Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity

    Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,Proceedings of the 60th Annual ...

  114. [126]

    Fp8 formats for deep learning

    Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, et al. Fp8 formats for deep learning. arXiv preprint arXiv:2209.05433, 2022

  115. [127]

    At which training stage does code data help LLMs reasoning?, 2023

    Yingwei Ma, Yue Liu, Yue Yu, Yuanliang Zhang, Yu Jiang, Changjian Wang, and Shanshan Li. At which training stage does code data help LLMs reasoning?, 2023. URL http: //arxiv.org/abs/2309.16298

  116. [128]

    Cognitive biases, task complexity, and result intepretability in large language models

    Mario Mina, Valle Ruiz-Fernández, Júlia Falcão, Luis Vasquez-Reina, and Aitor Gonzalez- Agirre. Cognitive biases, task complexity, and result intepretability in large language models. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Ste...

  117. [129]

    Large language models: A survey

    Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. Large language models: A survey. arXiv preprint arXiv:2402.06196, 2024

  118. [130]

    Guerreiro, Ricardo Rei, Duarte M

    Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, and André F. T. Martins. EuroLLM: Multil...

  119. [131]

    Model cards for model reporting

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency , pages 220–229, 2019

  120. [132]

    Mixed precision training

    Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. Mixed precision training. arXiv preprint arXiv:1710.03740, 2017

  121. [133]

    Efficient large-scale language model training on gpu clusters using megatron-lm

    Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, et al. Efficient large-scale language model training on gpu clusters using megatron-lm. In Proceedin...

  122. [134]

    Extending off-the- shelf NER systems to personal information detection in dialogues with a virtual agent: Findings from a real-life use case

    Mario Mina, Carlos Rodríguez, Aitor Gonzalez-Agirre, and Marta Villegas. Extending off-the- shelf NER systems to personal information detection in dialogues with a virtual agent: Findings from a real-life use case. In Elena V olodina, David Alfter, Simon Dobnik, Therese Lind- ...

  123. [135]

    H100 Tensor Core GPU Architecture Overview

    NVIDIA. H100 Tensor Core GPU Architecture Overview. https://resources.nvidia. com/en-us-tensor-core , 2022

  124. [136]

    Polish parliamentary corpus, 2018

    Maciej Ogrodniczuk. Polish parliamentary corpus, 2018. URL https://api. semanticscholar.org/CorpusID:235134113

  125. [137]

    Beyond scale: The diversity coefficient as a data quality metric for variability in natural language data

    Brando Miranda, Alycia Lee, Sudharsan Sundar, Allison Casasola, and Sanmi Koyejo. Beyond scale: The diversity coefficient as a data quality metric for variability in natural language data. ArXiv, 2023. doi: 10.48550/ARXIV .2306.13840. URL https://arxiv.org/abs/2306. 13840. Pub...

  126. [138]

    Towards an open platform for legal information

    Malte Ostendorff, Till Blume, and Saskia Ostendorff. Towards an open platform for legal information. In Proceedings of the ACM/IEEE Joint Conference on Digital Libraries in 2020, JCDL ’20, pages 385–388. Association for Computing Machinery, 2020. ISBN 978-1- 4503-7585-6. doi: ...

  127. [139]

    Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel

    Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. Scaling data-constrained language models, 2023. URL http://arxiv.org/abs/2305.16264. 54

  128. [140]

    Word embeddings from large-scale greek web content

    Stamatis Outsios, Konstantinos Skianis, Polykarpos Meladianos, Christos Xypolopoulos, and Michalis Vazirgiannis. Word embeddings from large-scale greek web content. ArXiv, 2018

  129. [141]

    Spruit, C

    NLLB Team, Marta Ruiz Costa-jussà, James Cross, Onur cCelebi, Maha Elbayad, Ken- neth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Mail- lard, Anna Sun, Skyler Wang, Guillaume Wenzek, Alison Youngblood, Bapi Akula, Loïc Barrault, Gabriel Mejia Gonz...

  130. [142]

    A CURATEd CATalog: Rethinking the extraction of pretraining corpora for mid-resourced languages

    Jorge Palomar-Giner, Jose Javier Saiz, Ferran Espuña, Mario Mina, Severino Da Dalt, Joan Llop, Malte Ostendorff, Pedro Ortiz Suarez, Georg Rehm, Aitor Gonzalez-Agirre, and Marta Villegas. A CURATEd CATalog: Rethinking the extraction of pretraining corpora for mid-resourced lan...

  131. [143]

    Multi-granular legal topic classification on greek legislation

    Christos Papaloukas, Ilias Chalkidis, Konstantinos Athinaios, Despina-Athanasia Pantazi, and Manolis Koubarakis. Multi-granular legal topic classification on greek legislation. In Proceedings of the Natural Legal Language Processing Workshop 2021, pages 63–75. As- sociation fo...

  132. [144]

    ChatML, 2022

    OpenAI. ChatML, 2022. URL https://github.com/openai/openai-python/blob/ e389823ba013a24b4c32ce38fa0bd87e6bccae94/chatml.md

  133. [145]

    REGULATION (EU) 2024/1689 (Artificial Intelligence Act)

    European Parliament and Council of the European Union. REGULATION (EU) 2024/1689 (Artificial Intelligence Act). Official Journal of the European Union , 2024-06-13. URL http://data.europa.eu/eli/reg/2024/1689/oj

  134. [146]

    LLM-datasets: An open framework for pretraining datasets of large language models

    Malte Ostendorff, Pedro Ortiz Suarez, Lucas Fonseca Lage, and Georg Rehm. LLM-datasets: An open framework for pretraining datasets of large language models. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=5RdIMlGLXL

  135. [148]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:2773...

  136. [149]

    Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language models as knowledge bases? In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in N...

  137. [150]

    Towards a common understanding of contributing factors for cross-lingual transfer in multilingual language models: A review, 2023

    Fred Philippy, Siwen Guo, and Shohreh Haddadan. Towards a common understanding of contributing factors for cross-lingual transfer in multilingual language models: A review, 2023. URL https://arxiv.org/abs/2305.16768. Version Number: 1

  138. [151]

    Bleu: a method for au- tomatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for au- tomatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics , ACL ’02, page 311–318, USA, 2002. As- sociation for Computati...

  139. [152]

    Online dpo: Online direct preference optimization with fast-slow chasing.arXiv preprint arXiv:2406.05534, 2024

    Biqing Qi, Pengfei Li, Fangyuan Li, Junqi Gao, Kaiyan Zhang, and Bowen Zhou. Online dpo: Online direct preference optimization with fast-slow chasing.arXiv preprint arXiv:2406.05534, 2024

  140. [153]

    Nemotron-4 15b technical report, 2024

    Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Mostofa Patwary, Sandeep Subrama- nian, Dan Su, Chen Zhu, Deepak Narayanan, Aastha Jhunjhunwala, Ayush Dattagupta, Vibhu Jawa, Jiwei Liu, Ameya Mahabaleshwarkar, Osvald Nitski, Annika Brundyn, James Maki, Miguel Martinez, J...

  141. [154]

    Language models are unsupervised multitask learners

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019

  142. [155]

    The FineWeb datasets: Decanting the web for the finest text data at scale, 2024

    Guilherme Penedo, Hynek Kydlíˇcek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro V on Werra, and Thomas Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale, 2024. URL http://arxiv.org/abs/2406.17557

  143. [156]

    Scaling language models: Methods, analysis & insights from training gopher

    Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021

  144. [157]

    Direct preference optimization: Your language model is secretly a reward model

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in N...

  145. [158]

    French contextualized word-embeddings with a sip of CaBeRnet: a new french balanced reference corpus

    Murielle Popa-Fabre, Pedro Javier Ortiz Suárez, Benoît Sagot, and Éric de la Clergerie. French contextualized word-embeddings with a sip of CaBeRnet: a new french balanced reference corpus. In Proceedings of the 8th Workshop on Challenges in the Management of Large Corpora, pa...

  146. [159]

    Rush, and Thomas Wolf

    Nazneen Rajani, Lewis Tunstall, Edward Beeching, Nathan Lambert, Alexander M. Rush, and Thomas Wolf. No robots. https://huggingface.co/datasets/HuggingFaceH4/ no_robots, 2023

  147. [160]

    Improving language understanding by generative pre-training, 2018

    Alec Radford. Improving language understanding by generative pre-training, 2018

  148. [161]

    Zero- infinity: Breaking the gpu memory wall for extreme scale deep learning

    Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. Zero- infinity: Breaking the gpu memory wall for extreme scale deep learning. In Proceedings of the international conference for high performance computing, networking, storage and analysis, pages ...

  149. [162]

    Compressive transformers for long-range sequence modelling

    Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. ArXiv, 2019. URL https: //arxiv.org/abs/1911.05507

  150. [163]

    Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20...

  151. [164]

    {Zero-offload}: Democratizing {billion-scale} model training

    Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. {Zero-offload}: Democratizing {billion-scale} model training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21) , pages 551–564, 2021

  152. [165]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023. URL https://arxiv.org/abs/1910.10683

  153. [166]

    Xstest: A test suite for identifying exaggerated safety behaviours in large language models

    Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. arXiv preprint arXiv:2308.01263, 2023

  154. [167]

    Zero: Memory opti- mizations toward training trillion parameter models

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory opti- mizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16. IEEE, 2020

  155. [168]

    Rainbow teaming: Open-ended generation of diverse adversarial prompts

    Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, et al. Rainbow teaming: Open-ended generation of diverse adversarial prompts. arXiv preprint arXiv:2402.16822, 2024

  156. [169]

    Searching for activation functions

    Prajit Ramachandran, Barret Zoph, and Quoc V Le. Searching for activation functions. arXiv preprint arXiv:1710.05941, 2017

  157. [170]

    Neural machine translation of rare words with subword units

    Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909, 2015

  158. [171]

    Andrew Schwartz, and Dirk Hovy

    Deven Santosh Shah, H. Andrew Schwartz, and Dirk Hovy. Predictive biases in natural language processing models: A conceptual framework and overview. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, page 5248–5264, Online, 2020. Associ...

  159. [172]

    Advancing neural encoding of portuguese with transformer albertina PT-*, 2023

    João Rodrigues, Luís Gomes, João Silva, António Branco, Rodrigo Santos, Henrique Lopes Cardoso, and Tomás Osório. Advancing neural encoding of portuguese with transformer albertina PT-*, 2023

  160. [173]

    Glu variants improve transformer, 2020

    Noam Shazeer. Glu variants improve transformer, 2020. URL https://arxiv.org/abs/ 2002.05202

  161. [174]

    The swedish culturomics gigaword CorpusThe swedish culturomics gigaword corpus, 2016

    Stian Rødven-Eide. The swedish culturomics gigaword CorpusThe swedish culturomics gigaword corpus, 2016. URL https://spraakbanken.gu.se/resurser/gigaword

  162. [175]

    The woman worked as a babysitter: On biases in language generation

    Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. The woman worked as a babysitter: On biases in language generation. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language ...

  163. [176]

    Criteria for the annotation of implicit stereotypes

    Wolfgang Schmeisser-Nieto, Montserrat Nofre, and Mariona Taulé. Criteria for the annotation of implicit stereotypes. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 753–762, 2022

  164. [177]

    Megatron-lm: Training multi-billion parameter language models using model parallelism

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019

  165. [179]

    BIGPATENT: A large-scale dataset for abstractive and coherent summarization

    Eva Sharma, Chen Li, and Lu Wang. BIGPATENT: A large-scale dataset for abstractive and coherent summarization. ArXiv, abs/1906.03741, 2019. URL http://arxiv.org/abs/ 1906.03741. 57

  166. [180]

    Manning, Andrew Ng, and Christopher Potts

    Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In David Yarowsky, Timothy Baldwin, Anna Korhonen, Karen Livescu, and Steven Bethard,...

  167. [181]

    Large language model alignment: A survey

    Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. Large language model alignment: A survey. arXiv preprint arXiv:2309.15025, 2023

  168. [182]

    Leon Strømberg-Derczynski, Manuel Ciosici, Rebekah Baglini, Morten H. Christiansen, Jacob Aarup Dalsgaard, Riccardo Fusaroli, Peter Juel Henrichsen, Rasmus Hvingelby, Andreas Kirkedal, Alex Speed Kjeldsen, Claus Ladefoged, Finn Årup Nielsen, Jens Madsen, Malte Lau Petersen, Jo...

  169. [183]

    Roformer: Enhanced transformer with rotary position embedding

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024

  170. [184]

    Language models are multilingual chain-of-thought reasoners

    Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush V osoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on L...

  171. [185]

    BLEU is not suitable for the evaluation of text simplification

    Elior Sulem, Omri Abend, and Ari Rappoport. BLEU is not suitable for the evaluation of text simplification. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii, editors, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, p...

  172. [186]

    Challenging big- bench tasks and whether chain-of-thought can solve them

    Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. Challenging big- bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022

  173. [187]

    Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemi´nski, Hakimeh ...

  174. [188]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024

  175. [189]

    peS2o (pretraining efficiently on s2orc) dataset, 2023

    Luca Soldaini and Kyle Lo. peS2o (pretraining efficiently on s2orc) dataset, 2023

  176. [190]

    Gemma 2: Improving open language models at a practical size, 2024

    Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. Gemma 2: Improving open language models at a practical size, 2024. URL https://arxiv. org/abs/2408.0011...

  177. [191]

    Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024

    Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Ziteng Wang, Rob Fergus, Yann LeCun, and Saining Xie. Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024....

  178. [192]

    Detecting personal information in training corpora: an analysis

    Nishant Subramani, Sasha Luccioni, Jesse Dodge, and Margaret Mitchell. Detecting personal information in training corpora: an analysis. In The Third Workshop on Trustworthy Natural Language Processing, pages 208–220, 2023. doi: 10.18653/v1/2023.trustnlp-1.18. 58

  179. [193]

    Llama 2: Open foundation and fine-tuned chat models

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023

  180. [194]

    Aya model: An in- struction finetuned open-access multilingual language model.arXiv preprint arXiv:2402.07827, 2024

    Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al. Aya model: An in- struction finetuned open-access multilingual language model.arXiv preprint arXiv:2402.07827, 2024

  181. [195]

    Hashimoto

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023

  182. [196]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017

  183. [197]

    Gemma: Open models based on gemini research and technology

    Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024

  184. [198]

    Introducing the CURLICAT corpora: Seven-language domain specific annotated corpora 59 from curated sources

    Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadi´c, Vanja Štefanec, Maciej Ogrodniczuk, Bartlomiej Nito´n, Piotr Pezik, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan, Dan Tufi{\textbackslash}textcommabelows, Radovan Garabík, Simon Krek, and Andraž Repar. Introducin...

  185. [199]

    The brwac corpus: A new open resource for brazilian portuguese

    Jorge A Wagner Filho, Rodrigo Wilkens, Marco Idiart, and Aline Villavicencio. The brwac corpus: A new open resource for brazilian portuguese. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), 2018

  186. [200]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timo- thée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023

  187. [201]

    Element-aware summarization with large language models: Expert-aligned evaluation and chain-of-thought method

    Yiming Wang, Zhuosheng Zhang, and Rui Wang. Element-aware summarization with large language models: Expert-aligned evaluation and chain-of-thought method. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association...

  188. [202]

    Aligning large language models with human: A survey

    Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu. Aligning large language models with human: A survey. arXiv preprint arXiv:2307.12966, 2023

  189. [203]

    DaNewsroom: A large-scale danish summarisation dataset

    Daniel Varab and Natalie Schluter. DaNewsroom: A large-scale danish summarisation dataset. In Proceedings of The 12th Language Resources and Evaluation Conference, pages 6731–

  190. [204]

    Mind the instructions: a holistic evaluation of consistency and interactions in prompt-based learning

    Lucas Weber, Elia Bruni, and Dieuwke Hupkes. Mind the instructions: a holistic evaluation of consistency and interactions in prompt-based learning. In Jing Jiang, David Reitter, and Shumin Deng, editors, Proceedings of the 27th Conference on Computational Natural Language Lear...

  191. [205]

    Emergent abilities of large language models

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022

  192. [206]

    Introducing v0

    Bertie Vidgen, Adarsh Agrawal, Ahmed M Ahmed, Victor Akinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Max Bartolo, et al. Introducing v0. 5 of the ai safety benchmark from mlcommons. arXiv preprint arXiv:2404.12241, 2024

  193. [207]

    Google’s neural machine translation system: Bridging the gap between human and machine translation

    Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016

  194. [208]

    Le, Tengyu Ma, and Adams Wei Yu

    Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy Liang, Quoc V . Le, Tengyu Ma, and Adams Wei Yu. DoReMi: Optimizing data mixtures speeds up language model pretraining, 2023. URL http://arxiv.org/abs/2305.10429

  195. [209]

    Neural machine translation with byte-level subwords

    Changhan Wang, Kyunghyun Cho, and Jiatao Gu. Neural machine translation with byte-level subwords. In Proceedings of the AAAI conference on artificial intelligence, volume 34-05, pages 9154–9160, 2020

  196. [210]

    Fung, Sha Li, Zixuan Huang, Xu Cao, Xingyao Wang, Yiquan Wang, Heng Ji, and Chengxiang Zhai

    Ke Yang, Jiateng Liu, John Wu, Chaoqi Yang, Yi R. Fung, Sha Li, Zixuan Huang, Xu Cao, Xingyao Wang, Yiquan Wang, Heng Ji, and Chengxiang Zhai. If llm is the wizard, then code is the wand: A survey on how code empowers large language models to serve as intelligent agents, 2024....

  197. [211]

    PAWS-X: A cross-lingual adversarial dataset for paraphrase identification

    Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Pr...

  198. [212]

    Lipton, and Yulia Tsvetkov

    Zirui Wang, Zachary C. Lipton, and Yulia Tsvetkov. On negative interference in multilingual models: Findings and a meta-learning treatment, 2020. URL http://arxiv.org/abs/2010. 03017

  199. [213]

    Low-resource languages jailbreak gpt-4

    Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. Low-resource languages jailbreak gpt-4. arXiv preprint arXiv:2310.02446, 2023

  200. [214]

    Yi: Open foundation models by 01

    Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652, 2024

  201. [215]

    Fundamental limitations of alignment in large language models

    Yotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine, and Amnon Shashua. Fundamental limitations of alignment in large language models. arXiv preprint arXiv:2304.11082, 2023

  202. [216]

    Root mean square layer normalization

    Biao Zhang and Rico Sennrich. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019

  203. [217]

    Opt: Open pre-trained transformer language models

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022

  204. [218]

    Qwen2 technical report

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024

  205. [219]

    A survey of large language models

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023

  206. [220]

    P Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023

  207. [221]

    A survey on large language model (llm) security and privacy: The good, the bad, and the ugly

    Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. A survey on large language model (llm) security and privacy: The good, the bad, and the ugly. High-Confidence Computing, page 100211, 2024

  208. [222]

    MoE-LPR: Multilingual extension of large language models through mixture-of-experts with language priors routing, 2024

    Hao Zhou, Zhijun Wang, Shujian Huang, Xin Huang, Xue Han, Junlan Feng, Chao Deng, Weihua Luo, and Jiajun Chen. MoE-LPR: Multilingual extension of large language models through mixture-of-experts with language priors routing, 2024. URL https://arxiv.org/ abs/2408.11396. Version...

  209. [223]

    Fine-tuning language models from human preferences

    Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019

  210. [224]

    Sigmoid loss for language image pre-training, 2023

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training, 2023. URL https://arxiv.org/abs/2303.15343

  211. [227]

    Calibrate before use: Improving few-shot performance of language models

    Tony Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, 2021. URL https://api.semanticscholar.org/CorpusID:231979430

  212. [230]

    Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x

    Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, et al. Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x. arXiv preprint arXiv:2303.17568, 2023

  213. [233]

    Corpus of academic slovene KAS 2.0, 2022

    Aleš Žagar, Matic Kavaš, Marko Robnik-Šikonja, Tomaž Erjavec, Darja Fišer, Nikola Ljubeši´c, Marko Ferme, Mladen Boroviˇc, Borko Boškoviˇc, Milan Ojsteršek, and Goran Hrovat. Corpus of academic slovene KAS 2.0, 2022. URL http://hdl.handle.net/11356/1448. 61 A Author Contributi...

  214. [234]

    Web-sourced datasets with some preprocessing available under permissive license

  215. [235]

    Domain-specific or language-specific raw crawls

  216. [236]

    CATalog)

    Manually curated data obtained through collaborators, data providers (by means of legal assignment agreements) or open source projects (e.g. CATalog). What mechanisms or procedures were used to collect the data (e.g., hardware apparatuses or sensors, manual human curation, sof...

  217. [237]

    We validate the data with a data integrity check, which ensures that the downloaded files are complete, uncorrupted and in the expected format and structure

    Open Direct Download: Data were obtained directly from publicly accessible sources, such as websites or repositories that provide open data downloads. We validate the data with a data integrity check, which ensures that the downloaded files are complete, uncorrupted and in the...

  218. [238]

    These scripts navigate web pages, extract relevant data and store it in a structured format

    Ad hoc scrapers or crawlers: Custom web scraping scripts or crawlers were used to extract data from various online sources where direct downloads were not available. These scripts navigate web pages, extract relevant data and store it in a structured format. We validate this m...

  219. [239]

    raw" data saved in addition to the preprocessed/cleaned/labeled data (e.g., to support unanticipated future uses)? If so, please provide a link or other access point to the

    Direct download via FTP, SFTP, API or S3: Some datasets were acquired using secure transfer protocols such as FTP (File Transfer Protocol), SFTP (Secure File Transfer Protocol), or API (Application Programming Interface) requests from cloud storage services such as Amazon S3. ...

  220. [240]

    Un taxi amarillo: está ubicado en la esquina superior izquierda de la imagen

  221. [241]

    Un autobús verde: está posicionado en el centro de la imagen

  222. [242]

    Un autobús azul: se encuentra en el lado derecho de la imagen

  223. [243]

    Un autobús blanco: está ubicado en el lado inferior derecho de la imagen

  224. [244]

    Un camión rojo: está situado en el lado inferior izquierdo de la imagen

  225. [245]

    Una bicicleta: está ubicada en la esquina superior izquierda de la imagen

  226. [246]

    Un scooter: está posicionado en la esquina superior derecha de la imagen

  227. [247]

    Barcelona Supercomputing Centre

    Un automóvil blanco: se encuentra en el lado derecho de la imagen. Grounding Figure 21: Grounding example in Spanish. Salamandra provides a numbered list to describe the colors and relative positions of each vehicle shown in the image. It misses the yellow car and misidentifie...

  228. [2019]

    doi: 10.18653/v1/D19-1339

    Association for Computational Linguistics. doi: 10.18653/v1/D19-1339. URL https: //aclanthology.org/D19-1339

  229. [2021]

    URL http://hdl.handle.net/11234/1-4635

  230. [2022]

    doi: 10.18653/v1/2022.bionlp-1.19

    Association for Computational Linguistics. doi: 10.18653/v1/2022.bionlp-1.19. URL https://aclanthology.org/2022.bionlp-1.19

  231. [2023]

    URL https://github.com/togethercomputer/RedPajama-Data

  232. [2024]

    Version Number: 2

    URL https://arxiv.org/abs/2410.03730. Version Number: 2

  233. [6739]

    ISBN 979-10-95546-34-4

    European Language Resources Association, 2020. ISBN 979-10-95546-34-4. URL https://www.aclweb.org/anthology/2020.lrec-1.831

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.