Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper establishes that large language models can be trained to competitive performance using only public domain and openly licensed text, and releases an 8TB dataset plus two 7B models as evidence.

desk verdict A large, reproducible open-license corpus with a clean code win, but the knowledge/reasoning headline is partly an artifact of tuning the mixture on the same benchmarks as the comparison. read the letter →

arxiv 2506.05209 v1 pith:L7QFMMII submitted 2025-06-05 cs.CL cs.LG

classification cs.CLcs.LG
keywords openlylicensedtextlanguagemodelpretrainingpublicdomaindatasetcurationdatalicensingcopyrightlicenselaundering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether a capable large language model can be built using only text that is in the public domain or under a genuinely open license. To answer it, the authors assembled the Common Pile v0.1, an 8TB corpus from 30 sources—research papers, code, government and legal records, books, wikis, educational materials, transcripts, and web pages—each vetted for license status. They then trained two 7-billion-parameter models, Comma v0.1-1T and Comma v0.1-2T, on 1 and 2 trillion tokens drawn from a filtered and reweighted version of the corpus. Both models perform competitively with budget-matched models trained on unlicensed text, such as the Llama 1 and 2 7B models, and a controlled small-scale comparison shows the corpus outperforms earlier openly licensed collections. The authors release the dataset, training mixture, code, and checkpoints, while flagging that license laundering could still have let some unlicensed text through.

What carries the argument

The load-bearing object is the Common Pile corpus itself, together with the Comma training mixture derived from it. The Common Pile is an 8TB collection assembled from 30 sources under strict sourcing rules: only content whose license meets the Open Definition 2.1 standard—free access, use, modification, and sharing for any purpose—and whose providers are believed to be the actual rights holders. The argument is carried by three mechanisms: license due diligence at the source level, including manual verification of web domains and exclusion of unreliable sources; a filtering and deduplication pipeline that turns raw text into training tokens; and per-source quality evaluation that sets mixing weights to up-weight high-quality sources while capping repetition. The final validation mechanism is a controlled comparison: identical small models trained on different corpora, followed by two full-scale 7B training runs.

What would settle it

Take a random sample of documents in the Common Pile, trace each to the stated rights holder, and check whether the open-license claim matches the holder's terms; if a substantial fraction are in-copyright or mislabeled, the claim that the corpus is openly licensed fails.

Watch

Extended reading notes

Core claim

The authors set out to determine whether LLM performance depends on unlicensed web text. They construct the Common Pile v0.1, an 8TB corpus from 30 sources, all selected under the Open Definition 2.1 standard for open licenses or public domain status. They then filter, deduplicate, and reweight this corpus into the Comma training mixture and train two 7-billion-parameter models on 1 and 2 trillion tokens. On standard knowledge, reasoning, and coding benchmarks, Comma v0.1-1T and Comma v0.1-2T are competitive with budget-matched models trained on unlicensed data such as Llama 1 and 2 7B, and a controlled 1.7B-parameter ablation shows the Common Pile outperforms prior open-license corpora. The authors also release the dataset, code, mixture, and checkpoints, while acknowledging that license laundering and changed license terms may have allowed some unlicensed text into the corpus.

Load-bearing premise

The central premise is that the license labels on the 30 sources faithfully reflect the rights of the people who posted them, so the corpus really is openly licensed.

Editorial extensions

If this is right

  • A 7B model trained entirely on openly licensed text can match or beat budget-matched models trained on unlicensed web text on several standard benchmarks, including knowledge and code tasks.
  • The 8TB Common Pile is, to the authors' knowledge, the largest openly licensed pretraining corpus, making open-license pretraining a practical option rather than a toy-scale exercise.
  • Releasing the dataset, mixture, code, and checkpoints allows others to reproduce and extend the results without scraping unlicensed data.
  • The controlled 1.7B-parameter ablation shows the Common Pile outperforms prior openly licensed corpora on the same budget, so corpus quality, not just license status, drives the result.
  • The 2T run repeats the 1T mixture and still stays competitive, which the paper attributes to the mixture's quality and a possible ceiling from heavy repetition.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the performance result and the licensing result are separable; even if some unlicensed text slipped in, the main contribution could still hold as a scale and curation demonstration, but the legal and ethical selling point would weaken.
  • Beyond the paper: the same sourcing-and-validation recipe could be transferred to non-English corpora, code-specific datasets, or multimodal data, where open licensing questions are equally pressing.
  • Beyond the paper: a direct test of the licensing premise would be to retrain on a provenance-verified subset and compare; the paper's own ablation removing one curated source suggests performance is fairly robust to source changes.
  • Beyond the paper: the reported growth of openly licensed text over time implies the constraint may become less binding as more contemporary content is released under open licenses.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces the Common Pile v0.1, an 8 TB collection of text that the authors argue is composed exclusively of public domain and openly licensed material drawn from 30 diverse sources. The authors apply filtering and deduplication to obtain the Comma dataset, tune source mixing weights using per-source 1.7B-parameter models evaluated on a set of early-signal benchmarks, and then train two 7B-parameter models, Comma v0.1-1T and Comma v0.1-2T, on 1T and 2T tokens respectively. The models are evaluated on knowledge, reasoning, and code benchmarks and compared against budget-matched models trained on unlicensed text, with claims that Comma is competitive with Llama 1/2, OLMo, and DeepSeekLLM. The paper also reports controlled 1.7B-scale ablations against earlier openly licensed corpora and several unlicensed baselines. The manuscript releases the dataset, preprocessing code, training mixture, and model checkpoints.

Significance. If the central claims hold, this is a substantial contribution: it provides the largest openly licensed pretraining corpus to date, along with trained models and a reproducible pipeline, and it presents evidence that open-license-only pretraining can approach the quality of models trained on unlicensed web text. The paper is commendably transparent about licensing caveats, includes controlled small-scale ablations, and ships machine-checkable artifacts (dataset, code, checkpoints). The main evidence for the headline claim, however, is weakened by the fact that the data mixture was tuned on the same evaluation benchmarks used to compare the final models against baselines, and by the use of a custom tokenizer that makes the '1 trillion tokens' budget not directly comparable with baseline tokenizers. These issues are addressable with additional experiments or re-analysis, which is why I recommend major revision rather than rejection.

major comments (3)
  1. [Section 4.2 and Section 4.4 (Tables 10-11)] The source mixing weights for the Comma dataset were selected by training per-source 1.7B models and evaluating them on the 'early signal' tasks from Penedo et al. (ARC, MMLU, HellaSwag, OBQA, CSQA, PIQA, SIQA). The final 7B models are then evaluated on almost exactly that benchmark suite, with only BoolQ and the code benchmarks added. Consequently, the comparisons against Llama 1/2, MPT, OLMo, and DeepSeekLLM are not clean: the Comma mixture was explicitly optimized for these benchmarks, while the baseline mixtures were not. This undermines the inference from 'Comma scores comparably' to 'openly licensed text is sufficient for performant LLMs.' The code results (HumanEval, MBPP) are less affected because those benchmarks were not used during mixture selection, but the knowledge/reasoning scores, which dominate the reported averages, are directly confounded. I recommend either evaluating on a held-out suite that was not used in any stage of mixture selection, or demonstrating through an ablation that mixture selection on this benchmark set does not materially affect the final comparisons.
  2. [Section 4.3 (controlled 1.7B experiments)] The controlled dataset-quality experiment in Section 4.3 shares the same circularity concern. The 'Comma dataset' used in this comparison is a filtered and reweighted mixture whose weights were selected using the same early-signal benchmarks on which the models are then evaluated. When comparing against OLC, Common Corpus, KL3M, the Pile, OSCAR, and FineWeb, the Comma model benefits not only from the underlying text but also from benchmark-specific tuning; the baseline datasets were not given an equivalent mixture-tuning step. The paper's statement that differences in performance 'stem primarily from the quality of each dataset' is therefore too strong. A fairer comparison would either apply the same per-source tuning procedure to each candidate corpus (using a held-out validation set) or compare unweighted versions of each corpus.
  3. [Section 4.4 (tokenization and budget matching)] The claim that Comma v0.1-1T and v0.1-2T are 'budget-matched' to Llama 1/2 and other baselines implicitly assumes that 1T (or 2T) tokens of Comma data correspond to the same training budget as 1T (or 2T) tokens of the baselines. However, Comma uses a custom BPE tokenizer with a vocabulary size of 64,000, whereas Llama 1/2, MPT, OLMo, and DeepSeekLLM use different tokenizers, typically with smaller vocabularies. If the Comma tokenizer is more efficient (fewer tokens per byte of text), then training on 1T Comma tokens may actually process more text than training on 1T tokens of a baseline, making the comparison generous to Comma. The paper does not report tokenizer efficiency (e.g., bytes per token, or the number of tokens per document for a fixed text sample), nor does it compare training FLOPs. This should be addressed by reporting a tokenizer-comparability metric or by comparing models trained with the same tokenizer on the same number of tokens.
minor comments (5)
  1. [Appendix B.3 (USPTO)] In the description of the USPTO source, the text says 'We include parents from the US Patents and Trademark Office' but this should read 'We include patents from the US Patents and Trademark Office.'
  2. [Appendix B.8 (PEPs)] The sentence 'There are been 661 PEPs published' contains a grammatical error and should be 'There have been 661 PEPs published.'
  3. [Table 9 caption] The caption states that removing DPI data has 'marginal affect' on dataset quality; this should be 'marginal effect.'
  4. [Appendix O] The text says 'the the training batch size is 8.3M ( 223) versus the 2.1M ( 221) tokens per step' with a doubled 'the' and missing superscript formatting for the exponents; these should be corrected for readability.
  5. [Tables 10-11] Abbreviations such as 'HS', 'HEval', and 'MBPP' are used without definition in the table captions or nearby text; defining them (HellaSwag, HumanEval, MBPP) would improve accessibility.

Circularity Check

1 steps flagged · score 5.0 of 10

Benchmark-tuned data mixture is evaluated on the same benchmark suite, making the headline competitive-performance claim partly constructed.

  1. fitted input called prediction [Section 4.2 (data mixing), Section 4.3 (controlled 1.7B ablations), Section 4.4 (Comma v0.1 evaluation)]
    "To determine mixing weights, we first trained per-source language models ... Based on the performance of these per-source models, we heuristically set mixing weights to up- and down-weight high- and low-performance sources. Each model was then evaluated using the set of 'early signal' tasks identified by Penedo et al.: ARC, MMLU, HellaSwag, OpenBookQA, CommonSenseQA, PIQA, and SocialIQA. We evaluate the Comma v0.1 models on ARC, MMLU, BoolQ, HellaSwag, OpenBookQA, CommonsenseQA, PIQA, and SIQA to probe world knowledge and reasoning."

    The mixing weights are fitted to per-source 1.7B models' scores on the 'early signal' tasks (ARC, MMLU, HellaSwag, OBQA, CSQA, PIQA, SIQA), and the 7B Comma models are then evaluated on almost exactly the same tasks plus BoolQ and code. The paper's central evidence that openly licensed data 'can be used as the foundation for competitive LLMs' therefore uses the same benchmark signal that determined the up/down-weighting of sources. The knowledge/reasoning columns in Tables 10-11 are not an independent test of the dataset: the mixture was explicitly chosen to make those tasks score well, so competitive performance on them is partly an artifact of selection rather than purely of the open-license constraint.

full rationale

The paper's derivation chain is otherwise self-contained: the dataset is built from sources with documented licensing choices, preprocessing and deduplication are standard, and the 7B models are compared against fixed external baselines that were not tuned on this benchmark suite. The main circularity is the shared benchmark signal between mixture selection and evaluation. This does not invalidate the dataset or the code results, but it materially weakens the inference from 'Comma scores well on these benchmarks' to 'openly licensed text alone supports competitive performance.' A held-out evaluation would settle whether the effect is real or an artifact of tuning. No other load-bearing self-citation or definitional equivalence was found.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The central claims depend on two classes of assumptions: the accuracy of license labels on 30 sources, and the absence of benchmark contamination. The mixture weights and filtering thresholds are free parameters tuned on the evaluation suite. No new theoretical entities are introduced.

free parameters (4)
  • Source mixing weights (Table 7) = Various repeats, e.g., peS2o 6x, USPTO 0.25x
    Chosen based on per-source 1.7B model performance on the same benchmark suite used for evaluation, i.e., fitted to the eval set.
  • Cool-down mixture weights (Table 8) = 37.5B token budget from 11 sources
    Hand-selected 'high-quality sources'; no independent criterion given.
  • Per-source filtering thresholds (Table 5) = e.g., language > 0.5, min length 100-700, toxicity > 0.1
    Manual thresholds per source; not derived from a principled optimization.
  • Weight decay = 0.2
    Selected because it gave 'slightly improved performance' on the benchmark suite (Section 4.3).
assumptions (3)
  • domain assumption The evaluated benchmarks are not present in the DPI-sourced training data
    Only Winogrande is explicitly excluded; other tasks (ARC, MMLU, HellaSwag, etc.) are assumed to be unseen without verification.
  • domain assumption Per-source model performance on the early-signal tasks predicts contribution to the mixture
    This assumption drives the heuristic mixing in Section 4.2.
  • domain assumption License metadata from each source is accurate and satisfies the Open Definition 2.1
    The validity of the 'openly licensed only' premise rests on this; caveat admitted in Section 2.1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text." pith.science (2026). https://pith.science/paper/L7QFMMII

@misc{pith2026250605209,
  author       = {Pith},
  title        = {Pith review of: The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/L7QFMMII}},
  note         = {Machine review of arXiv:2506.05209}
}
read the original abstract

Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement and ethical concerns. Training LLMs on openly licensed text presents a first step towards addressing these issues, but prior data collection efforts have yielded datasets too small or low-quality to produce performant LLMs. To address this gap, we collect, curate, and release the Common Pile v0.1, an eight terabyte collection of openly licensed text designed for LLM pretraining. The Common Pile comprises content from 30 sources that span diverse domains including research papers, code, books, encyclopedias, educational materials, audio transcripts, and more. Crucially, we validate our efforts by training two 7 billion parameter LLMs on text from the Common Pile: Comma v0.1-1T and Comma v0.1-2T, trained on 1 and 2 trillion tokens respectively. Both models attain competitive performance to LLMs trained on unlicensed text with similar computational budgets, such as Llama 1 and 2 7B. In addition to releasing the Common Pile v0.1 itself, we also release the code used in its creation as well as the training mixture and checkpoints for the Comma v0.1 models.

Figures

Figures reproduced from arXiv: 2506.05209 by the authors.

Figure 1
Figure 1. The Common Pile is an 8TB dataset of openly licensed text curated from 30 diverse sources. The sources comprising the Common Pile are shown above, categorized by textual domain. LLMs, often without compensation to the creators of this content. Recent estimates suggest that compensating the authors of pre-training data, even at conservatively low wage rates, would cost billions of US dollars [82]. While copyright exe… view at source ↗
Figure 2
Figure 2. The Common Pile consistently outperforms other openly licensed corpora as a pre￾training dataset. Following the setup from Penedo et al. [132], we train and evaluate 1.7B parameter models on 28B tokens of data from each dataset. Stars denote benchmarks on which the model trained using the Common Pile outperforms all other models. a lower threshold to filter out low-quality code than was used for SmolLM2, resulting i… view at source ↗
Figure 3
Figure 3. Compared to models trained with similar resources (7 billion parameters, 1 trillion tokens), Comma v0.1-1T is the strongest model on several standard benchmarks. To contextual￾ize these results, we include Qwen3 8B (trained on 36 trillion tokens) as a “current best-practices” upper bound. Stars denote benchmarks on which Comma v0.1-1T outperforms all other compute￾matched models (i.e., all models other than Qwen3). … view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Comma v0.1-2T is also competitive with budget-matched models (7 billion parameters, 2 trillion tokens) trained on unlicensed data. We additionally include Qwen3 8B as a higher budget upper bound. Stars denote benchmarks where Comma v0.1-2T outperforms budget-matched mo…
Figure 5
Figure 5. Figure 5: Author contributions to this work. Large squares indicate a major contribution and small squares indicate a supporting contribution. B Detailed Description of Sources Below, we give a more in-depth overview of the sources that make up the Common Pile, including specifi…
Figure 6
Figure 6. Figure 6: The amount of openly licensed text grows steadily over time. We visualize the cumulative proportion of data created up to various cutoff dates for sources in the Common Pile with reliable creation date metadata. This includes all sources except for the Caselaw Access P…
Figure 7
Figure 7. Figure 7: A model trained on the Comma dataset consistently outperforms models trained on other corpora of openly licensed text and outperforms the Pile on all but two tasks. We train identical 1.7B parameter models on 28B tokens from each dataset following Penedo et al. [132]. …

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Applying the Chinese Wall Reverse Engineering Technique to Large Language Model Code Editing

    cs.SE 2025-07 conditional novelty 3.0 of 10

    Using Gemini 2.5 Pro to annotate code with edit instructions improved Comma v0.1 1T's CanItEdit pass@20 from 20.00 to 33.33 and Starcoder2 Instruct's pass@1 from 35.10 to 42.05.

Reference graph

Works this paper leans on

206 extracted references · 10 canonical work pages · cited by 1 Pith paper

  1. [1]

    Code § 102

    17 U.S. Code § 102. Subject matter of copyright: In general, December 1990. URL https: //www.law.cornell.edu/uscode/text/17/102

  2. [2]

    Code § 105

    17 U.S. Code § 105. Subject matter of copyright: United States Government works, December

  3. [3]

    Efficient online data mixing for language model pre-training

    Alon Albalak, Liangming Pan, Colin Raffel, and William Yang Wang. Efficient online data mixing for language model pre-training. arXiv preprint arXiv:2312.02406, 2023

  4. [4]

    A survey on data selection for language models

    Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. A survey on data selection for language models. Transactions on Machine Learning Research, 2024

  5. [5]

    Smollm2: When smol goes big – data-centric training of a small language model, 2025

    Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíˇcek, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von Wer...

  6. [6]

    On the cross-lingual transferability of monolingual representations

    Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. On the cross-lingual transferability of monolingual representations. In Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, 2019

  7. [7]

    To code, or not to code? exploring impact of code in pre-training

    Viraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot, Ivan Zhang, Acyr Locatelli, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker. To code, or not to code? exploring impact of code in pre-training. arXiv preprint arXiv:2408.10914, 2024

  8. [8]

    Program synthesis with large language models

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021

Show all 206 references
  1. [9]

    Towards best practices for open datasets for llm training, 2025

    Stefan Baack, Stella Biderman, Kasia Odrozek, Aviya Skowron, Ayah Bdeir, Jillian Bom- marito, Jennifer Ding, Maximilian Gahntz, Paul Keller, Pierre-Carl Langlais, Greg Lindahl, Sebastian Majstorovic, Nik Marda, Guilherme Penedo, Maarten Van Segbroeck, Jennifer Wang, Leandro vo...

  2. [10]

    Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction

    Adrien Barbaresi. Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction. In Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natu...

  3. [11]

    Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp

    Max Bartolo, A. Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp. Beat the ai: Investigating adversarial human annotation for reading comprehension. Transactions of the Association for Computational Linguistics, 8:662–678, 2020

  4. [12]

    Stable lm 2 1.6 b technical report

    Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, et al. Stable lm 2 1.6 b technical report. arXiv preprint arXiv:2402.17834, 2024

  5. [13]

    Chou, Roy Frostig, and Percy Liang

    Jonathan Berant, A. Chou, Roy Frostig, and Percy Liang. Semantic parsing on freebase from question-answer pairs. In Conference on Empirical Methods in Natural Language Processing, pages 1533–1544, 2013

  6. [14]

    Elastic ChatNoir: Search Engine for the ClueWeb and the Common Crawl

    Janek Bevendorff, Benno Stein, Matthias Hagen, and Martin Potthast. Elastic ChatNoir: Search Engine for the ClueWeb and the Common Crawl. In Leif Azzopardi, Allan Hanbury, Gabriella Pasi, and Benjamin Piwowarski, editors, Advances in Information Retrieval. 40th European Confer...

  7. [15]

    FastWARC: Optimizing Large-Scale Web Archive Analytics

    Janek Bevendorff, Martin Potthast, and Benno Stein. FastWARC: Optimizing Large-Scale Web Archive Analytics. In Andreas Wagner, Christian Guetl, Michael Granitzer, and Stefan V oigt, editors,3rd International Symposium on Open Search Technology (OSSYM 2021) . International Open...

  8. [16]

    Deepseek llm: Scaling open-source language models with longtermism

    Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024

  9. [17]

    Emergent and predictable memorization in large language models

    Stella Biderman, Usvsn Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raff. Emergent and predictable memorization in large language models. Advances in Neural Information Processing Systems, 36:28072–28090, 2023

  10. [18]

    Pythia: A suite for analyzing large language models across training and scaling, 2023

    Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. Pythia: A suite for analyzing large language mode...

  11. [19]

    Piqa: Reasoning about physical commonsense in natural language, 2019

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. Piqa: Reasoning about physical commonsense in natural language, 2019. URL https://arxiv.org/abs/ 1911.11641

  12. [20]

    License List (version 15), 2025

    Blue Oak Council. License List (version 15), 2025. URL https://blueoakcouncil.org/ list

  13. [21]

    e-snli: Nat- ural language inference with natural language explanations

    Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. e-snli: Nat- ural language inference with natural language explanations. In Neural Information Processing Systems, pages 9560–9572, 2018

  14. [22]

    Quantifying memorization across neural language models

    Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models. In The Eleventh International Conference on Learning Representations, 2022

  15. [23]

    Hartung, M

    Ilias Chalkidis, Abhik Jana, D. Hartung, M. Bommarito, Ion Androutsopoulos, D. Katz, and Nikolaos Aletras. Lexglue: A benchmark dataset for legal language understanding in english. In Annual Meeting of the Association for Computational Linguistics, pages 4310–4330, 2021

  16. [24]

    URL https://chatgptiseatingtheworld.com

    Chat GPT Is Eating the World, 2024. URL https://chatgptiseatingtheworld.com. 12

  17. [25]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott G...

  18. [26]

    HybridQA: A dataset of multi-hop question answering over tabular and textual data

    Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. HybridQA: A dataset of multi-hop question answering over tabular and textual data. In Trevor Cohn, Yulan He, and Yang Liu, editors,Findings of the Association for Computational Linguistics: EMN...

  19. [27]

    Logic2text: High-fidelity natural language generation from logical forms

    Zhiyu Chen, Wenhu Chen, Hanwen Zha, Xiyou Zhou, Yunkai Zhang, Sairam Sundaresan, and William Yang Wang. Logic2text: High-fidelity natural language generation from logical forms. ArXiv, abs/2004.14579, 2020

  20. [28]

    What is your data worth to gpt? llm-scale data valuation with influence functions, 2024

    Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, and Eric Xing. What is your data worth to gpt? llm-scale data valuation with influen...

  21. [29]

    Quac: Question answering in context

    Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. Quac: Question answering in context. In Conference on Empirical Methods in Natural Language Processing, pages 2174–2184, 2018

  22. [30]

    Toxic comment classification challenge

    cjadams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, nithum, and Will Cukierski. Toxic comment classification challenge. https://kaggle.com/competitions/ jigsaw-toxic-comment-classification-challenge , 2017. Kaggle

  23. [31]

    Boolq: Exploring the surprising difficulty of natural yes/no questions

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019

  24. [32]

    What does BERT look at? an analysis of BERT’s attention

    Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. What does BERT look at? an analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, 2019

  25. [33]

    Think you have solved question answering? try arc, the ai2 reasoning challenge

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018

  26. [34]

    Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman

    Karl Cobbe, V . Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. ArXiv, abs/2110.14168, 2021

  27. [35]

    CC0 1.0 Universal (CC0 1.0) Public Domain Dedication, 2025

    Creative Commons. CC0 1.0 Universal (CC0 1.0) Public Domain Dedication, 2025. URL https://creativecommons.org/publicdomain/zero/1.0/

  28. [36]

    Creative Commons Attribution 4.0 International License § 2(a)(1)(A),

    Creative Commons. Creative Commons Attribution 4.0 International License § 2(a)(1)(A),

  29. [37]

    Creative Commons Attribution 4.0 International License § 3(a)(1)(A)(i),

    Creative Commons. Creative Commons Attribution 4.0 International License § 3(a)(1)(A)(i),

  30. [38]

    Public Domain Mark 1.0, 2025

    Creative Commons. Public Domain Mark 1.0, 2025. URL https://creativecommons. org/publicdomain/mark/1. 13

  31. [39]

    Creative Commons Attribution-ShareAlike 4.0 International License,

    Creative Commons. Creative Commons Attribution-ShareAlike 4.0 International License,

  32. [40]

    URL https://creativecommons.org/licenses/by/4.0/legalcode

  33. [41]

    Feder Cooper, Aaron Gokaslan, Amy B

    A. Feder Cooper, Aaron Gokaslan, Amy B. Cyphert, Christopher De Sa, Mark A. Lemley, Daniel E. Ho, and Percy Liang. Extracting memorized pieces of (copyrighted) books from open-weight language models. arXiv preprint arXiv:2505.12546, 2025

  34. [42]

    A span-extraction dataset for chinese machine reading comprehension

    Yiming Cui, Ting Liu, Li Xiao, Zhipeng Chen, Wentao Ma, Wanxiang Che, Shijin Wang, and Guoping Hu. A span-extraction dataset for chinese machine reading comprehension. In EMNLP-IJCNLP, pages 5882–5888, 2019

  35. [43]

    URL https://creativecommons.org/licenses/by-sa/4.0/

  36. [44]

    Feder Cooper and James Grimmelmann

    A. Feder Cooper and James Grimmelmann. The Files are in the Computer: Copyright, Memorization, and Generative AI. arXiv preprint arXiv:2404.12590, 2024

  37. [45]

    Few-nerd: A few-shot named entity recognition dataset

    Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, and Zhiyuan Liu. Few-nerd: A few-shot named entity recognition dataset. ArXiv, abs/2105.07464, 2021

  38. [46]

    Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs

    Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In North American Chapter of the Association for Computational Linguistics , pages 2368–2378, 2019

  39. [47]

    Tracking state changes in procedural text: a challenge dataset and models for process paragraph comprehen- sion

    Bhavana Dalvi, Lifu Huang, Niket Tandon, Wen tau Yih, and Peter Clark. Tracking state changes in procedural text: a challenge dataset and models for process paragraph comprehen- sion. In North American Chapter of the Association for Computational Linguistics , pages 1595–1604, 2018

  40. [48]

    Liu, Ana Marasovi´c, Noah A

    Pradeep Dasigi, Nelson F. Liu, Ana Marasovi´c, Noah A. Smith, and Matt Gardner. Quoref: A reading comprehension dataset with questions requiring coreferential reasoning. In Conference on Empirical Methods in Natural Language Processing, volume abs/1908.05803, 2019

  41. [49]

    Where’s my head? definition, data set, and models for numeric fused-head identification and resolution

    Yanai Elazar and Yoav Goldberg. Where’s my head? definition, data set, and models for numeric fused-head identification and resolution. Transactions of the Association for Compu- tational Linguistics, 7:519–535, 2019

  42. [50]

    Measuring causal effects of data statistics on language model’sfactual’predictions

    Yanai Elazar, Nora Kassner, Shauli Ravfogel, Amir Feder, Abhilasha Ravichander, Marius Mosbach, Yonatan Belinkov, Hinrich Schütze, and Yoav Goldberg. Measuring causal effects of data statistics on language model’sfactual’predictions. arXiv preprint arXiv:2207.14251, 2022

  43. [51]

    Do membership inference attacks work on large language models? arXiv preprint arXiv:2402.07841, 2024

    Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettle- moyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. Do membership inference attacks work on large language models? arXiv preprint arXiv:2402.07841, 2024

  44. [52]

    Dumitrescu, Petru Rebeja, Beáta L ˝orincz, Mihaela G ˘aman, M

    S. Dumitrescu, Petru Rebeja, Beáta L ˝orincz, Mihaela G ˘aman, M. Ilie, Andrei Pruteanu, Adriana Stan, Luciana Morogan, Traian Rebedea, and Sebastian Ruder. Liro: Benchmark and leaderboard for romanian language tasks. In NeurIPS Datasets and Benchmarks, 2021

  45. [53]

    Iirc: A dataset of incomplete information reading comprehension questions

    James Ferguson, Matt Gardner, Tushar Khot, and Pradeep Dasigi. Iirc: A dataset of incomplete information reading comprehension questions. In Conference on Empirical Methods in Natural Language Processing, pages 1137–1147, 2020

  46. [54]

    Nancy Fulda, Nathan Tibbetts, Zachary Brown, and D. Wingate. Harvesting common-sense navigational knowledge for robotics from uncurated text corpora. In Conference on Robot Learning, pages 525–534, 2017. 14

  47. [55]

    Hwang, Maxwell Forbes, and Yejin Choi

    Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. Moral stories: Situated reasoning about norms, intents, actions, and their consequences. ArXiv, abs/2012.15738, 2020

  48. [56]

    Datasets, documents, and repetitions: The practicalities of unequal data quality

    Alex Fang, Hadi Pouransari, Matt Jordan, Alexander Toshev, Vaishaal Shankar, Ludwig Schmidt, and Tom Gunter. Datasets, documents, and repetitions: The practicalities of unequal data quality. arXiv preprint arXiv:2503.07879, 2025

  49. [57]

    The pile: An 800gb dataset of diverse text for language modeling, 2020

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling, 2020. URL https: //arxiv.org/abs/2101.00027

  50. [58]

    Openllama: An open reproduction of llama, May 2023

    Xinyang Geng and Hao Liu. Openllama: An open reproduction of llama, May 2023. URL https://github.com/openlm-research/open_llama

  51. [59]

    A new algorithm for data compression

    Philip Gage. A new algorithm for data compression. The C Users Journal archive, 12:23–38,

  52. [60]

    Breaking nli systems with sentences that require simple lexical inferences

    Max Glockner, Vered Shwartz, and Yoav Goldberg. Breaking nli systems with sentences that require simple lexical inferences. ArXiv, abs/1805.02266, 2018

  53. [61]

    N. Gale, G. Heath, E. Cameron, S. Rashid, and S. Redwood. Using the framework method for the analysis of qualitative data in multi-disciplinary health research. BMC Medical Research Methodology, 13:117 – 117, 2013

  54. [62]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravanku- mar, Artem Korenev, A...

  55. [63]

    Grobid. Grobid. https://github.com/kermitt2/grobid, 2008–2025

  56. [64]

    Roth, and Jonathan Berant

    Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, D. Roth, and Jonathan Berant. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361, 2021

  57. [65]

    Olmes: A standard for language model evaluations, 2025

    Yuling Gu, Oyvind Tafjord, Bailey Kuehl, Dany Haddad, Jesse Dodge, and Hannaneh Hajishirzi. Olmes: A standard for language model evaluations, 2025. URL https: //arxiv.org/abs/2406.08446

  58. [66]

    Feder Cooper, Jasmine Collins, Landan Seguin, Austin Jacobson, Mihir Patel, Jonathan Frankle, Cory Stephenson, and V olodymyr Kuleshov

    Aaron Gokaslan, A. Feder Cooper, Jasmine Collins, Landan Seguin, Austin Jacobson, Mihir Patel, Jonathan Frankle, Cory Stephenson, and V olodymyr Kuleshov. CommonCanvas: Open Diffusion Models Trained on Creative-Commons Images. In Proceedings of the IEEE/CVF Conference on Compu...

  59. [67]

    Disfl- qa: A benchmark dataset for understanding disfluencies in question answering

    Aditya Gupta, Jiacheng Xu, Shyam Upadhyay, Diyi Yang, and Manaal Faruqui. Disfl- qa: A benchmark dataset for understanding disfluencies in question answering. ArXiv, abs/2106.04016, 2021

  60. [68]

    C4Corpus: Multilingual web-size corpus with free license

    Ivan Habernal, Omnia Zayed, and Iryna Gurevych. C4Corpus: Multilingual web-size corpus with free license. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, and Stelios...

  61. [69]

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkin- son, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Ja...

  62. [70]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020

  63. [71]

    Wiki-40b: Multilingual language model dataset

    Mandy Guo, Zihang Dai, Denny Vrande ˇci´c, and Rami Al-Rfou. Wiki-40b: Multilingual language model dataset. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 2440–2452, 2020

  64. [72]

    Universal language model fine-tuning for text classi- fication

    Jeremy Howard and Sebastian Ruder. Universal language model fine-tuning for text classi- fication. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 2018

  65. [73]

    Minicpm: Unveiling the potential of small language models with scalable training strategies

    Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, et al. Minicpm: Unveiling the potential of small language models with scalable training strategies. arXiv preprint arXiv:2404.06395, 2024

  66. [74]

    URL https://huggingface.co/datasets/ PleIAs/common_corpus

    HuggingFace: Common Corpus, 2025. URL https://huggingface.co/datasets/ PleIAs/common_corpus

  67. [75]

    AI Training and Copyright Infringement: Solutions from Asia, Octo- ber 2024

    Seth Hays. AI Training and Copyright Infringement: Solutions from Asia, Octo- ber 2024. URL https://www.techpolicy.press/ai-training-and-copyright- infringement-solutions-from-asia/

  68. [76]

    Simple and scalable strategies to continually pre-train large language models

    Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish. Simple and scalable strategies to continually pre-train large language models. arXiv preprint arXiv:2403.08763, 2024

  69. [77]

    Cuad: An expert-annotated nlp dataset for legal contract review

    Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. Cuad: An expert-annotated nlp dataset for legal contract review. ArXiv, abs/2103.06268, 2021

  70. [78]

    The kl3m data project: Copyright-clean training resources for large language models, 2025

    Michael J Bommarito II, Jillian Bommarito, and Daniel Martin Katz. The kl3m data project: Copyright-clean training resources for large language models, 2025. URL https://arxiv. org/abs/2504.07854

  71. [79]

    Model AI Governance Framework for Generative AI: Fostering a Trusted Ecosys- tem, May 2024

    Infocomm Media Development Authority of Singapore (IMDA), Aicadium, and AI Verify Foundation. Model AI Governance Framework for Generative AI: Fostering a Trusted Ecosys- tem, May 2024. URL https://aiverifyfoundation.sg/wp-content/uploads/2024/ 05/Model-AI-Governance-Framework...

  72. [80]

    Analysing and reclassifying open access informa- tion in OpenAlex, 2023

    Najko Jahn, Nick Haupka, and Anne Hobert. Analysing and reclassifying open access informa- tion in OpenAlex, 2023. URL https://subugoe.github.io/scholcomm_analytics/ posts/oalex_oa_status/?utm_source=chatgpt.com

  73. [81]

    Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi

    Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi. Comet-atomic 2020: On symbolic and neural commonsense knowledge graphs. In AAAI Conference on Artificial Intelligence, pages 6384–6392, 2020. 17

  74. [82]

    Position: The most expensive part of an llm should be its training data

    Nikhil Kandpal and Colin Raffel. Position: The most expensive part of an llm should be its training data. arXiv preprint arXiv:2504.12427, 2025

  75. [83]

    Google patents public data

    IFI CLAIMS Patent Services and Google. Google patents public data. https://patents. google.com/, 2023. Licensed under a Creative Commons Attribution 4.0 International License

  76. [84]

    Large language models struggle to learn long-tail knowledge

    Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning, pages 15696–15707. PMLR, 2023

  77. [85]

    When choosing plausible alternatives, clever hans can be clever

    Pride Kavumba, Naoya Inoue, Benjamin Heinzerling, Keshav Singh, Paul Reisert, and Kentaro Inui. When choosing plausible alternatives, clever hans can be clever. ArXiv, abs/1911.00225, 2019

  78. [86]

    Scitail: A textual entailment dataset from science question answering

    Tushar Khot, Ashish Sabharwal, and Peter Clark. Scitail: A textual entailment dataset from science question answering. In AAAI Conference on Artificial Intelligence, pages 5189–5197, 2018

  79. [87]

    Bag of tricks for efficient text classification

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 427–431. Associatio...

  80. [88]

    Andreas Kopf, Yannic Kilcher, Dimitri von Rutte, Sotiris Anagnostidis, Zhi Rui Tam, K. Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Rich’ard Nagyfi, ES Shahul, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and A. ...

  81. [89]

    Deduplicating training data mitigates privacy risks in language models, 2022

    Nikhil Kandpal, Eric Wallace, and Colin Raffel. Deduplicating training data mitigates privacy risks in language models, 2022. URL https://arxiv.org/abs/2202.06539

  82. [90]

    Faisal Ladhak, Esin Durmus, Claire Cardie, and K. McKeown. Wikilingua: A new benchmark dataset for multilingual abstractive summarization. ArXiv, abs/2010.03093, 2020

  83. [91]

    Releasing Common Corpus: the largest public domain dataset for training LLMs, 2024

    Pierre-Carl Langlais. Releasing Common Corpus: the largest public domain dataset for training LLMs, 2024. URL https://huggingface.co/blog/Pclanglais/common-corpus

  84. [92]

    AI White Paper 2024: New Strate- gies in Stage II, Toward the world’s most AI-friendly country, April 2024

    LDP Headquarters for the Promotion of Digital Society and Project Team on the Evolution and Implementation of AIs. AI White Paper 2024: New Strate- gies in Stage II, Toward the world’s most AI-friendly country, April 2024. URL https://aiverifyfoundation.sg/wp-content/uploads/2...

  85. [93]

    Graham, F.Q

    Rodney Michael Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, Miles Crawford, Doug Downey, Jason Dunkelberger, Oren Etzioni, Rob Evans, Sergey Feldman, Joseph Go...

  86. [94]

    Deduplicating training data makes language models better, 2022

    Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better, 2022. URL https://arxiv.org/abs/2107.06499

  87. [95]

    Kwiatkowski, J

    T. Kwiatkowski, J. Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, D. Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc V . Le, and Slav Petrov....

  88. [96]

    Feder Cooper, and James Grimmelmann

    Katherine Lee, A. Feder Cooper, and James Grimmelmann. Talkin’ ’Bout AI Generation: Copyright and the Generative-AI Supply Chain. arXiv preprint arXiv:2309.08133, 2023

  89. [97]

    Kontokostas, Pablo N

    Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, D. Kontokostas, Pablo N. Mendes, Sebastian Hellmann, M. Morsey, Patrick van Kleef, S. Auer, and Christian Bizer. Dbpedia - a large-scale, multilingual knowledge base extracted from wikipedia. Semantic Web, 6:167–195, 2015

  90. [98]

    Levesque, E

    H. Levesque, E. Davis, and L. Morgenstern. The winograd schema challenge. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning, 2011

  91. [99]

    Lebret, David Grangier, and Michael Auli

    R. Lebret, David Grangier, and Michael Auli. Neural text generation from structured data with application to the biography domain. In Conference on Empirical Methods in Natural Language Processing, pages 1203–1213, 2016

  92. [100]

    Xin Li and D. Roth. Learning question classifiers. In International Conference on Computational Linguistics, pages 1–7, 2002

  93. [101]

    Feder Cooper, James Grimmelmann, and Daphne Ippolito

    Katherine Lee, A. Feder Cooper, James Grimmelmann, and Daphne Ippolito. AI and Law: The Next Generation. SSRN, 2023. http://dx.doi.org/10.2139/ssrn.4580739

  94. [102]

    Testing the ability of language models to interpret figurative language

    Emmy Liu, Chenxuan Cui, Kenneth Zheng, and Graham Neubig. Testing the ability of language models to interpret figurative language. ArXiv, abs/2204.12632, 2022. 19

  95. [103]

    Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, Richard Fan, Yi Gu, Victor Miller, Yonghao Zhuang, Guowei He, Haonan Li, Fajri Koto, Liping Tang, Nikhil Ranjan, Zhiqiang Shen, Xuguang Ren, Rob...

  96. [104]

    S2ORC: The semantic scholar open research corpus

    Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. S2ORC: The semantic scholar open research corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 4969–4983, Online, July 2020. Association for Computational...

  97. [105]

    Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar

    Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan ...

  98. [106]

    URL https://arxiv.org/abs/2406.11794

  99. [107]

    Consent in crisis: The rapid decline of the AI data commons

    Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderinwale, William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh, Manuel Cherep, Ahmad Anis, An Dinh, Caroline Chitongo, Da ...

  100. [108]

    Lin, Jacob Hilton, and Owain Evans

    Stephanie C. Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. In Annual Meeting of the Association for Computational Linguistics, pages 3214–3252, 2021

  101. [109]

    Bridging the data provenance gap across text, speech and video

    Shayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary, Joanna Materzynska, William Brannon, Robert Mahari, Naana Obeng-Marnu, Manan Dey, Mohammed Hamdy, et al. Bridging the data provenance gap across text, speech and video. arXiv preprint arXiv:2412.17847, 2024

  102. [110]

    A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity

    Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, et al. A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity. In Proceedings of th...

  103. [111]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations , 2019. URL https: //openreview.net/forum?id=Bkg6RiCqY7

  104. [112]

    The responsible foundation model development cheatsheet: A review of tools & resources

    Shayne Longpre, Stella Biderman, Alon Albalak, Hailey Schoelkopf, Daniel McDuff, Sayash Kapoor, Kevin Klyman, Kyle Lo, Gabriel Ilharco, Nay San, et al. The responsible foundation model development cheatsheet: A review of tools & resources. Transactions on Machine Learning Rese...

  105. [113]

    A large-scale audit of dataset licensing and attribution in AI

    Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, Xinyi (Alexis) Wu, Enrico Shippole, Kurt Bollacker, Tongshuang Wu, Luis Villa, Sandy Pentland, and Sara Hooker. A la...

  106. [114]

    Discit ergo est: Training data provenance and fair use

    Robert Mahari and Shayne Longpre. Discit ergo est: Training data provenance and fair use. Robert Mahari and Shayne Longpre, Discit ergo est: Training Data Provenance And Fair Use, Dynamics of Generative AI (ed. Thibault Schrepel & Volker Stocker), Network Law Review, Winter, 2023

  107. [115]

    Data authenticity, consent, & provenance for ai are all broken: what will it take to fix them? arXiv preprint arXiv:2404.12691, 2024

    Shayne Longpre, Robert Mahari, Naana Obeng-Marnu, William Brannon, Tobin South, Katy Gero, Sandy Pentland, and Jad Kabbara. Data authenticity, consent, & provenance for ai are all broken: what will it take to fix them? arXiv preprint arXiv:2404.12691, 2024

  108. [116]

    Stephen Merity, Caiming Xiong, James Bradbury, and R. Socher. Pointer sentinel mixture models. ArXiv, abs/1609.07843, 2016

  109. [117]

    Gray, The Google Books Team, Joseph P

    Jean-Baptiste Michel, Yuan Kui Shen, Aviva Presser Aiden, Adrian Veres, Matthew K. Gray, The Google Books Team, Joseph P. Pickett, Dale Hoiberg, Dan Clancy, Peter Norvig, Jon Orwant, Steven Pinker, Martin A. Nowak, and Erez Lieberman Aiden. Quantitative analysis of culture usi...

  110. [118]

    Can a suit of armor conduct electricity? a new dataset for open book question answering

    Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018

  111. [119]

    i’d rather just go to bed

    Annie Louis, D. Roth, and Filip Radlinski. “i’d rather just go to bed”’: Understanding indirect answers. In Conference on Empirical Methods in Natural Language Processing , volume abs/2010.03450, 2020

  112. [120]

    Starcoder 2 and the stack v2: The next generation, 2024

    Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Z...

  113. [121]

    Scaling data-constrained language models

    Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Alek- sandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A Raffel. Scaling data-constrained language models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advance...

  114. [122]

    Large text compression benchmark, 2011

    Matt Mahoney. Large text compression benchmark, 2011

  115. [123]

    The e2e dataset: New challenges for end-to-end generation

    Jekaterina Novikova, Ondrej Dusek, and Verena Rieser. The e2e dataset: New challenges for end-to-end generation. ArXiv, abs/1706.09254, 2017

  116. [124]

    Ananiadou

    Tomoko Ohta, Sampo Pyysalo, Junichi Tsujii, and S. Ananiadou. Open-domain anatomical entity mention detection. In Annual Meeting of the Association for Computational Linguistics, pages 27–36, 2012

  117. [125]

    Zhang, Eunsol Choi, and Greg Durrett

    Yasumasa Onoe, Michael J.Q. Zhang, Eunsol Choi, and Greg Durrett. Creak: A dataset for commonsense reasoning over entity knowledge. ArXiv, abs/2109.01653, 2021

  118. [126]

    Smith, and Luke Zettlemoyer

    Sewon Min, Suchin Gururangan, Eric Wallace, Weijia Shi, Hannaneh Hajishirzi, Noah A. Smith, and Luke Zettlemoyer. SILO language models: Isolating legal risk in a nonparametric datastore. In The Twelfth International Conference on Learning Representations, 2024. URL https://ope...

  119. [127]

    Scaling data-constrained language models

    Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. Scaling data-constrained language models. Advances in Neural Information Processing Systems, 36, 2023

  120. [128]

    Privacy auditing of large language models

    Ashwinee Panda, Xinyu Tang, Milad Nasr, Christopher A Choquette-Choo, and Prateek Mittal. Privacy auditing of large language models. arXiv preprint arXiv:2503.06808, 2025

  121. [129]

    Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. Crows-pairs: A challenge dataset for measuring social biases in masked language models. In Conference on Empirical Methods in Natural Language Processing, pages 1953–1967, 2020

  122. [130]

    Directive (eu) 2019/790,

    European Parliament and Council of the European Union. Directive (eu) 2019/790,

  123. [131]

    Parser for uk parliament proceedings

    ParlParse. Parser for uk parliament proceedings. https://parser.theyworkforyou.com/,

  124. [132]

    The FineWeb datasets: Decanting the web for the finest text data at scale

    Guilherme Penedo, Hynek Kydlí ˇcek, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro V on Werra, and Thomas Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale. Advances in Neural Information Processing Systems, 37, 2024

  125. [133]

    URL https://openalex.org

    OpenAlex, 2025. URL https://openalex.org. 21

  126. [134]

    Khudanpur

    Vassil Panayotov, Guoguo Chen, Daniel Povey, and S. Khudanpur. Librispeech: An asr corpus based on public domain audio books. 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5206–5210, 2015

  127. [135]

    Dynasent: A dynamic benchmark for sentiment analysis

    Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. Dynasent: A dynamic benchmark for sentiment analysis. ArXiv, abs/2012.15349, 2020

  128. [136]

    Trak: Attributing model behavior at scale, 2023

    Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry. Trak: Attributing model behavior at scale, 2023. URL https://arxiv.org/abs/2303.14186

  129. [137]

    Language models are unsupervised multitask learners, 2019

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners, 2019

  130. [138]

    Robust speech recognition via large-scale weak supervision, 2022

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision, 2022. URL https://arxiv.org/abs/2212.04356

  131. [139]

    Balog, B

    Filip Radlinski, K. Balog, B. Byrne, and K. Krishnamoorthi. Coached conversational preference elicitation: A case study in understanding movie preferences. In SIGDIAL Conferences, pages 353–360, 2019

  132. [140]

    Accessed: 2025-05-09

  133. [141]

    Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hen- nigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendrick...

  134. [142]

    Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019

    Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019

  135. [143]

    Ponti, Goran Glavavs, Olga Majewska, Qianchu Liu, Ivan Vulic, and A

    E. Ponti, Goran Glavavs, Olga Majewska, Qianchu Liu, Ivan Vulic, and A. Korhonen. Xcopa: A multilingual dataset for causal commonsense reasoning. In Conference on Empirical Methods in Natural Language Processing, pages 2362–2376, 2020

  136. [144]

    Nazneen Rajani, Bryan McCann, Caiming Xiong, and R. Socher. Explain yourself! leveraging language models for commonsense reasoning. ArXiv, abs/1906.02361, 2019

  137. [145]

    Improving language understanding by generative pre-training, 2018

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training, 2018

  138. [146]

    Squad: 100,000+ questions for machine comprehension of text

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. In Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, 2016

  139. [147]

    Know what you don’t know: Unanswerable questions for squad

    Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don’t know: Unanswerable questions for squad. In Annual Meeting of the Association for Computational Linguistics , volume abs/1806.03822, 2018

  140. [148]

    Schema-guided dialogue state tracking task at dstc8

    Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. Schema-guided dialogue state tracking task at dstc8. ArXiv, abs/2002.01359, 2020

  141. [149]

    Compressive transformers for long-range sequence modelling

    Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. arXiv preprint, 2019. URL https://arxiv.org/abs/1911.05507

  142. [150]

    Schmidt, Yuval Pinter, and Chris Tanner

    Varshini Reddy, Craig W. Schmidt, Yuval Pinter, and Chris Tanner. How much is enough? the diminishing returns of tokenization training data, 2025. URL https://arxiv.org/abs/2502.20273

  143. [151]

    Thang, N

    Hammam Riza, Michael Purwoadi, Gunarso, Teduh Uliniansyah, Aw Ai Ti, Sharifah Mahani Aljunied, Luong Chi Mai, V . Thang, N. Thai, Vichet Chea, Rapid Sun, Sethserey Sam, Sopheap Seng, K. Soe, K. Nwet, M. Utiyama, and Chenchen Ding. Introduction of the asian language treebank. I...

  144. [152]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140), 2020

  145. [153]

    The size of datasets used to train language models doubles approximately every seven months, 2024

    Robi Rahman and David Owen. The size of datasets used to train language models doubles approximately every seven months, 2024. URL https://epoch.ai/data- insights/dataset-size-trend. Accessed: 2025-05-08

  146. [154]

    A primer in bertology: What we know about how BERT works

    Anna Rogers, Olga Kovaleva, and Anna Rumshisky. A primer in bertology: What we know about how BERT works. Transactions of the association for computational linguistics, 8, 2021

  147. [155]

    Outsider oversight: Designing a third party audit ecosystem for ai governance

    Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, and Daniel Ho. Outsider oversight: Designing a third party audit ecosystem for ai governance. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, pages 557–571, 2022

  148. [156]

    Matthew Sag and Peter K. Yu. The globalization of copyright exceptions for ai train- ing. Emory Law Journal , 74, 2025. doi: http://dx.doi.org/10.2139/ssrn.4976393. URL https://ssrn.com/abstract=4976393

  149. [157]

    Conjnli: Natural language inference over conjunctive sentences

    Swarnadeep Saha, Yixin Nie, and Mohit Bansal. Conjnli: Natural language inference over conjunctive sentences. In Conference on Empirical Methods in Natural Language Processing, pages 8240–8252, 2020

  150. [158]

    Winogrande

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande. Communications of the ACM, 64:99 – 106, 2019

  151. [159]

    Condaqa: A contrastive reading comprehension dataset for reasoning about negation

    Abhilasha Ravichander, Matt Gardner, and Ana Marasovi´c. Condaqa: A contrastive reading comprehension dataset for reasoning about negation. ArXiv, abs/2211.00295, 2022

  152. [160]

    Socialiqa: Commonsense reasoning about social interactions, 2019

    Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. Socialiqa: Commonsense reasoning about social interactions, 2019. URL https://arxiv.org/abs/1904.09728

  153. [161]

    Sboev, A

    A. Sboev, A. Naumov, and R. Rybka. Data-driven model for emotion detection in russian texts. In BICAAI, pages 637–642, 2020

  154. [162]

    How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020

    Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020

  155. [163]

    Getting closer to ai complete question answering: A set of prerequisite real tasks

    Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. Getting closer to ai complete question answering: A set of prerequisite real tasks. In AAAI Conference on Artificial Intelligence, pages 8722–8731, 2020

  156. [164]

    Hausknecht

    Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew J. Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning. ArXiv, abs/2010.03768, 2020

  157. [165]

    Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A

    Rachel Rudinger, Vered Shwartz, Jena D. Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A. Smith, and Yejin Choi. Thinking like a skeptic: Defeasible inference in natural language. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the As- sociation f...

  158. [166]

    peS2o (Pretraining Efficiently on S2ORC) Dataset

    Luca Soldaini and Kyle Lo. peS2o (Pretraining Efficiently on S2ORC) Dataset. Technical report, Allen Institute for AI, 2023. ODC-By, https://github.com/allenai/pes2o

  159. [167]

    Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A

    Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Mu...

  160. [168]

    Smith, and Luke Zettlemoyer

    Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. Evaluating gender bias in machine translation. ArXiv, abs/1906.00591, 2019

  161. [169]

    Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures

    Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures. In 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7). Leibniz-Institut für Deutsche Sprache, 2019. 24

  162. [170]

    How to Think About Remedies in the Generative AI Copyright Cases

    Pamela Samuelson. How to Think About Remedies in the Generative AI Copyright Cases. Lawfare, February 2024. URL https://www.lawfaremedia.org/article/how-to- think-about-remedies-in-the-generative-ai-copyright-cases

  163. [171]

    Commonsenseqa: A question answering challenge targeting commonsense knowledge

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. arXiv preprint arXiv:1811.00937, 2018

  164. [172]

    Niket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson, and E. Hovy. A dataset for tracking entities in open domain procedural text. ArXiv, abs/2011.08092, 2020

  165. [173]

    Barzilay

    Tal Schuster, Adam Fisch, and R. Barzilay. Get your vitamin c! robust fact verification with contrastive evidence. In North American Chapter of the Association for Computational Linguistics, pages 624–643, 2021

  166. [174]

    Emily Sheng and David C. Uthus. Investigating societal biases in a poetry composition system. ArXiv, abs/2011.02686, 2020

  167. [175]

    Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023

    MosaicML NLP Team. Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023. URL www.mosaicml.com/blog/mpt-7b. Accessed: 2023-05-05

  168. [176]

    Shivalika Singh, Freddie Vargus, Daniel Dsouza, B"orje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemi’nski, Hakimeh...

  169. [177]

    Fever: a large-scale dataset for fact extraction and verification

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. Fever: a large-scale dataset for fact extraction and verification. ArXiv, abs/1803.05355, 2018

  170. [178]

    Mad- dison

    Anvith Thudi, Evianne Rovers, Yangjun Ruan, Tristan Thrush, and Chris J. Mad- dison. Mixmin: Finding data mixtures via convex minimization, 2025. URL https://arxiv.org/abs/2502.10510

  171. [179]

    Llama: Open and efficient foundation language models, 2023

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language...

  172. [180]

    Llama 2: Open foundation and fine-tuned chat models

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023

  173. [181]

    Quartz: An open-domain dataset of qualitative relationship questions

    Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. Quartz: An open-domain dataset of qualitative relationship questions. In Conference on Empirical Methods in Natural Language Processing, volume abs/1909.03553, 2019

  174. [182]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017

  175. [183]

    Meta Lingua: A minimal PyTorch LLM training library, 2024

    Mathurin Videau, Badr Youbi Idrissi, Daniel Haziza, Luca Wehrstedt, Jade Copet, Olivier Teytaud, and David Lopez-Paz. Meta Lingua: A minimal PyTorch LLM training library, 2024. URL https://github.com/facebookresearch/lingua

  176. [184]

    Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, and Eric P. Xing. Txt360: A top-quality llm pre-training dataset requires t...

  177. [185]

    Choudhury

    Ishan Tarunesh, Somak Aditya, and M. Choudhury. Trusting roberta over bert: Insights from checklisting the natural language inference task. ArXiv, abs/2107.07229, 2021

  178. [186]

    Resolving gendered ambiguous pronouns with bert

    Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. Resolving gendered ambiguous pronouns with bert. ArXiv, abs/1906.01161, 2019

  179. [187]

    Qwen3, April 2025

    Qwen Team. Qwen3, April 2025. URL https://qwenlm.github.io/blog/qwen3/

  180. [188]

    Organize the web: Constructing domains enhances pre-training data curation

    Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen, and Luca Soldaini. Organize the web: Constructing domains enhances pre-training data curation. arXiv preprint arXiv:2502.10341, 2025

  181. [189]

    Doremi: Optimizing data mixtures speeds up language model pretraining

    Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy S Liang, Quoc V Le, Tengyu Ma, and Adams Wei Yu. Doremi: Optimizing data mixtures speeds up language model pretraining. Advances in Neural Information Processing Systems, 36, 2023

  182. [190]

    Cohen, R

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, R. Salakhutdinov, and Christopher D. Manning. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing , pages 2369–2380, 2018

  183. [191]

    OpenAI prepares to fight for its life as legal troubles mount

    Cat Zakrzewski, Nitasha Tiku, and Elizabeth Dwoskin. OpenAI prepares to fight for its life as legal troubles mount. The Washington Post, 2024

  184. [192]

    Open parliament license

    UK Parliament. Open parliament license. https://www.parliament.uk/site- information/copyright-parliament/open-parliament-licence/ , Unknown. Accessed: 2025-05-09

  185. [193]

    MAP-Neo: Highly capable and transparent bilingual large language model series

    Ge Zhang, Scott Qu, Jiaheng Liu, Chenchen Zhang, Chenghua Lin, Chou Leuang Yu, Danny Pan, Esther Cheng, Jie Liu, Qunshu Lin, Raven Yuan, Tuney Zheng, Wei Pang, Xinrun Du, Yiming Liang, Yinghao Ma, Yizhi Li, Ziyang Ma, Bill Lin, Emmanouil Benetos, Huan Yang, Junting Zhou, Kaiji...

  186. [194]

    Winowhy: A deep diagnosis of essential commonsense knowledge for answering winograd schema challenge

    Hongming Zhang, Xinran Zhao, and Yangqiu Song. Winowhy: A deep diagnosis of essential commonsense knowledge for answering winograd schema challenge. In Annual Meeting of the Association for Computational Linguistics, pages 5736–5745, 2020

  187. [195]

    Helpsteer: Multi-attribute helpfulness dataset for steerlm

    Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Polak Scowcroft, Neel Kant, Aidan Swope, and Oleksii Kuchaiev. Helpsteer: Multi-attribute helpfulness dataset for steerlm. ArXiv, abs/2311.09528, 2023

  188. [196]

    Redpajama: an open dataset for training large language models, 2024

    Maurice Weber, Daniel Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Ré, Irina Rish, and Ce Zhang. Redpaja...

  189. [198]

    Le, Andrew M

    Wei Wei, Quoc V . Le, Andrew M. Dai, and Jia Li. Airdialogue: An environment for goal-oriented dialogue research. In Conference on Empirical Methods in Natural Language Processing, pages 3844–3854, 2018

  190. [203]

    Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019

  191. [206]

    Wildchat: 1m chatGPT interaction logs in the wild

    Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wildchat: 1m chatGPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Bl8u7ZRlbM

  192. [207]

    accepted answer

    Ben Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, and D. Roth. Temporal reasoning on implicit events from distant supervision. ArXiv, abs/2010.12753, 2020. 26 Appendix Table of Contents A Contributions 28 B Detailed Description of Sources 28 B.1 Scientific ...

  193. [1994]

    URL https://api.semanticscholar.org/CorpusID:59804030

  194. [2016]

    URL https://aclanthology

    European Language Resources Association (ELRA). URL https://aclanthology. org/L16-1146/

  195. [2019]

    URL https://eur-lex.europa.eu/legal-content/EN/ALL/?uri=CELEX: 32019L0790#art_3

  196. [2020]

    doi: 10.18653/v1/2020.findings-emnlp.418

    Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.418. URL https://aclanthology.org/2020.findings-emnlp.418/. 23

  197. [2021]

    URL https://api.semanticscholar.org/CorpusID:245353475

  198. [2024]

    URL https://www.law.cornell.edu/uscode/text/17/105

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.