REVIEW 3 major objections 5 minor 1 cited by
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper establishes that large language models can be trained to competitive performance using only public domain and openly licensed text, and releases an 8TB dataset plus two 7B models as evidence.
desk verdict A large, reproducible open-license corpus with a clean code win, but the knowledge/reasoning headline is partly an artifact of tuning the mixture on the same benchmarks as the comparison. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Common Pile corpus itself, together with the Comma training mixture derived from it. The Common Pile is an 8TB collection assembled from 30 sources under strict sourcing rules: only content whose license meets the Open Definition 2.1 standard—free access, use, modification, and sharing for any purpose—and whose providers are believed to be the actual rights holders. The argument is carried by three mechanisms: license due diligence at the source level, including manual verification of web domains and exclusion of unreliable sources; a filtering and deduplication pipeline that turns raw text into training tokens; and per-source quality evaluation that sets mixing weights to up-weight high-quality sources while capping repetition. The final validation mechanism is a controlled comparison: identical small models trained on different corpora, followed by two full-scale 7B training runs.
What would settle it
Take a random sample of documents in the Common Pile, trace each to the stated rights holder, and check whether the open-license claim matches the holder's terms; if a substantial fraction are in-copyright or mislabeled, the claim that the corpus is openly licensed fails.
Extended reading notes
Core claim
The authors set out to determine whether LLM performance depends on unlicensed web text. They construct the Common Pile v0.1, an 8TB corpus from 30 sources, all selected under the Open Definition 2.1 standard for open licenses or public domain status. They then filter, deduplicate, and reweight this corpus into the Comma training mixture and train two 7-billion-parameter models on 1 and 2 trillion tokens. On standard knowledge, reasoning, and coding benchmarks, Comma v0.1-1T and Comma v0.1-2T are competitive with budget-matched models trained on unlicensed data such as Llama 1 and 2 7B, and a controlled 1.7B-parameter ablation shows the Common Pile outperforms prior open-license corpora. The authors also release the dataset, code, mixture, and checkpoints, while acknowledging that license laundering and changed license terms may have allowed some unlicensed text into the corpus.
Load-bearing premise
The central premise is that the license labels on the 30 sources faithfully reflect the rights of the people who posted them, so the corpus really is openly licensed.
Editorial extensions
If this is right
- A 7B model trained entirely on openly licensed text can match or beat budget-matched models trained on unlicensed web text on several standard benchmarks, including knowledge and code tasks.
- The 8TB Common Pile is, to the authors' knowledge, the largest openly licensed pretraining corpus, making open-license pretraining a practical option rather than a toy-scale exercise.
- Releasing the dataset, mixture, code, and checkpoints allows others to reproduce and extend the results without scraping unlicensed data.
- The controlled 1.7B-parameter ablation shows the Common Pile outperforms prior openly licensed corpora on the same budget, so corpus quality, not just license status, drives the result.
- The 2T run repeats the 1T mixture and still stays competitive, which the paper attributes to the mixture's quality and a possible ceiling from heavy repetition.
Reading between the lines
- Beyond the paper: the performance result and the licensing result are separable; even if some unlicensed text slipped in, the main contribution could still hold as a scale and curation demonstration, but the legal and ethical selling point would weaken.
- Beyond the paper: the same sourcing-and-validation recipe could be transferred to non-English corpora, code-specific datasets, or multimodal data, where open licensing questions are equally pressing.
- Beyond the paper: a direct test of the licensing premise would be to retrain on a provenance-verified subset and compare; the paper's own ablation removing one curated source suggests performance is fairly robust to source changes.
- Beyond the paper: the reported growth of openly licensed text over time implies the constraint may become less binding as more contemporary content is released under open licenses.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Common Pile v0.1, an 8 TB collection of text that the authors argue is composed exclusively of public domain and openly licensed material drawn from 30 diverse sources. The authors apply filtering and deduplication to obtain the Comma dataset, tune source mixing weights using per-source 1.7B-parameter models evaluated on a set of early-signal benchmarks, and then train two 7B-parameter models, Comma v0.1-1T and Comma v0.1-2T, on 1T and 2T tokens respectively. The models are evaluated on knowledge, reasoning, and code benchmarks and compared against budget-matched models trained on unlicensed text, with claims that Comma is competitive with Llama 1/2, OLMo, and DeepSeekLLM. The paper also reports controlled 1.7B-scale ablations against earlier openly licensed corpora and several unlicensed baselines. The manuscript releases the dataset, preprocessing code, training mixture, and model checkpoints.
Significance. If the central claims hold, this is a substantial contribution: it provides the largest openly licensed pretraining corpus to date, along with trained models and a reproducible pipeline, and it presents evidence that open-license-only pretraining can approach the quality of models trained on unlicensed web text. The paper is commendably transparent about licensing caveats, includes controlled small-scale ablations, and ships machine-checkable artifacts (dataset, code, checkpoints). The main evidence for the headline claim, however, is weakened by the fact that the data mixture was tuned on the same evaluation benchmarks used to compare the final models against baselines, and by the use of a custom tokenizer that makes the '1 trillion tokens' budget not directly comparable with baseline tokenizers. These issues are addressable with additional experiments or re-analysis, which is why I recommend major revision rather than rejection.
major comments (3)
- [Section 4.2 and Section 4.4 (Tables 10-11)] The source mixing weights for the Comma dataset were selected by training per-source 1.7B models and evaluating them on the 'early signal' tasks from Penedo et al. (ARC, MMLU, HellaSwag, OBQA, CSQA, PIQA, SIQA). The final 7B models are then evaluated on almost exactly that benchmark suite, with only BoolQ and the code benchmarks added. Consequently, the comparisons against Llama 1/2, MPT, OLMo, and DeepSeekLLM are not clean: the Comma mixture was explicitly optimized for these benchmarks, while the baseline mixtures were not. This undermines the inference from 'Comma scores comparably' to 'openly licensed text is sufficient for performant LLMs.' The code results (HumanEval, MBPP) are less affected because those benchmarks were not used during mixture selection, but the knowledge/reasoning scores, which dominate the reported averages, are directly confounded. I recommend either evaluating on a held-out suite that was not used in any stage of mixture selection, or demonstrating through an ablation that mixture selection on this benchmark set does not materially affect the final comparisons.
- [Section 4.3 (controlled 1.7B experiments)] The controlled dataset-quality experiment in Section 4.3 shares the same circularity concern. The 'Comma dataset' used in this comparison is a filtered and reweighted mixture whose weights were selected using the same early-signal benchmarks on which the models are then evaluated. When comparing against OLC, Common Corpus, KL3M, the Pile, OSCAR, and FineWeb, the Comma model benefits not only from the underlying text but also from benchmark-specific tuning; the baseline datasets were not given an equivalent mixture-tuning step. The paper's statement that differences in performance 'stem primarily from the quality of each dataset' is therefore too strong. A fairer comparison would either apply the same per-source tuning procedure to each candidate corpus (using a held-out validation set) or compare unweighted versions of each corpus.
- [Section 4.4 (tokenization and budget matching)] The claim that Comma v0.1-1T and v0.1-2T are 'budget-matched' to Llama 1/2 and other baselines implicitly assumes that 1T (or 2T) tokens of Comma data correspond to the same training budget as 1T (or 2T) tokens of the baselines. However, Comma uses a custom BPE tokenizer with a vocabulary size of 64,000, whereas Llama 1/2, MPT, OLMo, and DeepSeekLLM use different tokenizers, typically with smaller vocabularies. If the Comma tokenizer is more efficient (fewer tokens per byte of text), then training on 1T Comma tokens may actually process more text than training on 1T tokens of a baseline, making the comparison generous to Comma. The paper does not report tokenizer efficiency (e.g., bytes per token, or the number of tokens per document for a fixed text sample), nor does it compare training FLOPs. This should be addressed by reporting a tokenizer-comparability metric or by comparing models trained with the same tokenizer on the same number of tokens.
minor comments (5)
- [Appendix B.3 (USPTO)] In the description of the USPTO source, the text says 'We include parents from the US Patents and Trademark Office' but this should read 'We include patents from the US Patents and Trademark Office.'
- [Appendix B.8 (PEPs)] The sentence 'There are been 661 PEPs published' contains a grammatical error and should be 'There have been 661 PEPs published.'
- [Table 9 caption] The caption states that removing DPI data has 'marginal affect' on dataset quality; this should be 'marginal effect.'
- [Appendix O] The text says 'the the training batch size is 8.3M ( 223) versus the 2.1M ( 221) tokens per step' with a doubled 'the' and missing superscript formatting for the exponents; these should be corrected for readability.
- [Tables 10-11] Abbreviations such as 'HS', 'HEval', and 'MBPP' are used without definition in the table captions or nearby text; defining them (HellaSwag, HumanEval, MBPP) would improve accessibility.
Circularity Check
Benchmark-tuned data mixture is evaluated on the same benchmark suite, making the headline competitive-performance claim partly constructed.
-
fitted input called prediction
[Section 4.2 (data mixing), Section 4.3 (controlled 1.7B ablations), Section 4.4 (Comma v0.1 evaluation)]
"To determine mixing weights, we first trained per-source language models ... Based on the performance of these per-source models, we heuristically set mixing weights to up- and down-weight high- and low-performance sources. Each model was then evaluated using the set of 'early signal' tasks identified by Penedo et al.: ARC, MMLU, HellaSwag, OpenBookQA, CommonSenseQA, PIQA, and SocialIQA. We evaluate the Comma v0.1 models on ARC, MMLU, BoolQ, HellaSwag, OpenBookQA, CommonsenseQA, PIQA, and SIQA to probe world knowledge and reasoning."
The mixing weights are fitted to per-source 1.7B models' scores on the 'early signal' tasks (ARC, MMLU, HellaSwag, OBQA, CSQA, PIQA, SIQA), and the 7B Comma models are then evaluated on almost exactly the same tasks plus BoolQ and code. The paper's central evidence that openly licensed data 'can be used as the foundation for competitive LLMs' therefore uses the same benchmark signal that determined the up/down-weighting of sources. The knowledge/reasoning columns in Tables 10-11 are not an independent test of the dataset: the mixture was explicitly chosen to make those tasks score well, so competitive performance on them is partly an artifact of selection rather than purely of the open-license constraint.
full rationale
The paper's derivation chain is otherwise self-contained: the dataset is built from sources with documented licensing choices, preprocessing and deduplication are standard, and the 7B models are compared against fixed external baselines that were not tuned on this benchmark suite. The main circularity is the shared benchmark signal between mixture selection and evaluation. This does not invalidate the dataset or the code results, but it materially weakens the inference from 'Comma scores well on these benchmarks' to 'openly licensed text alone supports competitive performance.' A held-out evaluation would settle whether the effect is real or an artifact of tuning. No other load-bearing self-citation or definitional equivalence was found.
Assumptions & free parameters
free parameters (4)
- Source mixing weights (Table 7) =
Various repeats, e.g., peS2o 6x, USPTO 0.25x
- Cool-down mixture weights (Table 8) =
37.5B token budget from 11 sources
- Per-source filtering thresholds (Table 5) =
e.g., language > 0.5, min length 100-700, toxicity > 0.1
- Weight decay =
0.2
assumptions (3)
- domain assumption The evaluated benchmarks are not present in the DPI-sourced training data
- domain assumption Per-source model performance on the early-signal tasks predicts contribution to the mixture
- domain assumption License metadata from each source is accurate and satisfies the Open Definition 2.1
Cite this review
Pith. "Pith review of The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text." pith.science (2026). https://pith.science/paper/L7QFMMII
@misc{pith2026250605209,
author = {Pith},
title = {Pith review of: The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text},
year = {2026},
howpublished = {\url{https://pith.science/paper/L7QFMMII}},
note = {Machine review of arXiv:2506.05209}
}
read the original abstract
Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement and ethical concerns. Training LLMs on openly licensed text presents a first step towards addressing these issues, but prior data collection efforts have yielded datasets too small or low-quality to produce performant LLMs. To address this gap, we collect, curate, and release the Common Pile v0.1, an eight terabyte collection of openly licensed text designed for LLM pretraining. The Common Pile comprises content from 30 sources that span diverse domains including research papers, code, books, encyclopedias, educational materials, audio transcripts, and more. Crucially, we validate our efforts by training two 7 billion parameter LLMs on text from the Common Pile: Comma v0.1-1T and Comma v0.1-2T, trained on 1 and 2 trillion tokens respectively. Both models attain competitive performance to LLMs trained on unlicensed text with similar computational budgets, such as Llama 1 and 2 7B. In addition to releasing the Common Pile v0.1 itself, we also release the code used in its creation as well as the training mixture and checkpoints for the Comma v0.1 models.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
Applying the Chinese Wall Reverse Engineering Technique to Large Language Model Code Editing
Using Gemini 2.5 Pro to annotate code with edit instructions improved Comma v0.1 1T's CanItEdit pass@20 from 20.00 to 33.33 and Starcoder2 Instruct's pass@1 from 35.10 to 42.05.
Reference graph
Works this paper leans on
-
[1]
Code § 102
17 U.S. Code § 102. Subject matter of copyright: In general, December 1990. URL https: //www.law.cornell.edu/uscode/text/17/102
1990
-
[2]
Code § 105
17 U.S. Code § 105. Subject matter of copyright: United States Government works, December
-
[3]
Efficient online data mixing for language model pre-training
Alon Albalak, Liangming Pan, Colin Raffel, and William Yang Wang. Efficient online data mixing for language model pre-training. arXiv preprint arXiv:2312.02406, 2023
arXiv 2023
-
[4]
A survey on data selection for language models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. A survey on data selection for language models. Transactions on Machine Learning Research, 2024
2024
-
[5]
Smollm2: When smol goes big – data-centric training of a small language model, 2025
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíˇcek, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von Wer...
arXiv 2025
-
[6]
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. On the cross-lingual transferability of monolingual representations. In Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, 2019
2019
-
[7]
To code, or not to code? exploring impact of code in pre-training
Viraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot, Ivan Zhang, Acyr Locatelli, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker. To code, or not to code? exploring impact of code in pre-training. arXiv preprint arXiv:2408.10914, 2024
arXiv 2024
-
[8]
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021
arXiv 2021
Show all 206 references
-
[9]
Towards best practices for open datasets for llm training, 2025
Stefan Baack, Stella Biderman, Kasia Odrozek, Aviya Skowron, Ayah Bdeir, Jillian Bom- marito, Jennifer Ding, Maximilian Gahntz, Paul Keller, Pierre-Carl Langlais, Greg Lindahl, Sebastian Majstorovic, Nik Marda, Guilherme Penedo, Maarten Van Segbroeck, Jennifer Wang, Leandro vo...
2025 arXiv
-
[10]
Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction
Adrien Barbaresi. Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction. In Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natu...
2021
-
[11]
Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp
Max Bartolo, A. Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp. Beat the ai: Investigating adversarial human annotation for reading comprehension. Transactions of the Association for Computational Linguistics, 8:662–678, 2020
2020
-
[12]
Stable lm 2 1.6 b technical report
Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, et al. Stable lm 2 1.6 b technical report. arXiv preprint arXiv:2402.17834, 2024
2024 arXiv
-
[13]
Chou, Roy Frostig, and Percy Liang
Jonathan Berant, A. Chou, Roy Frostig, and Percy Liang. Semantic parsing on freebase from question-answer pairs. In Conference on Empirical Methods in Natural Language Processing, pages 1533–1544, 2013
2013
-
[14]
Elastic ChatNoir: Search Engine for the ClueWeb and the Common Crawl
Janek Bevendorff, Benno Stein, Matthias Hagen, and Martin Potthast. Elastic ChatNoir: Search Engine for the ClueWeb and the Common Crawl. In Leif Azzopardi, Allan Hanbury, Gabriella Pasi, and Benjamin Piwowarski, editors, Advances in Information Retrieval. 40th European Confer...
2018
-
[15]
FastWARC: Optimizing Large-Scale Web Archive Analytics
Janek Bevendorff, Martin Potthast, and Benno Stein. FastWARC: Optimizing Large-Scale Web Archive Analytics. In Andreas Wagner, Christian Guetl, Michael Granitzer, and Stefan V oigt, editors,3rd International Symposium on Open Search Technology (OSSYM 2021) . International Open...
2021
-
[16]
Deepseek llm: Scaling open-source language models with longtermism
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024
2024 arXiv
-
[17]
Emergent and predictable memorization in large language models
Stella Biderman, Usvsn Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raff. Emergent and predictable memorization in large language models. Advances in Neural Information Processing Systems, 36:28072–28090, 2023
2023
-
[18]
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. Pythia: A suite for analyzing large language mode...
2023
-
[19]
Piqa: Reasoning about physical commonsense in natural language, 2019
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. Piqa: Reasoning about physical commonsense in natural language, 2019. URL https://arxiv.org/abs/ 1911.11641
2019 arXiv
-
[20]
License List (version 15), 2025
Blue Oak Council. License List (version 15), 2025. URL https://blueoakcouncil.org/ list
2025
-
[21]
e-snli: Nat- ural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. e-snli: Nat- ural language inference with natural language explanations. In Neural Information Processing Systems, pages 9560–9572, 2018
2018
-
[22]
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models. In The Eleventh International Conference on Learning Representations, 2022
2022
-
[23]
Hartung, M
Ilias Chalkidis, Abhik Jana, D. Hartung, M. Bommarito, Ion Androutsopoulos, D. Katz, and Nikolaos Aletras. Lexglue: A benchmark dataset for legal language understanding in english. In Annual Meeting of the Association for Computational Linguistics, pages 4310–4330, 2021
2021
-
[24]
URL https://chatgptiseatingtheworld.com
Chat GPT Is Eating the World, 2024. URL https://chatgptiseatingtheworld.com. 12
2024
-
[25]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott G...
2021
-
[26]
HybridQA: A dataset of multi-hop question answering over tabular and textual data
Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. HybridQA: A dataset of multi-hop question answering over tabular and textual data. In Trevor Cohn, Yulan He, and Yang Liu, editors,Findings of the Association for Computational Linguistics: EMN...
2020 doi
-
[27]
Logic2text: High-fidelity natural language generation from logical forms
Zhiyu Chen, Wenhu Chen, Hanwen Zha, Xiyou Zhou, Yunkai Zhang, Sairam Sundaresan, and William Yang Wang. Logic2text: High-fidelity natural language generation from logical forms. ArXiv, abs/2004.14579, 2020
2004 arXiv
-
[28]
What is your data worth to gpt? llm-scale data valuation with influence functions, 2024
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, and Eric Xing. What is your data worth to gpt? llm-scale data valuation with influen...
2024 arXiv
-
[29]
Quac: Question answering in context
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. Quac: Question answering in context. In Conference on Empirical Methods in Natural Language Processing, pages 2174–2184, 2018
2018
-
[30]
Toxic comment classification challenge
cjadams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, nithum, and Will Cukierski. Toxic comment classification challenge. https://kaggle.com/competitions/ jigsaw-toxic-comment-classification-challenge , 2017. Kaggle
2017
-
[31]
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019
1905 arXiv
-
[32]
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. What does BERT look at? an analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, 2019
2019
-
[33]
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018
2018 arXiv
-
[34]
Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman
Karl Cobbe, V . Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. ArXiv, abs/2110.14168, 2021
-
[35]
CC0 1.0 Universal (CC0 1.0) Public Domain Dedication, 2025
Creative Commons. CC0 1.0 Universal (CC0 1.0) Public Domain Dedication, 2025. URL https://creativecommons.org/publicdomain/zero/1.0/
2025
-
[36]
Creative Commons Attribution 4.0 International License § 2(a)(1)(A),
Creative Commons. Creative Commons Attribution 4.0 International License § 2(a)(1)(A),
-
[37]
Creative Commons Attribution 4.0 International License § 3(a)(1)(A)(i),
Creative Commons. Creative Commons Attribution 4.0 International License § 3(a)(1)(A)(i),
-
[38]
Public Domain Mark 1.0, 2025
Creative Commons. Public Domain Mark 1.0, 2025. URL https://creativecommons. org/publicdomain/mark/1. 13
2025
-
[39]
Creative Commons Attribution-ShareAlike 4.0 International License,
Creative Commons. Creative Commons Attribution-ShareAlike 4.0 International License,
-
[40]
URL https://creativecommons.org/licenses/by/4.0/legalcode
-
[41]
Feder Cooper, Aaron Gokaslan, Amy B
A. Feder Cooper, Aaron Gokaslan, Amy B. Cyphert, Christopher De Sa, Mark A. Lemley, Daniel E. Ho, and Percy Liang. Extracting memorized pieces of (copyrighted) books from open-weight language models. arXiv preprint arXiv:2505.12546, 2025
2025 arXiv
-
[42]
A span-extraction dataset for chinese machine reading comprehension
Yiming Cui, Ting Liu, Li Xiao, Zhipeng Chen, Wentao Ma, Wanxiang Che, Shijin Wang, and Guoping Hu. A span-extraction dataset for chinese machine reading comprehension. In EMNLP-IJCNLP, pages 5882–5888, 2019
2019
-
[43]
URL https://creativecommons.org/licenses/by-sa/4.0/
-
[44]
Feder Cooper and James Grimmelmann
A. Feder Cooper and James Grimmelmann. The Files are in the Computer: Copyright, Memorization, and Generative AI. arXiv preprint arXiv:2404.12590, 2024
2024 arXiv
-
[45]
Few-nerd: A few-shot named entity recognition dataset
Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, and Zhiyuan Liu. Few-nerd: A few-shot named entity recognition dataset. ArXiv, abs/2105.07464, 2021
2021 arXiv
-
[46]
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In North American Chapter of the Association for Computational Linguistics , pages 2368–2378, 2019
2019
-
[47]
Tracking state changes in procedural text: a challenge dataset and models for process paragraph comprehen- sion
Bhavana Dalvi, Lifu Huang, Niket Tandon, Wen tau Yih, and Peter Clark. Tracking state changes in procedural text: a challenge dataset and models for process paragraph comprehen- sion. In North American Chapter of the Association for Computational Linguistics , pages 1595–1604, 2018
2018
-
[48]
Liu, Ana Marasovi´c, Noah A
Pradeep Dasigi, Nelson F. Liu, Ana Marasovi´c, Noah A. Smith, and Matt Gardner. Quoref: A reading comprehension dataset with questions requiring coreferential reasoning. In Conference on Empirical Methods in Natural Language Processing, volume abs/1908.05803, 2019
1908 arXiv
-
[49]
Where’s my head? definition, data set, and models for numeric fused-head identification and resolution
Yanai Elazar and Yoav Goldberg. Where’s my head? definition, data set, and models for numeric fused-head identification and resolution. Transactions of the Association for Compu- tational Linguistics, 7:519–535, 2019
2019
-
[50]
Measuring causal effects of data statistics on language model’sfactual’predictions
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Amir Feder, Abhilasha Ravichander, Marius Mosbach, Yonatan Belinkov, Hinrich Schütze, and Yoav Goldberg. Measuring causal effects of data statistics on language model’sfactual’predictions. arXiv preprint arXiv:2207.14251, 2022
2022 arXiv
-
[51]
Do membership inference attacks work on large language models? arXiv preprint arXiv:2402.07841, 2024
Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettle- moyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. Do membership inference attacks work on large language models? arXiv preprint arXiv:2402.07841, 2024
2024 arXiv
-
[52]
Dumitrescu, Petru Rebeja, Beáta L ˝orincz, Mihaela G ˘aman, M
S. Dumitrescu, Petru Rebeja, Beáta L ˝orincz, Mihaela G ˘aman, M. Ilie, Andrei Pruteanu, Adriana Stan, Luciana Morogan, Traian Rebedea, and Sebastian Ruder. Liro: Benchmark and leaderboard for romanian language tasks. In NeurIPS Datasets and Benchmarks, 2021
2021
-
[53]
Iirc: A dataset of incomplete information reading comprehension questions
James Ferguson, Matt Gardner, Tushar Khot, and Pradeep Dasigi. Iirc: A dataset of incomplete information reading comprehension questions. In Conference on Empirical Methods in Natural Language Processing, pages 1137–1147, 2020
2020
-
[54]
Nancy Fulda, Nathan Tibbetts, Zachary Brown, and D. Wingate. Harvesting common-sense navigational knowledge for robotics from uncurated text corpora. In Conference on Robot Learning, pages 525–534, 2017. 14
2017
-
[55]
Hwang, Maxwell Forbes, and Yejin Choi
Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. Moral stories: Situated reasoning about norms, intents, actions, and their consequences. ArXiv, abs/2012.15738, 2020
2012 arXiv
-
[56]
Datasets, documents, and repetitions: The practicalities of unequal data quality
Alex Fang, Hadi Pouransari, Matt Jordan, Alexander Toshev, Vaishaal Shankar, Ludwig Schmidt, and Tom Gunter. Datasets, documents, and repetitions: The practicalities of unequal data quality. arXiv preprint arXiv:2503.07879, 2025
2025
-
[57]
The pile: An 800gb dataset of diverse text for language modeling, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling, 2020. URL https: //arxiv.org/abs/2101.00027
2020 arXiv
-
[58]
Openllama: An open reproduction of llama, May 2023
Xinyang Geng and Hao Liu. Openllama: An open reproduction of llama, May 2023. URL https://github.com/openlm-research/open_llama
2023
-
[59]
A new algorithm for data compression
Philip Gage. A new algorithm for data compression. The C Users Journal archive, 12:23–38,
-
[60]
Breaking nli systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. Breaking nli systems with sentences that require simple lexical inferences. ArXiv, abs/1805.02266, 2018
2018 arXiv
-
[61]
N. Gale, G. Heath, E. Cameron, S. Rashid, and S. Redwood. Using the framework method for the analysis of qualitative data in multi-disciplinary health research. BMC Medical Research Methodology, 13:117 – 117, 2013
2013
-
[62]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravanku- mar, Artem Korenev, A...
2024 arXiv
-
[63]
Grobid. Grobid. https://github.com/kermitt2/grobid, 2008–2025
2008
-
[64]
Roth, and Jonathan Berant
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, D. Roth, and Jonathan Berant. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361, 2021
2021
-
[65]
Olmes: A standard for language model evaluations, 2025
Yuling Gu, Oyvind Tafjord, Bailey Kuehl, Dany Haddad, Jesse Dodge, and Hannaneh Hajishirzi. Olmes: A standard for language model evaluations, 2025. URL https: //arxiv.org/abs/2406.08446
2025 arXiv
-
[66]
Feder Cooper, Jasmine Collins, Landan Seguin, Austin Jacobson, Mihir Patel, Jonathan Frankle, Cory Stephenson, and V olodymyr Kuleshov
Aaron Gokaslan, A. Feder Cooper, Jasmine Collins, Landan Seguin, Austin Jacobson, Mihir Patel, Jonathan Frankle, Cory Stephenson, and V olodymyr Kuleshov. CommonCanvas: Open Diffusion Models Trained on Creative-Commons Images. In Proceedings of the IEEE/CVF Conference on Compu...
2024
-
[67]
Disfl- qa: A benchmark dataset for understanding disfluencies in question answering
Aditya Gupta, Jiacheng Xu, Shyam Upadhyay, Diyi Yang, and Manaal Faruqui. Disfl- qa: A benchmark dataset for understanding disfluencies in question answering. ArXiv, abs/2106.04016, 2021
2021 arXiv
-
[68]
C4Corpus: Multilingual web-size corpus with free license
Ivan Habernal, Omnia Zayed, and Iryna Gurevych. C4Corpus: Multilingual web-size corpus with free license. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, and Stelios...
-
[69]
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkin- son, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Ja...
2024 arXiv
-
[70]
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020
2009 arXiv
-
[71]
Wiki-40b: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrande ˇci´c, and Rami Al-Rfou. Wiki-40b: Multilingual language model dataset. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 2440–2452, 2020
2020
-
[72]
Universal language model fine-tuning for text classi- fication
Jeremy Howard and Sebastian Ruder. Universal language model fine-tuning for text classi- fication. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 2018
2018
-
[73]
Minicpm: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, et al. Minicpm: Unveiling the potential of small language models with scalable training strategies. arXiv preprint arXiv:2404.06395, 2024
2024 arXiv
-
[74]
URL https://huggingface.co/datasets/ PleIAs/common_corpus
HuggingFace: Common Corpus, 2025. URL https://huggingface.co/datasets/ PleIAs/common_corpus
2025
-
[75]
AI Training and Copyright Infringement: Solutions from Asia, Octo- ber 2024
Seth Hays. AI Training and Copyright Infringement: Solutions from Asia, Octo- ber 2024. URL https://www.techpolicy.press/ai-training-and-copyright- infringement-solutions-from-asia/
2024
-
[76]
Simple and scalable strategies to continually pre-train large language models
Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish. Simple and scalable strategies to continually pre-train large language models. arXiv preprint arXiv:2403.08763, 2024
2024 arXiv
-
[77]
Cuad: An expert-annotated nlp dataset for legal contract review
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. Cuad: An expert-annotated nlp dataset for legal contract review. ArXiv, abs/2103.06268, 2021
2021 arXiv
-
[78]
The kl3m data project: Copyright-clean training resources for large language models, 2025
Michael J Bommarito II, Jillian Bommarito, and Daniel Martin Katz. The kl3m data project: Copyright-clean training resources for large language models, 2025. URL https://arxiv. org/abs/2504.07854
2025 arXiv
-
[79]
Model AI Governance Framework for Generative AI: Fostering a Trusted Ecosys- tem, May 2024
Infocomm Media Development Authority of Singapore (IMDA), Aicadium, and AI Verify Foundation. Model AI Governance Framework for Generative AI: Fostering a Trusted Ecosys- tem, May 2024. URL https://aiverifyfoundation.sg/wp-content/uploads/2024/ 05/Model-AI-Governance-Framework...
2024
-
[80]
Analysing and reclassifying open access informa- tion in OpenAlex, 2023
Najko Jahn, Nick Haupka, and Anne Hobert. Analysing and reclassifying open access informa- tion in OpenAlex, 2023. URL https://subugoe.github.io/scholcomm_analytics/ posts/oalex_oa_status/?utm_source=chatgpt.com
2023
-
[81]
Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi
Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi. Comet-atomic 2020: On symbolic and neural commonsense knowledge graphs. In AAAI Conference on Artificial Intelligence, pages 6384–6392, 2020. 17
2020
-
[82]
Position: The most expensive part of an llm should be its training data
Nikhil Kandpal and Colin Raffel. Position: The most expensive part of an llm should be its training data. arXiv preprint arXiv:2504.12427, 2025
2025 arXiv
-
[83]
Google patents public data
IFI CLAIMS Patent Services and Google. Google patents public data. https://patents. google.com/, 2023. Licensed under a Creative Commons Attribution 4.0 International License
2023
-
[84]
Large language models struggle to learn long-tail knowledge
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning, pages 15696–15707. PMLR, 2023
2023
-
[85]
When choosing plausible alternatives, clever hans can be clever
Pride Kavumba, Naoya Inoue, Benjamin Heinzerling, Keshav Singh, Paul Reisert, and Kentaro Inui. When choosing plausible alternatives, clever hans can be clever. ArXiv, abs/1911.00225, 2019
1911 arXiv
-
[86]
Scitail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark. Scitail: A textual entailment dataset from science question answering. In AAAI Conference on Artificial Intelligence, pages 5189–5197, 2018
2018
-
[87]
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 427–431. Associatio...
2017
-
[88]
Andreas Kopf, Yannic Kilcher, Dimitri von Rutte, Sotiris Anagnostidis, Zhi Rui Tam, K. Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Rich’ard Nagyfi, ES Shahul, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and A. ...
2023 arXiv
-
[89]
Deduplicating training data mitigates privacy risks in language models, 2022
Nikhil Kandpal, Eric Wallace, and Colin Raffel. Deduplicating training data mitigates privacy risks in language models, 2022. URL https://arxiv.org/abs/2202.06539
2022 arXiv
-
[90]
Faisal Ladhak, Esin Durmus, Claire Cardie, and K. McKeown. Wikilingua: A new benchmark dataset for multilingual abstractive summarization. ArXiv, abs/2010.03093, 2020
2010 arXiv
-
[91]
Releasing Common Corpus: the largest public domain dataset for training LLMs, 2024
Pierre-Carl Langlais. Releasing Common Corpus: the largest public domain dataset for training LLMs, 2024. URL https://huggingface.co/blog/Pclanglais/common-corpus
2024
-
[92]
AI White Paper 2024: New Strate- gies in Stage II, Toward the world’s most AI-friendly country, April 2024
LDP Headquarters for the Promotion of Digital Society and Project Team on the Evolution and Implementation of AIs. AI White Paper 2024: New Strate- gies in Stage II, Toward the world’s most AI-friendly country, April 2024. URL https://aiverifyfoundation.sg/wp-content/uploads/2...
2024
-
[93]
Graham, F.Q
Rodney Michael Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, Miles Crawford, Doug Downey, Jason Dunkelberger, Oren Etzioni, Rob Evans, Sergey Feldman, Joseph Go...
2023 arXiv
-
[94]
Deduplicating training data makes language models better, 2022
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better, 2022. URL https://arxiv.org/abs/2107.06499
2022 arXiv
-
[95]
Kwiatkowski, J
T. Kwiatkowski, J. Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, D. Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc V . Le, and Slav Petrov....
2019
-
[96]
Feder Cooper, and James Grimmelmann
Katherine Lee, A. Feder Cooper, and James Grimmelmann. Talkin’ ’Bout AI Generation: Copyright and the Generative-AI Supply Chain. arXiv preprint arXiv:2309.08133, 2023
2023 arXiv
-
[97]
Kontokostas, Pablo N
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, D. Kontokostas, Pablo N. Mendes, Sebastian Hellmann, M. Morsey, Patrick van Kleef, S. Auer, and Christian Bizer. Dbpedia - a large-scale, multilingual knowledge base extracted from wikipedia. Semantic Web, 6:167–195, 2015
2015
-
[98]
Levesque, E
H. Levesque, E. Davis, and L. Morgenstern. The winograd schema challenge. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning, 2011
2011
-
[99]
Lebret, David Grangier, and Michael Auli
R. Lebret, David Grangier, and Michael Auli. Neural text generation from structured data with application to the biography domain. In Conference on Empirical Methods in Natural Language Processing, pages 1203–1213, 2016
2016
-
[100]
Xin Li and D. Roth. Learning question classifiers. In International Conference on Computational Linguistics, pages 1–7, 2002
2002
-
[101]
Feder Cooper, James Grimmelmann, and Daphne Ippolito
Katherine Lee, A. Feder Cooper, James Grimmelmann, and Daphne Ippolito. AI and Law: The Next Generation. SSRN, 2023. http://dx.doi.org/10.2139/ssrn.4580739
2023 doi
-
[102]
Testing the ability of language models to interpret figurative language
Emmy Liu, Chenxuan Cui, Kenneth Zheng, and Graham Neubig. Testing the ability of language models to interpret figurative language. ArXiv, abs/2204.12632, 2022. 19
2022 arXiv
-
[103]
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, Richard Fan, Yi Gu, Victor Miller, Yonghao Zhuang, Guowei He, Haonan Li, Fajri Koto, Liping Tang, Nikhil Ranjan, Zhiqiang Shen, Xuguang Ren, Rob...
2023 arXiv
-
[104]
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. S2ORC: The semantic scholar open research corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 4969–4983, Online, July 2020. Association for Computational...
2020 doi
-
[105]
Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan ...
-
[106]
URL https://arxiv.org/abs/2406.11794
-
[107]
Consent in crisis: The rapid decline of the AI data commons
Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderinwale, William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh, Manuel Cherep, Ahmad Anis, An Dinh, Caroline Chitongo, Da ...
2024
-
[108]
Lin, Jacob Hilton, and Owain Evans
Stephanie C. Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. In Annual Meeting of the Association for Computational Linguistics, pages 3214–3252, 2021
2021
-
[109]
Bridging the data provenance gap across text, speech and video
Shayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary, Joanna Materzynska, William Brannon, Robert Mahari, Naana Obeng-Marnu, Manan Dey, Mohammed Hamdy, et al. Bridging the data provenance gap across text, speech and video. arXiv preprint arXiv:2412.17847, 2024
2024 arXiv
-
[110]
A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, et al. A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity. In Proceedings of th...
2024
-
[111]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations , 2019. URL https: //openreview.net/forum?id=Bkg6RiCqY7
2019
-
[112]
The responsible foundation model development cheatsheet: A review of tools & resources
Shayne Longpre, Stella Biderman, Alon Albalak, Hailey Schoelkopf, Daniel McDuff, Sayash Kapoor, Kevin Klyman, Kyle Lo, Gabriel Ilharco, Nay San, et al. The responsible foundation model development cheatsheet: A review of tools & resources. Transactions on Machine Learning Rese...
2024
-
[113]
A large-scale audit of dataset licensing and attribution in AI
Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, Xinyi (Alexis) Wu, Enrico Shippole, Kurt Bollacker, Tongshuang Wu, Luis Villa, Sandy Pentland, and Sara Hooker. A la...
2024
-
[114]
Discit ergo est: Training data provenance and fair use
Robert Mahari and Shayne Longpre. Discit ergo est: Training data provenance and fair use. Robert Mahari and Shayne Longpre, Discit ergo est: Training Data Provenance And Fair Use, Dynamics of Generative AI (ed. Thibault Schrepel & Volker Stocker), Network Law Review, Winter, 2023
2023
-
[115]
Data authenticity, consent, & provenance for ai are all broken: what will it take to fix them? arXiv preprint arXiv:2404.12691, 2024
Shayne Longpre, Robert Mahari, Naana Obeng-Marnu, William Brannon, Tobin South, Katy Gero, Sandy Pentland, and Jad Kabbara. Data authenticity, consent, & provenance for ai are all broken: what will it take to fix them? arXiv preprint arXiv:2404.12691, 2024
2024 arXiv
-
[116]
Stephen Merity, Caiming Xiong, James Bradbury, and R. Socher. Pointer sentinel mixture models. ArXiv, abs/1609.07843, 2016
2016 arXiv
-
[117]
Gray, The Google Books Team, Joseph P
Jean-Baptiste Michel, Yuan Kui Shen, Aviva Presser Aiden, Adrian Veres, Matthew K. Gray, The Google Books Team, Joseph P. Pickett, Dale Hoiberg, Dan Clancy, Peter Norvig, Jon Orwant, Steven Pinker, Martin A. Nowak, and Erez Lieberman Aiden. Quantitative analysis of culture usi...
2011 doi
-
[118]
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018
2018 arXiv
-
[119]
i’d rather just go to bed
Annie Louis, D. Roth, and Filip Radlinski. “i’d rather just go to bed”’: Understanding indirect answers. In Conference on Empirical Methods in Natural Language Processing , volume abs/2010.03450, 2020
2010 arXiv
-
[120]
Starcoder 2 and the stack v2: The next generation, 2024
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Z...
2024
-
[121]
Scaling data-constrained language models
Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Alek- sandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A Raffel. Scaling data-constrained language models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advance...
2023
-
[122]
Large text compression benchmark, 2011
Matt Mahoney. Large text compression benchmark, 2011
2011
-
[123]
The e2e dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondrej Dusek, and Verena Rieser. The e2e dataset: New challenges for end-to-end generation. ArXiv, abs/1706.09254, 2017
2017 arXiv
-
[124]
Ananiadou
Tomoko Ohta, Sampo Pyysalo, Junichi Tsujii, and S. Ananiadou. Open-domain anatomical entity mention detection. In Annual Meeting of the Association for Computational Linguistics, pages 27–36, 2012
2012
-
[125]
Zhang, Eunsol Choi, and Greg Durrett
Yasumasa Onoe, Michael J.Q. Zhang, Eunsol Choi, and Greg Durrett. Creak: A dataset for commonsense reasoning over entity knowledge. ArXiv, abs/2109.01653, 2021
2021 arXiv
-
[126]
Smith, and Luke Zettlemoyer
Sewon Min, Suchin Gururangan, Eric Wallace, Weijia Shi, Hannaneh Hajishirzi, Noah A. Smith, and Luke Zettlemoyer. SILO language models: Isolating legal risk in a nonparametric datastore. In The Twelfth International Conference on Learning Representations, 2024. URL https://ope...
2024
-
[127]
Scaling data-constrained language models
Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. Scaling data-constrained language models. Advances in Neural Information Processing Systems, 36, 2023
2023
-
[128]
Privacy auditing of large language models
Ashwinee Panda, Xinyu Tang, Milad Nasr, Christopher A Choquette-Choo, and Prateek Mittal. Privacy auditing of large language models. arXiv preprint arXiv:2503.06808, 2025
2025 arXiv
-
[129]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. Crows-pairs: A challenge dataset for measuring social biases in masked language models. In Conference on Empirical Methods in Natural Language Processing, pages 1953–1967, 2020
1953
-
[130]
Directive (eu) 2019/790,
European Parliament and Council of the European Union. Directive (eu) 2019/790,
2019
-
[131]
Parser for uk parliament proceedings
ParlParse. Parser for uk parliament proceedings. https://parser.theyworkforyou.com/,
-
[132]
The FineWeb datasets: Decanting the web for the finest text data at scale
Guilherme Penedo, Hynek Kydlí ˇcek, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro V on Werra, and Thomas Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale. Advances in Neural Information Processing Systems, 37, 2024
2024
-
[133]
URL https://openalex.org
OpenAlex, 2025. URL https://openalex.org. 21
2025
-
[134]
Khudanpur
Vassil Panayotov, Guoguo Chen, Daniel Povey, and S. Khudanpur. Librispeech: An asr corpus based on public domain audio books. 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5206–5210, 2015
2015
-
[135]
Dynasent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. Dynasent: A dynamic benchmark for sentiment analysis. ArXiv, abs/2012.15349, 2020
2012 arXiv
-
[136]
Trak: Attributing model behavior at scale, 2023
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry. Trak: Attributing model behavior at scale, 2023. URL https://arxiv.org/abs/2303.14186
2023 arXiv
-
[137]
Language models are unsupervised multitask learners, 2019
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners, 2019
2019
-
[138]
Robust speech recognition via large-scale weak supervision, 2022
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision, 2022. URL https://arxiv.org/abs/2212.04356
2022 arXiv
-
[139]
Balog, B
Filip Radlinski, K. Balog, B. Byrne, and K. Krishnamoorthi. Coached conversational preference elicitation: A case study in understanding movie preferences. In SIGDIAL Conferences, pages 353–360, 2019
2019
-
[140]
Accessed: 2025-05-09
2025
-
[141]
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hen- nigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendrick...
-
[142]
Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019
2019
-
[143]
Ponti, Goran Glavavs, Olga Majewska, Qianchu Liu, Ivan Vulic, and A
E. Ponti, Goran Glavavs, Olga Majewska, Qianchu Liu, Ivan Vulic, and A. Korhonen. Xcopa: A multilingual dataset for causal commonsense reasoning. In Conference on Empirical Methods in Natural Language Processing, pages 2362–2376, 2020
2020
-
[144]
Nazneen Rajani, Bryan McCann, Caiming Xiong, and R. Socher. Explain yourself! leveraging language models for commonsense reasoning. ArXiv, abs/1906.02361, 2019
1906 arXiv
-
[145]
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training, 2018
2018
-
[146]
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. In Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, 2016
2016
-
[147]
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don’t know: Unanswerable questions for squad. In Annual Meeting of the Association for Computational Linguistics , volume abs/1806.03822, 2018
2018 arXiv
-
[148]
Schema-guided dialogue state tracking task at dstc8
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. Schema-guided dialogue state tracking task at dstc8. ArXiv, abs/2002.01359, 2020
2002 arXiv
-
[149]
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. arXiv preprint, 2019. URL https://arxiv.org/abs/1911.05507
2019 arXiv
-
[150]
Schmidt, Yuval Pinter, and Chris Tanner
Varshini Reddy, Craig W. Schmidt, Yuval Pinter, and Chris Tanner. How much is enough? the diminishing returns of tokenization training data, 2025. URL https://arxiv.org/abs/2502.20273
2025 arXiv
-
[151]
Thang, N
Hammam Riza, Michael Purwoadi, Gunarso, Teduh Uliniansyah, Aw Ai Ti, Sharifah Mahani Aljunied, Luong Chi Mai, V . Thang, N. Thai, Vichet Chea, Rapid Sun, Sethserey Sam, Sopheap Seng, K. Soe, K. Nwet, M. Utiyama, and Chenchen Ding. Introduction of the asian language treebank. I...
2016
-
[152]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140), 2020
2020
-
[153]
The size of datasets used to train language models doubles approximately every seven months, 2024
Robi Rahman and David Owen. The size of datasets used to train language models doubles approximately every seven months, 2024. URL https://epoch.ai/data- insights/dataset-size-trend. Accessed: 2025-05-08
2024
-
[154]
A primer in bertology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. A primer in bertology: What we know about how BERT works. Transactions of the association for computational linguistics, 8, 2021
2021
-
[155]
Outsider oversight: Designing a third party audit ecosystem for ai governance
Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, and Daniel Ho. Outsider oversight: Designing a third party audit ecosystem for ai governance. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, pages 557–571, 2022
2022
-
[156]
Matthew Sag and Peter K. Yu. The globalization of copyright exceptions for ai train- ing. Emory Law Journal , 74, 2025. doi: http://dx.doi.org/10.2139/ssrn.4976393. URL https://ssrn.com/abstract=4976393
2025 doi
-
[157]
Conjnli: Natural language inference over conjunctive sentences
Swarnadeep Saha, Yixin Nie, and Mohit Bansal. Conjnli: Natural language inference over conjunctive sentences. In Conference on Empirical Methods in Natural Language Processing, pages 8240–8252, 2020
2020
-
[158]
Winogrande
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande. Communications of the ACM, 64:99 – 106, 2019
2019
-
[159]
Condaqa: A contrastive reading comprehension dataset for reasoning about negation
Abhilasha Ravichander, Matt Gardner, and Ana Marasovi´c. Condaqa: A contrastive reading comprehension dataset for reasoning about negation. ArXiv, abs/2211.00295, 2022
2022 arXiv
-
[160]
Socialiqa: Commonsense reasoning about social interactions, 2019
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. Socialiqa: Commonsense reasoning about social interactions, 2019. URL https://arxiv.org/abs/1904.09728
2019 arXiv
-
[161]
Sboev, A
A. Sboev, A. Naumov, and R. Rybka. Data-driven model for emotion detection in russian texts. In BICAAI, pages 637–642, 2020
2020
-
[162]
How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020
Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020
2020
-
[163]
Getting closer to ai complete question answering: A set of prerequisite real tasks
Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. Getting closer to ai complete question answering: A set of prerequisite real tasks. In AAAI Conference on Artificial Intelligence, pages 8722–8731, 2020
2020
-
[164]
Hausknecht
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew J. Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning. ArXiv, abs/2010.03768, 2020
2010 arXiv
-
[165]
Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A
Rachel Rudinger, Vered Shwartz, Jena D. Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A. Smith, and Yejin Choi. Thinking like a skeptic: Defeasible inference in natural language. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the As- sociation f...
2020
-
[166]
peS2o (Pretraining Efficiently on S2ORC) Dataset
Luca Soldaini and Kyle Lo. peS2o (Pretraining Efficiently on S2ORC) Dataset. Technical report, Allen Institute for AI, 2023. ODC-By, https://github.com/allenai/pes2o
2023
-
[167]
Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Mu...
2024
-
[168]
Smith, and Luke Zettlemoyer
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. Evaluating gender bias in machine translation. ArXiv, abs/1906.00591, 2019
1906 arXiv
-
[169]
Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures. In 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7). Leibniz-Institut für Deutsche Sprache, 2019. 24
2019
-
[170]
How to Think About Remedies in the Generative AI Copyright Cases
Pamela Samuelson. How to Think About Remedies in the Generative AI Copyright Cases. Lawfare, February 2024. URL https://www.lawfaremedia.org/article/how-to- think-about-remedies-in-the-generative-ai-copyright-cases
2024
-
[171]
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. arXiv preprint arXiv:1811.00937, 2018
2018 arXiv
-
[172]
Niket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson, and E. Hovy. A dataset for tracking entities in open domain procedural text. ArXiv, abs/2011.08092, 2020
2011 arXiv
-
[173]
Barzilay
Tal Schuster, Adam Fisch, and R. Barzilay. Get your vitamin c! robust fact verification with contrastive evidence. In North American Chapter of the Association for Computational Linguistics, pages 624–643, 2021
2021
-
[174]
Emily Sheng and David C. Uthus. Investigating societal biases in a poetry composition system. ArXiv, abs/2011.02686, 2020
2011 arXiv
-
[175]
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML NLP Team. Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023. URL www.mosaicml.com/blog/mpt-7b. Accessed: 2023-05-05
2023
-
[176]
Shivalika Singh, Freddie Vargus, Daniel Dsouza, B"orje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemi’nski, Hakimeh...
2024 arXiv
-
[177]
Fever: a large-scale dataset for fact extraction and verification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. Fever: a large-scale dataset for fact extraction and verification. ArXiv, abs/1803.05355, 2018
2018 arXiv
-
[178]
Mad- dison
Anvith Thudi, Evianne Rovers, Yangjun Ruan, Tristan Thrush, and Chris J. Mad- dison. Mixmin: Finding data mixtures via convex minimization, 2025. URL https://arxiv.org/abs/2502.10510
2025
-
[179]
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language...
2023 arXiv
-
[180]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[181]
Quartz: An open-domain dataset of qualitative relationship questions
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. Quartz: An open-domain dataset of qualitative relationship questions. In Conference on Empirical Methods in Natural Language Processing, volume abs/1909.03553, 2019
1909 arXiv
-
[182]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017
2017
-
[183]
Meta Lingua: A minimal PyTorch LLM training library, 2024
Mathurin Videau, Badr Youbi Idrissi, Daniel Haziza, Luca Wehrstedt, Jade Copet, Olivier Teytaud, and David Lopez-Paz. Meta Lingua: A minimal PyTorch LLM training library, 2024. URL https://github.com/facebookresearch/lingua
2024
-
[184]
Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, and Eric P. Xing. Txt360: A top-quality llm pre-training dataset requires t...
2024
-
[185]
Choudhury
Ishan Tarunesh, Somak Aditya, and M. Choudhury. Trusting roberta over bert: Insights from checklisting the natural language inference task. ArXiv, abs/2107.07229, 2021
2021 arXiv
-
[186]
Resolving gendered ambiguous pronouns with bert
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. Resolving gendered ambiguous pronouns with bert. ArXiv, abs/1906.01161, 2019
1906 arXiv
-
[187]
Qwen3, April 2025
Qwen Team. Qwen3, April 2025. URL https://qwenlm.github.io/blog/qwen3/
2025
-
[188]
Organize the web: Constructing domains enhances pre-training data curation
Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen, and Luca Soldaini. Organize the web: Constructing domains enhances pre-training data curation. arXiv preprint arXiv:2502.10341, 2025
2025 arXiv
-
[189]
Doremi: Optimizing data mixtures speeds up language model pretraining
Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy S Liang, Quoc V Le, Tengyu Ma, and Adams Wei Yu. Doremi: Optimizing data mixtures speeds up language model pretraining. Advances in Neural Information Processing Systems, 36, 2023
2023
-
[190]
Cohen, R
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, R. Salakhutdinov, and Christopher D. Manning. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing , pages 2369–2380, 2018
2018
-
[191]
OpenAI prepares to fight for its life as legal troubles mount
Cat Zakrzewski, Nitasha Tiku, and Elizabeth Dwoskin. OpenAI prepares to fight for its life as legal troubles mount. The Washington Post, 2024
2024
-
[192]
Open parliament license
UK Parliament. Open parliament license. https://www.parliament.uk/site- information/copyright-parliament/open-parliament-licence/ , Unknown. Accessed: 2025-05-09
2025
-
[193]
MAP-Neo: Highly capable and transparent bilingual large language model series
Ge Zhang, Scott Qu, Jiaheng Liu, Chenchen Zhang, Chenghua Lin, Chou Leuang Yu, Danny Pan, Esther Cheng, Jie Liu, Qunshu Lin, Raven Yuan, Tuney Zheng, Wei Pang, Xinrun Du, Yiming Liang, Yinghao Ma, Yizhi Li, Ziyang Ma, Bill Lin, Emmanouil Benetos, Huan Yang, Junting Zhou, Kaiji...
2024 arXiv
-
[194]
Winowhy: A deep diagnosis of essential commonsense knowledge for answering winograd schema challenge
Hongming Zhang, Xinran Zhao, and Yangqiu Song. Winowhy: A deep diagnosis of essential commonsense knowledge for answering winograd schema challenge. In Annual Meeting of the Association for Computational Linguistics, pages 5736–5745, 2020
2020
-
[195]
Helpsteer: Multi-attribute helpfulness dataset for steerlm
Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Polak Scowcroft, Neel Kant, Aidan Swope, and Oleksii Kuchaiev. Helpsteer: Multi-attribute helpfulness dataset for steerlm. ArXiv, abs/2311.09528, 2023
2023 arXiv
-
[196]
Redpajama: an open dataset for training large language models, 2024
Maurice Weber, Daniel Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Ré, Irina Rish, and Ce Zhang. Redpaja...
2024 arXiv
-
[198]
Le, Andrew M
Wei Wei, Quoc V . Le, Andrew M. Dai, and Jia Li. Airdialogue: An environment for goal-oriented dialogue research. In Conference on Empirical Methods in Natural Language Processing, pages 3844–3854, 2018
2018
-
[203]
Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019
1905 arXiv
-
[206]
Wildchat: 1m chatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wildchat: 1m chatGPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Bl8u7ZRlbM
2024
-
[207]
accepted answer
Ben Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, and D. Roth. Temporal reasoning on implicit events from distant supervision. ArXiv, abs/2010.12753, 2020. 26 Appendix Table of Contents A Contributions 28 B Detailed Description of Sources 28 B.1 Scientific ...
2010 arXiv
-
[1994]
URL https://api.semanticscholar.org/CorpusID:59804030
-
[2016]
URL https://aclanthology
European Language Resources Association (ELRA). URL https://aclanthology. org/L16-1146/
-
[2019]
URL https://eur-lex.europa.eu/legal-content/EN/ALL/?uri=CELEX: 32019L0790#art_3
-
[2020]
doi: 10.18653/v1/2020.findings-emnlp.418
Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.418. URL https://aclanthology.org/2020.findings-emnlp.418/. 23
2020 doi
-
[2021]
URL https://api.semanticscholar.org/CorpusID:245353475
-
[2024]
URL https://www.law.cornell.edu/uscode/text/17/105
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.