REVIEW 4 major objections 3 minor 3 cited by
LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch
T0 review · 4 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper reports that K2 Diamond, a 65-billion-parameter model trained completely from scratch on 1.4 trillion tokens, surpasses LLaMA-65B and rivals Llama2-70B on standard benchmarks, and it releases the code, data sequence, logs, and…
desk verdict Genuinely valuable 65B open-release, but the exact-reproducibility claim is undercut by the paper's own rollback description and incomplete checkpoint uploads. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the auditable training record, not any single architectural innovation. The architecture follows LLaMA-65B; the claimed novelty is that every component that usually stays private is published: the TxT360 data curation pipeline with rich per-document metadata, the exact data mix and order (47% RefinedWeb, code via StarCoder, math boosted to 40% in stage 2, repetition and truncation rules), a two-stage curriculum, a 4D-parallelism setup (8-way tensor, 4-way pipeline, 15-way data parallel on 480 A100s) with a 2,040-token global batch, and a checkpoint after each data chunk (380 saved, 140 currently shared), plus training logs and incident reports on the two malignant loss spikes. This record is what would let a reader verify the training-from-scratch claim and study training dynamics.
What would settle it
Compare the hash of each released checkpoint against the training logs and the evaluation gallery outputs, then re-run the reported evaluation on the released final checkpoint and check whether the scores in the paper's benchmark table reproduce; a mismatch, or an inability to trace the released data sequence through to the published checkpoints, would falsify the central transparency claim.
Extended reading notes
Core claim
K2 Diamond is a 65-billion-parameter dense Transformer (80 layers, hidden size 8,192, no grouped-query attention) trained from scratch on 1.4 trillion tokens in two stages: a major stage at 2,048-token context and a long-context stage at 8,192 tokens using RoPE theta scaling. The paper reports an average score of 57.20 across 21 benchmarks, against 53.77 for LLaMA-65B and 57.11 for Llama2-70B, with its largest margins on code (HumanEval pass@1 of 32.0 versus 22.8 and 30.0) and medical QA, while using about 35% fewer FLOPs than Llama2-70B. It claims this makes K2 Diamond the first fully open-source LLM at this scale, releasing code, the exact ordering of training data, training logs, evaluation galleries, and checkpoints from during training.
Load-bearing premise
The load-bearing premise is that the released code, data sequence, logs, and checkpoints are exactly the ones that produced the reported K2 Diamond scores; the paper's own footnotes say 380 checkpoints were saved but only 120 of the stage-1 checkpoints are uploaded so far, so the complete record is not yet available.
Editorial extensions
If this is right
- At 65B parameters, a fully open-source LLM can reportedly reach performance comparable to a leading open-weight recipe: the paper's 21-benchmark average is 57.20 for K2 Diamond, versus 53.77 for LLaMA-65B and 57.11 for Llama2-70B.
- A two-stage curriculum on 1.4 trillion tokens is reported as sufficient to rival a model trained on 2 trillion tokens, giving evidence about data efficiency at the 65B scale.
- Public loss-spike checkpoints and incident logs allow researchers to study why some spikes are benign and others destructive, using data rather than anecdote.
- The longitudinal evaluations across 120 checkpoints give a fine-grained record of when reasoning, coding, medical, and bias-related behaviors appear or disappear during pretraining.
- The Apache 2.0 license and released fine-tuning recipes make K2 a practical base for distillation, function calling, and domain adaptation without use restrictions.
Reading between the lines
- If the released artifacts are exact, this makes a frontier-scale 'training from scratch' claim auditable for the first time, and it lets resource-constrained groups study scaling behavior without paying the full compute cost.
- The paper's own checkpoint footnote suggests a practical tension in full-transparency releases: at more than 100GB per checkpoint, storage costs cap what can actually be shared, so complete openness at this scale may require new formats, incremental weight formats, or long-term archival commitments.
- A natural testable extension is to use K2's released data sequence to pretrain a smaller model and check whether the same loss-spike and capability-acquisition patterns reproduce, which would separate model-scale effects from data-order effects.
- The reported disappearance of some early abilities (e.g., certain GSM8K questions answered correctly at intermediate checkpoints but not at the end) invites a concrete follow-up: determine whether this reflects data-order interference, evaluation noise, or genuine forgetting during continued pretraining.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports the LLM360 K2 project and its first model, K2 Diamond, a 65B-parameter transformer pretrained from scratch on roughly 1.4T tokens. The authors claim that K2 Diamond surpasses LLaMA-65B and rivals Llama2-70B on a 21-benchmark suite while using fewer FLOPs and tokens, and that the project is fully reproducible under the LLM360 360-degree open-source framework: code, data mix, exact data sequence, checkpoints, logs, and incident records are released. The paper also documents the TxT360 data curation pipeline, pretraining infrastructure and parallelism choices, loss-spike incidents and rollbacks, post-training and safety fine-tuning, and a qualitative longitudinal capability analysis. The central contribution is the transparency claim: a 65B model trained from scratch with a complete audit trail of data, code, checkpoints, and logs.
Significance. If the reproducibility claim holds, K2 Diamond would be a valuable community resource: it is the largest fully open-source 65B-scale model described to date, it re-evaluates comparison baselines with a shared harness, and it documents training incidents (loss spikes, hardware failures) and energy/carbon overhead in unusual detail. The longitudinal evaluations across 120 checkpoints and the release of evaluation galleries are genuinely useful for studying capability acquisition. These strengths are substantial and should be credited. However, the paper's own footnotes and incident descriptions reveal gaps between the advertised 'exact data sequence' + 'all checkpoints' and what is actually released, and the headline FLOPs-efficiency claim excludes the extra 30 days spent on loss-spike recovery. Because the openness/reproducibility claim is the paper's central contribution, these issues are load-bearing rather than cosmetic.
major comments (4)
- [Sec. 3 and Sec. 4.4] The central reproducibility claim is undermined by an unspecified relation between the released data sequence and the data actually consumed after the two malignant loss spikes. Section 3 promises 'the exact sequence of training data, segmented into chunks corresponding to each checkpoint,' and Section 4.4 states that the authors 'restarted training from earlier checkpoints' and 'allowed training to proceed for a few steps after each spike.' The paper never states whether the released 360 chunks correspond to one sequential pass through the intended order, whether spike-containing chunks were replayed, partially skipped, or re-ordered, or how checkpoint numbering (1..360) maps to the post-rollback runs, which are said to reside in separate repositories. Without this mapping, a third party cannot reconstruct K2 Diamond from the released artifacts, so the exact-sequence claim needs either a precise specification of the rollback/replay policy or a restatement of what the released sequence actually represents.
- [Sec. 3 vs. Sec. 4.3, footnote 4] The checkpoint availability statements are internally inconsistent. Section 3 lists 'Model Checkpoints: 140 intermediate model checkpoints, evenly distributed across stage 1 (120 checkpoints) and stage 2 (20 checkpoints),' while Section 4.3 says 360 and 20 chunks produce 'a total of 380 K2 Diamond checkpoints,' and footnote 4 says only 120 of the 380 are shared. The reader cannot tell how many checkpoints exist, how many are accessible, or whether the 140 number is a typo. Given that the paper advertises 'all intermediate model checkpoints saved during training' as part of the transparency claim, this discrepancy and the acknowledged storage-limited upload need to be resolved and clearly stated.
- [Abstract, Sec. 1, and Power Consumption section] The headline claim of 'an approximately 35% reduction in FLOPs compared to Llama2-70B' and the statement that K2 Diamond 'requires fewer FLOPs' exclude the additional 30 days of compute spent handling loss spikes, which the paper itself reports as an extra 129.3 MWh on top of the 430.8 MWh base training. Since loss-spike rollbacks involve recomputation from earlier checkpoints, the total FLOPs consumed by the project are materially higher than the useful-training FLOPs implied by 1.4T tokens. The paper should either report the efficiency claim net of this overhead or clearly separate 'training FLOPs to the final checkpoint' from 'total FLOPs expended,' and should qualify the abstract accordingly.
- [Sec. 7, Table 15] The benchmark comparison that supports the 'surpasses LLaMA-65B and rivals Llama2-70B' claim is reported without error bars, confidence intervals, or significance tests. For example, the overall average scores are 57.20 (K2 Diamond) versus 57.11 (Llama2-70B), a difference of 0.09 points, and individual benchmarks show both favorable and unfavorable gaps. Because the evaluation harness is deterministic given the exact prompts and decoding settings, the authors could provide multi-seed or resampling-based variance estimates, or at minimum state the number of runs and the stability of the differences. As written, the central empirical comparison rests on averages that may be within run-to-run noise.
minor comments (3)
- [Table 1] The CommonCrawl cut-off date is listed as '2024-30,' which is not a valid month-day value; this is likely a typo and should be corrected.
- [Sec. 4.4 and Abstract] The paper says the model is trained on '1.4 trillion tokens' in the abstract, but Section 4.4 describes a major stage of 1.4T tokens plus a long-context stage of 69.3B tokens. Clarify whether the 1.4T figure already includes the long-context stage or whether the total is approximately 1.47T.
- [Sec. 8.3 and Table 18] The 'emergent and disappearing abilities' analysis uses a 90%-correct cutoff in Bucket 6 and averages over only 20 checkpoints per bucket; the authors do acknowledge that these are preliminary observations, but the presentation would benefit from stating the number of questions examined and the instability of the per-question estimates, since many questions show near-zero frequency in most buckets.
Circularity Check
Safety evaluation reuses safety-SFT training benchmarks (DoNotAnswer, AdvBench, MITRE) without documented splits; headline capability claims remain externally grounded.
-
fitted input called prediction
[Risks and Mitigation (after Section 10); Section 6.2.2 and Table 8]
"This process incorporates datasets such as DoNotAnswer, AdvBench, and MITRE, alongside custom-built and region-specific prompts to address nuanced risks, including cybersecurity and culturally sensitive content. To evaluate safety alignment, we rigorously tested the model before and after fine-tuning using diverse benchmarks."
The paper names MITRE as a safety-SFT training source, then reports the model's MITRE score rising from 3.20% to 57.30% as evidence of improved safety alignment. No train/eval split is documented for MITRE, and the paper elsewhere shows it knows how to make such an exclusion: 'AttackSerious ... but not included in the SFT data.' Without a stated held-out split, the MITRE gain is at least partly memorization of the test set, so this evaluation result reduces to the training input by construction.
-
fitted input called prediction
[Section 5.1 (Building the K2 Chat Baseline); Table 5; Section 6.2.1 and Table 8]
"These included 2,700 samples from the Do-Not-Answer dataset (Wang et al., 2024) and a small set of UAE culture-related prompts to ensure region-specific alignment."
DoNotAnswer appears both in the SFT construction (2,700 samples, with Table 5 listing 1,839 'Do_Not_Answer_for_FT' entries) and as an evaluation benchmark (939 prompts, Section 6.2.1). The paper never states that the evaluation prompts were withheld from the safety-SFT data. The reported improvement from 67.94% to 87.65% on DoNotAnswer is therefore partly fitted: the model was fine-tuned on the same benchmark it is then scored on.
1 more flagged steps
-
fitted input called prediction
[Section 6.1.2 (Few-shots Attack); Section 6.2.1 and Table 8]
"we add 3 randomly sampled harmful question-answer pairs from AdvBench (Zou et al., 2023) and add them before the direct attack prompts."
AdvBench harmful behaviors are used to build few-shot demonstrations in the SFT data, and the same AdvBench dataset (500 harmful strings) is used as a safety evaluation benchmark. With no disclosed split between the AdvBench items used in SFT and those used in evaluation, the reported gain from 52.12% to 81.73% is partly a training-set-memorization artifact rather than an independent measure of robustness.
full rationale
The paper's headline capability comparisons (K2 Diamond vs LLaMA-65B / Llama2-70B, Table 15) are external-benchmark results produced with the same evaluation code for all models; no parameter is fitted to those benchmarks, and the data mix/hyperparameters are reported rather than reverse-engineered from the scores. The central 'trained from scratch' claim is an empirical account, not a derivation, so it is not circular in the sense of this review. The circularity found is localized to the safety fine-tuning section: several benchmarks used for safety evaluation (DoNotAnswer, AdvBench, MITRE) are also named as safety-SFT training data, and the paper does not document any held-out split for those datasets. Consequently the reported safety-score improvements are partly by construction. The loss-spike rollback ambiguity and the partial checkpoint uploads (footnote 4/10) are real reproducibility/transparency limitations—the released data sequence may not exactly match the consumed sequence after rollbacks—but they are consistency gaps, not circular reasoning. Self-citations to LLM360, TxT360, and Crystal are present but not load-bearing for the external benchmark numbers. Overall circularity is partial and confined to an auxiliary claim.
Assumptions & free parameters
free parameters (5)
- data_mix_proportions =
47% RefinedWeb, 1% math in stage 1, 40% math in stage 2, 10-30% code, various repetition counts
- RoPE_theta =
10,000 in stage 1; 500,000 in stage 2
- learning_rate_schedule =
1.5e-4 to 1.5e-5 cosine in stage 1; 1e-4 to 0 linear in stage 2
- global_batch_size =
2040 sequences (4M tokens)
- loss_spike_rollback_policy =
restart from earlier checkpoints during two malignant spikes
assumptions (4)
- domain assumption The LLaMA-65B architecture is a sound base for a 65B model.
- domain assumption The public datasets listed are sufficient to train a competitive 65B model.
- domain assumption The lm-evaluation-harness settings give comparable scores across K2 and the baseline models.
- standard math AdamW with the stated hyperparameters and gradient clipping converges as expected.
Cite this review
Pith. "Pith review of LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch." pith.science (2026). https://pith.science/paper/SV4BEA7K
@misc{pith2026250107124,
author = {Pith},
title = {Pith review of: LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch},
year = {2026},
howpublished = {\url{https://pith.science/paper/SV4BEA7K}},
note = {Machine review of arXiv:2501.07124}
}
read the original abstract
We detail the training of the LLM360 K2-65B model, scaling up our 360-degree OPEN SOURCE approach to the largest and most powerful models under project LLM360. While open-source LLMs continue to advance, the answer to "How are the largest LLMs trained?" remains unclear within the community. The implementation details for such high-capacity models are often protected due to business considerations associated with their high cost. This lack of transparency prevents LLM researchers from leveraging valuable insights from prior experience, e.g., "What are the best practices for addressing loss spikes?" The LLM360 K2 project addresses this gap by providing full transparency and access to resources accumulated during the training of LLMs at the largest scale. This report highlights key elements of the K2 project, including our first model, K2 DIAMOND, a 65 billion-parameter LLM that surpasses LLaMA-65B and rivals LLaMA2-70B, while requiring fewer FLOPs and tokens. We detail the implementation steps and present a longitudinal analysis of K2 DIAMOND's capabilities throughout its training process. We also outline ongoing projects such as TXT360, setting the stage for future models in the series. By offering previously unavailable resources, the K2 project also resonates with the 360-degree OPEN SOURCE principles of transparency, reproducibility, and accessibility, which we believe are vital in the era of resource-intensive AI research.
Figures
Figures from the paper (17 more)
Forward citations
Cited by 3 Pith papers
-
Fluid Language Model Benchmarking
Fluid Benchmarking, combining IRT-based ability estimation with Fisher-information-based adaptive item selection, improves LM evaluation across efficiency, validity, variance, and saturation in pretraining settings.
-
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
LibriSpeech and Common Voice evaluation sentences leak into the Pile, and controlled LLM pretraining experiments show that contamination biases output probabilities even when error rates barely change.
-
CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion
CollabEval turns model evaluation into low-rank matrix completion and uses the imputations as control variates to cut CI width and MSE at fixed annotation budget while preserving unbiasedness.
Reference graph
Works this paper leans on
-
[1]
01-ai/yi: A series of large language models trained from scratch by developers @01-ai, 2023
01.ai. 01-ai/yi: A series of large language models trained from scratch by developers @01-ai, 2023. URL https://github.com/01-ai/Yi
2023
-
[2]
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024
arXiv 2024
-
[3]
Nemotron-4 340b technical report
Bo Adler, Niket Agarwal, Ashwath Aithal, Dong H Anh, Pallab Bhattacharya, Annika Brundyn, Jared Casper, Bryan Catanzaro, Sharon Clay, Jonathan Cohen, et al. Nemotron-4 340b technical report. arXiv preprint arXiv:2406.11704, 2024
arXiv 2024
-
[4]
Santacoder: don't reach for the stars! arXiv preprint arXiv:2301.03988, 2023
Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, et al. Santacoder: don't reach for the stars! arXiv preprint arXiv:2301.03988, 2023
arXiv 2023
-
[5]
Smollm - blazingly fast and remarkably powerful
Loubna Ben Allal, Anton Lozhkov, and Elie Bakouch. Smollm - blazingly fast and remarkably powerful. https://huggingface.co/blog/smollm, 2024. Accessed: 2024-09-12
2024
-
[6]
The falcon series of open language models, 2023
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. The falcon series of open language models, 2023
2023
-
[7]
Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019
2019
-
[8]
GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch , 9 2023
Alex Andonian, Quentin Anthony, Stella Biderman, Sid Black, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Michael Pieler, Jason Phang, Shivanshu Purohit, Hailey Schoelkopf, Dashiell Stander, Tri Songz, Curt Tigges, Benjamin Thérien, Phil Wang, and Samuel Weinbach. GPT-NeoX: Large Scale Autoregressive Lan...
2023
Show all 165 references
-
[9]
Palm 2 technical report
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023
2023 arXiv
-
[10]
Model card for claude 3
Anthropic. Model card for claude 3. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf, 2024. Accessed: 2024-09-01
2024
-
[11]
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021
2021 arXiv
-
[12]
Jiang, Jia Deng, Stella Biderman, and Sean Welleck
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics, 2023
2023
-
[13]
Qwen technical report
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023
2023 arXiv
-
[14]
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan H...
2022 arXiv
-
[15]
Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction
Adrien Barbaresi. Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction . In Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Nat...
2021
-
[16]
Efficient training of language models to fill in the middle
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. Efficient training of language models to fill in the middle. arXiv preprint arXiv:2207.14255, 2022
2022 arXiv
-
[17]
Infinity instruct
Beijing Academy of Artificial Intelligence (BAAI) . Infinity instruct. arXiv preprint arXiv:2406.XXXX, 2024
2024
-
[18]
Stable lm 2 1.6 b technical report
Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, et al. Stable lm 2 1.6 b technical report. arXiv preprint arXiv:2402.17834, 2024
2024 arXiv
-
[19]
A framework for the evaluation of code generation models
Loubna Ben Allal, Niklas Muennighoff, Logesh Kumar Umapathi, Ben Lipkin, and Leandro von Werra. A framework for the evaluation of code generation models. https://github.com/bigcode-project/bigcode-evaluation-harness, 2022
2022
-
[20]
Red-teaming large language models using chain of utterances for safety-alignment, 2023
Rishabh Bhardwaj and Soujanya Poria. Red-teaming large language models using chain of utterances for safety-alignment, 2023. URL https://arxiv.org/abs/2308.09662
2023 arXiv
-
[21]
Purple llama cyberseceval: A secure coding benchmark for language models, 2023
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, Sasha Frolov, Ravi Prakash Giri, Dhaval Kapil, Yiannis Kozyrakis, David LeBlanc, James Milazzo, Aleksandar Straumann...
2023 arXiv
-
[23]
Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions, 2024
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions, 2024. URL https://arxiv.org/abs/2309.07875
2024 arXiv
-
[24]
Emergent and predictable memorization in large language models
Stella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raf. Emergent and predictable memorization in large language models. arXiv preprint arXiv:2304.11158, 2023 a
2023 arXiv
-
[25]
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In Intern...
2023
-
[26]
Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal...
2024
-
[27]
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. Piqa: Reasoning about physical commonsense in natural language. In Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020
2020
-
[28]
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow , March 2021
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow , March 2021. URL https://doi.org/10.5281/zenodo.5297715. If you use this software, please cite it using these metadata
2021 doi
-
[29]
Gpt-neox-20b: An open-source autoregressive language model, 2022
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. Gpt-neox-20b: An ope...
2022
-
[30]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...
1901
-
[31]
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021
2021 arXiv
-
[32]
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks, 2023 a . URL https://arxiv.org/abs/2211.12588
2023 arXiv
-
[33]
Berkay Celik
Yufan Chen, Arjun Arunasalam, and Z. Berkay Celik. Can large language models provide security & privacy advice? measuring the ability of llms to refute misconceptions. In Proceedings of the 39th Annual Computer Security Applications Conference, ACSAC '23, pp.\ 366–378, New Yor...
2023
-
[34]
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24 0 (240): 0 1--113, 2023
2023
-
[35]
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1, 2018
2018 arXiv
-
[36]
Claude 2.1 model card
Claude. Claude 2.1 model card. Technical report, Claude Inc., 2023. URL https://claude.ai/model-card/claude-2-1
2023
-
[37]
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021
-
[38]
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models
Damai Dai, Chengqi Deng, Chenggang Zhao, RX Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y Wu, et al. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. arXiv preprint arXiv:2401.06066, 2024
2024 arXiv
-
[39]
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691, 2023
2023 arXiv
-
[40]
Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality
Tri Dao and Albert Gu. Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality. In International Conference on Machine Learning (ICML), 2024
2024
-
[41]
Mpt-7b: How we scaled to train the world's largest open-source model, 2024
Databricks . Mpt-7b: How we scaled to train the world's largest open-source model, 2024. URL https://www.databricks.com/blog/mpt-7b. Accessed: 2024-09-09
2024
-
[42]
Deepseek llm: Scaling open-source language models with longtermism
DeepSeek-AI. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024. URL https://github.com/deepseek-ai/DeepSeek-LLM
2024 arXiv
-
[43]
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
Nolan Dey, Gurpreet Gosal, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, Joel Hestness, et al. Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster. arXiv preprint arXiv:2304.03208, 2023
2023 arXiv
-
[44]
Build it break it fix it for dialogue safety: Robustness from adversarial human attack, 2019
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. Build it break it fix it for dialogue safety: Robustness from adversarial human attack, 2019. URL https://arxiv.org/abs/1908.06083
2019 arXiv
-
[45]
Measuring the carbon intensity of ai in cloud instances
Jesse Dodge, Taylor Prewitt, Remi Tachet des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A Smith, Nicole DeCario, and Will Buchanan. Measuring the carbon intensity of ai in cloud instances. In Proceedings of the 2022 ACM conference on fair...
2022
-
[47]
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23 0 (120): 0 1--39, 2022
2022
-
[48]
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfi...
2022 arXiv
-
[49]
The P ile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The P ile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020
2020 arXiv
-
[50]
A framework for few-shot language model evaluation, 12 2023
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...
2023
-
[51]
Openllama: An open reproduction of llama, May 2023
Xinyang Geng and Hao Liu. Openllama: An open reproduction of llama, May 2023. URL https://github.com/openlm-research/open_llama
2023
-
[52]
Chatglm: A family of large language models from glm-130b to glm-4 all tools
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, et al. Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793, 2024
2024 arXiv
-
[53]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[54]
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. Olmo: Accelerating the science of language models. arXiv preprint arXiv:2402.00838, 2024
2024 arXiv
-
[55]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023
2023 arXiv
-
[56]
Textbooks are all you need
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio C \'e sar Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al. Textbooks are all you need. arXiv preprint arXiv:2306.11644, 2023
2023 arXiv
-
[57]
Krass*, Lucia Zheng, Neel Guha, Christopher D
Peter Henderson*, Mark S. Krass*, Lucia Zheng, Neel Guha, Christopher D. Manning, Dan Jurafsky, and Daniel E. Ho. Pile of law: Learning responsible data filtering from the law and a 256gb open-source legal dataset, 2022. URL https://arxiv.org/abs/2207.00220
2022 arXiv
-
[58]
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021
2021
-
[59]
Scaling laws and interpretability of learning from repeated data, 2022
Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, Scott Johnston, Ben Mann, Chris Olah, Catherine Olsson, Dario Amodei, Nicholas Joseph, Jared Kaplan, and Sam McCandlish. Scaling l...
2022
-
[60]
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al. Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in neural information processing systems, 32, 2019
2019
-
[61]
Mistral 7b
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023
-
[62]
Mixtral of experts
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024
2024 arXiv
-
[63]
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. arXiv preprint arXiv:2009.13081, 2020
2009 arXiv
-
[64]
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natur...
2019
-
[65]
Prosocialdialog: A prosocial backbone for conversational agents, 2022
Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, and Maarten Sap. Prosocialdialog: A prosocial backbone for conversational agents, 2022. URL https://arxiv.org/abs/2205.12688
2022 arXiv
-
[66]
Reducing activation recomputation in large transformer models
Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems, 5, 2023
2023
-
[67]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems...
2023
-
[68]
Race: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. Race: Large-scale reading comprehension dataset from examinations. arXiv preprint arXiv:1704.04683, 2017
2017 arXiv
-
[69]
Amp: Automatically finding model parallel strategies with heterogeneity awareness
Dacheng Li, Hongyi Wang, Eric Xing, and Hao Zhang. Amp: Automatically finding model parallel strategies with heterogeneity awareness. Advances in Neural Information Processing Systems, 35: 0 6630--6639, 2022
2022
-
[70]
Datacomp-lm: In search of the next generation of training sets for language models
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, et al. Datacomp-lm: In search of the next generation of training sets for language models. arXiv preprint arXiv:2406.11794, 2024
2024 arXiv
-
[71]
Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161, 2023 a
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161, 2023 a
2023 arXiv
-
[72]
Deepinception: Hypnotize large language model to be jailbreaker
Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. Deepinception: Hypnotize large language model to be jailbreaker. arXiv preprint arXiv:2311.03191, 2023 b
2023 arXiv
-
[73]
Textbooks are all you need ii: phi-1.5 technical report
Yuanzhi Li, S \'e bastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee. Textbooks are all you need ii: phi-1.5 technical report. arXiv preprint arXiv:2309.05463, 2023 c
2023 arXiv
-
[74]
Against the achilles' heel: A survey on red teaming for generative models, 2024
Lizhi Lin, Honglin Mu, Zenan Zhai, Minghan Wang, Yuxia Wang, Renxi Wang, Junjie Gao, Yixuan Zhang, Wanxiang Che, Timothy Baldwin, Xudong Han, and Haonan Li. Against the achilles' heel: A survey on red teaming for generative models, 2024
2024
-
[75]
Truthfulqa: Measuring how models mimic human falsehoods, 2021
Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods, 2021
2021
-
[76]
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation, 2023
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang. Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation, 2023. URL https://arxiv.org/abs/2310.17389
2023 arXiv
-
[77]
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, et al. Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model. arXiv preprint arXiv:2405.04434, 2024 a
2024 arXiv
-
[78]
Goal-oriented prompt attack and safety evaluation for llms
Chengyuan Liu, Fubang Zhao, Lizhi Qing, Yangyang Kang, Changlong Sun, Kun Kuang, and Fei Wu. Goal-oriented prompt attack and safety evaluation for llms. arXiv e-prints, pp.\ arXiv--2309, 2023 a
2023
-
[79]
Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding
Hanmeng Liu, Jian Liu, Leyang Cui, Zhiyang Teng, Nan Duan, Ming Zhou, and Yue Zhang. Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31: 0 2947--2962, 2023 b . doi:10.1109/...
2023
-
[80]
Prompt injection attack against llm-integrated applications
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, et al. Prompt injection attack against llm-integrated applications. arXiv preprint arXiv:2306.05499, 2023 c
2023 arXiv
-
[81]
Llm360: Towards fully transparent open-source llms
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, et al. Llm360: Towards fully transparent open-source llms. arXiv preprint arXiv:2312.06550, 2023 d
2023 arXiv
-
[82]
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, Richard Fan, Yi Gu, Victor Miller, Yonghao Zhuang, Guowei He, Haonan Li, Fajri Koto, Liping Tang, Nikhil Ranjan, Zhiqiang Shen, Roberto Iriondo,...
2024
-
[83]
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Michael Kinney, and Daniel S. Weld. S2orc: The semantic scholar open research corpus. In Annual Meeting of the Association for Computational Linguistics, 2020. URL https://api.semanticscholar.org/CorpusID:215416146
2020
-
[84]
The flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688, 2023
2023 arXiv
-
[85]
Starcoder 2 and the stack v2: The next generation
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, et al. Starcoder 2 and the stack v2: The next generation. arXiv preprint arXiv:2402.19173, 2024
2024 arXiv
-
[86]
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024. URL https://arxiv.org...
2024 arXiv
-
[87]
Introducing meta llama 3: The most capable openly available llm to date, 2024
Meta AI . Introducing meta llama 3: The most capable openly available llm to date, 2024. URL https://ai.meta.com/blog/meta-llama-3/
2024
-
[88]
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In EMNLP, 2018
2018
-
[89]
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786, 2022
2022 arXiv
-
[90]
Octopack: Instruction tuning code large language models
Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro Von Werra, and Shayne Longpre. Octopack: Instruction tuning code large language models. arXiv preprint arXiv:2308.07124, 2023
2023 arXiv
-
[91]
Olmoe: Open mixture-of-experts language models
Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, et al. Olmoe: Open mixture-of-experts language models. arXiv preprint arXiv:2409.02060, 2024
2024 arXiv
-
[92]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online, November 2020. Associa...
2020
-
[93]
Pipedream: Generalized pipeline parallelism for dnn training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia. Pipedream: Generalized pipeline parallelism for dnn training. In Proceedings of the 27th ACM symposium on operating systems principles, p...
2019
-
[94]
Efficient large-scale language model training on gpu clusters using megatron-lm
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, et al. Efficient large-scale language model training on gpu clusters using megatron-lm. In Proceedin...
2021
-
[95]
Gpt-4 technical report, 2023
OpenAI. Gpt-4 technical report, 2023
2023
-
[96]
Grok-1: Explainable ai system
XAI Organization. Grok-1: Explainable ai system. https://github.com/xai-org/grok-1, 2023. Accessed: 2024-09-12
2023
-
[97]
Reka core, flash, and edge: A series of powerful multimodal language models
Aitor Ormazabal, Che Zheng, Cyprien de Masson d'Autume, Dani Yogatama, Deyu Fu, Donovan Ong, Eric Chen, Eugenie Lamprecht, Hai Pham, Isaac Ong, et al. Reka core, flash, and edge: A series of powerful multimodal language models. arXiv preprint arXiv:2404.12387, 2024
2024 arXiv
-
[98]
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Gerardo Flores, George H Chen, Tom Pollard, Joyce C Ho, and Tristan Naumann (eds.), Proceedings of the Conference...
2022
-
[99]
Nemotron-4 15b technical report
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Mostofa Patwary, Sandeep Subramanian, Dan Su, Chen Zhu, Deepak Narayanan, Aastha Jhunjhunwala, Ayush Dattagupta, et al. Nemotron-4 15b technical report. arXiv preprint arXiv:2402.16819, 2024
2024 arXiv
-
[100]
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman. Bbq: A hand-built bias benchmark for question answering, 2022. URL https://arxiv.org/abs/2110.08193
2022 arXiv
-
[101]
Openwebmath: An open dataset of high-quality mathematical web text, 2023
Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, and Jimmy Ba. Openwebmath: An open dataset of high-quality mathematical web text, 2023
2023
-
[102]
Carbon emissions and large neural network training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021
2021 arXiv
-
[103]
The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only. arXiv pr...
2023 arXiv
-
[104]
The fineweb datasets: Decanting the web for the finest text data at scale, 2024
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. The fineweb datasets: Decanting the web for the finest text data at scale, 2024. URL https://arxiv.org/abs/2406.17557
2024 arXiv
-
[105]
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, et al. Rwkv: Reinventing rnns for the transformer era. arXiv preprint arXiv:2305.13048, 2023
2023 arXiv
-
[106]
Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence
Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Teddy Ferdinan, Haowen Hou, Przemys aw Kazienko, et al. Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence. arXiv preprint arXiv:2404.05892, 2024
2024 arXiv
-
[107]
Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models, 2024
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao. Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models, 2024
2024
-
[108]
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019
2019
-
[109]
Aart: Ai-assisted red-teaming with diverse data generation for new llm-powered applications, 2023
Bhaktipriya Radharapu, Kevin Robinson, Lora Aroyo, and Preethi Lahoti. Aart: Ai-assisted red-teaming with diverse data generation for new llm-powered applications, 2023. URL https://arxiv.org/abs/2311.08592
2023 arXiv
-
[110]
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks,...
2022 arXiv
-
[111]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023. URL https://arxiv.org/abs/1910.10683
2023 arXiv
-
[112]
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 11 2019. URL https://arxiv.org/abs/1908.10084
2019 arXiv
-
[113]
Winogrande: an adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: an adversarial winograd schema challenge at scale. Commun. ACM, 64 0 (9): 0 99–106, aug 2021. ISSN 0001-0782. doi:10.1145/3474381. URL https://doi.org/10.1145/3474381
2021 doi
-
[114]
What language model to train if you have one million gpu hours? arXiv preprint arXiv:2210.15424, 2022
Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M Saiful Bari, Stella Biderman, Hady Elsahar, Niklas Muennighoff, Jason Phang, et al. What language model to train if you have one million gpu hours? arXiv preprint arXiv:2210.15424, 2022
-
[115]
Scalable and transferable black-box jailbreaks for language models via persona modulation
Rusheb Shah, Soroush Pour, Arush Tagade, Stephen Casper, Javier Rando, et al. Scalable and transferable black-box jailbreaks for language models via persona modulation. arXiv preprint arXiv:2311.03348, 2023
2023 arXiv
-
[116]
``Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. ``Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models . In ACM SIGSAC Conference on Computer and Communications Security (CCS) . ACM, 2024 a
2024
-
[117]
Jetmoe: Reaching llama2 performance with 0.1 m dollars
Yikang Shen, Zhen Guo, Tianle Cai, and Zengyi Qin. Jetmoe: Reaching llama2 performance with 0.1 m dollars. arXiv preprint arXiv:2404.07413, 2024 b
2024 arXiv
-
[118]
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019
1909 arXiv
-
[119]
Snowflake arctic: The best llm for enterprise ai - efficiently intelligent, truly open, 2024
AI Research Team Snowflake. Snowflake arctic: The best llm for enterprise ai - efficiently intelligent, truly open, 2024. URL https://www.snowflake.com/blog/arctic-open-efficient-foundation-language-models-snowflake/. Accessed on May 28, 2024
2024
-
[120]
peS2o (Pretraining Efficiently on S2ORC) Dataset
Luca Soldaini and Kyle Lo. peS2o (Pretraining Efficiently on S2ORC) Dataset . Technical report, Allen Institute for AI , 2023. ODC-By, https://github.com/allenai/pes2o
2023
-
[121]
Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Mu...
2024
-
[122]
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew Mccallum. Energy and policy considerations for deep learning in nlp. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 3645--3650, 2019
2019
-
[123]
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864, 2021
2021 arXiv
-
[124]
Le, Ed H
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them, 2022
2022
-
[125]
Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, and Eric P. Xing. Txt360: A top-quality llm pre-training dataset requires t...
2024
-
[126]
Xing, and Zhengzhong Liu, 2024 a
Tianhua Tao, Junbo Li, Bowen Tan, Hongyi Wang, William Marshall, Bhargav M Kanakiya, Joel Hestness, Natalia Vassilieva, Zhiqiang Shen, Eric P. Xing, and Zhengzhong Liu, 2024 a . URL https://huggingface.co/LLM360/CrystalCoder
2024
-
[127]
Xing, and Zhengzhong Liu
Tianhua Tao, Junbo Li, Bowen Tan, Hongyi Wang, William Marshall, Bhargav M Kanakiya, Joel Hestness, Natalia Vassilieva, Zhiqiang Shen, Eric P. Xing, and Zhengzhong Liu. Crystal: Illuminating LLM abilities on language and code. In First Conference on Language Modeling, 2024 b
2024
-
[128]
Hashimoto
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023
2023
-
[129]
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023
2023 arXiv
-
[130]
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi \`e re, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024 a
2024 arXiv
-
[131]
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, L \'e onard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ram \'e , et al. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024 b
2024 arXiv
-
[132]
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML NLP Team. Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023. URL www.mosaicml.com/blog/mpt-7b. Accessed: 2023-05-05
2023
-
[133]
Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023
Teknium. Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023. URL https://huggingface.co/datasets/teknium/OpenHermes-2.5
2023
-
[134]
Common crawl web corpus
The Common Crawl Team . Common crawl web corpus. http://commoncrawl.org, 2024
2024
-
[135]
Redpajama-incite-7b-base, 2023 a
Together Computer . Redpajama-incite-7b-base, 2023 a . URL https://huggingface.co/togethercomputer/RedPajama-INCITE-7B-Base
2023
-
[136]
Redpajama: an open dataset for training large language models, October 2023 b
Together Computer . Redpajama: an open dataset for training large language models, October 2023 b . URL https://github.com/togethercomputer/RedPajama-Data
2023
-
[137]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023 a
2023 arXiv
-
[138]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023 b
2023 arXiv
-
[139]
Tensor trust: Interpretable prompt injection attacks from an online game, 2023
Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell. Tensor trust: Interpretable prompt injection attacks from an online game, 2023. URL https:/...
2023 arXiv
-
[140]
Bypassing the safety training of open-source LLM s with priming attacks
Jason Vega, Isha Chaudhary, Changming Xu, and Gagandeep Singh. Bypassing the safety training of open-source LLM s with priming attacks. In The Second Tiny Papers Track at ICLR 2024, 2024. URL https://openreview.net/forum?id=nz8Byp7ep6
2024
-
[141]
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model . https://github.com/kingoflolz/mesh-transformer-jax, May 2021
2021
-
[142]
Do-not-answer: A dataset for evaluating safeguards in llms, 2023
Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. Do-not-answer: A dataset for evaluating safeguards in llms, 2023. URL https://arxiv.org/abs/2308.13387
2023 arXiv
-
[143]
Do-not-answer: Evaluating safeguards in LLM s
Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. Do-not-answer: Evaluating safeguards in LLM s. In Yvette Graham and Matthew Purver (eds.), Findings of the Association for Computational Linguistics: EACL 2024, pp.\ 896--911, St. Julian ' s, Malta, March 2...
2024
-
[144]
Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024
2024
-
[145]
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. In International Conference on Learning Representations
-
[146]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models, 2022
2022
-
[147]
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023. URL https://arxiv.org/abs/2201.11903
2023 arXiv
-
[148]
Bloom: A 176b-parameter open-access multilingual language model
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili \'c , Daniel Hesslow, Roman Castagn \'e , Alexandra Sasha Luccioni, Fran c ois Yvon, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.0...
2022 arXiv
-
[149]
Sustainable ai: Environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al. Sustainable ai: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4: 0 795--...
2022
-
[150]
Cognitive overload: Jailbreaking large language models with overloaded logical thinking
Nan Xu, Fei Wang, Ben Zhou, Bangzheng Li, Chaowei Xiao, and Muhao Chen. Cognitive overload: Jailbreaking large language models with overloaded logical thinking. In Kevin Duh, Helena Gomez, and Steven Bethard (eds.), Findings of the Association for Computational Linguistics: NA...
2024
-
[151]
Baichuan 2: Open large-scale language models
Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, et al. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305, 2023
2023 arXiv
-
[152]
Qwen2 technical report
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024
2024 arXiv
-
[153]
Bloom+ 1: Adding language support to bloom for zero-shot prompting
Zheng-Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, et al. Bloom+ 1: Adding language support to bloom for zero-shot prompting. arXiv preprint arXiv:2212.09...
2022 arXiv
-
[154]
Yi: Open foundation models by 01
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652, 2024
2024 arXiv
-
[155]
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. arXiv preprint arXiv:2308.06463, 2023
2023 arXiv
-
[156]
Mammoth: Building math generalist models through hybrid instruction tuning
Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mammoth: Building math generalist models through hybrid instruction tuning. arXiv preprint arXiv:2309.05653, 2023
2023 arXiv
-
[157]
Hellaswag: Can a machine really finish your sentence? In Annual Meeting of the Association for Computational Linguistics, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Annual Meeting of the Association for Computational Linguistics, 2019. URL https://api.semanticscholar.org/CorpusID:159041722
2019
-
[158]
Glm-130b: An open bilingual pre-trained model
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Peng Zhang, Yuxiao Dong, and Jie Tang. Glm-130b: An open bilingual pre-trained model. arXiv prepr...
-
[159]
Map-neo: Highly capable and transparent bilingual large language model series
Ge Zhang, Scott Qu, Jiaheng Liu, Chenchen Zhang, Chenghua Lin, Chou Leuang Yu, Danny Pan, Esther Cheng, Jie Liu, Qunshu Lin, et al. Map-neo: Highly capable and transparent bilingual large language model series. arXiv preprint arXiv:2405.19327, 2024
2024 arXiv
-
[160]
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022
2022 arXiv
-
[161]
Alpa: Automating inter-and \ Intra-Operator \ parallelism for distributed deep learning
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P Xing, et al. Alpa: Automating inter-and \ Intra-Operator \ parallelism for distributed deep learning. In 16th USENIX Symposium on Operating Systems ...
2022
-
[162]
P Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
2023
-
[163]
Opencodeinterpreter: Integrating code generation with execution and refinement
Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue. Opencodeinterpreter: Integrating code generation with execution and refinement. arXiv preprint arXiv:2402.14658, 2024
2024 arXiv
-
[164]
Jiuzhang3.0: Efficiently improving mathematical reasoning by training small data synthesis models
Kun Zhou, Beichen Zhang, Jiapeng Wang, Zhipeng Chen, Wayne Xin Zhao, Jing Sha, Zhichao Sheng, Shijin Wang, and Ji-Rong Wen. Jiuzhang3.0: Efficiently improving mathematical reasoning by training small data synthesis models. 2024
2024
-
[165]
Astraios: Parameter-efficient instruction tuning code large language models
Terry Yue Zhuo, Armel Zebaze, Nitchakarn Suppattarachai, Leandro von Werra, Harm de Vries, Qian Liu, and Niklas Muennighoff. Astraios: Parameter-efficient instruction tuning code large language models. arXiv preprint arXiv:2401.00788, 2024
2024 arXiv
-
[166]
Zico Kolter, and Matt Fredrikson
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models, 2023. URL https://arxiv.org/abs/2307.15043
2023 arXiv
-
[167]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.