REVIEW 14 cited by
Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
We study recent research advances that improve large language models through efficient pre-training and scaling, and open datasets and tools. We combine these advances to introduce Cerebras-GPT, a family of open compute-optimal language models scaled from 111M to 13B parameters. We train Cerebras-GPT models on the Eleuther Pile dataset following DeepMind Chinchilla scaling rules for efficient pre-training (highest accuracy for a given compute budget). We characterize the predictable power-law scaling and compare Cerebras-GPT with other publicly-available models to show all Cerebras-GPT models have state-of-the-art training efficiency on both pre-training and downstream objectives. We describe our learnings including how Maximal Update Parameterization ($\mu$P) can further improve large model scaling, improving accuracy and hyperparameter predictability at scale. We release our pre-trained models and code, making this paper the first open and reproducible work comparing compute-optimal model scaling to models trained on fixed dataset sizes. Cerebras-GPT models are available on HuggingFace: https://huggingface.co/cerebras.
Forward citations
Cited by 14 Pith papers
-
How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
Vendi Score and scaling-law objectives belong to the class of matrix spectral functions, which are submodular, enabling efficient greedy selection of training data that outperforms random subsets in predicting held-ou...
-
Maximal Update Parametrization and Zero-Shot Hyperparameter Transfer for Fourier Neural Operators
The authors derive and test a parametrization for FNOs that keeps optimal hyperparameters stable as the number of Fourier modes grows, enabling zero-shot transfer from small proxy models to near-billion-parameter models.
-
Basis Transformers for Multi-Task Tabular Regression
Basis transformers beat fine-tuned LLMs on 34 multi-task tabular regression datasets while using five times fewer parameters and no data preprocessing.
-
TokenPrint: A Calibrated Token-Space Fingerprint for Language-Model Provenance
A token-overlap fingerprint on fixed probes tracks model lineage and shared training data without weight access, with calibrated similarity levels across 32 models.
-
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
Falcon-H1 reports competitive benchmark scores for a 0.5B to 34B family of parallel hybrid attention/Mamba-2 models, claiming 2x to 4x parameter efficiency versus dense transformers.
-
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters
Sharing one MLP layer's weights across several layers plus low-rank adapters recovers most of a pretrained LLM's quality with a fraction of the storage and faster phone inference.
-
MuLoCo: Muon is a practical inner optimizer for DiLoCo
Using Muon instead of AdamW inside DiLoCo improves worker scaling and critical batch size for LLM pre-training across 150M to 15B parameters.
-
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models
An attention variant that compresses Key heads harder than Value heads and widens Query heads yields up to 33.36% faster long-context attention, powering a system-domain LLM that reportedly outperforms GPT-4 on the ne...
-
LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch
K2 Diamond is a fully open 65B-parameter LLM that reaches Llama 2 70B-level performance on standard benchmarks.
-
Foundational Large Language Models for Materials Research
Domain-adapted LLaMA models (LLaMat) outperform commercial LLMs on materials NLP and structured extraction tasks and generate M3GNet-predicted stable crystals, with LLaMA-2-based variants beating LLaMA-3-based ones.
-
TensorSLM: Energy-efficient Embedding Compression of Sub-billion Parameter Language Models on Low-end Devices
TensorSLM applies per-vector tensor-train SVD to compress SLM token embeddings training-free, showing competitive task performance at roughly 2x embedding compression on Raspberry Pi with an estimated, pre-decoder ene...
-
Get Experience from Practice: LLM Agents with Record & Replay
AgentRR is a proposed paradigm that records agent traces, generalizes them into multi-level experiences, and replays them under safety checks to make LLM agents cheaper, faster, and more reliable.
-
A Probabilistic WxChallenge Proposal
Two optional WxChallenge games let players bet confidence credits on ensemble-based thresholds or bins, with scores based on information gain over the baseline.
-
YuLan-Mini: An Open Data-efficient Language Model
A 2.42B-parameter base model trained on 1.08T tokens matches or beats several industry baselines trained on 7T to 18T tokens across math, code, and general benchmarks.
Discussion (0). Continue with ORCID to comment.