REVIEW 13 cited by
Understanding Emergent Abilities of Language Models from the Loss Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent studies have put into question the belief that emergent abilities in language models are exclusive to large models. This skepticism arises from two observations: 1) smaller models can also exhibit high performance on emergent abilities and 2) there is doubt on the discontinuous metrics used to measure these abilities. In this paper, we propose to study emergent abilities in the lens of pre-training loss, instead of model size or training compute. We demonstrate that the Transformer models with the same pre-training loss, but different model and data sizes, generate the same performance on various downstream tasks, with a fixed data corpus, tokenization, and model architecture. We also discover that a model exhibits emergent abilities on certain tasks -- regardless of the continuity of metrics -- when its pre-training loss falls below a specific threshold. Before reaching this threshold, its performance remains at the level of random guessing. This inspires us to redefine emergent abilities as those that manifest in models with lower pre-training losses, highlighting that these abilities cannot be predicted by merely extrapolating the performance trends of models with higher pre-training losses.
Forward citations
Cited by 13 Pith papers
-
Predicting Emergent Capabilities by Finetuning
Finetuning small models shifts the point where capability emerges, and extrapolating this shift to the low-data limit predicts few-shot emergence up to about 4x the compute in advance.
-
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
Soup-of-Experts pretrains a shared parameter bank and many expert vectors, plus a router, so a small specialist language model can be instantiated instantly from any domain-weight mixture without retraining.
-
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
A training pipeline called T1 scales RL for LLM reasoning via diverse sampling and rewards, and claims that longer allowed generations directly improve math accuracy without verifiers.
-
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
RLHF training on LLMs shows diminishing returns from more response samples, larger reward models, and larger policy models, so it scales less efficiently than pretraining.
-
Loss-to-Loss Prediction: Scaling Laws for All Datasets
Losses of models trained on different datasets are related by shifted power laws, enabling translation of scaling laws and prediction of downstream performance from a few runs.
-
AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model
AMix-1, a 1.7B-parameter Bayesian Flow Network protein model conditioned on MSA profiles and refined by an evolutionary test-time scaling loop, produced AmeR variants with up to 50x wild-type activity in wet-lab tests.
-
Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law
A fitted token-level loss weighting, optimized against known model accuracies, predicts held-out downstream task performance more accurately than mean validation loss on five of six benchmarks.
-
Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need
The paper claims that proxy tasks selected by cross-model performance correlation and small-model variance ratios can predict LLM tool-use capability rankings at early training stages.
-
ChronoLLM: A Framework for Customizing Large Language Model for Digital Twins generalization based on PyChrono
Fine-tuning LLMs on PyChrono-specific data improves their success rate at generating runnable simulation code from about 40% to about 85%, compared to prompting general models.
-
Optimizing Sequential Recommendation Models with Scaling Laws and Approximate Entropy
The authors propose a 'Performance Law' for sequential recommendation models that predicts HR and NDCG from model layers, embedding dimension, and number of tokens divided by Approximate Entropy, then uses the fitted ...
-
Foundation Models for Astrophysics
Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...
-
Foundations of GenIR
A survey chapter proposing that generative AI reshapes information access through two paradigms, information generation and information synthesis.
- Transformers Struggle to Learn to Search
Discussion (0). Continue with ORCID to comment.