Pith. sign in

REVIEW 13 cited by

Understanding Emergent Abilities of Language Models from the Loss Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.15796 v3 pith:HFJFFBM2 submitted 2024-03-23 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords abilitiesmodelsemergentpre-traininglossmodelperformancedata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent studies have put into question the belief that emergent abilities in language models are exclusive to large models. This skepticism arises from two observations: 1) smaller models can also exhibit high performance on emergent abilities and 2) there is doubt on the discontinuous metrics used to measure these abilities. In this paper, we propose to study emergent abilities in the lens of pre-training loss, instead of model size or training compute. We demonstrate that the Transformer models with the same pre-training loss, but different model and data sizes, generate the same performance on various downstream tasks, with a fixed data corpus, tokenization, and model architecture. We also discover that a model exhibits emergent abilities on certain tasks -- regardless of the continuity of metrics -- when its pre-training loss falls below a specific threshold. Before reaching this threshold, its performance remains at the level of random guessing. This inspires us to redefine emergent abilities as those that manifest in models with lower pre-training losses, highlighting that these abilities cannot be predicted by merely extrapolating the performance trends of models with higher pre-training losses.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predicting Emergent Capabilities by Finetuning

    cs.LG 2024-11 conditional novelty 7.0 of 10

    Finetuning small models shifts the point where capability emerges, and extrapolating this shift to the low-data limit predicts few-shot emergence up to about 4x the compute in advance.

  2. Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Soup-of-Experts pretrains a shared parameter bank and many expert vectors, plus a router, so a small specialist language model can be instantiated instantly from any domain-weight mixture without retraining.

  3. T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A training pipeline called T1 scales RL for LLM reasoning via diverse sampling and rewards, and claims that longer allowed generations directly improve math accuracy without verifiers.

  4. Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

    cs.CL 2024-12 conditional novelty 6.0 of 10

    RLHF training on LLMs shows diminishing returns from more response samples, larger reward models, and larger policy models, so it scales less efficiently than pretraining.

  5. Loss-to-Loss Prediction: Scaling Laws for All Datasets

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Losses of models trained on different datasets are related by shifted power laws, enabling translation of scaling laws and prediction of downstream performance from a few runs.

  6. AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

    q-bio.BM 2025-07 conditional novelty 5.0 of 10

    AMix-1, a 1.7B-parameter Bayesian Flow Network protein model conditioned on MSA profiles and refined by an evolutionary test-time scaling loop, produced AmeR variants with up to 50x wild-type activity in wet-lab tests.

  7. Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A fitted token-level loss weighting, optimized against known model accuracies, predicts held-out downstream task performance more accurately than mean validation loss on five of six benchmarks.

  8. Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need

    cs.CL 2024-12 reject novelty 5.0 of 10

    The paper claims that proxy tasks selected by cross-model performance correlation and small-model variance ratios can predict LLM tool-use capability rankings at early training stages.

  9. ChronoLLM: A Framework for Customizing Large Language Model for Digital Twins generalization based on PyChrono

    cs.SE 2025-01 conditional novelty 4.0 of 10

    Fine-tuning LLMs on PyChrono-specific data improves their success rate at generating runnable simulation code from about 40% to about 85%, compared to prompting general models.

  10. Optimizing Sequential Recommendation Models with Scaling Laws and Approximate Entropy

    cs.AI 2024-11 reject novelty 4.0 of 10

    The authors propose a 'Performance Law' for sequential recommendation models that predicts HR and NDCG from model layers, embedding dimension, and number of tokens divided by Approximate Entropy, then uses the fitted ...

  11. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0 of 10

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...

  12. Foundations of GenIR

    cs.IR 2025-01 unverdicted novelty 1.0 of 10

    A survey chapter proposing that generative AI reshapes information access through two paradigms, information generation and information synthesis.

  13. Transformers Struggle to Learn to Search

    cs.CL 2024-12

Pith tools