Pith. sign in

REVIEW 14 cited by

Emergent Abilities in Large Language Models: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.05788 v2 pith:23C3AXTJ submitted 2025-02-28 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords abilitiesemergentmodelsreasoninglargetheylanguageleading
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Language Models (LLMs) are leading a new technological revolution as one of the most promising research streams toward artificial general intelligence. The scaling of these models, accomplished by increasing the number of parameters and the magnitude of the training datasets, has been linked to various so-called emergent abilities that were previously unobserved. These emergent abilities, ranging from advanced reasoning and in-context learning to coding and problem-solving, have sparked an intense scientific debate: Are they truly emergent, or do they simply depend on external factors, such as training dynamics, the type of problems, or the chosen metric? What underlying mechanism causes them? Despite their transformative potential, emergent abilities remain poorly understood, leading to misconceptions about their definition, nature, predictability, and implications. In this work, we shed light on emergent abilities by conducting a comprehensive review of the phenomenon, addressing both its scientific underpinnings and real-world consequences. We first critically analyze existing definitions, exposing inconsistencies in conceptualizing emergent abilities. We then explore the conditions under which these abilities appear, evaluating the role of scaling laws, task complexity, pre-training loss, quantization, and prompting strategies. Our review extends beyond traditional LLMs and includes Large Reasoning Models (LRMs), which leverage reinforcement learning and inference-time search to amplify reasoning and self-reflection. However, emergence is not inherently positive. As AI systems gain autonomous reasoning capabilities, they also develop harmful behaviors, including deception, manipulation, and reward hacking. We highlight growing concerns about safety and governance, emphasizing the need for better evaluation frameworks and regulatory oversight.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

    cs.CV 2025-11 unverdicted novelty 8.0 of 10

    MVI-Bench supplies the first taxonomy and dataset focused on misleading visual inputs to measure LVLM robustness, with tests on 18 models revealing clear weaknesses.

  2. Market Design for AI: Beyond the Copyright Binary

    econ.TH 2026-06 unverdicted novelty 7.0 of 10

    In a stylized model, AI-firm monopsony and correlated creator content produce an 'originality penalty' and a dynamic 'curse of precision'; a two-part-tariff data intermediary restores the social optimum.

  3. Vision Language Models Cannot Reason About Physical Transformation

    cs.AI 2026-03 accept novelty 6.5 of 10

    Current VLMs cannot maintain transformation-invariant representations of number, length, volume or size and instead rely on textual invariance priors that reverse on matched non-conserving controls.

  4. EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new Persian-Islamic trustworthiness benchmark ranks Claude highest and Qwen lowest across eight LLMs and finds safety is the weakest dimension.

  5. Artificial or Human Intelligence?

    econ.TH 2025-09 conditional novelty 6.0 of 10

    A theoretical model shows that AI's sharp capability cutoff and hallucination rate create a discontinuous gap in student ability around the AI frontier, and that AI-free assignments can correct students' overestimatio...

  6. The Other Mind: How Language Models Exhibit Human Temporal Cognition

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.

  7. Conversational AI as a Catalyst for Informal Learning: An Empirical Large-Scale Study on LLM Use in Everyday Learning

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Most adults in a German Prolific sample report using large language models for informal learning, with four distinct learner profiles emerging from their usage patterns.

  8. Lost in Context: Addressing Context Anxiety in Large Language Models

    cs.AI 2026-05 reject novelty 5.0 of 10

    Context anxiety — abandoning solvable tasks over perceived token limits — is measurable and reducible by fine-tuning on anxiety-free reasoning traces, but the paper's causal mechanism is not actually tested.

  9. Large Language Models and Emergence: A Complex Systems Perspective

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A perspective paper arguing that LLM emergence claims are incomplete without evidence of internal coarse-grained representations, and that LLMs have not shown emergent intelligence.

  10. Quantum-Structured World Models (QSWMs) for Predictive Latent Dynamics

    cs.LG 2026-08 conditional novelty 4.0 of 10

    A quantum-inspired world model with complex-valued latents beats matched classical baselines on one-step cellular-automaton prediction, but its advantage decays in long-horizon rollout.

  11. A Large Language Model-Driven Agent-Based Modeling Framework with Multi-Round Communication for Simulating Vaccine Opinion Dynamics

    cs.MA 2026-07 conditional novelty 4.0 of 10

    An LLM-driven agent-based model with multi-round dialogue reproduces non-linear social influence patterns in vaccination opinion dynamics, with memory increasing resistance and prompt diversity increasing adoption.

  12. TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization

    cs.SD 2025-08 reject novelty 4.0 of 10

    TinyMusician distills MusicGen and applies hand-picked mixed-precision quantization to make a 1.04 GB on-device music generator, but the headline '93% quality, 55% smaller' claims conflict with the paper's own tables.

  13. A Survey on Latent Reasoning

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey that organizes latent reasoning methods into vertical recurrence, horizontal recurrence, and infinite-depth diffusion, arguing that silent reasoning can beat explicit chain-of-thought.

  14. Not All Explanations for Deep Learning Phenomena Are Equally Valuable

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A position paper arguing that narrow, puzzle-solving explanations of deep learning edge case phenomena are low-value, and that these phenomena should instead be used to stress-test broad explanatory theories.

Pith tools