REVIEW 14 cited by
Emergent Abilities in Large Language Models: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large Language Models (LLMs) are leading a new technological revolution as one of the most promising research streams toward artificial general intelligence. The scaling of these models, accomplished by increasing the number of parameters and the magnitude of the training datasets, has been linked to various so-called emergent abilities that were previously unobserved. These emergent abilities, ranging from advanced reasoning and in-context learning to coding and problem-solving, have sparked an intense scientific debate: Are they truly emergent, or do they simply depend on external factors, such as training dynamics, the type of problems, or the chosen metric? What underlying mechanism causes them? Despite their transformative potential, emergent abilities remain poorly understood, leading to misconceptions about their definition, nature, predictability, and implications. In this work, we shed light on emergent abilities by conducting a comprehensive review of the phenomenon, addressing both its scientific underpinnings and real-world consequences. We first critically analyze existing definitions, exposing inconsistencies in conceptualizing emergent abilities. We then explore the conditions under which these abilities appear, evaluating the role of scaling laws, task complexity, pre-training loss, quantization, and prompting strategies. Our review extends beyond traditional LLMs and includes Large Reasoning Models (LRMs), which leverage reinforcement learning and inference-time search to amplify reasoning and self-reflection. However, emergence is not inherently positive. As AI systems gain autonomous reasoning capabilities, they also develop harmful behaviors, including deception, manipulation, and reward hacking. We highlight growing concerns about safety and governance, emphasizing the need for better evaluation frameworks and regulatory oversight.
Forward citations
Cited by 14 Pith papers
-
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs
MVI-Bench supplies the first taxonomy and dataset focused on misleading visual inputs to measure LVLM robustness, with tests on 18 models revealing clear weaknesses.
-
Market Design for AI: Beyond the Copyright Binary
In a stylized model, AI-firm monopsony and correlated creator content produce an 'originality penalty' and a dynamic 'curse of precision'; a two-part-tariff data intermediary restores the social optimum.
-
Vision Language Models Cannot Reason About Physical Transformation
Current VLMs cannot maintain transformation-invariant representations of number, length, volume or size and instead rely on textual invariance priors that reverse on matched non-conserving controls.
-
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
A new Persian-Islamic trustworthiness benchmark ranks Claude highest and Qwen lowest across eight LLMs and finds safety is the weakest dimension.
-
Artificial or Human Intelligence?
A theoretical model shows that AI's sharp capability cutoff and hallucination rate create a discontinuous gap in student ability around the AI frontier, and that AI-free assignments can correct students' overestimatio...
-
The Other Mind: How Language Models Exhibit Human Temporal Cognition
Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.
-
Conversational AI as a Catalyst for Informal Learning: An Empirical Large-Scale Study on LLM Use in Everyday Learning
Most adults in a German Prolific sample report using large language models for informal learning, with four distinct learner profiles emerging from their usage patterns.
-
Lost in Context: Addressing Context Anxiety in Large Language Models
Context anxiety — abandoning solvable tasks over perceived token limits — is measurable and reducible by fine-tuning on anxiety-free reasoning traces, but the paper's causal mechanism is not actually tested.
-
Large Language Models and Emergence: A Complex Systems Perspective
A perspective paper arguing that LLM emergence claims are incomplete without evidence of internal coarse-grained representations, and that LLMs have not shown emergent intelligence.
-
Quantum-Structured World Models (QSWMs) for Predictive Latent Dynamics
A quantum-inspired world model with complex-valued latents beats matched classical baselines on one-step cellular-automaton prediction, but its advantage decays in long-horizon rollout.
-
A Large Language Model-Driven Agent-Based Modeling Framework with Multi-Round Communication for Simulating Vaccine Opinion Dynamics
An LLM-driven agent-based model with multi-round dialogue reproduces non-linear social influence patterns in vaccination opinion dynamics, with memory increasing resistance and prompt diversity increasing adoption.
-
TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
TinyMusician distills MusicGen and applies hand-picked mixed-precision quantization to make a 1.04 GB on-device music generator, but the headline '93% quality, 55% smaller' claims conflict with the paper's own tables.
-
A Survey on Latent Reasoning
A survey that organizes latent reasoning methods into vertical recurrence, horizontal recurrence, and infinite-depth diffusion, arguing that silent reasoning can beat explicit chain-of-thought.
-
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
A position paper arguing that narrow, puzzle-solving explanations of deep learning edge case phenomena are low-value, and that these phenomena should instead be used to stress-test broad explanatory theories.
Discussion (0). Continue with ORCID to comment.