REVIEW 8 cited by
What is the Role of Small Models in the LLM Era: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have made significant progress in advancing artificial general intelligence (AGI), leading to the development of increasingly large models such as GPT-4 and LLaMA-405B. However, scaling up model sizes results in exponentially higher computational costs and energy consumption, making these models impractical for academic researchers and businesses with limited resources. At the same time, Small Models (SMs) are frequently used in practical settings, although their significance is currently underestimated. This raises important questions about the role of small models in the era of LLMs, a topic that has received limited attention in prior research. In this work, we systematically examine the relationship between LLMs and SMs from two key perspectives: Collaboration and Competition. We hope this survey provides valuable insights for practitioners, fostering a deeper understanding of the contribution of small models and promoting more efficient use of computational resources. The code is available at https://github.com/tigerchen52/role_of_small_models
Forward citations
Cited by 8 Pith papers
-
AI Propaganda factories with language models
Small language models sustain political personas and become more ideologically extreme when replying to counter-arguments, according to a language-model judge.
-
Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints
A modified DPO loss with a hinge margin improves small LLM alignment on AlpacaEval by about 2 points over the APO-zero baseline.
-
Model compression using knowledge distillation with integrated gradients
Overlaying teacher integrated-gradient maps on 10% of training images improves knowledge-distilled students by about one percentage point on CIFAR-10 at 4.1x compression.
-
LightRouter: Towards Efficient LLM Collaboration with Minimal Overhead
LightRouter uses short preview outputs to filter a pool of LLMs down to two, then aggregates their full responses, beating ensemble baselines and matching costlier models.
-
Energy-Aware Code Generation with LLMs: Benchmarking Small vs. Large Language Models for Sustainable AI Programming
On 150 LeetCode problems, GPT-4.0 and DeepSeek-Reasoner beat three 3B-parameter models on correctness and speed; the 52% energy-efficiency claim counts any of three SLMs on correct outputs, not a per-model advantage.
-
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
A survey of multi-LLM systems in edge computing, covering architectures, enabling technologies, trust mechanisms, applications, and open datasets for edge general intelligence.
-
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
The authors propose a competition with new scoring metrics to find benchmarks that give clean early-training signals for small language models, and show MMLU-var outperforms MMLU as a baseline.
-
Towards a Small Language Model Lifecycle Framework
A synthesis proposing a modular lifecycle framework that organizes Small Language Model development and deployment into interconnected main, optional, and cross-cutting components.
Discussion (0). Sign in to comment.