Pith. sign in

REVIEW 2 cited by

The Cost of Down-Scaling Language Models: Fact Recall Deteriorates before In-Context Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.04680 v1 pith:LUS63OLP submitted 2023-10-07 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords scalingin-contextmodelcapabilitiesfactlearningrecallcore
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How does scaling the number of parameters in large language models (LLMs) affect their core capabilities? We study two natural scaling techniques -- weight pruning and simply training a smaller or larger model, which we refer to as dense scaling -- and their effects on two core capabilities of LLMs: (a) recalling facts presented during pre-training and (b) processing information presented in-context during inference. By curating a suite of tasks that help disentangle these two capabilities, we find a striking difference in how these two abilities evolve due to scaling. Reducing the model size by more than 30\% (via either scaling approach) significantly decreases the ability to recall facts seen in pre-training. Yet, a 60--70\% reduction largely preserves the various ways the model can process in-context information, ranging from retrieving answers from a long context to learning parameterized functions from in-context exemplars. The fact that both dense scaling and weight pruning exhibit this behavior suggests that scaling model size has an inherently disparate effect on fact recall and in-context learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Pruning attention layers in five LLMs across eight datasets maintains accuracy but degrades faithfulness and calibration.

  2. The Structural Attention Tax: How Retrieval Format Hijacks In-Context Learning Independent of Content

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Knowledge graph triples capture 2-3x more attention per token than equivalent natural language due to structural patterns, compressing demonstration attention by up to 42% independent of semantic relevance.

Pith tools