REVIEW 7 cited by
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Despite the recent observation that large language models (LLMs) can store substantial factual knowledge, there is a limited understanding of the mechanisms of how they acquire factual knowledge through pretraining. This work addresses this gap by studying how LLMs acquire factual knowledge during pretraining. The findings reveal several important insights into the dynamics of factual knowledge acquisition during pretraining. First, counterintuitively, we observe that pretraining on more data shows no significant improvement in the model's capability to acquire and maintain factual knowledge. Next, there is a power-law relationship between training steps and forgetting of memorization and generalization of factual knowledge, and LLMs trained with duplicated training data exhibit faster forgetting. Third, training LLMs with larger batch sizes can enhance the models' robustness to forgetting. Overall, our observations suggest that factual knowledge acquisition in LLM pretraining occurs by progressively increasing the probability of factual knowledge presented in the pretraining data at each step. However, this increase is diluted by subsequent forgetting. Based on this interpretation, we demonstrate that we can provide plausible explanations for recently observed behaviors of LLMs, such as the poor performance of LLMs on long-tail knowledge and the benefits of deduplicating the pretraining corpus.
Forward citations
Cited by 7 Pith papers
-
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
A shortcut-filtered benchmark shows LLMs genuinely compose facts internally for country-bridge queries (over 80% for the best models) but almost never for year-bridge queries (about 5-6%).
-
Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory
LLMs' source-attribution ability is not fixed: it flips with conversational memory structure, and corrective feedback can invert judgments or sever confidence from accuracy.
-
FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
FPEdit uses knowledge editing with a promote-suppress objective to embed robust, stealthy natural-language fingerprints into LLMs, achieving 94 to 100 percent retention after fine-tuning while preserving benchmark per...
-
MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability
A pre-training task called RAMP, where models practice searching to fill masked text spans, improves downstream agentic open-domain QA performance across Qwen and LLaMA models.
-
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
A fine-tuned 14B LLM judge, trained with scenario-based prompts and controlled instruction generation, approaches GPT-4's human-agreement performance, and the paper documents why scaling distillation data can fail.
-
Episodic memory in AI agents poses risks that should be studied and mitigated
Episodic memory in AI agents could enable both safety benefits and significant new risks, and developers should adopt principles that keep memories interpretable, user-controllable, detachable, and not editable by the...
-
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
AttriBoT combines caching, hierarchical pruning, and smaller proxy models to approximate leave-one-out context attribution with a >300x speedup and little loss in faithfulness.
Discussion (0). Continue with ORCID to comment.