Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:07:17.187777Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2411.18700.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:07:17.187777Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8e94b6cc-9e36-4656-b103-5f0a93eb967b · outbound
On the Effectiveness of Incremental Training of Large Language Models Language Models are Few-Shot Learners
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c78e031-6c87-441f-bcb1-f6f298f0ad23 · outbound
On the Effectiveness of Incremental Training of Large Language Models BERT Rediscovers the Classical NLP Pipeline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b64c5ac2-0d07-4d4f-8f8b-1ebb53944b65 · outbound
On the Effectiveness of Incremental Training of Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14598270-c0ff-4c62-ac3d-440bb049dfb9 · outbound
On the Effectiveness of Incremental Training of Large Language Models Scaling Laws for Neural Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12439348-02dc-434e-ac1b-0217059f06b4 · outbound
On the Effectiveness of Incremental Training of Large Language Models Training Compute-Optimal Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980ea5dc-d5f2-4381-a8f7-06924a65c315 · outbound
On the Effectiveness of Incremental Training of Large Language Models The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25ee6e12-78f2-4a44-9998-5e19b622a658 · outbound
On the Effectiveness of Incremental Training of Large Language Models A fast learning algorithm for deep belief nets,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8ab18b5-2a3a-4326-a982-14b1dbc99b3d · outbound
On the Effectiveness of Incremental Training of Large Language Models Greedy layer- wise training of deep networks,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c8871e9c-a191-41c6-bfd5-3736c3830568 · outbound
On the Effectiveness of Incremental Training of Large Language Models The cascade-correlation learning architec- ture,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ee8b7273-9829-42aa-aa6e-efd1c33d98ab · outbound
On the Effectiveness of Incremental Training of Large Language Models Improving language models by retrieving from trillions of tokens
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac1514ad-4586-4dcf-b5a2-75d08f926f6e · outbound
On the Effectiveness of Incremental Training of Large Language Models Analyzing hidden representations in end- to-end automatic speech recognition systems,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation def03a05-a250-46df-94d8-c585e739b61d · outbound
On the Effectiveness of Incremental Training of Large Language Models The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35a86151-da4d-4fd4-a6ed-57636614eb5d · outbound
On the Effectiveness of Incremental Training of Large Language Models Transformer Feed-Forward Layers Are Key-Value Memories
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc3fc014-5cb3-4fef-9662-7f08d269cbe8 · outbound
On the Effectiveness of Incremental Training of Large Language Models Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 254cd807-f22d-46a5-b661-d8cd9626cad5 · outbound
On the Effectiveness of Incremental Training of Large Language Models Visualizing and understanding convolutional networks,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe957c28-c9de-4056-bd4b-9c6c96da2ad2 · outbound
On the Effectiveness of Incremental Training of Large Language Models Qwen2 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08b23097-01bd-4b1a-8f60-41a3e2c3e273 · outbound
On the Effectiveness of Incremental Training of Large Language Models Why does unsuper- vised pre-training help deep learning?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 53a07be4-1a81-4739-ad62-02493786a540 · outbound
On the Effectiveness of Incremental Training of Large Language Models The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b543ff2-8fd8-4520-affc-87eec956b3e2 · outbound
On the Effectiveness of Incremental Training of Large Language Models Distilling the Knowledge in a Neural Network
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b32f80-6923-4e83-a35f-758e6b4c2edc · outbound
On the Effectiveness of Incremental Training of Large Language Models Mixed Precision Training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04acfd1d-844b-4814-b545-67c71f54d203 · outbound
On the Effectiveness of Incremental Training of Large Language Models Universal Language Model Fine-tuning for Text Classification
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b95f1d8f-9be0-4227-913a-e7951f3a350c · outbound
On the Effectiveness of Incremental Training of Large Language Models Progressive Neural Networks
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2c69e01-9f59-46fd-94e0-52803e39fe34 · outbound
On the Effectiveness of Incremental Training of Large Language Models An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd30d1d6-9297-42ec-b819-f5ceff22f5c8 · outbound
On the Effectiveness of Incremental Training of Large Language Models How transferable are features in deep neural networks?
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ea9ea4-19ef-46ed-9985-cfc563d25274 · outbound
On the Effectiveness of Incremental Training of Large Language Models Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a86c7fe5-b006-41f8-a950-cbd05c080d48 · outbound
On the Effectiveness of Incremental Training of Large Language Models Fineweb- edu,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 46459acf-7dc4-488a-8c36-842f54d2ee83 · outbound
On the Effectiveness of Incremental Training of Large Language Models Decoupled Weight Decay Regularization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a60cc8-8eb3-4df5-8417-d28617c884ca · outbound
On the Effectiveness of Incremental Training of Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc76df30-b020-4876-a902-00052a076e06 · outbound
On the Effectiveness of Incremental Training of Large Language Models Deep feedforward net- works,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.