Pith. sign in

Paper Citation Record · LEDGER

On the Effectiveness of Incremental Training of Large Language Models

As of 13 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2411.18700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18700 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:07:17.187777Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e94b6cc-9e36-4656-b103-5f0a93eb967b · outbound

This paper cites Language Models are Few-Shot Learners.

On the Effectiveness of Incremental Training of Large Language Models Language Models are Few-Shot Learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.072102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.072102Z digest=sha256:4861d386b613b9bf13218dac312b1c8455759caccd58b2b82eae70e940d7049e

Observation 0c78e031-6c87-441f-bcb1-f6f298f0ad23 · outbound

This paper cites BERT Rediscovers the Classical NLP Pipeline.

On the Effectiveness of Incremental Training of Large Language Models BERT Rediscovers the Classical NLP Pipeline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.077285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.077285Z digest=sha256:8e12b8f4a4a68dde7d01af005090cf2c4be7f8d3d1f57064e4dd2dbb7eec0c9a

Observation b64c5ac2-0d07-4d4f-8f8b-1ebb53944b65 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

On the Effectiveness of Incremental Training of Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.081467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.081467Z digest=sha256:78c7f277356ad59e785a66e75203d6b9533827f1ad841da61f77c07c65157eee

Observation 14598270-c0ff-4c62-ac3d-440bb049dfb9 · outbound

This paper cites Scaling Laws for Neural Language Models.

On the Effectiveness of Incremental Training of Large Language Models Scaling Laws for Neural Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.085994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.085994Z digest=sha256:5ea35ecfaddb0f227306af7b209ce03fbec22a23819112dca4ae7a9223728a4c

Observation 12439348-02dc-434e-ac1b-0217059f06b4 · outbound

This paper cites Training Compute-Optimal Large Language Models.

On the Effectiveness of Incremental Training of Large Language Models Training Compute-Optimal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.090479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.090479Z digest=sha256:236a0a0e57ef13ea82dda7170c7f85936c9ac25ebba89f89f7a099f67200b555

Observation 980ea5dc-d5f2-4381-a8f7-06924a65c315 · outbound

This paper cites The Llama 3 Herd of Models.

On the Effectiveness of Incremental Training of Large Language Models The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.094732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.094732Z digest=sha256:421cf08d1681aea0d6d7c7037226c1d9347ec957e0ec43ab2dbadbd9f5d27e1e

Observation 25ee6e12-78f2-4a44-9998-5e19b622a658 · outbound

This paper cites A fast learning algorithm for deep belief nets,.

On the Effectiveness of Incremental Training of Large Language Models A fast learning algorithm for deep belief nets,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.099131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.099131Z digest=sha256:249d8d402ad64fc89a35d32102062957c4502650378c9b39e2d713e99b4241fe

Observation d8ab18b5-2a3a-4326-a982-14b1dbc99b3d · outbound

This paper cites Greedy layer- wise training of deep networks,.

On the Effectiveness of Incremental Training of Large Language Models Greedy layer- wise training of deep networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:07:17.523820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:07:17.103345Z digest=sha256:77c74f58b8a94dcfcfa391bab12c5453f31867445a5802873e0b513851c93b84

Observation c8871e9c-a191-41c6-bfd5-3736c3830568 · outbound

This paper cites The cascade-correlation learning architec- ture,.

On the Effectiveness of Incremental Training of Large Language Models The cascade-correlation learning architec- ture,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:07:17.510236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:07:17.107182Z digest=sha256:228174a3b927a972e4426edefa282de4a5f799577dca38dd076b7dfc32c1a968

Observation ee8b7273-9829-42aa-aa6e-efd1c33d98ab · outbound

This paper cites Improving language models by retrieving from trillions of tokens.

On the Effectiveness of Incremental Training of Large Language Models Improving language models by retrieving from trillions of tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.110794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.110794Z digest=sha256:09cc7e3766e10337072bf7ca480cac62e27b0d3e9ac2dd67b989bd9e860ddf98

Observation ac1514ad-4586-4dcf-b5a2-75d08f926f6e · outbound

This paper cites Analyzing hidden representations in end- to-end automatic speech recognition systems,.

On the Effectiveness of Incremental Training of Large Language Models Analyzing hidden representations in end- to-end automatic speech recognition systems,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:07:17.496896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:07:17.114888Z digest=sha256:41c6ed175e14af29ee378ff863f8aac686dde063dd4d67e33ab129a55ea36af1

Observation def03a05-a250-46df-94d8-c585e739b61d · outbound

This paper cites The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives.

On the Effectiveness of Incremental Training of Large Language Models The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.118607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.118607Z digest=sha256:a90f119994576cf5e8a649c9468113085520f3c7468f2951ef14e870adab36e4

Observation 35a86151-da4d-4fd4-a6ed-57636614eb5d · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

On the Effectiveness of Incremental Training of Large Language Models Transformer Feed-Forward Layers Are Key-Value Memories

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.122588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.122588Z digest=sha256:c0cbc4f9b50b61aed54d3d4bbf72409603326f70692b17d02fa71ee0edcb0297

Observation dc3fc014-5cb3-4fef-9662-7f08d269cbe8 · outbound

This paper cites Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability,.

On the Effectiveness of Incremental Training of Large Language Models Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:07:17.485310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:07:17.126445Z digest=sha256:f2517226f3fc2c2620acdc26d197fb1b31248534fa4fa1b8e62d4605ab9f6907

Observation 254cd807-f22d-46a5-b661-d8cd9626cad5 · outbound

This paper cites Visualizing and understanding convolutional networks,.

On the Effectiveness of Incremental Training of Large Language Models Visualizing and understanding convolutional networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:07:17.472343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:07:17.129855Z digest=sha256:2338bfe59884196760477447758ea43b08e0fe136b76175aae97d1ef7ce8509e

Observation fe957c28-c9de-4056-bd4b-9c6c96da2ad2 · outbound

This paper cites Qwen2 Technical Report.

On the Effectiveness of Incremental Training of Large Language Models Qwen2 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.133462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.133462Z digest=sha256:5a881ac392968ee585cb2dcf420dbc7a4722ea97e2cc0f7f1cd48ebf202d0ff2

Observation 08b23097-01bd-4b1a-8f60-41a3e2c3e273 · outbound

This paper cites Why does unsuper- vised pre-training help deep learning?.

On the Effectiveness of Incremental Training of Large Language Models Why does unsuper- vised pre-training help deep learning?

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:07:17.459738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:07:17.137562Z digest=sha256:228911502b175edf2cf088d0999fd51ac9ac336b8e3dd92797dfb16b2e870f4d

Observation 53a07be4-1a81-4739-ad62-02493786a540 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

On the Effectiveness of Incremental Training of Large Language Models The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.141100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.141100Z digest=sha256:90aec1e4d8c5f917d1678091c1140a74ce4fe071844a23b900c513c1e1ca7929

Observation 8b543ff2-8fd8-4520-affc-87eec956b3e2 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

On the Effectiveness of Incremental Training of Large Language Models Distilling the Knowledge in a Neural Network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.145022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.145022Z digest=sha256:93f80c47d518dfd0273e9eba8303a73a8815d3762cb00f1569e155eba9daace7

Observation b3b32f80-6923-4e83-a35f-758e6b4c2edc · outbound

This paper cites Mixed Precision Training.

On the Effectiveness of Incremental Training of Large Language Models Mixed Precision Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.148936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.148936Z digest=sha256:257dc1118f5ecf4541be5d803ab96a9bf1d8c97086df7e560f2f18efe05df34e

Observation 04acfd1d-844b-4814-b545-67c71f54d203 · outbound

This paper cites Universal Language Model Fine-tuning for Text Classification.

On the Effectiveness of Incremental Training of Large Language Models Universal Language Model Fine-tuning for Text Classification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.152917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.152917Z digest=sha256:98b185493fe7678599a1aecbc218fd0e480a2e8d9732b8beafa629049fb8b4d2

Observation b95f1d8f-9be0-4227-913a-e7951f3a350c · outbound

This paper cites Progressive Neural Networks.

On the Effectiveness of Incremental Training of Large Language Models Progressive Neural Networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.157078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.157078Z digest=sha256:c7eaf6f74e534e0d465f0d136c31d85cd42f2918509a7fd38acf120a8a8cf24d

Observation d2c69e01-9f59-46fd-94e0-52803e39fe34 · outbound

This paper cites An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks.

On the Effectiveness of Incremental Training of Large Language Models An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.161309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.161309Z digest=sha256:4b7455e6c054904e307ffb38e4ca8e78c8d210ebd59278b7dff6475898424970

Observation bd30d1d6-9297-42ec-b819-f5ceff22f5c8 · outbound

This paper cites How transferable are features in deep neural networks?.

On the Effectiveness of Incremental Training of Large Language Models How transferable are features in deep neural networks?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.166435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.166435Z digest=sha256:a2d894e43ae5e78b785f1ae75cabd5031b4bc8229a437633211fb85b22f01241

Observation f5ea9ea4-19ef-46ed-9985-cfc563d25274 · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

On the Effectiveness of Incremental Training of Large Language Models Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.170606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.170606Z digest=sha256:ce5b2f7898e68a28d86b14e0148592dafe74a7c2bc306df16222e85803ef0ded

Observation a86c7fe5-b006-41f8-a950-cbd05c080d48 · outbound

This paper cites Fineweb- edu,.

On the Effectiveness of Incremental Training of Large Language Models Fineweb- edu,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:07:17.439285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:07:17.174737Z digest=sha256:710db423d921f837e71374ae42715c0f4407d0c9dd2804ffd7b3571bdb4b0463

Observation 46459acf-7dc4-488a-8c36-842f54d2ee83 · outbound

This paper cites Decoupled Weight Decay Regularization.

On the Effectiveness of Incremental Training of Large Language Models Decoupled Weight Decay Regularization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.178636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.178636Z digest=sha256:ac589465eab3836b3032e61f2576b15f0c4e4c835f98912df5d7c69724d92452

Observation 87a60cc8-8eb3-4df5-8417-d28617c884ca · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

On the Effectiveness of Incremental Training of Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:17.183757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:17.183757Z digest=sha256:2c84ab2900de2f57bc6ec590790e666d864f741dff841aaae8b10d59c525d955

Observation bc76df30-b020-4876-a902-00052a076e06 · outbound

This paper cites Deep feedforward net- works,.

On the Effectiveness of Incremental Training of Large Language Models Deep feedforward net- works,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:07:17.426541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:07:17.187777Z digest=sha256:43e323263671dad504edcbffb77aaafa161589bc03a12409f285121dddc0afa1

Pith citing papers

No inbound Pith citation observations are available.