Pith. sign in

Paper Citation Record · LEDGER

Drop Dropout on Single-Epoch Language Model Pretraining

As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2505.24788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24788 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:18:56.233445Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T05:38:12.882914Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T05:41:02.181700Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec892921-0533-4864-8a44-1ec55ded0f7b · outbound

This paper cites InProceedings of the 2021 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 5484–5495, Online and Punta Cana, Dominican Republic.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 2021 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 5484–5495, Online and Punta Cana, Dominican Republic

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.303745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:54.922866Z digest=sha256:863a7fee9315fadc347e24916e94d166510c4f4d8940b6804c64971b1072a291

Observation d0fc1467-7b46-4b9d-8151-76d043095499 · outbound

This paper cites Association for Computational Linguistics.

Drop Dropout on Single-Epoch Language Model Pretraining Association for Computational Linguistics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.917281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:55.307375Z digest=sha256:b33b4ad4267492114feba32bc8366a30add6d643373efa47602a0b28dc1c4ef5

Observation 8b9aed26-2221-4d6c-8020-ac3c1db84b21 · outbound

This paper cites an unresolved cited work.

Drop Dropout on Single-Epoch Language Model Pretraining Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:18:57.569992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:55.555633Z digest=sha256:f5882a43bb76a840510f531b04fe6da8ff0ce4f314e7e045b8368aa66befb10d

Observation 9c845021-90ce-4769-b02e-54015936c838 · outbound

This paper cites InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.370618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:55.697998Z digest=sha256:739eccae3538cc2b3121c9311cf432a6a59b067528eff0d8860d55a2a4182c49

Observation 60612f92-7e0d-45af-a219-1073cb261708 · outbound

This paper cites An adapted version of the official evaluation script was used to obtain the dev-slice results reported in this work.

Drop Dropout on Single-Epoch Language Model Pretraining An adapted version of the official evaluation script was used to obtain the dev-slice results reported in this work

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.523414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:56.183766Z digest=sha256:5c62161988f758fc6d98514209737f850395e312e7af2e3d1d5e4f314d7ffaae

Observation 97e2a51e-cf1b-4b21-9de7-94b0fdac0da8 · outbound

This paper cites When MLP dropout is used, p= 0.1.

Drop Dropout on Single-Epoch Language Model Pretraining When MLP dropout is used, p= 0.1

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.207353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:55.880575Z digest=sha256:16c67fe969dfdb381736fd16db9cb84a8ad8972f877a68e5ef1c65a507cd52b7

Observation 61ab845e-e8a5-4776-80d5-d0e8cf75666c · outbound

This paper cites Batching was done sequentially with the Pytorch Data Loader, sequence lengths are capped at 512 tokens.

Drop Dropout on Single-Epoch Language Model Pretraining Batching was done sequentially with the Pytorch Data Loader, sequence lengths are capped at 512 tokens

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.970449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:55.962071Z digest=sha256:fa9fa75ae0fdb7b8e81efad0be2d07502926d26dc32453065c13443b2ea283c3

Observation 20f42294-0e67-46a4-8297-ab6a97aa1076 · outbound

This paper cites Optimization was donewith regulariza- tionusing AdamW (Loshchilov and Hutter,.

Drop Dropout on Single-Epoch Language Model Pretraining Optimization was donewith regulariza- tionusing AdamW (Loshchilov and Hutter,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.712805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:56.060676Z digest=sha256:d46485d476f7e620d0339cc5de981b0df713f51cf4eccd67c2789b2042f88165

Observation 6d2db1f3-db5b-493c-9e77-02120fb5abc2 · outbound

This paper cites Batch size was set to 128, and dropout rate was set to 0.15 regardless of whether pretrain- ing the BERT model used dropout consistent with previous approaches.

Drop Dropout on Single-Epoch Language Model Pretraining Batch size was set to 128, and dropout rate was set to 0.15 regardless of whether pretrain- ing the BERT model used dropout consistent with previous approaches

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.414228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:56.233445Z digest=sha256:f2fe09480ea25f0c03372d27419f8dbd35af021e04bf6f00d835eb961626fd12

Observation c64bc1a1-f446-49a1-8444-2caa456e0456 · outbound

This paper cites Improving neural networks by preventing co-adaptation of feature detectors.

Drop Dropout on Single-Epoch Language Model Pretraining Improving neural networks by preventing co-adaptation of feature detectors

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.233415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.233415Z digest=sha256:db548f3378eb09783c26e61bd635d32d6b52738ace1c477c40c1273669b7ad04

Observation 24b0913e-3420-446b-9874-bdd1cad2b9f2 · outbound

This paper cites Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al.

Drop Dropout on Single-Epoch Language Model Pretraining Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.722046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:55.402456Z digest=sha256:92f3e34dae84738eb9460a30ac04955b0718bc57619873e4f40742a61119dde0

Observation 8daa68c7-a906-4998-9fb2-58914a0172b7 · outbound

This paper cites InProceedings of the 8th Workshop on Cognitive Modeling and Com- putational Linguistics (CMCL 2018), pages 10–18, Salt Lake City, Utah.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 8th Workshop on Cognitive Modeling and Com- putational Linguistics (CMCL 2018), pages 10–18, Salt Lake City, Utah

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.117199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:55.139696Z digest=sha256:064e82fc565f537674d40799c7540ae3ce155b59623d54a25a71959249248843

Observation a843f7c7-d1b9-4eec-8491-b4946b423fa7 · outbound

This paper cites an unresolved cited work.

Drop Dropout on Single-Epoch Language Model Pretraining Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:18:58.481051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:54.555900Z digest=sha256:4fd1327ed27b5b06740a3146bef6757fbfeda06ced91255713131c42716b824a

Observation 856c6ebb-6a4b-4186-ae65-14f8d09836ee · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Drop Dropout on Single-Epoch Language Model Pretraining The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:54.703706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:54.703706Z digest=sha256:631c7bf319ebe93b3f2d425957e60113475773404691c0f6e5c3e6d217f37c3f

Observation 5d8ec641-5d65-4c91-a58c-9eddb65b716e · outbound

This paper cites InPro- ceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 6491– 6506, Online and Punta Cana, Dominican Republic.

Drop Dropout on Single-Epoch Language Model Pretraining InPro- ceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 6491– 6506, Online and Punta Cana, Dominican Republic

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.666912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:18:54.415114Z digest=sha256:6868c28122467696d98a6d2aa1e8b3e1eb6ceecd9dc8756d9da8f387c01c7918

Observation 3fe71e64-6957-4a44-bb46-2f501760d989 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Drop Dropout on Single-Epoch Language Model Pretraining LLaMA: Open and Efficient Foundation Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.499541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.499541Z digest=sha256:18007007f19bbc59adbcf5262dbb43e99157e27cdce67457769be9be5a0e99e8

Observation 3f795873-d19c-46f8-9103-0ee8f35cf793 · outbound

This paper cites ReFT: Representation Finetuning for Language Models.

Drop Dropout on Single-Epoch Language Model Pretraining ReFT: Representation Finetuning for Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.792408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.792408Z digest=sha256:15fd7098f1fd4dc757ec74264d2e92fa359c5afc8ebca26fb00c92460c13bff1

Pith citing papers

Observation 09c0af84-9f63-4978-80fb-afcb9cce6080 · inbound

Language models recognize dropout and Gaussian noise applied to their activations cites this paper.

Language models recognize dropout and Gaussian noise applied to their activations Drop Dropout on Single-Epoch Language Model Pretraining

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.183113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:38:12.882914Z digest=sha256:763c23d4ff2ba7f4cd4862b1f48c81dfc5dd4eee75ee5a4e0e2a467155baba18