Pith. sign in

Paper Citation Record · LEDGER

Drop Dropout on Single-Epoch Language Model Pretraining

As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2505.24788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24788 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:18:56.233445Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T05:38:12.882914Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T05:41:02.181700Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec892921-0533-4864-8a44-1ec55ded0f7b · outbound

This paper cites InProceedings of the 2021 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 5484–5495, Online and Punta Cana, Dominican Republic.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 2021 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 5484–5495, Online and Punta Cana, Dominican Republic

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.303745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:54.922866Z digest=sha256:d2bdf90ea0df33593139e9629f0970c8f69057ead539118a6d40ee335c4b27b3

Observation d0fc1467-7b46-4b9d-8151-76d043095499 · outbound

This paper cites Association for Computational Linguistics.

Drop Dropout on Single-Epoch Language Model Pretraining Association for Computational Linguistics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.917281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:55.307375Z digest=sha256:63c8a1a44c1c9dc3bef2a77481fa0ed713fa8bc6c362ff34811ec42553bacb2f

Observation 8b9aed26-2221-4d6c-8020-ac3c1db84b21 · outbound

This paper cites an unresolved cited work.

Drop Dropout on Single-Epoch Language Model Pretraining Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:18:57.569992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:55.555633Z digest=sha256:86840dfa9d0190425955fdd39cdb28f2aec43ad59087bfa09dba0c2fcff584cd

Observation 9c845021-90ce-4769-b02e-54015936c838 · outbound

This paper cites InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.370618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:55.697998Z digest=sha256:7bc1a510676cbb4781a3b25b5314b585d9efb8c16eeba65471b6408cf718bb34

Observation 60612f92-7e0d-45af-a219-1073cb261708 · outbound

This paper cites An adapted version of the official evaluation script was used to obtain the dev-slice results reported in this work.

Drop Dropout on Single-Epoch Language Model Pretraining An adapted version of the official evaluation script was used to obtain the dev-slice results reported in this work

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.523414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:56.183766Z digest=sha256:1990c28c302dba03e5d0f4c7607b23dc945b7ba4a3d2aca96f8079b2daefdcce

Observation 97e2a51e-cf1b-4b21-9de7-94b0fdac0da8 · outbound

This paper cites When MLP dropout is used, p= 0.1.

Drop Dropout on Single-Epoch Language Model Pretraining When MLP dropout is used, p= 0.1

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.207353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:55.880575Z digest=sha256:8ea5f0c883482b7bcfd75d059e8e485ac14de99bb9d91997c7a890e7e9a1abba

Observation 61ab845e-e8a5-4776-80d5-d0e8cf75666c · outbound

This paper cites Batching was done sequentially with the Pytorch Data Loader, sequence lengths are capped at 512 tokens.

Drop Dropout on Single-Epoch Language Model Pretraining Batching was done sequentially with the Pytorch Data Loader, sequence lengths are capped at 512 tokens

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.970449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:55.962071Z digest=sha256:5b01c90daa495c787f3c579bca4fe7082701e147d96606422344a7897229ddde

Observation 20f42294-0e67-46a4-8297-ab6a97aa1076 · outbound

This paper cites Optimization was donewith regulariza- tionusing AdamW (Loshchilov and Hutter,.

Drop Dropout on Single-Epoch Language Model Pretraining Optimization was donewith regulariza- tionusing AdamW (Loshchilov and Hutter,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.712805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:56.060676Z digest=sha256:25bbdbcc6c375c6fb1584d9682be24fcce5159cc532e42c2ac26a169594de3cb

Observation 6d2db1f3-db5b-493c-9e77-02120fb5abc2 · outbound

This paper cites Batch size was set to 128, and dropout rate was set to 0.15 regardless of whether pretrain- ing the BERT model used dropout consistent with previous approaches.

Drop Dropout on Single-Epoch Language Model Pretraining Batch size was set to 128, and dropout rate was set to 0.15 regardless of whether pretrain- ing the BERT model used dropout consistent with previous approaches

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.414228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:56.233445Z digest=sha256:8e45cede956bfc711e234f29b5ef58eccea6acff5114f4f060f8baf4b62390a5

Observation c64bc1a1-f446-49a1-8444-2caa456e0456 · outbound

This paper cites Improving neural networks by preventing co-adaptation of feature detectors.

Drop Dropout on Single-Epoch Language Model Pretraining Improving neural networks by preventing co-adaptation of feature detectors

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.233415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.233415Z digest=sha256:db548f3378eb09783c26e61bd635d32d6b52738ace1c477c40c1273669b7ad04

Observation 24b0913e-3420-446b-9874-bdd1cad2b9f2 · outbound

This paper cites Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al.

Drop Dropout on Single-Epoch Language Model Pretraining Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.722046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:55.402456Z digest=sha256:07a5eeff089e3f37789cdb70dd92df61b6dc5af65cf3be545ad9a08b6364b258

Observation 8daa68c7-a906-4998-9fb2-58914a0172b7 · outbound

This paper cites InProceedings of the 8th Workshop on Cognitive Modeling and Com- putational Linguistics (CMCL 2018), pages 10–18, Salt Lake City, Utah.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 8th Workshop on Cognitive Modeling and Com- putational Linguistics (CMCL 2018), pages 10–18, Salt Lake City, Utah

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.117199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:55.139696Z digest=sha256:616f9449da42aa7fc1453d109bfd0ca077db1eb5e31d8693ca4d86d9b3b1ea68

Observation a843f7c7-d1b9-4eec-8491-b4946b423fa7 · outbound

This paper cites an unresolved cited work.

Drop Dropout on Single-Epoch Language Model Pretraining Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:18:58.481051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:54.555900Z digest=sha256:2095e1c8192110bef9a78a6873e8ccfc5a65b0fe71dcd19b6e91a0c630516cdf

Observation 856c6ebb-6a4b-4186-ae65-14f8d09836ee · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Drop Dropout on Single-Epoch Language Model Pretraining The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:54.703706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:54.703706Z digest=sha256:631c7bf319ebe93b3f2d425957e60113475773404691c0f6e5c3e6d217f37c3f

Observation 5d8ec641-5d65-4c91-a58c-9eddb65b716e · outbound

This paper cites InPro- ceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 6491– 6506, Online and Punta Cana, Dominican Republic.

Drop Dropout on Single-Epoch Language Model Pretraining InPro- ceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 6491– 6506, Online and Punta Cana, Dominican Republic

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.666912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:18:54.415114Z digest=sha256:24bb60290087be34f0c50715783e47b2a0bef333e4adc20686dae646e66b8de2

Observation 3fe71e64-6957-4a44-bb46-2f501760d989 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Drop Dropout on Single-Epoch Language Model Pretraining LLaMA: Open and Efficient Foundation Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.499541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.499541Z digest=sha256:18007007f19bbc59adbcf5262dbb43e99157e27cdce67457769be9be5a0e99e8

Observation 3f795873-d19c-46f8-9103-0ee8f35cf793 · outbound

This paper cites ReFT: Representation Finetuning for Language Models.

Drop Dropout on Single-Epoch Language Model Pretraining ReFT: Representation Finetuning for Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.792408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.792408Z digest=sha256:15fd7098f1fd4dc757ec74264d2e92fa359c5afc8ebca26fb00c92460c13bff1

Pith citing papers

Observation 09c0af84-9f63-4978-80fb-afcb9cce6080 · inbound

Language models recognize dropout and Gaussian noise applied to their activations cites this paper.

Language models recognize dropout and Gaussian noise applied to their activations Drop Dropout on Single-Epoch Language Model Pretraining

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.183113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:38:12.882914Z digest=sha256:dffc250cd84d7b2b4a12e296610b1fa43e1cab4868acc1eb0d8fee662e252568