Pith. sign in

Paper Citation Record · LEDGER

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 6 inbound Pith citation observations for arXiv:2506.12543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12543 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:30.649181Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:38:48.446865Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a21bd626-5831-4ece-acb5-d4840f36f7b9 · outbound

This paper cites Disentangling Adaptive Gradient Methods from Learning Rates.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Disentangling Adaptive Gradient Methods from Learning Rates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.766980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.766980Z digest=sha256:32d98243705f6874f2293f23b07775eed2d0a951677f4d3fd06a2d090b0374df

Observation 49d4b4c9-e468-4c4f-b591-c01e47235889 · outbound

This paper cites As before, the gap decreases the longer we train, and SGD can eventually outperform Adam.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling As before, the gap decreases the longer we train, and SGD can eventually outperform Adam

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:31.048280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:50:30.645921Z digest=sha256:a2f50406a9cfad37821da11e3c7f7929c71dc2d378a1888eaab0596854a08b03

Observation 42de01fd-13d3-4930-a9bc-9a2017e833c3 · outbound

This paper cites On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:30.995734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:50:30.238100Z digest=sha256:ba9974e8e9b0054ade7eec114e39a4a747c983973d3926368bb0e178aa6130cf

Observation bbf646a2-602d-4105-bf75-187a47e1043a · outbound

This paper cites Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.555421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.555421Z digest=sha256:fdedfbcd8b790fb6fc79b551a7f8254434cdcb69f80f4fa87ae718083f62e5d8

Observation faba8271-5cc9-4dd9-a404-e1ec57e7b9b5 · outbound

This paper cites How Does Adaptive Optimization Impact Local Neural Network Geometry?.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How Does Adaptive Optimization Impact Local Neural Network Geometry?

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:50:30.948721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:50:30.563102Z digest=sha256:6c0d3b2aca68267d3ca471039bbb564ff677372b045e5b557aea414412fad7e3

Observation 3ed2fc01-02dd-4f7e-831e-dbf28261e37e · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.570637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.570637Z digest=sha256:f725edbe3d8e1445e0b007c82cdc1d4b9095d1cf11e03800365df38fc3eede52

Observation 93b8ace5-b1e8-41e7-976f-b4ee1f3a29da · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.584903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.584903Z digest=sha256:56e68f6bfdc35aa118e3a8188293640f7b61f30f7ea77084b3256491b2f435c2

Observation e4f2a746-919f-4dfd-9dad-d458551d844e · outbound

This paper cites Muon is Scalable for LLM Training.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Muon is Scalable for LLM Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.594639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.594639Z digest=sha256:d3361bb74ce961cee38946f11cb92313bd30c05a639146872673ac5cde1c7c2f

Observation cf808663-47b5-4c1a-9afc-7e27f355ba34 · outbound

This paper cites On the SDEs and Scaling Rules for Adaptive Gradient Algorithms.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling On the SDEs and Scaling Rules for Adaptive Gradient Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.606742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.606742Z digest=sha256:14a6473220e4bd737910fd3fca2286ea7fc0cd0cbbc0ac708386c9d48cac2be9

Observation 3c264e59-9a2c-40dd-b301-bfb40bb9ea61 · outbound

This paper cites Orvieto and R.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Orvieto and R

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.612924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.612924Z digest=sha256:6d7cf50e10da62c010d42f7599e5987e6c2c49fb5c4f8ad7b000f2af9d61aa35

Observation c90daf0d-ba2c-4c9c-8d1c-cfba25a472cb · outbound

This paper cites Toward Understanding Why Adam Converges Faster Than SGD for Transformers.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.615997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.615997Z digest=sha256:2af8b3ef740614c60d1f35ade2f4b1182e7b16116a4972086f07dc8671277333

Observation 69e7fee2-9ad5-4733-b8d8-8f7515dc23f4 · outbound

This paper cites arXiv:2502.00213 [cs].

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling arXiv:2502.00213 [cs]

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.624491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.624491Z digest=sha256:5276d8050afe9ac2757af9edd8013a2406e44fbad1f9b8b75b52f3436949ac39

Observation ff2174f1-3d25-4dac-8c73-a60ed35c049d · outbound

This paper cites Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.627480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.627480Z digest=sha256:78eb417272f7a2726d5831765adab957e6d2907e61d93f89586194c4904db65f

Observation 37a5bffc-f69f-4701-ab14-eb9cf9c9450d · outbound

This paper cites How Does Critical Batch Size Scale in Pre-training?.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How Does Critical Batch Size Scale in Pre-training?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.630441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.630441Z digest=sha256:19abe11b52c6e63d068bc9ea32d32a19bf35150ba19348b57ce2c1743cacd625

Observation f7dcaf01-f68f-4aad-a8d0-86ec4ad73dc9 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Why Transformers Need Adam: A Hessian Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.633781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.633781Z digest=sha256:ce65dca09bd88c125d296e3a4c0bd462e973aac94f4e2c98412138bbff8303bb

Observation d6d36155-b1fa-4519-8f90-785c1d93fa03 · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Deconstructing What Makes a Good Optimizer for Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.636520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.636520Z digest=sha256:858be04de0095def03d6a962e215d90f7d5e769fa80e69a6f07118a92ecba996

Observation a976dd82-e9e3-4506-b999-0f76593a6436 · outbound

This paper cites an unresolved cited work.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:31.065026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:50:30.639985Z digest=sha256:ccb789f0d528d22102062eadf6568be0f50d529166d7768e89cb6bf9cf1dd2f6

Observation 08b0145f-91e0-48b2-ae3e-3a5593e047d9 · outbound

This paper cites [2024], and uses the codebase of Orvieto and Gower [2025].

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling [2024], and uses the codebase of Orvieto and Gower [2025]

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:31.039339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:50:30.649181Z digest=sha256:536fd46395370988621f9807c49166b174bc20c4f85931acb1cc1cb8a1de5bc1

Observation f34f1c0d-9218-4ca8-aec6-2a8131c3d5ec · outbound

This paper cites an unresolved cited work.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Unresolved cited work

Reference 1024

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:31.056492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:50:30.642968Z digest=sha256:dd4e94db4207c66089114655abbcd9c50c1463d115f275722b1c0d877c45b358

Observation f176b608-2526-4295-95c2-1538cb7575ee · outbound

This paper cites Transformers without Tears: Improving the Normalization of Self-Attention.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Transformers without Tears: Improving the Normalization of Self-Attention

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.609473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.609473Z digest=sha256:7a9ae29102558c7bfa8a80e2962e9eb359ff492063bc39145ed6e035285267be

Observation 6799d912-962d-42ed-bda7-de130b44a4a0 · outbound

This paper cites How to Fine-Tune Vision Models with SGD.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How to Fine-Tune Vision Models with SGD

Reference 2014

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:30.926338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:50:30.576108Z digest=sha256:e6ef4af870c783a8b2b33e7a0543994f939990518f9c2bb81e1626e0aeff1623

Observation c597f601-3722-465a-851c-2e2ab7b14dcb · outbound

This paper cites The Llama 3 Herd of Models.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling The Llama 3 Herd of Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.548431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.548431Z digest=sha256:5074e02c514b16ffb2275ad33b5583118900113396fa16e40fa6560e73da4619

Observation c4d8fd51-e618-4b44-9367-f940273b414c · outbound

This paper cites signSGD: Compressed Optimisation for Non-Convex Problems.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling signSGD: Compressed Optimisation for Non-Convex Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.980013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.980013Z digest=sha256:ff8ddffb1894ea38abb2a39bfe41353b27fa3e8b70cf276229f8b88b9e0ee6ff

Observation ef191c64-b70d-4e05-a88f-27d993caa9c9 · outbound

This paper cites Practical Efficiency of Muon for Pretraining.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Practical Efficiency of Muon for Pretraining

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.618825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.618825Z digest=sha256:7b22640769035c6fa95c21de1f78c1ff25d55ddcafae780e585fdf74e1ed9089

Observation 4f04113b-7173-4087-a8fc-853e94d5a18b · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.464304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.464304Z digest=sha256:015dd7357590453bbfb968a16124da63aeaeb8464ae2d295b538d84757802d3c

Observation 72b5ec18-fdd9-4bf8-b615-731d1aa0231f · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.603923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.603923Z digest=sha256:0337dacab9014ed725fb718dc1851d2e951bb93ae30f237a4542bce7abb2d04a

Observation a959df74-d31d-472b-b5f2-150c36d047b8 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.348876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.348876Z digest=sha256:46837a6d6d11214cb87f0f0bc394c185485e808f1b6dc0a24e5fc2fcc126dc16

Observation 6cec26bc-75cf-40f6-8c17-d216703fe868 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.096081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.096081Z digest=sha256:b961a52e0d87d58d93e8c3146bbdd1528165bb237fa58c663d40545ec364be70

Observation 3c56be0f-8610-4504-8510-4e9111b6f761 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.838100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.838100Z digest=sha256:063fec530be8006f8daff7c13d38caeea707ef765dcee1a12e055d320255f56c

Observation 0da4f519-d61b-47a1-8033-3430a73bbbff · outbound

This paper cites GLU Variants Improve Transformer.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling GLU Variants Improve Transformer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.621589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.621589Z digest=sha256:996aa38ba3bca2f65aec03359d124c2a4cf49e22c791ac79bca79e3f82afaa74

Pith citing papers

Observation 34bb4621-ccf2-4792-bf5c-1e3ccfebbc71 · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:48.446865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:48.446865Z digest=sha256:4cb753f15742f85ef0546451f4670d9f30ceb461a65f01d20fd1807245c39432

Observation 0d414ffe-531d-4c19-a57a-892ac77a246f · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.656518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:219db54bca81ab77d51bfa335c01d53dd2f64635a947f2eab5f5e4a6a54e9f7e

Observation 01bd3cbc-d651-4a8c-bf5a-ad2e192dce3d · inbound

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates cites this paper.

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.427606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T13:34:09.129379Z digest=sha256:500f4d7ac6de046857721da8ccd9479192c6db903860d8aacd76baa080b45d96

Observation 53449721-9569-4766-a9a1-050f989768bf · inbound

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics cites this paper.

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.363127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T08:06:52.309619Z digest=sha256:dbcd91ff4b8345336f5872402857e11fd478f12bd189dc4ac3b21b8cb8cf4fcf

Observation 3aa9feae-5033-46d6-8a2e-f130184609b1 · inbound

A Theoretical and Experimental Study of a Novel Adaptive Learning Algorithm cites this paper.

A Theoretical and Experimental Study of a Novel Adaptive Learning Algorithm Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:23:16.439125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T09:18:31.526770Z digest=sha256:a1655aa8d0287f0cb6352c021da6a2aac0155ad8c846e2e2062f9ce5206b926d

Observation e9e7170e-6db5-4851-b0e2-92c553ad4910 · inbound

Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior cites this paper.

Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-26T08:49:14.928649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:45:34.884703Z digest=sha256:f3b549f856ae60197cd21d913d914173215065933f4035e00081ab922c2fd07a