Pith. sign in

Paper Citation Record · LEDGER

Large Batch Training of Convolutional Networks

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 64 inbound Pith citation observations for arXiv:1708.03888.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1708.03888 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 64 of 64 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:59:27.531983Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

509
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9cb22c49-6e9f-4d33-bfb5-b261ac09eba2 · inbound

Large Batch Optimization for Deep Learning: Training BERT in 76 minutes cites this paper.

Large Batch Optimization for Deep Learning: Training BERT in 76 minutes Large Batch Training of Convolutional Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:39:00.060875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T21:38:59.970003Z digest=sha256:d32cd83e1fa7f6f29cfa11cec21f0c4aea63b8434c6cc49c07a8f43507e69cb1

Observation 1311239e-a50a-407a-8c4c-9e63160af736 · inbound

Gradient Noise Convolution (GNC): Smoothing Loss Function for Distributed Large-Batch SGD cites this paper.

Gradient Noise Convolution (GNC): Smoothing Loss Function for Distributed Large-Batch SGD Large Batch Training of Convolutional Networks

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-25T15:45:59.730986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T15:41:27.232807Z digest=sha256:0d16b055ede84c6d54aae5aaf81fb61e3ba9a481dad70a4138e6777a89e62978

Observation 93bf0d58-8017-47f6-93a3-71a7943b59e0 · inbound

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models cites this paper.

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models Large Batch Training of Convolutional Networks

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:24:35.887384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:24:35.826882Z digest=sha256:08ace3a6633300b47e7def95ef3760e03c5a69dacb2b06f2372855acf9ebb0cf

Observation bff25c17-cc58-4d49-a48b-490f377e4a08 · inbound

Solving Rubik's Cube with a Robot Hand cites this paper.

Solving Rubik's Cube with a Robot Hand Large Batch Training of Convolutional Networks

Reference 118

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:38:28.978739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:38:28.621842Z digest=sha256:37b8cc7ba4ab9167ae08547dab2cf8d9112c533abb74eb45f607a9ca2a54f81c

Observation 94dedb8c-ae06-426b-8262-7c30943c2e5e · inbound

A Simple Framework for Contrastive Learning of Visual Representations cites this paper.

A Simple Framework for Contrastive Learning of Visual Representations Large Batch Training of Convolutional Networks

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:31:53.353223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T18:31:53.283456Z digest=sha256:eb3a85ca072354f47442b81e8bae254a5d8aa6b90de6c9e7f234b6e939e8e8fe

Observation 529719b8-9929-496a-bdf8-7c9b5c15a5f9 · inbound

Scaling Laws for Transfer cites this paper.

Scaling Laws for Transfer Large Batch Training of Convolutional Networks

Reference 155

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:58:13.495889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T00:58:13.116663Z digest=sha256:f38240cba55698f1eaaf5138ed288ff1976742cec2fd5a94256f23cdfdd7776e

Observation 43b6046d-a27a-4921-811b-d222bf9b1490 · inbound

VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning cites this paper.

VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning Large Batch Training of Convolutional Networks

Reference 105

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T15:29:28.894899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T15:29:28.776822Z digest=sha256:a091139a72edf22703df20264c0338bdd5c5fd4c603cfce2a8d8effa128598ee

Observation a773b300-cc4d-49df-ab69-af2a7dac7bdb · inbound

Masked Autoencoders Are Scalable Vision Learners cites this paper.

Masked Autoencoders Are Scalable Vision Learners Large Batch Training of Convolutional Networks

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:53:57.559273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:53:57.519348Z digest=sha256:2330df7a88f9a3cf8fd38d3767b62a4b43c33c39ee746a5d70c1dfca6c42adb7

Observation 3f7fc28c-706e-4cac-9485-d546d62ff8ec · inbound

A General Language Assistant as a Laboratory for Alignment cites this paper.

A General Language Assistant as a Laboratory for Alignment Large Batch Training of Convolutional Networks

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:22:58.919879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T14:22:57.925354Z digest=sha256:cbb09fa6d6e038bc5170705a441096a48b616b20f56354be03d57dfa2e3d2ea2

Observation d28d8083-a7af-49c7-8c93-f289aed0591e · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know Large Batch Training of Convolutional Networks

Reference 275

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:42:48.032070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:78ae61ee8935f48d994e569ad38c5560a1294187c901844ebc153fff1a5ac5dc

Observation 3d146861-c72a-4e03-8cbe-6e256bb37de8 · inbound

Vision Transformers Need Registers cites this paper.

Vision Transformers Need Registers Large Batch Training of Convolutional Networks

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:41:38.263721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T09:41:37.937046Z digest=sha256:4dd5224a1b19ee01f7eb85e742840e3f794b56333ef4b2a7ea007f74529501c9

Observation b35ec985-2ec0-4cae-a89d-2802e85bd033 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video Large Batch Training of Convolutional Networks

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T12:40:23.973006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:af1e5f3ee71782d73236413359316f28db22c33ae670bdc0d54cae606b517b65

Observation 407ae796-a0b5-493e-9aad-65210c420069 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video Large Batch Training of Convolutional Networks

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:23.801636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:b67555708e60b2fe2fe03b44407f4a6ef7970fb5977fc5012471b3b7c591a1f6

Observation d15cf685-93f5-4056-9c13-7d31334ba6f6 · inbound

Training Deep Learning Models with Norm-Constrained LMOs cites this paper.

Training Deep Learning Models with Norm-Constrained LMOs Large Batch Training of Convolutional Networks

Reference 215

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:22:37.139631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T21:22:36.870292Z digest=sha256:9cab72e7f75d7e285287e5ac174b75184c8ded690bf4dc647c3ddef007364e78

Observation 0cb11246-b373-4c03-b683-aeadd8269646 · inbound

A Physics-Inspired Optimizer: Velocity Regularized Adam cites this paper.

A Physics-Inspired Optimizer: Velocity Regularized Adam Large Batch Training of Convolutional Networks

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T14:34:54.229546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T14:34:37.468180Z digest=sha256:108bd2d77214da64a4e6c35e5a266778a111cdd3f0a8bc73a53263153862e97e

Observation a1f2a9c6-63bc-434c-b70e-d4e9b35d6ff6 · inbound

Energy Consumption in Parallel Neural Network Training cites this paper.

Energy Consumption in Parallel Neural Network Training Large Batch Training of Convolutional Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:59:27.531983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:59:27.531983Z digest=sha256:bb2f78ce059535f32a5bb4d56f7436a046e96dbdc6d762cc46130fb4ca5fe2b6

Observation 6b69d579-96dc-4258-9a3f-02eab965d4c2 · inbound

On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning'' cites this paper.

On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning'' Large Batch Training of Convolutional Networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:24.631571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:34:24.631571Z digest=sha256:c983605df3327ec264dfce0404c02999c7a5758ac5e203afd58d6e5a68fbed9c

Observation b9edcc1e-5cd5-493e-bc73-98be6fd1e637 · inbound

Closed-Form Last Layer Optimization cites this paper.

Closed-Form Last Layer Optimization Large Batch Training of Convolutional Networks

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T10:01:13.454551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T09:58:50.815850Z digest=sha256:95a1db7292e20cd260f3e4ecb0a12455f1c915812b70306794c0a00d47b729fe

Observation ff8fd775-01d0-4378-ac00-102e61f6f019 · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Large Batch Training of Convolutional Networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:46.832837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:46.832837Z digest=sha256:d6fbb606116135ba5a5a5d1d3e9d0db4e7db9a0199174e109d96fee40fb7cda9

Observation 458c65d6-4e3f-4e86-bf90-5147e89a4716 · inbound

Self-Supervised Learning with a Multi-Task Latent Space Objective cites this paper.

Self-Supervised Learning with a Multi-Task Latent Space Objective Large Batch Training of Convolutional Networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:20.720337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:08:20.720337Z digest=sha256:04918558c4dc708baea39636e1dcd4cb832fc7be22b15b7d8a496d6d3ebf7f31

Observation 86c4113e-abc7-4574-9849-6901a4b1bc8f · inbound

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation cites this paper.

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation Large Batch Training of Convolutional Networks

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:20:06.963846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:15:57.019720Z digest=sha256:e3721febc9c7e3c15b532c1ecc7a156903c89df5678a1fcd988f8a54d016a035

Observation 3dc82b7b-110e-4a83-8a0c-fc2a11bfcd7c · inbound

AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models cites this paper.

AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models Large Batch Training of Convolutional Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T23:54:54.872017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:54:54.872017Z digest=sha256:36bff7ffd11c8f33977225e8b669697cc628854e0a91f1915efc4d397a28ccfd

Observation 14291e2f-f549-4b9e-96b7-ea0014d8683c · inbound

Communication-Efficient Gluon in Federated Learning cites this paper.

Communication-Efficient Gluon in Federated Learning Large Batch Training of Convolutional Networks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:03.189861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T15:16:15.739456Z digest=sha256:f0f99cc7d9e3a5f5f17c8256843531d99a8f1adeb1937249c03ba91aba39c0da

Observation 05852816-1d02-49e7-9687-aca73fbcff49 · inbound

When Losses Align: Gradient-Based Composite Loss Weighting for Efficient Pretraining cites this paper.

When Losses Align: Gradient-Based Composite Loss Weighting for Efficient Pretraining Large Batch Training of Convolutional Networks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:54.991024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:11:37.361024Z digest=sha256:15d8c1ee50f1593ef93efa3255fec1b50d3f75bc6f4a31f260f946c751a94f13

Observation 18b279e6-a61d-4ca9-b15a-30667080af38 · inbound

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling cites this paper.

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling Large Batch Training of Convolutional Networks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:55.478901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:47:08.766380Z digest=sha256:093329881cbc1fe7867c9eabc66ec57f357fb5106bf6997191cdd6043f2a24ee

Observation be0c4285-9882-4b3f-b5a2-3b5d56e794b2 · inbound

Consolidation-Expansion Operator Mechanics:A Unified Framework for Adaptive Learning cites this paper.

Consolidation-Expansion Operator Mechanics:A Unified Framework for Adaptive Learning Large Batch Training of Convolutional Networks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:21:18.615137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T03:20:27.751171Z digest=sha256:1e195a90005a2a1bd397370f3855b34cd2899d69f055b7e8f1ddb9d3795ea8af

Observation c27bcd73-5a8f-48c3-812b-743f7a7382b6 · inbound

Consolidation-Expansion Operator Mechanics:A Unified Framework for Adaptive Learning cites this paper.

Consolidation-Expansion Operator Mechanics:A Unified Framework for Adaptive Learning Large Batch Training of Convolutional Networks

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:59:27.892999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T20:55:08.166991Z digest=sha256:d2c663d691107cac0b41e8bd608a337ae72eb7e7069bb462ec108b1cddb03ea4

Observation c4c6cdf8-01eb-43f2-8a68-cc5bf37c183e · inbound

ShardTensor: Domain Parallelism for Scientific Machine Learning cites this paper.

ShardTensor: Domain Parallelism for Scientific Machine Learning Large Batch Training of Convolutional Networks

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:09.069725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T03:03:32.316700Z digest=sha256:ba547ea09671a312cdd11d39a0895fc5dfd19f8a593fde718e4a801cb77b3145

Observation 054a41a5-49a8-4afa-a1a1-181b840f75cf · inbound

Information theoretic underpinning of self-supervised learning by clustering cites this paper.

Information theoretic underpinning of self-supervised learning by clustering Large Batch Training of Convolutional Networks

Reference 154

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:22:23.144321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T06:22:03.439595Z digest=sha256:5376d3c9f89b6278d85074940d16e36aa313113f55d7d4010e3ff6273fed5f22

Observation af435a0d-41de-4723-964e-87085347eeed · inbound

Convergence of difference inclusions via a diameter criterion cites this paper.

Convergence of difference inclusions via a diameter criterion Large Batch Training of Convolutional Networks

Reference 136

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T02:23:31.753896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T02:21:30.228735Z digest=sha256:d6f03e6dc8734d90de45f7f8e5951e08ed2c9a8ec5b8b6cb7d30316db32dee82

Observation 5992d4d7-63d5-4d7c-873c-51f9aa8fee84 · inbound

Rethinking Neural Network Learning Rates: A Stackelberg Perspective cites this paper.

Rethinking Neural Network Learning Rates: A Stackelberg Perspective Large Batch Training of Convolutional Networks

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:43:06.680507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T14:42:45.648114Z digest=sha256:8554888d35a16de80e12713697fb15337a43665b3186e85fa0dd076aabaea505

Observation 800d998a-0b4c-42c8-b31c-11643c796af4 · inbound

Rethinking Neural Network Learning Rates: A Stackelberg Perspective cites this paper.

Rethinking Neural Network Learning Rates: A Stackelberg Perspective Large Batch Training of Convolutional Networks

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:55:01.066572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T19:53:45.109784Z digest=sha256:e03e53db28b68adb8e6488659e91209bcd59961f972a117656d6fa0d661620b7

Observation 8ff62883-6c78-4b02-abf8-2d2bcaab760d · inbound

Accelerated Gradient Descent for Faster Convergence with Minimal Overhead cites this paper.

Accelerated Gradient Descent for Faster Convergence with Minimal Overhead Large Batch Training of Convolutional Networks

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T21:23:44.641716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T21:19:33.691048Z digest=sha256:4e9fa757a71813b811d81e14fd02e3bc24d628f275addc3bdb39c5af877ccc95

Observation 623c9667-4669-4c30-842f-f3ba9ee69d3a · inbound

PEIRA: Learning Predictive Encoders through Inter-View Regressor Alignment cites this paper.

PEIRA: Learning Predictive Encoders through Inter-View Regressor Alignment Large Batch Training of Convolutional Networks

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:43:19.538540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T13:41:52.182460Z digest=sha256:ed143168dc915e084dfb09c8a97e0c8a43cc6c5ccc1437872f17f368c94e9d63

Observation c9aab725-fe4d-4406-a151-afae85ca14fc · inbound

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates cites this paper.

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates Large Batch Training of Convolutional Networks

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:38:19.434014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T13:34:09.129379Z digest=sha256:213bc700c3267692398079e9f7845d164549ec7bb39b5db0754ba353cbfb3083

Observation ec4d4292-9a21-4a62-8103-029687114936 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Large Batch Training of Convolutional Networks

Reference 163

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:38:11.041872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T09:34:45.186929Z digest=sha256:68c7bccb7285297f10ad3718304dbd44a0b957f1f7793de3e880ee5bf86e5b1a

Observation 5f476e18-4e91-49d4-ba50-ad5f99d69804 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Large Batch Training of Convolutional Networks

Reference 166

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.321169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T18:42:01.854481Z digest=sha256:1de94fdc42d61946644dea734920f9efbd6a20e5a2c97b042ebc850e5abd5ae6

Observation dbfbd077-a323-4578-9cf2-f32a130551c6 · inbound

Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise cites this paper.

Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise Large Batch Training of Convolutional Networks

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T08:53:10.394204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T08:48:42.019359Z digest=sha256:493d92e777e16c19b385bd5b69f036507acf612c1148f3ea33b45c271aff1467

Observation 7bb80b6d-5e53-4ad4-8955-0d2c87b7758e · inbound

Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise cites this paper.

Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise Large Batch Training of Convolutional Networks

Reference 135

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.574590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T18:27:40.390908Z digest=sha256:bf57334f34712bc23f9f7fd7d9374db40ebb28f38c6e95359de26017a9554dc6

Observation 838e2945-7c01-423e-9910-ec61d0994e08 · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling Large Batch Training of Convolutional Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:24:27.783480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:21:15.933174Z digest=sha256:edd37ffa4eb2fc53ddda1c288b938c5c9db4165515077ea4a10543ecaf3901c5

Observation f0fdad98-32cb-48ae-9555-684aa36d4bf5 · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling Large Batch Training of Convolutional Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:24:57.645906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:17:59.698127Z digest=sha256:24ff5d50ce16d5aa8fad68ce9ab3ccd95b0e8df87e63a973aaa40e72cfe014ce

Observation 8121dd3c-5625-4345-8ac2-a40633c461a1 · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Large Batch Training of Convolutional Networks

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:44:42.670722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T07:44:16.677054Z digest=sha256:f97dd1b21bb875c939619b7858cbcb51c1e5170c293882383e5ccb36628ef3a6

Observation e9a303f1-aa0d-4451-af4e-933cde503842 · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Large Batch Training of Convolutional Networks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.606238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:40d7b1ef2ac362487ba0fe81ee3a3f8078ce12b3d4eaf8ae99a8ef1551ab4df5

Observation 7d506ace-a146-4498-8e16-6eba52ff7877 · inbound

From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression cites this paper.

From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression Large Batch Training of Convolutional Networks

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:14:45.590648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T14:09:52.456340Z digest=sha256:ce8c36b834cb122b44cfccb2c9db6c5806305dedfcb84642c4b739a06a580f26

Observation cc81bd36-f1a5-4a6c-b1d1-240da1e3096b · inbound

Unified Neural Scaling Laws cites this paper.

Unified Neural Scaling Laws Large Batch Training of Convolutional Networks

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:44:03.347192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:56:43.393302Z digest=sha256:43f7e0ba72c39a7c651b618338b95a3f3b05e8e783185a192ac0f943010b3d65

Observation 02230bd6-51b6-4260-8d26-e487bbd9f6c0 · inbound

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training cites this paper.

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training Large Batch Training of Convolutional Networks

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:23:54.190890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T19:15:49.229099Z digest=sha256:7920599a60e670acf0f7709b0f2291bd9acd6133df45ab9fe51bd069c5ae5c3d

Observation a9e6114c-9e22-4647-9159-63b926f1bb1f · inbound

Gradient Perturbation: Learning to Perturb Gradients for Adaptive Training cites this paper.

Gradient Perturbation: Learning to Perturb Gradients for Adaptive Training Large Batch Training of Convolutional Networks

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T09:13:16.168939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T09:08:26.043300Z digest=sha256:7f121b5bf13621c17a6c36a231ebc4102f36d4d2547403b8715d441dc4eaac42

Observation a637fcf7-7352-472b-86de-0e33f0fab09f · inbound

Continual Visual and Verbal Learning Through a Child's Egocentric Input cites this paper.

Continual Visual and Verbal Learning Through a Child's Egocentric Input Large Batch Training of Convolutional Networks

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T07:46:46.279821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T06:40:26.958891Z digest=sha256:5ad4696f074a092f4916aa20a646410f079ba7e7b317b512ec9cfede2de309fd

Observation b3e8831b-0721-4f40-8f52-44bc08e9875e · inbound

SMI: Efficient Self-Supervised Learning via Mutual-Information-Inspired Dependency Optimization cites this paper.

SMI: Efficient Self-Supervised Learning via Mutual-Information-Inspired Dependency Optimization Large Batch Training of Convolutional Networks

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:37:24.775934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T19:41:51.921338Z digest=sha256:b93067c171c43c5abf028d99acbd2e99d4dac3d85c89fd5a8e4744357595fe3d

Observation 4e7db590-e0f2-498b-9b0d-8bca223e2b6b · inbound

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework cites this paper.

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework Large Batch Training of Convolutional Networks

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-27T18:31:07.894671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T18:26:32.883834Z digest=sha256:bd19ff12627a4485a1639a47ebceb337cd1767df3050b8bda74541f576860ed1

Observation 86d8f0e1-d7ac-4474-afad-f88e45eab2ea · inbound

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework cites this paper.

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework Large Batch Training of Convolutional Networks

Reference 129

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:07:27.215803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T18:26:32.883834Z digest=sha256:413cf9ea874a3c27e045e41a4b3fe532648b02ff89a2ca6054fb00075a38fcdb

Observation 18a77217-3ce0-4da7-b72d-1d87a00024c9 · inbound

RCAP: Robust, Class-Aware, Probabilistic Dynamic Dataset Pruning cites this paper.

RCAP: Robust, Class-Aware, Probabilistic Dynamic Dataset Pruning Large Batch Training of Convolutional Networks

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:17:57.247646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T10:13:13.967151Z digest=sha256:32eb70a0d23ca7595a29be047e328afa6b8bf968017b732c0d4171a5b611906c

Observation 417fe08a-3981-4f24-9677-20114c735241 · inbound

LEAP: Layer-skipping Efficiency via Adaptive Progression for Vision Transformer Distillation cites this paper.

LEAP: Layer-skipping Efficiency via Adaptive Progression for Vision Transformer Distillation Large Batch Training of Convolutional Networks

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:16.777414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:07:20.133335Z digest=sha256:a1be8a4b4209417a4e7e64f567edf1989277a191325d4ebb6876506cc6fbb19a

Observation adaab7df-b61d-44d3-9c83-034cc7b6ce0a · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Large Batch Training of Convolutional Networks

Reference 162

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:59:44.836737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:d1f80b7e3a1c2e0cf0d6244e67272b6a96ed83bfc0a6747a5e8f36d284f3eb1b

Observation 03497dd4-779e-458f-80a0-5f5623bb5e80 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Large Batch Training of Convolutional Networks

Reference 162

Resolution
verified exact
local_arxiv, observed 2026-07-01T18:55:59.454775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:284346a9f2e9aae29ca535695ee9cba97a008d4d8032aa030c44f501ae25eae1

Observation 92437153-2dd8-4055-a73c-30c1fa44165d · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Large Batch Training of Convolutional Networks

Reference 125

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T20:30:07.628097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:79c47c30e379f26722dfd346b31fba5eda9f5be5025526d9d690509e1a8246df

Observation f2ec9f3d-cd85-43b1-8b4a-1cd7605b4452 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Large Batch Training of Convolutional Networks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:13.181331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:13.181331Z digest=sha256:1b5865f334f3f50a195dae5a1bfe2ed9972dc4e64ef356113ed3b99feda4e30b

Observation e144ee98-9f2e-41df-8229-d3b47777bab2 · inbound

Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules cites this paper.

Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules Large Batch Training of Convolutional Networks

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:21.500959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T07:31:11.050545Z digest=sha256:84344e3fc6fbfc5e8d72e02afa2d0cf3e1cfbf12c9b649999a6d6bfb6fd74faf

Observation 3753eee2-05cc-4415-96a8-1515526e9942 · inbound

LionVote: Per-Layer Learning Rate Adaptation for Lion cites this paper.

LionVote: Per-Layer Learning Rate Adaptation for Lion Large Batch Training of Convolutional Networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T04:17:58.962415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T04:17:58.962415Z digest=sha256:b119317d4a8166f46c42f3f89914f21bf1a907536fd4758ebabfcdcd5b53ae06

Observation 3474dcef-4995-47ea-8416-90180d1f769b · inbound

M+Adam: Low-Precision Training via Additive-Multiplicative Optimization cites this paper.

M+Adam: Low-Precision Training via Additive-Multiplicative Optimization Large Batch Training of Convolutional Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T10:29:19.000404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:29:19.000404Z digest=sha256:4e66c28265021615ff8232ce5807357e7084d3f8df82ff9dbe7972e1b9b669f7

Observation c1fa8d58-4a0a-457c-9ab4-62addecdba3d · inbound

M+Adam: Low-Precision Training via Additive-Multiplicative Optimization cites this paper.

M+Adam: Low-Precision Training via Additive-Multiplicative Optimization Large Batch Training of Convolutional Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.134837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.134837Z digest=sha256:89542e2e09b8f5f5d5cfbdaa5976a3646c9f282c102d6c0b5ebd2a27fea88945

Observation 9dc59862-e3ab-48d3-892d-784b945a2ac8 · inbound

PsiLogic: Chaos-Aware Active Cancellation for Adam with a Fair Cross-Domain Benchmark cites this paper.

PsiLogic: Chaos-Aware Active Cancellation for Adam with a Fair Cross-Domain Benchmark Large Batch Training of Convolutional Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:42:05.942052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:42:05.942052Z digest=sha256:49ded3f77d2fe85baf1eb2323395ab9a7689497166b374e8fa9d09f2a2122a3b

Observation b23984c7-5cce-4ad0-a718-a220fb885fa6 · inbound

Backpropagation-Free Trunk Training via the Split Forward Gradients cites this paper.

Backpropagation-Free Trunk Training via the Split Forward Gradients Large Batch Training of Convolutional Networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T20:29:59.024974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:29:59.024974Z digest=sha256:75496dc61522a940fa73c4a525b5616527303432c9738a919c6dca01952b9af7

Observation f33cc214-6673-4722-a1a7-a122be0ab08c · inbound

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales cites this paper.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Large Batch Training of Convolutional Networks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.783035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.783035Z digest=sha256:899fad370800344421ace6dba01f2f341ce735fe4f4f0da5c16f3c7673165114