Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:01:50.972685Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2412.07210.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:01:50.972685Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 02128430-a0c8-433c-b6c1-e08388b5d020 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models A Survey on Data Selection for Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4bf485-785d-4396-ba82-b1bf31081284 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d98eb2e-c506-4273-925a-d7fe5ab15716 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models On the choice of learning rate for local sgd
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14d78ca3-c140-47f1-a970-16db4580e558 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Multi-level local sgd: Distributed sgd for heterogeneous hierarchical networks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 46dcee5e-64ac-49b9-b133-795e94cc48d0 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Accelerating gossip sgd with periodic global averaging
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fcfce43c-30ce-47b1-b3e4-3757cb80497e · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Large scale distributed deep networks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c81e9e-f3a7-4fe7-b8b9-3873153fb9b1 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Local sgd optimizes overparameterized neural networks in polynomial time
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aa79819e-02f6-440b-a216-0172499fa719 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models DiLoCo: Distributed Low-Communication Training of Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9394a52-6502-4d38-ba64-9fc2be638d8c · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Lighteval: A lightweight framework for llm evaluation, 2023
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a1c683a-31d5-4702-951f-bbc5fccf2fe6 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Why (and when) does local sgd generalize better than sgd? In The Eleventh International Conference on Learning Representations, 2022
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0b8eb900-25d4-4c18-8814-707fb99df4cf · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Tighter theory for local sgd on identical and heterogeneous data
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718a2e4c-9738-408c-9825-981da2d171bb · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Lyra: Elastic scheduling for deep learning clusters
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b91b25c0-4ad3-4c4f-bbbd-506c791394f9 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb03af3c-12d8-4f80-9530-fbae14c399d8 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Asynchronous decentralized parallel stochastic gradient descent
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3aaeea4a-5341-4ff0-97ea-e73c94af78b4 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Don't use large mini-batches, use local sgd
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d7f7e6da-14fb-424c-9952-4a4bb991f624 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Asynchronous Local-SGD Training for Language Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb954cd-310c-4ff0-a903-4f7fd6b02e9b · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Decoupled weight decay regularization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c03838-4df0-43db-b82b-baf48a5eeb8f · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Fineweb-edu, May 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33cf37c2-dd16-4a7b-9155-9c0e52e045e9 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Efficient large-scale language model training on gpu clusters using megatron-lm
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c134f586-12a4-4025-8745-71529614a710 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee196b0d-9be9-4519-b27e-151060818add · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Federated learning with buffered asynchronous aggregation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 402b8a00-7355-4d75-8cd8-3bc4b70e84a0 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Opencompass: A universal evaluation platform for foundation models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dcca39fd-0815-409d-b51c-a1f84329abd5 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Local SGD Accelerates Convergence by Exploiting Second Order Information of the Loss Function
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0fba23ee-ceb2-4089-acd0-0f544e2fe9ec · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ba9c49-fe2a-45ca-876b-0007b2bac3e2 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Zero: Memory optimizations toward training trillion parameter models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650c8030-1a74-4be4-bcca-5a2cb55d060d · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models A stochastic approximation method
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56d8a16-03b5-4e22-865b-c6b256ddb323 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Stl-sgd: Speeding up local sgd with stagewise communication period
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3f737601-00b6-433f-854b-036adf1f6f44 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f2f3474-5f05-467f-acce-516eac130318 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Local SGD With a Communication Overhead Depending Only on the Number of Workers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f599164a-05f3-402b-b91e-2fec05b75b4b · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Efficient distributed training with full communication-computation overlap
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e9e18d58-e543-435e-85d7-c2ef250c48aa · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models On the importance of initialization and momentum in deep learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c1516d-2e87-443a-bcfb-3f6f0513f9a8 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Self-Influence Guided Data Reweighting for Language Model Pre-training
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb8b5939-9f9a-468c-a7ac-5d23561428aa · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2727d3e7-8285-4ff4-8eb5-5ac8833701cf · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Adaptive communication strategies to achieve the best error-runtime trade-off in local-update sgd
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a7ceaba7-f6dd-4e7c-9f38-9d177a0db455 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Slowmo: Improving communication-efficient distributed sgd with slow momentum
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f8e5596-19d1-4b43-924b-38750029f303 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Asynchronous Federated Optimization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da53f5d-c929-4cc7-8635-f020e0f7f443 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Tuning large neural networks via zero-shot hyperparameter transfer
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aaddf191-10df-4f23-be96-55fec23447dd · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9161b451-eb35-4362-bbfe-755d63b54514 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Parallel SGD: When does averaging help?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf03a536-3ad2-4452-a072-bb519bfe0223 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Timelyfl: Heterogeneity-aware asynchronous federated learning with adaptive partial training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b42961bd-82ab-416f-b6b9-55078b33ae54 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models write newline
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd775047-53a3-401a-8f8d-c9018ae64769 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models @esa (Ref
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4366e7aa-4907-4662-9ac9-b97745479cf6 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9090612e-c3c9-442c-925c-a00bb6bf3263 · outbound
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.