Pith. sign in

Paper Citation Record · LEDGER

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models

As of 18 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2412.07210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07210 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:01:50.972685Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact3
  • verified fuzzy17
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02128430-a0c8-433c-b6c1-e08388b5d020 · outbound

This paper cites A Survey on Data Selection for Language Models.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models A Survey on Data Selection for Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.697707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.697707Z digest=sha256:cb87d334a21144988b7d5e3ea7d9d9e7c23648dd6bb1a6b9e39727c4d832bdaa

Observation bf4bf485-785d-4396-ba82-b1bf31081284 · outbound

This paper cites Qwen Technical Report.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.704928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.704928Z digest=sha256:3b41f9a5808a3aba3d250a4f38c033e380f7832fa72b343d2a8772adb184ce4e

Observation 6d98eb2e-c506-4273-925a-d7fe5ab15716 · outbound

This paper cites On the choice of learning rate for local sgd.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models On the choice of learning rate for local sgd

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:17.255885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.711299Z digest=sha256:01f0f7d5f38ecd4297e0bddb673e52aec7ea5d1c5ba36870c926b76b4c177f12

Observation 14d78ca3-c140-47f1-a970-16db4580e558 · outbound

This paper cites Multi-level local sgd: Distributed sgd for heterogeneous hierarchical networks.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Multi-level local sgd: Distributed sgd for heterogeneous hierarchical networks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:17.234512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.718130Z digest=sha256:721f027258a8767649bc603e0df5b6153f2c33b8e1b3205eef89662fda57788e

Observation 46dcee5e-64ac-49b9-b133-795e94cc48d0 · outbound

This paper cites Accelerating gossip sgd with periodic global averaging.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Accelerating gossip sgd with periodic global averaging

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:17.212707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.725557Z digest=sha256:d7b15f3812131fbb2db6e6b386a50e6c215b63b0e7308a4ec08550a67d87aff1

Observation fcfce43c-30ce-47b1-b3e4-3757cb80497e · outbound

This paper cites Large scale distributed deep networks.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Large scale distributed deep networks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.731733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.731733Z digest=sha256:d54f1fe92d24fa8449269eb3fe923a40a605bc50b7a220ce9cedc9c418a0e85d

Observation f5c81e9e-f3a7-4fe7-b8b9-3873153fb9b1 · outbound

This paper cites Local sgd optimizes overparameterized neural networks in polynomial time.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Local sgd optimizes overparameterized neural networks in polynomial time

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:17.170440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.741566Z digest=sha256:78df05e95e12ba53d97d9d749f6fbd3449672cf352953861b6ce5a5193437a86

Observation aa79819e-02f6-440b-a216-0172499fa719 · outbound

This paper cites DiLoCo: Distributed Low-Communication Training of Language Models.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models DiLoCo: Distributed Low-Communication Training of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.746824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.746824Z digest=sha256:eb3c53f271e51fae956f4bec7daa79627eb465cb16a0de8e58eca9809f1ef4d0

Observation c9394a52-6502-4d38-ba64-9fc2be638d8c · outbound

This paper cites Lighteval: A lightweight framework for llm evaluation, 2023.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Lighteval: A lightweight framework for llm evaluation, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.754095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.754095Z digest=sha256:e8e53845db8dd385f9a1366d98500411d73bdca2438a2999a3a71ddd2e731610

Observation 1a1c683a-31d5-4702-951f-bbc5fccf2fe6 · outbound

This paper cites Why (and when) does local sgd generalize better than sgd? In The Eleventh International Conference on Learning Representations, 2022.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Why (and when) does local sgd generalize better than sgd? In The Eleventh International Conference on Learning Representations, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:17.122929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.759345Z digest=sha256:2b99ba1c394c6a16b9de44283cf0bad1091f806edea9c97c8168c31bfc53c784

Observation 0b8eb900-25d4-4c18-8814-707fb99df4cf · outbound

This paper cites Tighter theory for local sgd on identical and heterogeneous data.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Tighter theory for local sgd on identical and heterogeneous data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.768713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.768713Z digest=sha256:3090a7adcd0b4374f17e21986aba8db2163de5979b8ca03753910596af7361e4

Observation 718a2e4c-9738-408c-9825-981da2d171bb · outbound

This paper cites Lyra: Elastic scheduling for deep learning clusters.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Lyra: Elastic scheduling for deep learning clusters

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:17.075927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.775914Z digest=sha256:d472d3b03a6748b65538ab139ab336d8f443b81e9d54f7fc06f93cc81ec98926

Observation b91b25c0-4ad3-4c4f-bbbd-506c791394f9 · outbound

This paper cites Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.781092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.781092Z digest=sha256:d6e30be5462fa694e7dbef7e2a2fd635abd6e2663418888dfd1d976d5a4a0777

Observation cb03af3c-12d8-4f80-9530-fbae14c399d8 · outbound

This paper cites Asynchronous decentralized parallel stochastic gradient descent.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Asynchronous decentralized parallel stochastic gradient descent

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:17.024482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.787834Z digest=sha256:58c0cfd0d274f4117f94a8d5b1d999b85f7129866ea6fb2e2b74815d0459d4c4

Observation 3aaeea4a-5341-4ff0-97ea-e73c94af78b4 · outbound

This paper cites Don't use large mini-batches, use local sgd.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Don't use large mini-batches, use local sgd

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:17.004193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.795388Z digest=sha256:ea5d29b01abbb5b747aedf276d914620d3d1a39800c0105c77b1bc679f55cd80

Observation d7f7e6da-14fb-424c-9952-4a4bb991f624 · outbound

This paper cites Asynchronous Local-SGD Training for Language Modeling.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Asynchronous Local-SGD Training for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.802686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.802686Z digest=sha256:f73682613e132f73d8aa758f95daaa7dbab6e3ce50cbfa63cf6e41ef9e6c7236

Observation cfb954cd-310c-4ff0-a903-4f7fd6b02e9b · outbound

This paper cites Decoupled weight decay regularization.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Decoupled weight decay regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.809059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.809059Z digest=sha256:1c8877d696c250932858e3516945924967d22a9386726fab5a635ab9c4911817

Observation 98c03838-4df0-43db-b82b-baf48a5eeb8f · outbound

This paper cites Fineweb-edu, May 2024.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Fineweb-edu, May 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.814182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.814182Z digest=sha256:45d7787ae0514ff2b126c23adea098ede79536b2096ec188ae07e82a0ebd8f94

Observation 33cf37c2-dd16-4a7b-9155-9c0e52e045e9 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.819444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.819444Z digest=sha256:ed72c90f3333875ec417f521dfa95290ca25f4bfe11ce93523c5df1a36cf3c21

Observation c134f586-12a4-4025-8745-71529614a710 · outbound

This paper cites an unresolved cited work.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.825187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.825187Z digest=sha256:24bd83a5fd575bb7fc42c1f8991b3afa3ef21754372399985298c080c8d52c40

Observation ee196b0d-9be9-4519-b27e-151060818add · outbound

This paper cites Federated learning with buffered asynchronous aggregation.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Federated learning with buffered asynchronous aggregation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:16.922045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.830487Z digest=sha256:d6bc66fd8612ede9ca531e52eb0ee7e5f3b83d21f49c896b06c43e0db2dae630

Observation 402b8a00-7355-4d75-8cd8-3bc4b70e84a0 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Opencompass: A universal evaluation platform for foundation models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:16.901960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.838036Z digest=sha256:f49481013af477a671b814493182df514b0626aea5cda33fe4c645b233cecf68

Observation dcca39fd-0815-409d-b51c-a1f84329abd5 · outbound

This paper cites Local SGD Accelerates Convergence by Exploiting Second Order Information of the Loss Function.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Local SGD Accelerates Convergence by Exploiting Second Order Information of the Loss Function

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:02:06.233883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.842956Z digest=sha256:2a6acce32c6bfb5053ac8b54e8580df3c2bce5b4abab41eb20c663f461a08fd1

Observation 0fba23ee-ceb2-4089-acd0-0f544e2fe9ec · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.849060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.849060Z digest=sha256:44b8fe50e066b2688af0013dfe6a388130afe66309851e4223f9aed656ebf8cb

Observation d2ba9c49-fe2a-45ca-876b-0007b2bac3e2 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Zero: Memory optimizations toward training trillion parameter models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.857149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.857149Z digest=sha256:ef8413b2e3e86427ac9efda955955e6724b63c180cdcbd1feb35c95eb7272ad9

Observation 650c8030-1a74-4be4-bcca-5a2cb55d060d · outbound

This paper cites A stochastic approximation method.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models A stochastic approximation method

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.863773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.863773Z digest=sha256:b7620395168c9a76f7bd38d2be9ed922f23b68470121bec005c1b2a6298b8c09

Observation a56d8a16-03b5-4e22-865b-c6b256ddb323 · outbound

This paper cites Stl-sgd: Speeding up local sgd with stagewise communication period.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Stl-sgd: Speeding up local sgd with stagewise communication period

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:16.855473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.869547Z digest=sha256:67139aa9d7e2207e56504aaf3e6765b5f893d7162c3bcd58841385c24b5e0543

Observation 3f737601-00b6-433f-854b-036adf1f6f44 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.875908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.875908Z digest=sha256:fa0ad40ebf577e2b6fcb3eb66dccf6b4005c32a1a0d55c5da680925d1e722eab

Observation 8f2f3474-5f05-467f-acce-516eac130318 · outbound

This paper cites Local SGD With a Communication Overhead Depending Only on the Number of Workers.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Local SGD With a Communication Overhead Depending Only on the Number of Workers

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:02:06.175013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.881039Z digest=sha256:50685207b7764e66b0b933f03abbd73642b320b78f5d9725c21ed5d51557f189

Observation f599164a-05f3-402b-b91e-2fec05b75b4b · outbound

This paper cites Efficient distributed training with full communication-computation overlap.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Efficient distributed training with full communication-computation overlap

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:16.830375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.886713Z digest=sha256:1c04fdf9850de769ffd90e11efc78b9fe2c4ce235e4c2880242bc01c3ecfa118

Observation e9e18d58-e543-435e-85d7-c2ef250c48aa · outbound

This paper cites On the importance of initialization and momentum in deep learning.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models On the importance of initialization and momentum in deep learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.892264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.892264Z digest=sha256:c9e2c1d09bd46c342d27c273e946451bc5e7494b2dfa4a67454f98fd21ecaa48

Observation 78c1516d-2e87-443a-bcfb-3f6f0513f9a8 · outbound

This paper cites Self-Influence Guided Data Reweighting for Language Model Pre-training.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Self-Influence Guided Data Reweighting for Language Model Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.897363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.897363Z digest=sha256:e88009ae0c65ba2a03921310e92cad01721d060fb9973b0ab68eec5a0d7ebf61

Observation fb8b5939-9f9a-468c-a7ac-5d23561428aa · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.902814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.902814Z digest=sha256:698349f619721c069240a82dc829cc4019d6a760ecf11795a295b6b3ff96363f

Observation 2727d3e7-8285-4ff4-8eb5-5ac8833701cf · outbound

This paper cites Adaptive communication strategies to achieve the best error-runtime trade-off in local-update sgd.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Adaptive communication strategies to achieve the best error-runtime trade-off in local-update sgd

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:16.794055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.908682Z digest=sha256:c479b3b58a290e05a47a27034349914309716263fdb7a40d92ef10aef8b39104

Observation a7ceaba7-f6dd-4e7c-9f38-9d177a0db455 · outbound

This paper cites Slowmo: Improving communication-efficient distributed sgd with slow momentum.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Slowmo: Improving communication-efficient distributed sgd with slow momentum

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:16.777749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.915047Z digest=sha256:cc64a481e1a64990b7282b473f8d1cd5b13230ac434827f4b54dcea1377023eb

Observation 9f8e5596-19d1-4b43-924b-38750029f303 · outbound

This paper cites Asynchronous Federated Optimization.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Asynchronous Federated Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.921207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.921207Z digest=sha256:0bcf6393f2966774de68dac594d884e4b6919ba52a729c33cdce3bc7fb0120e8

Observation 6da53f5d-c929-4cc7-8635-f020e0f7f443 · outbound

This paper cites Tuning large neural networks via zero-shot hyperparameter transfer.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Tuning large neural networks via zero-shot hyperparameter transfer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:16.759690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.927403Z digest=sha256:273f38175e14e792bf0e8fdd128dd1b742bd1a0523888cc5cfa770fd0d02ff55

Observation aaddf191-10df-4f23-be96-55fec23447dd · outbound

This paper cites Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:16.739467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.932676Z digest=sha256:0399ff60b1ca38a5d124fe37fd419d72a2c3f622ed1e75fbffb138cb17766cf9

Observation 9161b451-eb35-4362-bbfe-755d63b54514 · outbound

This paper cites Parallel SGD: When does averaging help?.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Parallel SGD: When does averaging help?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.937643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.937643Z digest=sha256:07fa9b94b80d0f1fa67325c088ece2e9de405da9ae6b5bbacc3aa486318d3932

Observation bf03a536-3ad2-4452-a072-bb519bfe0223 · outbound

This paper cites Timelyfl: Heterogeneity-aware asynchronous federated learning with adaptive partial training.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Timelyfl: Heterogeneity-aware asynchronous federated learning with adaptive partial training

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:16.713224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.942747Z digest=sha256:fc7c798ecdc5735b0076eebcbaf43319a85b9fdbad3724c6790cd7ba48fa342e

Observation b42961bd-82ab-416f-b6b9-55078b33ae54 · outbound

This paper cites write newline.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models write newline

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.949594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.949594Z digest=sha256:82eae466a85b7afac46adb38b95567b4c37124b7ae1ed9a783b7e8e8cac8dee7

Observation cd775047-53a3-401a-8f8d-c9018ae64769 · outbound

This paper cites @esa (Ref.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models @esa (Ref

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.960059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.960059Z digest=sha256:e241b6a548b8a15ca964cb079a3c8135e8677ae01893ee31c85f4cbc040f5f99

Observation 4366e7aa-4907-4662-9ac9-b97745479cf6 · outbound

This paper cites an unresolved cited work.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.965732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.965732Z digest=sha256:bff7c2d3ba22886f3ccdaed5819874f8113e1e71f8ef053e24ef42ff3b787abb

Observation 9090612e-c3c9-442c-925c-a00bb6bf3263 · outbound

This paper cites an unresolved cited work.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models Unresolved cited work

Reference 44

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:02:06.074504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:01:50.972685Z digest=sha256:0f0edd1dea8f252195c77200f837416b6f439e48fa881ca538d80e38590dbdb4

Pith citing papers

No inbound Pith citation observations are available.