Pith. sign in

Paper Citation Record · LEDGER

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining

As of 23 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.07494.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07494 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T09:15:12.214083Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact18
  • verified fuzzy5
  • unresolved0
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2059980-4aad-474d-9191-13547e0f19a3 · outbound

This paper cites Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.522575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:92295f36b38f285e9299b5ebd51ffb2e9969264d2ca705ecb9558dbfc0f5b83a

Observation 25c15d21-ac6a-4772-999b-aa7130eb0aa7 · outbound

This paper cites Bandwidth optimal all-reduce algorithms for clusters of workstations,.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Bandwidth optimal all-reduce algorithms for clusters of workstations,

Reference 2

Resolution
verified exact
doi, observed 2026-07-09T09:16:06.530439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:17b6faa842ac20fdf08a9768687e60692420635a975874b857306fa4d54f6acb

Observation d0b93274-9f9b-475e-b361-3f92cbd9b5b7 · outbound

This paper cites InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(Vancouver, BC, Canada)(ASPLOS 2023).

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(Vancouver, BC, Canada)(ASPLOS 2023)

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T09:16:06.475475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:2a968236b552eb69dbf5473cd2b7807ecaabb69ab82ba0238722e65c57affa74

Observation 0a318c2e-7a83-4f3b-8fbf-d728b7ed2f33 · outbound

This paper cites FP8-LM: Training FP8 Large Language Models.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining FP8-LM: Training FP8 Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.500582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:9b5fd1b7ebff4608b3b1cd361d557c120c49088165b67f76e819cb1ffb69fa98

Observation ddf31843-2c99-487e-b4f2-f2f4ac1398c0 · outbound

This paper cites COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.523377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:e9b995fa4a5a13c4cfac957b24cb0c34a10504db2028282a8e94d363489e1d9b

Observation 6c557088-e4b3-49a4-bcc5-9c1618a6957d · outbound

This paper cites Decoupled Weight Decay Regularization,.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Decoupled Weight Decay Regularization,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T09:16:06.711251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:511f39bb93b45600b7083fe648930a884336572c5f1256803f2a12f8fd2273a4

Observation c8ba83ef-4f75-4bbc-bd4e-b3b4c2d95e2a · outbound

This paper cites Decoupled Weight Decay Regularization.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Decoupled Weight Decay Regularization

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T09:16:06.503181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:3c7a4f63b49742faefd44ce28227b1f8909b76c8e4013cca3ee2b091a66fc42b

Observation 14715872-799e-4a17-9d7f-200112564f06 · outbound

This paper cites SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.535859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:f5d6612d8dc3e080c2eacf14be67fddc848df5e00fcfebee21492241ef85438e

Observation 93309472-6084-4e32-bf80-1856cef73652 · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks,.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Muon: An optimizer for hidden layers in neural networks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T09:16:06.715190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:bb57f41858949ab95e88c913bf5c9ed564452d4964fc0e3475053d7c3b66a16c

Observation 028df0d8-96d5-43b3-95c1-2f594ef0e20b · outbound

This paper cites Available: https://kellerjordan.github.io/posts/muon/.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Available: https://kellerjordan.github.io/posts/muon/

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T09:16:06.712955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:a71385721adde0d56d039b002159889cdc259271179e8ed260b1180ed928b400

Observation 63e731fd-9438-4aac-b9ac-457c38815966 · outbound

This paper cites Amari,Information Geometry and Its Applications.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Amari,Information Geometry and Its Applications

Reference 11

Resolution
malformed identifier
raw_fallback, observed 2026-07-09T09:16:06.717572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:dcdfd0793925e888de08f88c1154337b245d2ff59826cfa547fcc236b2ff63ea

Observation 5fe4aa30-baa6-4314-b6f0-26c35db697a5 · outbound

This paper cites Natural Gradient Works Efficiently in Learning , journal =.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Natural Gradient Works Efficiently in Learning , journal =

Reference 12

Resolution
verified exact
doi, observed 2026-07-09T09:16:06.528219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T08:38:13.253949+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:a4734aa1cdcf4504a2c0d7595ab35cb3d8c5ea93ec69e9260e1c0a08e26edb4e

Observation 8782c345-28f4-4a7e-9dd9-83fdb7b5ad87 · outbound

This paper cites Optimizing Neural Networks with Kronecker-factored Approximate Curvature.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Optimizing Neural Networks with Kronecker-factored Approximate Curvature

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.515307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:2b6a30f88b32dfd48899fd0ce2aadf21080daf09725528156c98de778cf3e19c

Observation 7ed52c58-f009-466d-870c-16cd34446578 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.525993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:82156de5a1e0b9c0834881c43fdfd4cf58067134023bca08c9f9e0e1089e2912

Observation f9b0c093-bb8a-4dfe-8673-35189ed11d3e · outbound

This paper cites Openwebtext cor- pus,.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Openwebtext cor- pus,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T09:16:06.705423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:34b40c7f1ec8f60004d4cc052bc98a94e3b2925b51daabe14524306e91ae5ec5

Observation 77bd7475-7b24-454f-81e8-3cc2d5443949 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.679697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:5a8a82d5f4aa0a771585df2472b7fd97c161716640017d5a017c51767cb86d52

Observation 5cf6404c-142b-481a-8678-195109daa2e2 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T09:16:06.482358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:d869bb63bcda1f48a4fc14e0e9ba50472be15c2e1eb03eddf8419eb72e9d0c40

Observation e4cca9e8-6a67-4820-8d2b-b420695220e3 · outbound

This paper cites GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.520887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:60adde8de02d52765fdfb72741b0a2acf5562b3a1b6496b4b7033a4d58e1e76e

Observation 49378994-0105-4fc3-be1b-69c070a6aebb · outbound

This paper cites Memory-Efficient Pipeline-Parallel DNN Training.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Memory-Efficient Pipeline-Parallel DNN Training

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.540962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:2ce00414401c22e44aac655d33710b17d630afc24087cf8bbedad3aebc80361a

Observation 8ede1d75-5111-4ccb-bd94-561dca7c8fda · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.530633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:ec6bc49485c180b2caa97d01507f0e64a63c170ce8c3883252395c633cfc9b41

Observation f72b09a6-2543-4ecb-aea5-a3711a6cba83 · outbound

This paper cites ZeRO++: Extremely Efficient Collective Communication for Giant Model Training.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining ZeRO++: Extremely Efficient Collective Communication for Giant Model Training

Reference 21

Resolution
malformed identifier
local_arxiv, observed 2026-07-09T09:16:06.676458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:f2f8542052ff3ef93aa3b57b475620e765d786e2926d92701ccb9d36f996789e

Observation c31c9677-d310-4239-ab2b-88cf985e9f0a · outbound

This paper cites Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T09:16:06.535810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:62d34c84527cfa530ad07a424278f579343b34167b146175eff12c52bf9cb836

Observation fde0184c-24a9-4b70-8164-6af20c5072b6 · outbound

This paper cites QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.525204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:fa8cbfd269e145d86bea586b729cb09f30e5f9b6f18aa5bbf8af99d3b52a8d02

Observation 384737c0-8c4b-4c26-b2e0-cc3911f32a6e · outbound

This paper cites Deep gradient compression: Reducing the communication bandwidth for distributed training,.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Deep gradient compression: Reducing the communication bandwidth for distributed training,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T09:16:06.708056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:f358b93efe208eae0cb476d8ddb24344c0772f23b8fe17814c10d79dd4be168e

Observation cd1c83f9-abf0-49e7-af9b-d7af0fb478c3 · outbound

This paper cites Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T09:16:06.510183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:2893d576cab02be1b1c6ff45e1501bc03116e535faede5bf53e1d84bbac74f99

Observation b8e742f2-10f4-40be-bcfc-b4d96c844f24 · outbound

This paper cites PowerSGD: Practical Low-Rank Gradient Compression for Distributed Optimization.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining PowerSGD: Practical Low-Rank Gradient Compression for Distributed Optimization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.518204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:fcc94a6d0c1d57c14985e9de7b4e90b0bf92f576c5e94f6ce34d238871452ba0

Observation 4c5ab36d-92f2-4a45-9c18-3f0e2ae59ea7 · outbound

This paper cites 1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining 1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T09:16:06.538354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:4eac5d952e9b695133f5af7705449526fb52581dbe5e1956d8ed7b7f7ad92ea4

Observation b7609279-ffb8-41e6-87ae-ae7f4e5522a6 · outbound

This paper cites FP8 Formats for Deep Learning.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining FP8 Formats for Deep Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.492807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:ae6429e552f08581883fc4f0e93094e85505e7b307ba757c854261029e9a989f

Observation 0223999c-3059-413c-9253-5e6703b9c96a · outbound

This paper cites LoCo: Low-Bit Communication Adaptor for Large-scale Model Training.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining LoCo: Low-Bit Communication Adaptor for Large-scale Model Training

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.532939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:cc97f1ed6857d3850c75bd111fb96e305fe10d928633f4843965a4534509d783

Observation 75ae97cc-eea2-4299-99a6-6ba24ebfe5e2 · outbound

This paper cites Mixed Precision Training.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Mixed Precision Training

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.471623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:47200bf27886aa6a2e157bc4cd84e0874a659c971e66421e11803c60d0e78dfe

Observation 58e148da-2af6-44f8-aa17-242246c20812 · outbound

This paper cites Training and inference of large language models using 8-bit floating point.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Training and inference of large language models using 8-bit floating point

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.538522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:81340c34396c9b5c1c4ddda411acd0b39121618033db46e64e383ad2dda54426

Pith citing papers

No inbound Pith citation observations are available.