Pith. sign in

Paper Citation Record · LEDGER

Scaling Exponents Across Parameterizations and Optimizers

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2407.05872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.05872 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:21:20.377243Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 28c08736-595d-4735-9602-dae2f7e9964a · inbound

GWT: Scalable Optimizer State Compression for Large Language Model Training cites this paper.

GWT: Scalable Optimizer State Compression for Large Language Model Training Scaling Exponents Across Parameterizations and Optimizers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:57:36.655891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:57:09.276224Z digest=sha256:565c6a5c0264f45bf3a8adf960d5f201a0d1e7094a4c5263d6e3c587bec153ff

Observation b70c9695-3fb7-44af-abc1-b35ac67cf3bb · inbound

Avoiding spurious sharpness minimization broadens applicability of SAM cites this paper.

Avoiding spurious sharpness minimization broadens applicability of SAM Scaling Exponents Across Parameterizations and Optimizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:21:20.377243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:21:20.377243Z digest=sha256:7148fdba8954865344519412f7029acc00065343919f0334a241b90b449d74fd

Observation a7e52749-50eb-49f8-8199-6388b3bf013b · inbound

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer cites this paper.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Scaling Exponents Across Parameterizations and Optimizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.874008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.874008Z digest=sha256:d2d029ee59f8e6712ff68cd940f8774e29f4f5cab52f0b68f653b6c64e772daa

Observation 54ecae26-47be-45b3-96f0-58a71145a502 · inbound

Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models cites this paper.

Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models Scaling Exponents Across Parameterizations and Optimizers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:30.972203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:04.110552Z digest=sha256:bb592f63ff458321a94838b444ea01bd47774faafdb8baa32dac2a2425553487

Observation 2d965bc0-1928-4efa-a0e7-9ba8e7c6afd4 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Scaling Exponents Across Parameterizations and Optimizers

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:39:41.136195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:f4c869f4afb0d09baa3140197f10f6684df6c4fa0159fb9a23e8e98aebc921bf

Observation f225e47c-d49a-40d2-8c6e-9d2916046295 · inbound

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning cites this paper.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Exponents Across Parameterizations and Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.364789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.364789Z digest=sha256:ea336750aeae765c49e0bdc00f0a4d1570b7d83e69d44257a800de7000f6f4ad

Observation 0231a66f-5898-4deb-a78f-3976de64ab0d · inbound

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks cites this paper.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Scaling Exponents Across Parameterizations and Optimizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.834417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.834417Z digest=sha256:25070917fc19fa832a5ba1ea3527f462d0de384d05ebf656354c9c875b5b2bf3

Observation e0385b0d-7ebe-4c00-8d84-3431dec3bbdc · inbound

Decoupled Relative Learning Rate Schedules cites this paper.

Decoupled Relative Learning Rate Schedules Scaling Exponents Across Parameterizations and Optimizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:14:11.824040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:14:11.824040Z digest=sha256:a9de46393baa8da4a688044e227e3d3499111d8c6e426dc798aaf8c313877a27

Observation 1af5ec08-b7dd-43b3-acae-d35d9369cca7 · inbound

Feature learning is decoupled from generalization in high capacity neural networks cites this paper.

Feature learning is decoupled from generalization in high capacity neural networks Scaling Exponents Across Parameterizations and Optimizers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:37.081920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:17:37.081920Z digest=sha256:4ab7c9f5cd2b75e1daed8025ddfc495d5659dba4f00fa8e231413ed6350d7ab8

Observation f93461aa-05fe-4da9-adf6-69f66ffbcde1 · inbound

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model cites this paper.

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Scaling Exponents Across Parameterizations and Optimizers

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:00:43.299777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T06:58:38.927268Z digest=sha256:49834ea2f1d2abdcbdd50a7ab0e446176cbf7d6fa615d79bfab266c4a7428ab0

Observation 9102af60-12fc-4a52-b523-61953e1188a0 · inbound

Parcae: Scaling Laws For Stable Looped Language Models cites this paper.

Parcae: Scaling Laws For Stable Looped Language Models Scaling Exponents Across Parameterizations and Optimizers

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:21:01.021466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:33:04.442462Z digest=sha256:99f517fa958058972cccc15979628facea747ab9c8e2a2d86abb5e92fbd1cc39

Observation a223976d-c983-44ac-85da-6472e97be7d7 · inbound

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions cites this paper.

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions Scaling Exponents Across Parameterizations and Optimizers

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:05:24.996003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:02:52.920735Z digest=sha256:658f1ae9e050fc5453cd86250fb7dc4b98da618a33fd5d022db9e7447eb6fec0

Observation 805efe36-d7e0-4353-96ba-7b0aba39edc1 · inbound

Learning Rate Transfer in Normalized Transformers cites this paper.

Learning Rate Transfer in Normalized Transformers Scaling Exponents Across Parameterizations and Optimizers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:51:28.083967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T09:01:45.365126Z digest=sha256:488d85d949e2ba63b0aaae7e3e3989d90ed72546c87adf5a65af3ded487d2e10

Observation 2e55164d-4c29-441a-a4e6-ec14cbc6e6b6 · inbound

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data cites this paper.

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data Scaling Exponents Across Parameterizations and Optimizers

Reference 141

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:57:21.512477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:56:38.042978Z digest=sha256:01240b768f2f4b99aeca8a366827e6152238cade8925e23eb73d2a3e268e5851

Observation d85931f2-99b1-4294-9198-7a845be45fa1 · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Scaling Exponents Across Parameterizations and Optimizers

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.837389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:92765cd13a6ff7161a0d69edf30086aab0e113c78bd2095209d5e3f9ed8b7442

Observation fa8b0839-ae38-430c-8d99-827db6b9e8fc · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention Scaling Exponents Across Parameterizations and Optimizers

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:39.842088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:b8af01ac3a658fb02d8ac8cd966af60d8093026500175e3737341f7edb0bc25d

Observation 08b3d673-2af3-4b48-a3b2-09dd84cd135e · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Scaling Exponents Across Parameterizations and Optimizers

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T15:39:33.195350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T15:35:51.654392Z digest=sha256:8258d6029a3b04cb37f79bcc0f8171fa4976d882ec3ff9d98c8e3cffbac93129

Observation d0848018-1c61-4497-a063-900a48ac9109 · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Scaling Exponents Across Parameterizations and Optimizers

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:57:25.317838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T21:51:13.457071Z digest=sha256:4f85312f54e6888872bca3ae6ddb0c66b5843926d676a5adb761f1975b208662

Observation 38c7f7b7-d4c8-4b42-8683-bb38aed6a0d7 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Exponents Across Parameterizations and Optimizers

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.676652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:90d6295597d5c09617c9795adc3141c930a6a4f1feafd610aae555a7c43969da

Observation 44235205-ee6f-4780-be28-54aab9ef7635 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Exponents Across Parameterizations and Optimizers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:06.198299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:06.198299Z digest=sha256:ca0dcbaaa6f3283a898ed476c70c6575813ef30ddb66ec6b7ef9627afe9f8fc3

Observation 0ca474ff-08db-4668-8d38-4b65e51d4f23 · inbound

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales cites this paper.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scaling Exponents Across Parameterizations and Optimizers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.238514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.238514Z digest=sha256:2f26039746860b96097522dcbb7e075680a84ddf31c1c7553bf2fb55ac508be8