Pith. sign in

Paper Citation Record · LEDGER

Scaling Exponents Across Parameterizations and Optimizers

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2407.05872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.05872 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:03:53.041974Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 28c08736-595d-4735-9602-dae2f7e9964a · inbound

GWT: Scalable Optimizer State Compression for Large Language Model Training cites this paper.

GWT: Scalable Optimizer State Compression for Large Language Model Training Scaling Exponents Across Parameterizations and Optimizers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:57:36.655891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T05:57:09.276224Z digest=sha256:735fb0ea1fe66536920a1573a3b3114fe35851028e9533c0cb56bcbe169e0dda

Observation b70c9695-3fb7-44af-abc1-b35ac67cf3bb · inbound

Avoiding spurious sharpness minimization broadens applicability of SAM cites this paper.

Avoiding spurious sharpness minimization broadens applicability of SAM Scaling Exponents Across Parameterizations and Optimizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:21:20.377243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:21:20.377243Z digest=sha256:fbf22e53da27a715e938fde6600301e4d1629ce26b70aa894e4b9e950acc5a69

Observation a7e52749-50eb-49f8-8199-6388b3bf013b · inbound

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer cites this paper.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Scaling Exponents Across Parameterizations and Optimizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.874008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.874008Z digest=sha256:57d047d4908bf83583a157da7d01e4855c783f2ea10c00ed7c717df56c0af46c

Observation 54ecae26-47be-45b3-96f0-58a71145a502 · inbound

Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models cites this paper.

Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models Scaling Exponents Across Parameterizations and Optimizers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:30.972203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T04:16:04.110552Z digest=sha256:fdab8484fc9d0a7de6b8016bcd6ed7f14e76f369d2c31321d1e1de4a09585c60

Observation 2d965bc0-1928-4efa-a0e7-9ba8e7c6afd4 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Scaling Exponents Across Parameterizations and Optimizers

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:39:41.136195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:1e12160a17050ff7a27af66d95e18c1f70e56ef97df2ee11f69250bcd8a75197

Observation 1936ef49-33e9-404f-ae58-372b5fffb8ac · inbound

Practical Efficiency of Muon for Pretraining cites this paper.

Practical Efficiency of Muon for Pretraining Scaling Exponents Across Parameterizations and Optimizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T01:03:53.041974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T01:03:53.041974Z digest=sha256:13a1fd9f0452371bb1b2a6513ae7936d359a0832f97b82f1e7cd531ffc1d820e

Observation f225e47c-d49a-40d2-8c6e-9d2916046295 · inbound

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning cites this paper.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Exponents Across Parameterizations and Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.364789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.364789Z digest=sha256:196b32507961f906f731cffa6c2c6448566fbb0a8ca23d3b42c92458f444bb43

Observation 3c8dd48f-a720-4c59-839c-43c18097175c · inbound

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size cites this paper.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Exponents Across Parameterizations and Optimizers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.469149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.469149Z digest=sha256:8ba73b0d704b0df5c84cdc83950fe5b40a267d9e16a1c2df63233f16e86766b1

Observation 0231a66f-5898-4deb-a78f-3976de64ab0d · inbound

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks cites this paper.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Scaling Exponents Across Parameterizations and Optimizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.834417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.834417Z digest=sha256:f95a499b3c958aff04dd20b22ef6bee380869bc0b58dd40df1adeddb00971f31

Observation e0385b0d-7ebe-4c00-8d84-3431dec3bbdc · inbound

Decoupled Relative Learning Rate Schedules cites this paper.

Decoupled Relative Learning Rate Schedules Scaling Exponents Across Parameterizations and Optimizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:14:11.824040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:14:11.824040Z digest=sha256:fdddba1fbdc6f1341a2ec7a9e6ee13d3f3e3a9fc6fffcb23927a0837554db406

Observation 1af5ec08-b7dd-43b3-acae-d35d9369cca7 · inbound

Feature learning is decoupled from generalization in high capacity neural networks cites this paper.

Feature learning is decoupled from generalization in high capacity neural networks Scaling Exponents Across Parameterizations and Optimizers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:37.081920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:17:37.081920Z digest=sha256:8f6e60d2b5620b0cc158e40159bbbd7b559125e91744ed2b6d28d3560e4a480f

Observation f93461aa-05fe-4da9-adf6-69f66ffbcde1 · inbound

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model cites this paper.

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Scaling Exponents Across Parameterizations and Optimizers

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:00:43.299777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T06:58:38.927268Z digest=sha256:200d1e49f44a070351e8aebbbd4dedaffb2225d2bc172ed99b26bf7d843c584e

Observation 9102af60-12fc-4a52-b523-61953e1188a0 · inbound

Parcae: Scaling Laws For Stable Looped Language Models cites this paper.

Parcae: Scaling Laws For Stable Looped Language Models Scaling Exponents Across Parameterizations and Optimizers

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:21:01.021466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:33:04.442462Z digest=sha256:d2a6b0184b7ff0e4be58352872325d8905b63cef9da64b9c096ed17ac318d677

Observation a223976d-c983-44ac-85da-6472e97be7d7 · inbound

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions cites this paper.

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions Scaling Exponents Across Parameterizations and Optimizers

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:05:24.996003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:02:52.920735Z digest=sha256:c58d02169805fb688abbb5875c5fbeab69ec00c6184480715b321ced5a076d0a

Observation 805efe36-d7e0-4353-96ba-7b0aba39edc1 · inbound

Learning Rate Transfer in Normalized Transformers cites this paper.

Learning Rate Transfer in Normalized Transformers Scaling Exponents Across Parameterizations and Optimizers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:51:28.083967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T09:01:45.365126Z digest=sha256:4934ee2695046653c236839de59ec9f06943a8a2a372a2c924f2f3febc8f4b0e

Observation 2e55164d-4c29-441a-a4e6-ec14cbc6e6b6 · inbound

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data cites this paper.

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data Scaling Exponents Across Parameterizations and Optimizers

Reference 141

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:57:21.512477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T05:56:38.042978Z digest=sha256:b0c494cc04e6c985a06d347004be8554fd08359f74cb1cad83f19da2675c5586

Observation d85931f2-99b1-4294-9198-7a845be45fa1 · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Scaling Exponents Across Parameterizations and Optimizers

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.837389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:c9423b119747f5359eca5eb436ed9fe19fe6e79397b78646c1e2a3aa6307f921

Observation fa8b0839-ae38-430c-8d99-827db6b9e8fc · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention Scaling Exponents Across Parameterizations and Optimizers

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:39.842088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:017b7782359c11b7506f746c46e0161ae86d8281c3837203bd62005be0399b4e

Observation 08b3d673-2af3-4b48-a3b2-09dd84cd135e · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Scaling Exponents Across Parameterizations and Optimizers

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T15:39:33.195350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T15:35:51.654392Z digest=sha256:c44ddffcce6eb8950b32c8794b772cbe2bfd6d43385ecc0b42973c41a80fbb7e

Observation d0848018-1c61-4497-a063-900a48ac9109 · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Scaling Exponents Across Parameterizations and Optimizers

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:57:25.317838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-02T21:51:13.457071Z digest=sha256:b589ef2041ffc487bbcfe4b9882ecbcb63cc3df4d78080fbe3f0961e13b64013

Observation 38c7f7b7-d4c8-4b42-8683-bb38aed6a0d7 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Exponents Across Parameterizations and Optimizers

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.676652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:1ede7abeadd9789a075c65c97e765d50d0a44a4050ab16c7ce5f0b26e3d92ccd

Observation 44235205-ee6f-4780-be28-54aab9ef7635 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Exponents Across Parameterizations and Optimizers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:06.198299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:06.198299Z digest=sha256:422f34f09ee5a1fb81e5320cad6c4d4433062e381c68cc381a0dfbe0b8341b86

Observation 0ca474ff-08db-4668-8d38-4b65e51d4f23 · inbound

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales cites this paper.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scaling Exponents Across Parameterizations and Optimizers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.238514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.238514Z digest=sha256:caf6212f8cb392132cf1227b6e8b4a2b957a83215bc5b456fe974dfdaec18408