Pith. sign in

Paper Citation Record · LEDGER

Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2503.04715.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.04715 v7

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:27:10.359838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5d54bad5-dbfe-43d1-a3f7-0e20e17efeee · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:05:47.673917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:da8c9c290814976445ea8fbdd9be345d906585ed7ed1abd46014b12c5becc8f8

Observation d0f628a8-daea-4fbf-8d14-e5bcf9243c97 · inbound

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning cites this paper.

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T05:27:10.359838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:27:10.359838Z digest=sha256:b1e5aebe84739c4ca038f5c2e4c83d7b63fb84b68a7535a4fcc1899c3c835463

Observation 8e25eb6b-12cf-4bb2-9d36-37a8aecf8ea8 · inbound

OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale cites this paper.

OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:14:00.828682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:14:00.828682Z digest=sha256:f4b0b45d004e243e369cdcb04a7aec66afb0dfba52c58c699936b043cd3c0f21

Observation aa01726a-0612-4346-824f-30a2507a8b6f · inbound

Rethinking Language Model Scaling under Transferable Hypersphere Optimization cites this paper.

Rethinking Language Model Scaling under Transferable Hypersphere Optimization Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:53:02.428845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:51:14.678941Z digest=sha256:ef2722f62822c9eebc102772c70688bd6e1f2324aaefd941f65ec223e662b044

Observation 89c6de2c-1130-4202-8f08-e3cd7811a13d · inbound

AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive cites this paper.

AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:17:02.023735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:15:58.175350Z digest=sha256:c9f5f2cad690da795cf5bea9034a5b4e6da99d247f6ec433a6ff9bc9cea71302

Observation c3d8b653-6d90-4b50-b188-c87cb9dda4be · inbound

AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive cites this paper.

AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:39:10.210596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:36:56.061187Z digest=sha256:629d6c618bf1a4fba8b55097ef8c9cfe28f8b903efda7839450a5dd7e41b3643

Observation c272cd62-9e05-4b36-84e8-3bb687330044 · inbound

BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks cites this paper.

BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:52:46.284908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T20:48:20.856172Z digest=sha256:64b0dd44c9e816aff6b56bef12cf12b84a7b516dd14ef48edecffd38584bcf77

Observation 931597dd-a8b5-4142-8954-839468e0a7bc · inbound

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate cites this paper.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.973859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:38b9d013f038b91828094333bd63143e6b4e006dbe7f80b81caaea35fbd71203

Observation 78f9b9c1-6564-4ea0-a71e-b01b07db99e1 · inbound

Staged Factorial Screening for Budget-Constrained Micro-Pretraining cites this paper.

Staged Factorial Screening for Budget-Constrained Micro-Pretraining Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:35.015335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T08:39:46.699786Z digest=sha256:2173e2c27b77784cf1216968914c8b8d868b3eab27b6525fe264f33a917ca68a

Observation a77351ce-5419-44df-8d6a-cae45a7be5d3 · inbound

Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training cites this paper.

Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:46:56.731010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T01:53:04.715108Z digest=sha256:fa62f861f9b01c0067525dc412078c175416dc5146005efcd9da529cc67ee1d7

Observation b36f3396-bd91-48a9-827c-ec445c2739c4 · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:17:25.643954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:4237e92e7b68f8d0e498f170c9b95375385825ef8665fe3277170fae306c66aa

Observation beb7a600-b152-4956-909c-b2676abab6d1 · inbound

MultiHashFormer: Hash-based Generative Language Models cites this paper.

MultiHashFormer: Hash-based Generative Language Models Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:23:05.586278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T04:13:05.082903Z digest=sha256:3fc2ad36b074e4b297bbc0cb18d15111b3bea3005661563fdad9e5a5932608c2

Observation 23542f01-bb32-401a-9b90-90b1e45102bb · inbound

On the Nonlinearity of Learning Rate Scaling for LLM Training cites this paper.

On the Nonlinearity of Learning Rate Scaling for LLM Training Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.775578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:b1ad4e3af736bfd5355fd82d7c3646cd47c01c4fabaa94f54301e77de9c0ba18

Observation 9e8febb8-fbee-4d59-87cf-6fc4f48cb2cf · inbound

How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size cites this paper.

How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:08:57.606538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-03T21:02:31.246432Z digest=sha256:13ee096cf27d38f1c836b5904dff51beca0ae64289a263a62ce7cbdda9ec3211

Observation f9a9a766-4339-4587-a696-5daaaf3a386d · inbound

Convolution for Large Language Models cites this paper.

Convolution for Large Language Models Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T15:33:19.545195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:33:19.545195Z digest=sha256:f76b8a8cf9a3057c4b1d8fdb013ae7ee61693df3bf69d0dc196dc9ca13e18ab0

Observation 92ff9c0d-a7ff-4c5c-9f4c-26e6bab97084 · inbound

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales cites this paper.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.820219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.820219Z digest=sha256:5d5c45426afe41e732549ed4519e244d927e41339fd8e55f7154eb6d6f126d0d

Observation 9a71e5f3-88c8-4657-8a6b-b9ed478aebca · inbound

LLM-Based Generative Retrieval for Snapchat Content Recommendation cites this paper.

LLM-Based Generative Retrieval for Snapchat Content Recommendation Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T01:29:04.360430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:29:04.360430Z digest=sha256:fe34580ecb6a05fae4ef5997d2c2b08cd0f399f1effc281950877d8b3530d17c