Pith. sign in

Paper Citation Record · LEDGER

Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2405.18392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.18392 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:32:05.909089Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5c6e8e3e-a02d-4d8b-959e-cbcb79d340f2 · inbound

Optimization Hyper-parameter Laws for Large Language Models cites this paper.

Optimization Hyper-parameter Laws for Large Language Models Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.861142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:e35dd48b78507c496775094834083a6230a812bb0a30c0ede66057624def676f

Observation df95fbbe-38b5-4452-b272-d08a71de1ef3 · inbound

PoM: Efficient Image and Video Generation with the Polynomial Mixer cites this paper.

PoM: Efficient Image and Video Generation with the Polynomial Mixer Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:47.060664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:47.060664Z digest=sha256:6c6734a40bb52e2ad2724d09ed8a0ffcb843ce7d1a373b646b9e2beac9ae11de

Observation 4912248d-357f-4082-a29c-8c9bb19db3b5 · inbound

INTELLECT-1 Technical Report cites this paper.

INTELLECT-1 Technical Report Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:26.398063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:44:26.398063Z digest=sha256:a14f413b76efeb6a90fe142c2fd358f933a1abdeba05e920aca5eac9a12c85e5

Observation 06527b53-fad6-4c37-9a85-92dbe8aa7470 · inbound

Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs cites this paper.

Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:18.736521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:18.736521Z digest=sha256:e8459d56a38088bc5b42a62f66dec808a3cc7cf4a398abc1c1255859a31651cc

Observation d8d53b71-3169-4cce-ae0e-ce8fad15b0d4 · inbound

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference cites this paper.

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:46:46.973215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T17:46:46.845424Z digest=sha256:edd700ecea3e8d4761f70f8c81384a9845c1f52bb1c23680150273fd438af7e7

Observation 9faea4ab-6f22-4e48-8bd1-ffc020328510 · inbound

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference cites this paper.

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:46:47.063757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T17:46:46.845424Z digest=sha256:8c61f08723c91119ae3bffa5e97bc4e84fae57e1712d612dae01e58352f05bf3

Observation 6a403ff7-89b4-4607-92e6-f7666643267d · inbound

YuLan-Mini: An Open Data-efficient Language Model cites this paper.

YuLan-Mini: An Open Data-efficient Language Model Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:17:55.601448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:17:55.601448Z digest=sha256:9e967905a578e516dff6deeaf56fc5b128765c631935aa66cece92434970cbec

Observation 112eb45b-9638-4c86-8dcf-ab089d0b216b · inbound

METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring cites this paper.

METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:23.584880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:23.584880Z digest=sha256:a75c8f58c362b33f9d9b1a2f251b58583450a4e1a35f48d38389751ace4f4817

Observation ed6f8b7d-04b4-425d-acf6-25bc03cc08c1 · inbound

Proxies for Distortion and Consistency with Applications for Real-World Image Restoration cites this paper.

Proxies for Distortion and Consistency with Applications for Real-World Image Restoration Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:35:42.346497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:35:42.346497Z digest=sha256:8638bd21845da95c15d2d4d5cd9ff94d404655206e6a9bc4938d4695fb49126f

Observation 8481af53-3f09-4cf1-b860-802cd22b8d79 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 177

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:30:02.959073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:5e56f0e9669edd66daec7f84c5407ad2713c2ad40c7b93ab0895ee84df1ec9fb

Observation 50ca0f07-5562-4a8f-9e7f-74178882b260 · inbound

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient cites this paper.

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:23.521445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:06:23.521445Z digest=sha256:64c62429dbb7e8d55e87f36955374c8a6fb37a7223d0873ce8dd2499ffa22683

Observation 0b1cf965-f320-43ba-be7d-75782d723954 · inbound

Trillion 7B Technical Report cites this paper.

Trillion 7B Technical Report Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:32:05.909089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:32:05.909089Z digest=sha256:4995d9b34131a0bd0e138ec151c30a6b9608bf7a10e963af89eb826d7e1c78db

Observation 42786b4f-a28b-49c1-8c5f-56eee43f816b · inbound

BioVFM-21M: Benchmarking and Scaling Self-Supervised Vision Foundation Models for Biomedical Image Analysis cites this paper.

BioVFM-21M: Benchmarking and Scaling Self-Supervised Vision Foundation Models for Biomedical Image Analysis Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:37:04.412573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:37:04.412573Z digest=sha256:d223a12a6a6818ca34ae31e605cea5490e43bce0739d1524f438cb8ee06c3bf6

Observation c1428323-a73c-4f23-8e36-d0a5ebcdb450 · inbound

BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP cites this paper.

BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:41.884536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:21:41.884536Z digest=sha256:1fb722dc9a0a67909a76bdefaba4ea85bbdead04e3c0e0280e07c8ad130a895f

Observation bbf646a2-602d-4105-bf75-187a47e1043a · inbound

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling cites this paper.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.555421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.555421Z digest=sha256:e33441a73e7b4e6ba711050ea5b7b8f5844f4cdbdf4991052c712623d39d0000

Observation a237938b-8a27-45c8-befd-83775802303a · inbound

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements cites this paper.

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:24.529827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:24.529827Z digest=sha256:eebff4e80d6d7ff471ffac4c3abbf8aabde5c32296c7d7be4e79de268498d0be

Observation a50d57de-612e-4e2e-b028-99df93492290 · inbound

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling cites this paper.

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.960524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:51.960524Z digest=sha256:774a185b71b720ad7452e0e2c48a9f565d636073e0155c770f9a041c21b13d04

Observation 25ab7824-863b-428a-92a9-768cb60b62eb · inbound

Analysis of Schedule-Free Nonconvex Optimization cites this paper.

Analysis of Schedule-Free Nonconvex Optimization Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:42:10.446190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:42:10.446190Z digest=sha256:760c7217477fc049c8dd291f8c5f62c382d7c79c3b114c5e541f1327a83112cb

Observation f919b332-930b-47ca-94f3-762fe9bce088 · inbound

Foundation Models for Discovery and Exploration in Chemical Space cites this paper.

Foundation Models for Discovery and Exploration in Chemical Space Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 281

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:52:24.910606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T05:52:10.848118Z digest=sha256:6ad6ea2d31c45a885f98ec54918a5a57c0417bbf058823671aebe1f5c02b4eb1

Observation faa0eda6-b091-4192-941c-e878bdf1ea6c · inbound

Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems cites this paper.

Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:56:02.214483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:22:25.943097Z digest=sha256:915bc399ac294b3802b570f59f8c819b2362e7e2bcf1524a08ecf38102d6f9f0

Observation 0d3a9710-b345-4969-9776-c3b3da0ec585 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:48:00.932636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:097d65465530b7a93725a5fa16a7798388a5b365d431abe605df622719dde404

Observation 1c0c4aa7-a990-4f60-8fe7-430a2084560b · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:27.369790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:79a87fe04dffefdab49d5c7cfe40930b25bc50b0dd053591443938a2af545d32

Observation bc7b8a2f-2565-494d-a659-9a50a5bd435c · inbound

Anytime Training with Schedule-Free Spectral Optimization cites this paper.

Anytime Training with Schedule-Free Spectral Optimization Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:40:24.418964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:38:16.958574Z digest=sha256:90693448c9bed6a6cb1dc25a134173e714a30868c1d9e330a4dd8001271a362c

Observation 96f3dcb0-b37b-470b-a982-57501ec69192 · inbound

Mellum2 Technical Report cites this paper.

Mellum2 Technical Report Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:02:46.500620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T22:58:35.397914Z digest=sha256:20cedadc4c37c1abd1d16265944ee314fa833af6964aa3c1472c1ebc064073e0

Observation 137949ea-f578-451a-b46b-ffe82b532b32 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.784468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:b8c1a121fe8a903236425059b5ee213c8745634458cd5363020376151417e507

Observation 605173fd-8a1f-4072-8768-fee7b95d5b2d · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:06.866381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:06.866381Z digest=sha256:030481160605249b6ea6b6d18afc3eacbbb3dfcb70b0b35eb519b6f6655e81be

Observation db6076a0-f5b2-49bf-89e4-fff2dbae6c71 · inbound

Scaling Point-in-Time Language Models cites this paper.

Scaling Point-in-Time Language Models Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T15:39:37.355927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:39:37.355927Z digest=sha256:1e0905afd5fa3842f0488d2cf6f7adbc560f90f3891f8a06a364e11d9b53453f

Observation 861e5366-811d-4e69-be36-4504f6cc5566 · inbound

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts cites this paper.

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T21:08:06.445085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:08:06.445085Z digest=sha256:c540201f59e1468646263a8e65bceed8ba9cfb7b4e1b32389054a86a135007a0