Pith. sign in

Paper Citation Record · LEDGER

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks

As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2507.02119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02119 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:48:55.792391Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T06:58:38.927268Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T07:00:43.255075Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f8507b7-d662-4d9a-a6f9-a37c495b3814 · outbound

This paper cites GPT-4 Technical Report.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:48.890013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:48.890013Z digest=sha256:8d9562fe16938c0fc7bd14f0ce2475e66ceea54e878883e8532ce43d58e4837d

Observation ff16d3b2-6874-4029-bc2a-03c5d1a16f45 · outbound

This paper cites and Fisher, D.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Fisher, D

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:03.332256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:48.984871Z digest=sha256:f4b44c7decfc92858faf42cb89e8787cda938c0bf0c82a7957d6fa155103c629

Observation d210d09d-bd03-4756-84ab-8877e7587d6b · outbound

This paper cites High dimensional analysis reveals conservative sharpening and a stochastic edge of stability.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks High dimensional analysis reveals conservative sharpening and a stochastic edge of stability

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.116497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.116497Z digest=sha256:b532c8cdaedb08cee618a1c632ec4dde508d938561abf7b5676ef1e10e12d2f3

Observation 803ba54a-c036-4b78-a89f-9aacfaba5bff · outbound

This paper cites Power lines: Scaling laws for weight decay and batch size in llm pre-training.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Power lines: Scaling laws for weight decay and batch size in llm pre-training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.230831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.230831Z digest=sha256:b76a2e71aaed297a796ca83d758b3e052b64fd1b0cbf8907af1c4de0176daa17

Observation 19d23896-26c3-4357-a7ee-da726da3d950 · outbound

This paper cites Finite size scaling analysis of ising model block distribution functions.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Finite size scaling analysis of ising model block distribution functions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:03.061062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:49.341620Z digest=sha256:39ae9873e15612f6fde1f20da2c643c3f53cbd3f68df8ac2f32d184faa9bb310

Observation 00ddcf25-616c-4d98-97c1-562523c70530 · outbound

This paper cites and Pehlevan, C.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Pehlevan, C

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:02.877934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:49.465906Z digest=sha256:a09cf0699a3503a14e792bb82eea52e15e907d2b737073352bb1ed2eb7d04614

Observation 7c5eee10-f120-4418-be25-ed004d944912 · outbound

This paper cites Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.634542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.634542Z digest=sha256:3a401d4a2139526e6ac9e778682821ac97be16919d48b0a55e6ed8d9533825ef

Observation 34cf55a9-2ec1-4ee3-979d-31fb7e50ae16 · outbound

This paper cites A Dynamical Model of Neural Scaling Laws.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks A Dynamical Model of Neural Scaling Laws

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.730886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.730886Z digest=sha256:87e9f518fc51d332a2b7f12c705e3a3a6bc5aeb81f41101093c837d460864ff4

Observation 14832c06-88b9-461a-8b12-2d859f17a77f · outbound

This paper cites How Feature Learning Can Improve Neural Scaling Laws.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks How Feature Learning Can Improve Neural Scaling Laws

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.856090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.856090Z digest=sha256:fb5bacc71225d84d3f2fba5bb55c2df6b3e501863871171ca8fcd8039a37409f

Observation 03d58112-1883-4f77-943c-d1f8987069cd · outbound

This paper cites Infinite limits of multi-head transformer dynamics.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Infinite limits of multi-head transformer dynamics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:02.507256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:49.979212Z digest=sha256:47fb1cdb47a4f245a3298c6ab41bb3040a9f5d8b1750c2470934dbd94424f0b3

Observation b2a49a53-e9c1-4dc0-bdee-154cb355f9c6 · outbound

This paper cites Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.108888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.108888Z digest=sha256:ede063d16f99b4e4f66281c82f9608436c6894ff3c23e5c26cdef4f359726a1e

Observation 4d05b138-7864-4521-a851-7c0cf4d60ba8 · outbound

This paper cites Adaptive Gradient Methods at the Edge of Stability.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Adaptive Gradient Methods at the Edge of Stability

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.237889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.237889Z digest=sha256:6077a68fbf546694046e092b835aa8fbf462b7f5b53a916650eb149fd3dc7051

Observation 6e76a5f3-8709-4694-8311-a01ccecbc2c8 · outbound

This paper cites M., Damian, A., Talwalkar, A., Kolter, Z., and Lee, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks M., Damian, A., Talwalkar, A., Kolter, Z., and Lee, J

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.344017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.344017Z digest=sha256:3c02e7d200aa7d676905e0238b1a3656100a72d6f003b284c5bc2602af6ddc10

Observation 80156d54-db79-46f4-8069-df4302dbccae · outbound

This paper cites Optimal learning rate schedules in high-dimensional non-convex optimization problems.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Optimal learning rate schedules in high-dimensional non-convex optimization problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.480589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.480589Z digest=sha256:dea4b9bca6063f26d3b75eb3fcb5f4cc0d4950fb97dc9ab897ff0b299c18db75

Observation 9502e4e6-1a62-40ee-8cb0-5ed9eaf6c731 · outbound

This paper cites C., Noci, L., Li, M., Bordelon, B., Bergsma, S., Pehlevan, C., Hanin, B., and Hestness, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks C., Noci, L., Li, M., Bordelon, B., Bergsma, S., Pehlevan, C., Hanin, B., and Hestness, J

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.592169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.592169Z digest=sha256:6f235ac14e595810e2a129be0f68d7af6a049f436cd678bdf742cab6654a660f

Observation ad86be4a-4417-4076-9c0f-e4185cd2952a · outbound

This paper cites Kingma, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Kingma, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:02.206565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:50.697194Z digest=sha256:b62bbd6f6a2695ce0cbf90d0547b78cf61c1c666f8203c051adc161e046f8362

Observation 0231a66f-5898-4deb-a78f-3976de64ab0d · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Scaling Exponents Across Parameterizations and Optimizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.834417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.834417Z digest=sha256:b8bc22cebb2e397590c1a79e380973db8da5c4efbbd9c9ed26b4a4202694350b

Observation a27cd3c4-c959-4c6c-ade8-6651aebda303 · outbound

This paper cites an unresolved cited work.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:49:01.969127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:50.982931Z digest=sha256:061f7a146c225a2fdb6a2608fc5e9ff1db007c61875eeabbf67b69aeeafbf873

Observation 9b472174-aad9-4b29-a6f6-280c995a9936 · outbound

This paper cites Monte Carlo methods in financial engineering, volume 53.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Monte Carlo methods in financial engineering, volume 53

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:01.602856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:51.135576Z digest=sha256:cb03ab001ee580655a84f7ba1e0f0e093f2d086509dca56637ddbd3c99348b7b

Observation 894f4596-0b96-4b61-a659-2679eddf39c2 · outbound

This paper cites and Fisher, D.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Fisher, D

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:01.275473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:51.288401Z digest=sha256:8fa236c6f36fab89aacb5b0f233268364264c2a849f76091d0bc753e1b496560

Observation 01e15fe3-4ed2-492b-8957-bad8c1daf31e · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Gaussian Error Linear Units (GELUs)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:51.439168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:51.439168Z digest=sha256:8e721d23c56658deab2adf934316224fcb8d22da9d9d27f03e1d54761c8b1abd

Observation 4e936d7a-dbe8-4e16-ba21-53b8d78e44a6 · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Deep Learning Scaling is Predictable, Empirically

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:51.541070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:51.541070Z digest=sha256:5dfee6fe27e37930b7f4da78a8ed63cb73d90ed0a23fa3f647b78f4814eeaa14

Observation a53d9abd-f9f5-46f0-9a61-462a98c5354a · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Training Compute-Optimal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:51.686276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:51.686276Z digest=sha256:5f34eacce41f45b21bd5c3badaf0e65a74f194fa35f6d19c256c94b2267f2779

Observation 6c3c0699-c49f-492c-a582-7a53930cd057 · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:51.813397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:51.813397Z digest=sha256:bee26aadde83f73bb066c52beb254f1e8d2991e146d56f48da5f46b5d516eb53

Observation ff42298c-b9e7-4d0e-9192-5040dfe2c84d · outbound

This paper cites F., Blundell, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks F., Blundell, J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:00.969218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:51.957254Z digest=sha256:015d9b36ddd8ef95d51b855deae085343a1b2239176cfb42ad9d11bd19c7681d

Observation 285b1657-f9f6-4e76-9a21-83e8c78a68f3 · outbound

This paper cites Stochastic modified equations and adaptive stochastic gradient algorithms.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Stochastic modified equations and adaptive stochastic gradient algorithms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:00.702563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.040916Z digest=sha256:b7aa5c65d1863b82c6d08e0d4b952674e11b06c15cd8785ab85db7891d11cd08

Observation 1541ff55-00bc-4e67-beb4-0ec9ee48de10 · outbound

This paper cites A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:48:56.444458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.152237Z digest=sha256:8d45a57b3bde8016a23e642bd98b06417ab8fbdac2bd2274d9b02e6b8768e2d9

Observation aaa4604b-ec14-4b6c-be1a-efdd3c1a479d · outbound

This paper cites On the sdes and scaling rules for adaptive gradient algorithms.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks On the sdes and scaling rules for adaptive gradient algorithms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:00.420054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.302674Z digest=sha256:e1de708e54524a1bb3577c12e9c8dd283a0d687af9b9cfe01f08c735ad98368a

Observation f77b1822-0677-4628-920b-ba8297a7df0a · outbound

This paper cites and Pastur, L.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Pastur, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:00.239937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.402249Z digest=sha256:85794681c0b5334d57deee56a7b6f1370234d881e0c975775812eb7575683fb5

Observation 3593f90c-ab38-4294-a77a-dbf7cb00324e · outbound

This paper cites An Empirical Model of Large-Batch Training.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks An Empirical Model of Large-Batch Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:52.527254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:52.527254Z digest=sha256:bbf30a6c585aed2476b47750447614e10f5fef92fb6b4261066ded468c77b3d5

Observation 9d7c7f74-d112-40fa-87e1-a8fce35df8b4 · outbound

This paper cites Y., Singh, S., Bhatele, A., Goldblum, M., Panda, A., and Goldstein, T.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Y., Singh, S., Bhatele, A., Goldblum, M., Panda, A., and Goldstein, T

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:52.668462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:52.668462Z digest=sha256:9fd5170571f70b7000d011988bb75abd95a999bb08b1eb227c2ac0e0cfe63734

Observation 04774ee9-505a-45a6-9571-c86f9dea0c0d · outbound

This paper cites The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:52.826035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:52.826035Z digest=sha256:2c76d7160dfead933b9275d738878dacf5bc59b3da1c5a2880b603682997c43b

Observation 32a46c09-50ef-497d-b2dd-898847fea43b · outbound

This paper cites Super consistency of neural network landscapes and learning rate transfer.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Super consistency of neural network landscapes and learning rate transfer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.986705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.964417Z digest=sha256:fdc59988af1f8ba9d5255ab0cd7d589978975cbecebd248f15f194f3f286798f

Observation f53c09ca-dfd3-44cb-a19f-41b65fc5e507 · outbound

This paper cites Sgd in the large: Average-case analysis, asymptotics, and stepsize criticality.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Sgd in the large: Average-case analysis, asymptotics, and stepsize criticality

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.812856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.062783Z digest=sha256:38929dd3cacda409fd87bf8f1da4afa326fcd2c83eca38b66a1ddabffd9f245b

Observation 6a579c77-06eb-48d0-a7fb-542382b48a89 · outbound

This paper cites Homogenization of sgd in high-dimensions: Exact dynamics and generalization properties.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Homogenization of sgd in high-dimensions: Exact dynamics and generalization properties

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.564827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.171825Z digest=sha256:45f1a64d8201c3e6babbc7d67a426b6d5e93508d205bc74f2e173a9dfca6a5d2

Observation 17e6d158-514a-450d-ad6c-12a4ecc631b6 · outbound

This paper cites 4+3 Phases of Compute-Optimal Neural Scaling Laws.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks 4+3 Phases of Compute-Optimal Neural Scaling Laws

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:53.293116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:53.293116Z digest=sha256:ef759e682efdea105be895a337eb24e37152fdcb51e2c3c48815fc58f7c95b36

Observation ae4e746d-a68b-4e00-a98f-f1a8a904de60 · outbound

This paper cites T., Agarwala, A., and Fisher, D.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks T., Agarwala, A., and Fisher, D

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.330413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.414115Z digest=sha256:85e3cd55efb2971537c7963c6c18f0bf50113bc6fbe63a462ad40d895a823eef

Observation 5be144fc-b5ef-41b9-bb18-d0db1e2f5d94 · outbound

This paper cites Reconciling Kaplan and Chinchilla Scaling Laws.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Reconciling Kaplan and Chinchilla Scaling Laws

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:53.523148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:53.523148Z digest=sha256:4da63801d9f801eafeaffb884fc5bb6cc8a6df5dbca7c7c7f4002321979643c7

Observation 81426342-a206-4235-9622-0e17160b76f0 · outbound

This paper cites J., Davison, M., Bhaya, D., and Fisher, D.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks J., Davison, M., Bhaya, D., and Fisher, D

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.064308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.640808Z digest=sha256:4b41dc5f5ab4266b43312a114605277ba54c9bc336ef894fcf33072de20d6dcb

Observation 524ba4f5-5eea-4805-bd48-9394a5ddabe6 · outbound

This paper cites The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:53.764373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:53.764373Z digest=sha256:7aa73a113d5c5eaba3c0af7b223fafb3e9174003aecd011d62eca5b55e3f9b7a

Observation 671b23b1-1324-45fa-af71-19fe0c84092b · outbound

This paper cites Ubiquitous abundance distribution of non-dominant plankton across the global ocean.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Ubiquitous abundance distribution of non-dominant plankton across the global ocean

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:58.772167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.834070Z digest=sha256:92dddd933fc2275971e6c63ac54b362468617680d744eb985fd592754bd82f08

Observation fa376d3b-72b2-4d46-8bf4-f6869b571edc · outbound

This paper cites and Kaplan, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Kaplan, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:58.494464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.926114Z digest=sha256:7b764eb5264f616ac2d9971eaab80b402876e26b611de4e337663d897543c5b3

Observation 4091f172-60bd-41e7-a799-b1009cb2f517 · outbound

This paper cites Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:48:56.060606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:54.043892Z digest=sha256:7d106e271294f37da1c99d1f33f4bcef600a25e6e21a083cc755495f44806393

Observation 0d53d82b-059b-4a33-be5a-6c88d9f6f97d · outbound

This paper cites Scaling Law with Learning Rate Annealing.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Scaling Law with Learning Rate Annealing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:54.159538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:54.159538Z digest=sha256:cb3bbd8cfdea06dc7fbdfe587b808d88d080c7126dcbbbbca985cf3b9962bf2a

Observation 8b73fa26-db25-4502-92c4-30c46691e5dc · outbound

This paper cites R., Geiler-Samerotte, K., H \'e rissant, L., Blundell, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks R., Geiler-Samerotte, K., H \'e rissant, L., Blundell, J

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:58.140034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:54.271428Z digest=sha256:9dfc31a8d7012741b324526f3a2f342fb73ce0c0e1cda794418a0fded312efba

Observation dbb367e9-4461-4699-9c01-b99aa9f5d739 · outbound

This paper cites Feature-learning networks are consistent across widths at realistic scales.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Feature-learning networks are consistent across widths at realistic scales

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:57.822552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:54.396692Z digest=sha256:f0dee006d4254938eb48fe39259134f0989e0d342340155221b568ac50966fb1

Observation 16c13459-0ce0-4da9-8d2f-dab0399daa4c · outbound

This paper cites How to set AdamW's weight decay as you scale model and dataset size.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks How to set AdamW's weight decay as you scale model and dataset size

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:54.538611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:54.538611Z digest=sha256:3ff0aa94d30074ca7052c6333c4ed0409d14da9ca52ece351dbfdb40addd92e4

Observation 8e56a2c6-68f4-468a-b2d3-e557308434b4 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:54.682273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:54.682273Z digest=sha256:03b16b1668ec72c37d703cebaba481f0267cec2a8650e95c7fdf95b82d02d360

Observation 675f7208-5acd-4a1d-bed0-13fa8463fd56 · outbound

This paper cites an unresolved cited work.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:48:57.542119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:54.825321Z digest=sha256:090e37dd0fc01ba9b1cea87cdc4a1c6991d1849b79ebeac06ce8f518c1374e4d

Observation e268f4b4-f837-49ba-a19f-7c5fdae47cf3 · outbound

This paper cites Small-scale proxies for large-scale Transformer training instabilities.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Small-scale proxies for large-scale Transformer training instabilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:54.971531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:54.971531Z digest=sha256:fcd44bac45d20111655c963c8982f43c7ea5931f38f8450d47fe7e1c5ccc245a

Observation daf49d49-f471-46bf-8352-787c0b0f9991 · outbound

This paper cites Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:55.089882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:55.089882Z digest=sha256:8cb5ad8eb742973b9c72dc7e87ea8b2cbe74216490834b7f4281101a9fabfdb7

Observation cf60c245-08c2-470e-b2e2-307e619503a5 · outbound

This paper cites and Hu, E.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Hu, E

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:57.192764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:55.214821Z digest=sha256:8674307a66929da12062d15f45053b7056d8dc8140ce372e6303cad12fde92c4

Observation 6907316e-e6d2-48a5-97e1-881565c4a1ec · outbound

This paper cites and Littwin, E.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Littwin, E

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:57.010570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:55.349785Z digest=sha256:300d555c2dcc5cb683354efa3f10c73098987c0753ac1feeacab2011e64a11f8

Observation b05f31e8-fb3c-44e1-b053-f8a9f689ea30 · outbound

This paper cites J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:56.863551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:48:55.507002Z digest=sha256:cbc8eeb815174dc9a83dfb309ed11180b7c1e0404f6a6d003b8d75215ceddbcb

Observation d6b96640-4afb-48c5-a528-1ee0a2da6836 · outbound

This paper cites and Sennrich, R.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Sennrich, R

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:55.635949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:55.635949Z digest=sha256:d7d328bfc7121ed836c0d854330a18702b7c77188b577c308bddd7938b6fc3c8

Observation ae15619c-a965-4477-a706-d53476ec9131 · outbound

This paper cites an unresolved cited work.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:55.792391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:55.792391Z digest=sha256:5a128b61e69f3e1ebc116ff71464b93296c568bf99aaa1620609cf22b8f8268c

Pith citing papers

Observation 41fd3f8c-d9fe-418e-b9d8-5b79dd298a3f · inbound

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model cites this paper.

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:00:43.256809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T06:58:38.927268Z digest=sha256:0faba64160e9e9b4ef1730719f81d2da0db77afe92e9e5a00921768551086f7d