Pith. sign in

Paper Citation Record · LEDGER

Scale Weight Decay and Train Better

As of 23 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 0 inbound Pith citation observations for arXiv:2607.23777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23777 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-30T12:53:41.205912Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved73
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 779190fd-96e4-45cc-b818-7c5e72291d75 · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.043061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.043061Z digest=sha256:c10adf1cbfde0d0d58d85a076f11c60cbee40ac9e0dc08c31fdffbb5779a2d0b

Observation 276f9969-64b7-42fd-9f5b-83ec6d6a3308 · outbound

This paper cites Scaling Laws for Neural Language Models.

Scale Weight Decay and Train Better Scaling Laws for Neural Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.226740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.226740Z digest=sha256:789765bd8f3d6b0822cbec545faa296d5040d8fcf13a4b390f9c74a7d1941af3

Observation 1ec2ef78-4e4b-4ff2-8dd9-ac917bbf00f5 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scale Weight Decay and Train Better Training Compute-Optimal Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.363542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.363542Z digest=sha256:32c0e6c9f6f0e195e99695d8a4ce8490c6cead031c6a6aef8b48648927648e53

Observation a77a1adf-f488-4b0f-83a5-4a20f16a1d13 · outbound

This paper cites Explaining neural scaling laws.

Scale Weight Decay and Train Better Explaining neural scaling laws

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.488793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.488793Z digest=sha256:8cddd476b770b7e1d355c160b461a827eeb0e47ec653bffa0a9b065e22c6975c

Observation 7cf1cc50-7468-421f-aad3-e21fb22704e6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Scale Weight Decay and Train Better An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.597117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.597117Z digest=sha256:a93a0447cc102037bb7dde008e656ea7f61a6a0ff3e481bb7be1a12583c7448a

Observation 55b13f6c-2c76-4e87-b39a-92a8567d0e1e · outbound

This paper cites Scaling Vision Transformers.

Scale Weight Decay and Train Better Scaling Vision Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.706869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.706869Z digest=sha256:724c36543619a5d813699c263b820e8e88571e0cc4126195003ef2066085f539

Observation 4fd4085d-e984-4019-ab2c-1d6c49ae3d4c · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.797871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.797871Z digest=sha256:597f83cfd9219c4b1f78f1be0cabdb128f2bdb45e324f3f0ed4b921701495769

Observation 8e874d25-fdcd-44f7-8324-a1de5c0199b4 · outbound

This paper cites Neural Scaling Laws in Robotics.

Scale Weight Decay and Train Better Neural Scaling Laws in Robotics

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.884696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.884696Z digest=sha256:b2de6d1d10edecef9a75b5f62201302dc3bdafecfbd025d81293e76175bd5fa9

Observation 11f0e4a3-670c-453c-8ddd-786047876460 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Scale Weight Decay and Train Better Kimi K2: Open Agentic Intelligence

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.891094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.891094Z digest=sha256:102b39a33a034b9e977dce6b0c2c0872c6d5e1e30bcc77d69436ea7025e6c1fc

Observation ccc70d7e-a03d-487c-8525-aa737e8f1e96 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

Scale Weight Decay and Train Better GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.895797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.895797Z digest=sha256:ff29265e1eefb64431768df418a35a6e385f4fe505a4c0db6204d2d93cb3238d

Observation d353d982-f150-4baa-9ad5-6b6b537de4e2 · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.900324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.900324Z digest=sha256:b2d4ff5f7c052fa26b54173942775f6c77e610c594c79db9eee8c61c8ef6d591

Observation 0669252d-c402-42c7-bd53-32a758d5c53c · outbound

This paper cites A Simple Weight Decay Can Improve Generalization.

Scale Weight Decay and Train Better A Simple Weight Decay Can Improve Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.905115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.905115Z digest=sha256:791e5360dda888b7aaeb6df1613018cdb643fd212756e042f4782a232e94e478

Observation b48b4df0-84dc-45e0-8e09-424743d8b37b · outbound

This paper cites Why Do We Need Weight Decay in Modern Deep Learning?.

Scale Weight Decay and Train Better Why Do We Need Weight Decay in Modern Deep Learning?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.909600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.909600Z digest=sha256:1256193f4d34b9c9275a24298b97b6785bc647fdd01533cfd75caef7d642fe21

Observation 2cfdcf2d-ac1d-4ba0-9c3e-9e5461e0649c · outbound

This paper cites Decoupled Weight Decay Regularization.

Scale Weight Decay and Train Better Decoupled Weight Decay Regularization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.915205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.915205Z digest=sha256:137ee8abe2375a6cab16910e53a95133c92b0d2677573ac8e107dc0ce73a6b36

Observation 9737d18c-75bd-4ee0-bc67-7445f196d449 · outbound

This paper cites Why Warmup the Learning Rate? Underlying Mechanisms and Improvements.

Scale Weight Decay and Train Better Why Warmup the Learning Rate? Underlying Mechanisms and Improvements

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.920881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.920881Z digest=sha256:efd14f641013d8a9dfca0e2fede160ced4a26c57504640a98ffdae6290d0a5d4

Observation 278735a2-2063-4ab2-9f53-c192fd02a7ac · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Scale Weight Decay and Train Better SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.926826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.926826Z digest=sha256:8d9dd9f697a8fe4e169b01f740853e539a7b72213534503d8692e58647d53dff

Observation 03e11547-df7e-45e9-961d-a968c1cc15e5 · outbound

This paper cites 2024.url:https : //kellerjordan.github.io/posts/muon/.

Scale Weight Decay and Train Better 2024.url:https : //kellerjordan.github.io/posts/muon/

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.930943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.930943Z digest=sha256:9cb37090c10d00ea26575f591807c7acdba8955790e753fcd7f968752040b1d1

Observation 8a26d2d7-3548-4363-826a-20840195459d · outbound

This paper cites Muon is Scalable for LLM Training.

Scale Weight Decay and Train Better Muon is Scalable for LLM Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.936018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.936018Z digest=sha256:08d90a90d5f025d8861056eef01dda0671f1d9e165de364877e5a3fc3a142ddb

Observation 3cd153e1-72dc-4c18-94d5-1872b22e2f1e · outbound

This paper cites A Stochastic Approximation Method.

Scale Weight Decay and Train Better A Stochastic Approximation Method

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.940143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.940143Z digest=sha256:c90c1ad6e5ce8b729b38ef44e4666c8fd0ead1a622db3d7a375655e2871d9578

Observation 5e669c20-a196-457f-8b62-74c44e0c2476 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Scale Weight Decay and Train Better Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.946548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.946548Z digest=sha256:e0f8c30750ee09fb8500fcf347e134787302c837c5c9a81acfae93dc9ce81ad6

Observation 18742985-3239-465c-a100-fa1b65c49831 · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

Scale Weight Decay and Train Better Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.950533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.950533Z digest=sha256:f24fd97eadaf5fbd81645e26c9ec57637459aba78a7507c60a14a0db350f7912

Observation cf743017-35ec-4b6e-a066-38f82f0bb173 · outbound

This paper cites Why Gradients Rapidly Increase Near the End of Training.

Scale Weight Decay and Train Better Why Gradients Rapidly Increase Near the End of Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.954445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.954445Z digest=sha256:9c55c0329d67dc188f0ad3cc769c0270f9570ecfb317a4cc86fea8797b9f6bc8

Observation 875cf34d-83ba-4843-8eb2-c3b3de5b968c · outbound

This paper cites Shampoo: Preconditioned Stochastic Tensor Optimization.

Scale Weight Decay and Train Better Shampoo: Preconditioned Stochastic Tensor Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.959117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.959117Z digest=sha256:7666c588e081627615eeb71349f3bff4ab07d2719f03c520209f9e90ca074f71

Observation 142487d7-9078-4488-8392-681cf80f7dc5 · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

Scale Weight Decay and Train Better Scalable Second Order Optimization for Deep Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.963582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.963582Z digest=sha256:0e717dbff4edf839d95980cbde455286a10f6c4246d88733fba8cf0fbb459dad

Observation fb89b922-ff20-4ccf-be67-3ca20a3dccef · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Scale Weight Decay and Train Better SOAP: Improving and Stabilizing Shampoo using Adam

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.967581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.967581Z digest=sha256:2c94619a5183a912a438dc53f1d5696bd4f0463d4c54504f7be1b6d80c48e550

Observation 98613942-ad1f-44a9-9718-ed96c96bd207 · outbound

This paper cites Aurora: A Leverage-Aware Spectral Optimizer.

Scale Weight Decay and Train Better Aurora: A Leverage-Aware Spectral Optimizer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.971261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.971261Z digest=sha256:07fc2e4c48f3ee97e7bf92b6d820ddf6f9908a10acafca648f50e5328c211fab

Observation ffc2807d-03c3-4f46-9f86-c9d5212623e6 · outbound

This paper cites Dion: Distributed Orthonormalized Updates.

Scale Weight Decay and Train Better Dion: Distributed Orthonormalized Updates

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.975269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.975269Z digest=sha256:c859d8b2272fd2a85cce810a70230f0a2a70ccef2058074c0e7c8fa7eadbf80f

Observation 483bd6a9-0653-4ad7-8ffd-2b40df48d89f · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.978836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.978836Z digest=sha256:49ed7a6b0080b09b14e4c58e4cb683cf456863c38590ddeeb1de34cfa602e7b1

Observation c9973faa-5435-47d3-a70a-ed5cc56f6bac · outbound

This paper cites The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm.

Scale Weight Decay and Train Better The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.986961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.986961Z digest=sha256:d0c43965ac8ccbfcc2d0d8de8a22552b377e8be71cf90e7a1d031dfe4bdec061

Observation 911aa333-4fbd-400d-93a8-4a6c81f74004 · outbound

This paper cites Anytime Training with Schedule-Free Spectral Optimization.

Scale Weight Decay and Train Better Anytime Training with Schedule-Free Spectral Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.992720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.992720Z digest=sha256:1e430da1729f16bfc92166b38f6b4d10c1c5dc57f3969ed82a6b614edbf6bf8e

Observation 2c7fb9f6-0bcf-42ec-be46-d5395a69548d · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.997411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.997411Z digest=sha256:f1e1bb0df130614a1cb8384142b49deb52e28d5d20ecaa5ed898089e1a7c0eea

Observation 43073ace-70a4-47b1-8997-4d7465993ec2 · outbound

This paper cites Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning.

Scale Weight Decay and Train Better Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.001602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.001602Z digest=sha256:df89ebbcba4af374e72d9ca8c54d39f3824711f36de44e70a17f1e11ddfef012

Observation a23e4123-d953-48da-9821-295eba574a0c · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.006795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.006795Z digest=sha256:c23a4ad623a6ba5d1b7a1dcf5d000830ad623a3d291f552ca05a0fe591eb5207

Observation 92e69bab-9d3e-4f8c-b25c-0c5f98a6dd57 · outbound

This paper cites Optimization Methods for Large-Scale Machine Learning.

Scale Weight Decay and Train Better Optimization Methods for Large-Scale Machine Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.012922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.012922Z digest=sha256:20c0e93a6890279d493f39fb1be4ebcbdde499c3f392d4a747dfc7b1619ffd82

Observation e04770ed-0bf2-400d-aa57-03816929c724 · outbound

This paper cites Muown: Row-Norm Control for Muon Optimization.

Scale Weight Decay and Train Better Muown: Row-Norm Control for Muon Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.017797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.017797Z digest=sha256:f5f9134e9ddbdf7093ee17b1bb1a6de126277b078515e38e4a5c128efb22dcfa

Observation cc72e5ca-0af6-449e-8e6a-a644cf546291 · outbound

This paper cites Convergence Bound and Critical Batch Size of Muon Optimizer.

Scale Weight Decay and Train Better Convergence Bound and Critical Batch Size of Muon Optimizer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.022151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.022151Z digest=sha256:17a2717bf686a2eedfbf19303176fa731d30f1db53f4bd8f859168c479a0a1fc

Observation 9df2c4f6-02af-4a7c-bc0a-45ede472414b · outbound

This paper cites On the Convergence Analysis of Muon.

Scale Weight Decay and Train Better On the Convergence Analysis of Muon

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.026143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.026143Z digest=sha256:5efacb268b8134baf95eb37a6fd77cc2120db64a2d736fe90f56574008658b5b

Observation e5627a78-50d9-4101-929b-57ece23b4f20 · outbound

This paper cites On the Convergence of Muon and Beyond.

Scale Weight Decay and Train Better On the Convergence of Muon and Beyond

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.030100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.030100Z digest=sha256:75d28e4d0cac193a7b25b88030c9b4b1d330271dbd00bfdd153b52f6092a8df6

Observation 3155f39f-83fc-4fd8-9db8-fd674e0bf77f · outbound

This paper cites Language Models are Unsupervised Multitask Learners.

Scale Weight Decay and Train Better Language Models are Unsupervised Multitask Learners

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.034235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.034235Z digest=sha256:d5a8ca37c4c6f6755cfa6580447b11f6f25a0309a4ad3fe32ba2b3cbc7a08981

Observation fd28eaa8-3601-43c0-8170-5269eafe5a5a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Scale Weight Decay and Train Better Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.037795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.037795Z digest=sha256:670db6f6945abb27a1864a8c97ed5997d7666b034c12857fe12ef0b7aa58a1e0

Observation 38cbf6cf-6290-4515-aa64-80b98712a8be · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Scale Weight Decay and Train Better Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.041793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.041793Z digest=sha256:ead1242917d5533546f75a653e43eb9470730448519ae6a739fb7235964dc20c

Observation 0a31867d-ee51-4a9c-a149-e080d835209c · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.045322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.045322Z digest=sha256:fe3b4234645085e83194f58b358362e573483788278d7312edbd5c6bf7a7546c

Observation db4d0e64-30a4-4959-871d-4543df27f422 · outbound

This paper cites Scaling Optimal LR Across Token Horizons.

Scale Weight Decay and Train Better Scaling Optimal LR Across Token Horizons

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.048646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.048646Z digest=sha256:4cb60c2e72671fb3599ef566cb3a88116e2c1d8125ee3013273c1942e9920602

Observation b4c0e474-cbe7-43f0-8789-3ee316ab8a79 · outbound

This paper cites L2 Regularization versus Batch and Weight Normalization.

Scale Weight Decay and Train Better L2 Regularization versus Batch and Weight Normalization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.060313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.060313Z digest=sha256:2b9639167913ab7aa06b2a7aabf0ad1fda1216a4896d8fd165dfe35e2c16a2d3

Observation 08e3f009-f43f-49fc-aaa1-048424f4be36 · outbound

This paper cites Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks.

Scale Weight Decay and Train Better Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.065516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.065516Z digest=sha256:257646468f5be335ab66cb9c6300fd07f213707c07233711074feb3cca633c66

Observation 98bd111d-c63b-49d3-bf56-722ec3e78ec2 · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.070299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.070299Z digest=sha256:ddae940a037a5df54297f60aa2f89e8d2058d02539be10ed6b4c95430e7fa899

Observation 3a7f9e62-8d0b-4e5d-b464-af15068857ed · outbound

This paper cites 2022.url: https://openreview.net/forum?id=J7V_4aauV6B.

Scale Weight Decay and Train Better 2022.url: https://openreview.net/forum?id=J7V_4aauV6B

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.074979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.074979Z digest=sha256:aab723ab67ea2d99bde790f7225ba9996cb0bf378f69ded41773db201565b543

Observation 2b2cd47a-2e82-4f08-8730-f15f62c23f53 · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.079462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.079462Z digest=sha256:4c2acd338b7777cf6559862913882ae272bf1d79b8f0594e677a6f6bf11ff870

Observation ead716dd-08e9-45f8-82e0-348633c9b71a · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.088487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.088487Z digest=sha256:6e3c88051e5955cc83c9f123e6b9c8373a0a4a06b1702154ec342b3e3fbb75ab

Observation 8e0c32fe-8577-4f5d-97ec-90c1cca4c89c · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Scale Weight Decay and Train Better Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.096060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.096060Z digest=sha256:ae78dbd59096b2f0275f99180130574f71be4802b12157df9adef066731c3533

Observation 3337b5f0-3d51-4517-a330-ccb02fa01e20 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Scale Weight Decay and Train Better Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.100513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.100513Z digest=sha256:bc5559fd9a369cdfe62531b619882d143da4298c71b294d56f725163bb19c350

Observation 130db0ba-753b-4621-8371-0c81a93b997d · outbound

This paper cites Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.

Scale Weight Decay and Train Better Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.105706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.105706Z digest=sha256:fa43b9e957c5e04d37eeb66afa9c1a1eab1f7aa96b0b0a88a93b00522b5d2e23

Observation 4be47bb7-04b2-48d5-b48f-175cc1858efc · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

Scale Weight Decay and Train Better Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.111628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.111628Z digest=sha256:1c62c590c2a027db9a84a9412bcdfae4d343d0fc62dfe25a8d8a186e03289709

Observation 6381bee1-28d0-4edd-926d-8d7797d0d13b · outbound

This paper cites Parallax: Parameterized Local Linear Attention for Language Modeling.

Scale Weight Decay and Train Better Parallax: Parameterized Local Linear Attention for Language Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.116534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.116534Z digest=sha256:3c53f74508c9c1e779618060e33f711282623fc15eb88be62bf8f963a3e18769

Observation 75dac705-d999-4a3f-bec5-0e797ecd4b76 · outbound

This paper cites Training language models to follow instructions with human feedback.

Scale Weight Decay and Train Better Training language models to follow instructions with human feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.121326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.121326Z digest=sha256:fa8a1f65e04d030dfbad73c390d4dc3e5a2f8a71277dacc29c87347db20b734c

Observation b7ce3610-a63b-4f44-94d1-14a6decf3732 · outbound

This paper cites Deep reinforcement learning from human preferences.

Scale Weight Decay and Train Better Deep reinforcement learning from human preferences

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.126522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.126522Z digest=sha256:16b807fc2b5b8f3a6b5b8ea1c237c9903b94610ed1ee91648f1c8a32b2cd24b0

Observation 87455bc6-5cb8-485e-929a-4c26d58a777c · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.130547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.130547Z digest=sha256:41f6a8948b1e94389eb9025ea9091bc3923ab05911e59a43a15ae3cdd5338c30

Observation ad191bf0-4ca6-4f5c-bb49-cf36bf670141 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Scale Weight Decay and Train Better DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.139360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.139360Z digest=sha256:a76a6168e0b195d17a44a4905f71318b4bfa07182dd613b4587d7f0f977029e9

Observation 620e6b72-6b4d-4a11-92a8-b313380f6b19 · outbound

This paper cites DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning.

Scale Weight Decay and Train Better DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-07-30T12:53:41.143145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.143145Z digest=sha256:c4648ccc5b9894771477e73bcfa9e9ea81e47db0c21cb7434c82fe07588a26ad

Observation 80569cfd-2968-4e13-864e-cd06e2a97b27 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

Scale Weight Decay and Train Better Scaling Laws for Fine-Grained Mixture of Experts

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.147044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.147044Z digest=sha256:0ebdbf9cf95b5428bf9a1a966f21de1f9257e2636d9effa7f888524b7a0b1179

Observation fea8ca95-1122-4e78-ab85-9d8c13c619ce · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Scale Weight Decay and Train Better Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.134632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.134632Z digest=sha256:d132b930bb39520da09eb2aea270404be0688da03c839fc82c7d14fe32e4d73c

Observation c8dc26db-bf8f-4bb6-b57e-7126535a1c92 · outbound

This paper cites Inference economics of language models.

Scale Weight Decay and Train Better Inference economics of language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.155145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.155145Z digest=sha256:56e2f8be0e3efad7b857cc3b4b5996db58ef60f79b1adbfaa02e7087bb6262f7

Observation 4b2f1fff-d1a5-4f20-9eda-f16568a08cfb · outbound

This paper cites Root Mean Square Layer Normalization.

Scale Weight Decay and Train Better Root Mean Square Layer Normalization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.160509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.160509Z digest=sha256:3e8208014c460bbaf3f4fcd91d2be0c005cd4de8b3aef073d3d777887a054c6e

Observation ed33a81e-32e0-4c9c-8d04-753f9a6eba21 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Scale Weight Decay and Train Better RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.164912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.164912Z digest=sha256:71ff1ff5b666e5f57dc6c1f43d8c2c20723d4f83e24544ee9a60b23bccfe8dbd

Observation 36607c74-c48e-4e93-9420-5b0dbde4f094 · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.150792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.150792Z digest=sha256:cf5ad8d999146e19d09191c94de85690ae29faa69de7ef9d6923ad3d07638b36

Observation 970c7596-d80c-40ac-8060-6449687b52f1 · outbound

This paper cites Benchmarking Optimizers for Large Language Model Pretraining.

Scale Weight Decay and Train Better Benchmarking Optimizers for Large Language Model Pretraining

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.174828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.174828Z digest=sha256:9c1ee164700681d519213d0ae64524ac86b82073fdef1201e67a4b2e3b3d7876

Observation 9c8d2a59-dea0-4030-bb23-dd38c5a17c02 · outbound

This paper cites A Spectral Condition for Feature Learning.

Scale Weight Decay and Train Better A Spectral Condition for Feature Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.180597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.180597Z digest=sha256:1e1f33dea6cc7dc0154de2ecc0f5069705f45e63447549c56c7e157662434104

Observation 1e09b8ad-c1b1-4069-adca-33ab7def3d1a · outbound

This paper cites Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks.

Scale Weight Decay and Train Better Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.186767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.186767Z digest=sha256:fca25e19526d18dad9189e6c99385ee83a6b7dba6cedb7325f0a99643fc02b2a

Observation 5c428c06-d042-42ca-8493-4167a46d629a · outbound

This paper cites GLU Variants Improve Transformer.

Scale Weight Decay and Train Better GLU Variants Improve Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.169794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.169794Z digest=sha256:be18d92219c1ce2564c0bfe5941306c20eb396f502a27094cc4fb86e0ccdf3cf

Observation f5b7f52a-9f64-4d35-80c5-a5db93db43e4 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Scale Weight Decay and Train Better PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.205912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.205912Z digest=sha256:d8c37cc2788d0adce0473469b86cb1542dc4a1674bb9279cce8ef1539211c05a

Observation 93edebdd-6373-46d8-a16b-c672ba690870 · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.191180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.191180Z digest=sha256:c5effda908fa466eaf0ab5d56dc49a6d3224b1bd1ca539800f2a4b4040d930bf

Observation ad9c477a-93a7-45b9-a38d-34e26d5816b1 · outbound

This paper cites Rethinking Language Model Scaling under Transferable Hypersphere Optimization.

Scale Weight Decay and Train Better Rethinking Language Model Scaling under Transferable Hypersphere Optimization

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.195671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.195671Z digest=sha256:ee1b20a267f3e01dbfa428e77c84f77ccf8ff0cb516afa72e827e5e00750902e

Observation ca591d46-0694-44a9-8a69-f1879540b5c5 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Scale Weight Decay and Train Better The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.130980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.130980Z digest=sha256:5b810c8c0b326a029f369d695c195e0f20768e2b5b46777b4d42fd486c930165

Observation 64874ec6-a8ef-4481-9818-fec04fc460fd · outbound

This paper cites an unresolved cited work.

Scale Weight Decay and Train Better Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.982908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.982908Z digest=sha256:7d660066d5a5fe83659f5cf4f9f40cc39f867f24dcde9e56ea98253e934715e0

Pith citing papers

No inbound Pith citation observations are available.