Pith. sign in

Paper Citation Record · LEDGER

Optimization Hyper-parameter Laws for Large Language Models

As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2409.04777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.04777 v4

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T20:45:31.427677Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:38:09.654884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T22:38:10.015046Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact39
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71b857a4-ad49-4b43-bbad-b05c7db2613b · outbound

This paper cites GPT-4 Technical Report.

Optimization Hyper-parameter Laws for Large Language Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.845085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:1d4a6d3d7cc7699f518f57e49ae5fd165dfe385fc8a80b6686823fc07521c0a1

Observation 899e10d2-4277-4c5c-8efb-a159e65415de · outbound

This paper cites Qwen Technical Report.

Optimization Hyper-parameter Laws for Large Language Models Qwen Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.840667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:e5cd086b90762b78d2c2886348e17bd7715139aec0cb5d6609d56afb4b41ecc0

Observation 9d989cec-7771-41d1-a130-01b1618bc8a9 · outbound

This paper cites Chinchilla Scaling: A replication attempt.

Optimization Hyper-parameter Laws for Large Language Models Chinchilla Scaling: A replication attempt

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.784575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:d7eab6d9e479ab7ea6e71b7072e0c9abee8591524b325f1f43c2622b69c2308d

Observation f25b8a45-79ba-4bf4-89b6-dcda6f86d320 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Optimization Hyper-parameter Laws for Large Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.804562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:d36d096d321d55975f123bf37c8c808b87755f4b66ffa9b35e31239057d5f54d

Observation 29fc9b2b-188c-4ba0-b630-2c8f205a7ad6 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Optimization Hyper-parameter Laws for Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.854973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:d678aeb9ebc10bf34b6d5305249d8a0a92407822d17eed7b1475094c1a735d8d

Observation acffcd64-5776-47f1-b1ba-7b4b07a12ea9 · outbound

This paper cites Stochastic Bregman Subgradient Methods for Nonsmooth Nonconvex Optimization Problems.

Optimization Hyper-parameter Laws for Large Language Models Stochastic Bregman Subgradient Methods for Nonsmooth Nonconvex Optimization Problems

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.814838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:091e4632813a9e974d7920b03f8872382864b64b453040ebd02e06bc13573533

Observation 99dd7265-15b3-4db5-bc31-bd10b8b39e89 · outbound

This paper cites Adam-family Methods with Decoupled Weight Decay in Deep Learning.

Optimization Hyper-parameter Laws for Large Language Models Adam-family Methods with Decoupled Weight Decay in Deep Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.883210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:b235e5992276f833dfb1889133bb0b8b437fd1b2052c5443543861c678cd1145

Observation 30383b20-e117-4d2a-a365-e88e3066d929 · outbound

This paper cites The Llama 3 Herd of Models.

Optimization Hyper-parameter Laws for Large Language Models The Llama 3 Herd of Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.867382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:73947696372d6b0c2c7a411092a6b065965af8f0e540183864988b81bd9ceca0

Observation 0d247a0a-72ab-4130-8a6f-053fc7e1a02d · outbound

This paper cites Sharpness-Aware Minimization for Efficiently Improving Generalization.

Optimization Hyper-parameter Laws for Large Language Models Sharpness-Aware Minimization for Efficiently Improving Generalization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.799977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:5c4d175bdf4b25c71984aab173f3d08ae797e534dbb24d9272833a1fa4be5e00

Observation f97c4029-899f-445f-bcd8-352a29f3a66b · outbound

This paper cites Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting.

Optimization Hyper-parameter Laws for Large Language Models Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.872547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:5d4d3b43fc505ceb4973e6be5b341f8d10b3629ab76cc9daaf03219a578fc880

Observation 536f8344-1450-4939-bc2f-fd99b41393a1 · outbound

This paper cites Accelerated Objective Gap and Gradient Norm Convergence for Gradient Descent via Long Steps.

Optimization Hyper-parameter Laws for Large Language Models Accelerated Objective Gap and Gradient Norm Convergence for Gradient Descent via Long Steps

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.820443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:e7d089fc2695b3df8e0b611c8512556e8348a135d46f5b1a054845d27e7eb2ac

Observation dec96b09-c71c-48d9-a0f1-9a0f712cd44f · outbound

This paper cites Efficient Continual Pre-training by Mitigating the Stability Gap.

Optimization Hyper-parameter Laws for Large Language Models Efficient Continual Pre-training by Mitigating the Stability Gap

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.778583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:d489e6ef95b69a1a09ad3ad92941a8216e9341981b7ae090454bcfd0c924c860

Observation 3feba15a-7bca-4398-a3a3-68a93ee0a0ec · outbound

This paper cites A Novel Convergence Analysis for Algorithms of the Adam Family.

Optimization Hyper-parameter Laws for Large Language Models A Novel Convergence Analysis for Algorithms of the Adam Family

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.831372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:c6bbff88a79a2715747fcb7f343b07172103a8d602d54c063b73519d3afe03ce

Observation 5c6e8e3e-a02d-4d8b-959e-cbcb79d340f2 · outbound

This paper cites Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations.

Optimization Hyper-parameter Laws for Large Language Models Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.861142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:f072d464622f3d421269b8e4d274fea63615d531832c645629fdc2dec7d92748

Observation 58294b13-000a-4652-8853-27cc490a9b5d · outbound

This paper cites Scaling Laws for Transfer.

Optimization Hyper-parameter Laws for Large Language Models Scaling Laws for Transfer

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.835986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:70df638d0c456214cbf735a9ee5f13aedaa6d80cb7c87d67df1d7ca74aa0eb86

Observation 6672f655-71d9-4a21-9daa-bfff810d6d1f · outbound

This paper cites Training Compute-Optimal Large Language Models.

Optimization Hyper-parameter Laws for Large Language Models Training Compute-Optimal Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.808999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:803277034ba38ca2c488662c0dee91e6b2b1309bf808181527e7a82fa0036ca0

Observation 778dbfa0-7a15-4bc5-93da-7da1272a3d8d · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Optimization Hyper-parameter Laws for Large Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.795489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:493f93888a16a866ccc9c0b97d4ee66cb7df1523f5f6cbd0e50abcdd82236372

Observation 9d1a3c41-6262-40ba-8d3a-7e2dcbd5aa77 · outbound

This paper cites Simple and Scalable Strategies to Continually Pre-train Large Language Models.

Optimization Hyper-parameter Laws for Large Language Models Simple and Scalable Strategies to Continually Pre-train Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.824942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:93931af8c73ea59219fdf8c48fbe9065c7b41e7291ce695b58c32703ec4987c9

Observation 2920b524-ea34-4253-a30c-c14534a6a0e9 · outbound

This paper cites Scaling laws for downstream task performance of large language models.

Optimization Hyper-parameter Laws for Large Language Models Scaling laws for downstream task performance of large language models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.878109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:3ffb04c1d9c3128c644d5e8960138389ecca02b8e199255b18d7cd1dbbc97688

Observation 2a59b6c9-4c31-43e5-81db-eb207d1d272d · outbound

This paper cites Three Factors Influencing Minima in SGD.

Optimization Hyper-parameter Laws for Large Language Models Three Factors Influencing Minima in SGD

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.790050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:e0847bd96cbb3b5482a4578460e5991a437916f6c9b3cafc41e9524579af0f42

Observation 2821de96-b257-4f92-8b91-6275d9357eb6 · outbound

This paper cites Rethinking Learning Rate Tuning in the Era of Large Language Models.

Optimization Hyper-parameter Laws for Large Language Models Rethinking Learning Rate Tuning in the Era of Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.850073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:ff962b1877b2af64477fecb3b3c7da6faf44803e2abc239e4fddb4fac96d823c

Observation cee8d0b5-bf8e-47eb-af8b-71f0d9956ee6 · outbound

This paper cites Scaling Laws for Neural Language Models.

Optimization Hyper-parameter Laws for Large Language Models Scaling Laws for Neural Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.927406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:afe7a716bb24a15b788563bebbe5cbf4d417585bac751254e49ad20da91a22cc

Observation 921063f1-6e6e-4ce5-8f3d-19c5d5d8f569 · outbound

This paper cites Continual Pre-training of Language Models.

Optimization Hyper-parameter Laws for Large Language Models Continual Pre-training of Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.948200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:090226b299e6ad15c9df61611da79a6aebd193df058ae858b1af14a6bc01341e

Observation 66668037-1846-4b0c-b002-81c88b430151 · outbound

This paper cites On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.

Optimization Hyper-parameter Laws for Large Language Models On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.922422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:107f0c5cc678ab19c26981888df12417b0feeff21278a82400cd88832d5195d8

Observation 7cb45d94-fae8-4d42-9c33-cacdb13bfa97 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Optimization Hyper-parameter Laws for Large Language Models Adam: A Method for Stochastic Optimization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.932968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:40a645f83307405c7c06d7dc17749c987175a321b276e4e1a8db3304cb3e6395

Observation 68927d72-0c6c-41a1-bb77-4ef54908852b · outbound

This paper cites Decoupled Weight Decay Regularization.

Optimization Hyper-parameter Laws for Large Language Models Decoupled Weight Decay Regularization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.937948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:57edab847a6bee612a7675ac227383ab20b33410de5999e000bcc2aa40d312e1

Observation 671738de-d351-4a27-b967-585eccb6ef81 · outbound

This paper cites Full Parameter Fine-tuning for Large Language Models with Limited Resources.

Optimization Hyper-parameter Laws for Large Language Models Full Parameter Fine-tuning for Large Language Models with Limited Resources

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.888996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:6483426d2ff1986913d554520c9ed1abeaa68f9a695681e008a944a37659d3ea

Observation 56eabdd4-9e51-49a5-9325-d982698991b5 · outbound

This paper cites An SDE Perspective on Stochastic Inertial Gradient Dynamics with Time-Dependent Viscosity and Geometric Damping.

Optimization Hyper-parameter Laws for Large Language Models An SDE Perspective on Stochastic Inertial Gradient Dynamics with Time-Dependent Viscosity and Geometric Damping

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.906059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:da6f024b5e85329baef398a56bdb5d4d63b549200c34eac8f2c36591f8ee5ba4

Observation 0ba8aaa7-da0c-4ab4-81cd-3341c05d465c · outbound

This paper cites Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models.

Optimization Hyper-parameter Laws for Large Language Models Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.911815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:b1b7a7d72bfa5a6ead9252cee186ec71a67be4bb759d601b97dc216bf88f8c27

Observation f3b64c03-6053-4dbe-a6dc-f44a35ed181a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Optimization Hyper-parameter Laws for Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.971633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:12c41a0334b63b65046bd4580fb0fce0e3b54dec274baa309deea73cb41a0f0c

Observation 3b89fc53-5457-4ee3-8787-2eabdaa4ee53 · outbound

This paper cites Rotaru, F.

Optimization Hyper-parameter Laws for Large Language Models Rotaru, F

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.894809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:5c728c35e9d89ce9e4ec3f46be82614de07989188ff9f7cf85281631195370db

Observation 2e2cb812-1518-49a7-bfff-a1f2760f19ec · outbound

This paper cites RepoFusion: Training Code Models to Understand Your Repository.

Optimization Hyper-parameter Laws for Large Language Models RepoFusion: Training Code Models to Understand Your Repository

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.953393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:ad7b7b18d9e8223bb943a89f5f7f3fe70ea9ab1011832d369676a3bb1190f779

Observation b262fb87-6456-4f4c-a566-632f6612805f · outbound

This paper cites An SDE perspective on stochastic convex optimization.

Optimization Hyper-parameter Laws for Large Language Models An SDE perspective on stochastic convex optimization

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.943058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:db724094da55d41515a0276466cf269c253982d99d52e2c0112a4a52706ec6ff

Observation 9f78d267-bcdb-418d-99d7-86ab56a043d3 · outbound

This paper cites An elementary proof of anti-concentration for degree two non-negative Gaussian polynomials.

Optimization Hyper-parameter Laws for Large Language Models An elementary proof of anti-concentration for degree two non-negative Gaussian polynomials

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.961159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:d4bdcceb87be08a278800ccbc3ef52d85b0188241d4e1d7d97ddfd76af631369

Observation 733bf6a5-6cd0-414c-a0a7-c47f13c2bb89 · outbound

This paper cites Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models.

Optimization Hyper-parameter Laws for Large Language Models Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.917769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:9c8485a462191eb667b1f7a63bcdb32a91b680c84e620fbdbdff684a98bba555

Observation 0f862a80-8a16-4d01-82d7-a0d843490679 · outbound

This paper cites Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization.

Optimization Hyper-parameter Laws for Large Language Models Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-08-11T01:19:49.772953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:70306cea75b8710695e4fbcf5673f04a7031f0c7849be5a9cf016a6de22cc73f

Observation 770af79a-1280-4f19-9b18-16f11a657b15 · outbound

This paper cites LoCo: Low-Bit Communication Adaptor for Large-scale Model Training.

Optimization Hyper-parameter Laws for Large Language Models LoCo: Low-Bit Communication Adaptor for Large-scale Model Training

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.976479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:fc5a39472685b6a63eaffe0344f710cf95094fa735b2f92918bf6d3c76b0b4f6

Observation 7c9da6ef-cfaa-4ab7-baab-3039495d370a · outbound

This paper cites Qwen2 Technical Report.

Optimization Hyper-parameter Laws for Large Language Models Qwen2 Technical Report

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.981300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:55d8632dcd100156887cd1fe75cb031f3d2681e417882f1651b1aa32b4217d67

Observation 6ea7a178-6fbc-4f44-8eed-c145524dd004 · outbound

This paper cites LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models.

Optimization Hyper-parameter Laws for Large Language Models LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:45:48.900983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:4af7363e1e00989d53afc0249ebfa1b3b655c052f34183224586dc26f6d9af8d

Pith citing papers

Observation 238e251c-24b1-4d59-82da-23d5d43a8f14 · inbound

Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning cites this paper.

Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning Optimization Hyper-parameter Laws for Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-04T22:38:10.019370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:38:09.654884Z digest=sha256:106320c59dc74e7b143c74e9dded97d289313b95bd2dfb80cbbec1825c8a9426