Pith. sign in

Paper Citation Record · LEDGER

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate

As of 23 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2605.21486.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.21486 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T05:01:27.451161Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact13
  • verified fuzzy34
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 586a96da-6d6d-4d44-9db4-2a79ab78f9e5 · outbound

This paper cites arXiv preprint arXiv:2601.10684 , year =.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate arXiv preprint arXiv:2601.10684 , year =

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:03:57.948549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:3874e20bd3c2696d3b0d18e685edfad989666afc018be51fac20442cc494cc0d

Observation a1d05a41-616c-404b-b64e-2b876bfdddbe · outbound

This paper cites Power lines: Scaling laws for weight decay and batch size in LLM pre-training.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Power lines: Scaling laws for weight decay and batch size in LLM pre-training

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.660494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:0662f47b14c922ee102917d93cc984291094437fadae5c313dd3ed7b4ebb08bb

Observation 7a263e0a-64e2-41d2-be0f-f3adc7563951 · outbound

This paper cites Scaling optimal LR across token horizons.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Scaling optimal LR across token horizons

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.643714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:0c4ab8758886b1ba34365784a3ae65a7fc3c0576614041d31dfd539f9b539ef6

Observation ed4ce618-d317-494a-b97d-f3298cf78f88 · outbound

This paper cites Self-consistent dynamical field theory of kernel evolution in wide neural networks.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Self-consistent dynamical field theory of kernel evolution in wide neural networks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.640010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:70dc59c22d621818579002de8d6b91f42daa2dfd89f101277d13583450865590

Observation 1c25a29d-c012-404b-b05f-571ab0cc1c65 · outbound

This paper cites Infinite limits of multi-head transformer dynamics.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Infinite limits of multi-head transformer dynamics

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.662895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:d52eb8edcef5480a8daad0c5a7b00cc4ab1078d630fb8fa8f9469056a07f4c42

Observation f037e25d-30d7-42de-ae5b-2b44120b27e0 · outbound

This paper cites Depthwise hyperparam- eter transfer in residual networks: Dynamics and scaling limit.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Depthwise hyperparam- eter transfer in residual networks: Dynamics and scaling limit

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.670796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:4e222191324c8599e4dc208ff829627824d6e7b40e43226b4e5461ebc45ea00a

Observation 302c5891-7b69-4413-9158-6fdef40f214b · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:03:57.951136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:bd753f662a10b71363fb363612f28ac0ead550d80043bb4d0ba637ebcb7b2c08

Observation 2d7c648a-5696-49b8-b54b-b269cf976d9f · outbound

This paper cites Don’t be lazy: Completep enables compute-efficient deep transformers.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Don’t be lazy: Completep enables compute-efficient deep transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.668839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:422586eaa6e8e8aab37e65831556af089b65ea526b19f8fe2e61515cea8cfc44

Observation 1d32c1bb-2139-42dc-8f2d-fa80a1568b2d · outbound

This paper cites Don’t be lazy: Completep enables compute-efficient deep transformers.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Don’t be lazy: Completep enables compute-efficient deep transformers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.968125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:d4928c93f5d68130ebf5d85e4c2ebf96b51f0dc03a94ced63992f6e42c71bc0c

Observation e5b9bc47-bec3-4567-86ba-5d3fa806b059 · outbound

This paper cites Sparse maximal update parameterization: A holistic approach to sparse training dynamics.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Sparse maximal update parameterization: A holistic approach to sparse training dynamics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.674209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:b3a460ea95f970638bb42a2758edf6f836e75f263cb75c32734e753e6396be42

Observation 26446428-f3be-4386-8359-d849220f101f · outbound

This paper cites Scaling exponents across parameterizations and optimizers.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Scaling exponents across parameterizations and optimizers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.676694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:b4d49b3e9a766031944b484cb99e395346b03ebb3024b47165a5c409fcae28c5

Observation f0f791cf-69e3-41cb-9a38-d657860c477a · outbound

This paper cites Understanding the mechanisms of fast hyperpa- rameter transfer.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Understanding the mechanisms of fast hyperpa- rameter transfer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.959670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:21edb323ab391f0ab6973eff056ccfcacf61460e4589a9428736eb22dcd38eae

Observation b19202a0-b40d-4f11-bef0-9735ce7c424d · outbound

This paper cites A loss curvature perspective on training instabilities of deep learning models.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate A loss curvature perspective on training instabilities of deep learning models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.656727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:0950272ccfd77656da76454f6a75f71b2ceb924be7e69e69d805e016af1440b1

Observation 4f7d4861-cf7c-402e-a6ab-c38b0e9d9ef5 · outbound

This paper cites $\boldsymbol{\mu}\mathbf{P^2}$: Effective sharpness aware minimization requires layerwise perturbation scaling.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate $\boldsymbol{\mu}\mathbf{P^2}$: Effective sharpness aware minimization requires layerwise perturbation scaling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.652917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:0576d5a9f8425a5a20589c187ee0c98ae676adc8cbcda210c98e66e46604f4fc

Observation ceeecce6-f2dc-4c8c-8262-1b0c501f6b54 · outbound

This paper cites A proof of learning rate transfer under $\mu$p.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate A proof of learning rate transfer under $\mu$p

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.654977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:54fe0e442fc4b0d4c0752b6bc2dd55152532f9ed3c091e677423ae176a55f624

Observation 73b515ef-7989-41b6-931b-51d9f8343192 · outbound

This paper cites Optimal embedding learning rate in llms: The effect of vocabulary size.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Optimal embedding learning rate in llms: The effect of vocabulary size

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.651223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:41f094f3c4cb571926d398cb8d47a5d173520eab27f80d6d883596fb464994e8

Observation b12fee4f-c0fe-4629-9cfd-5c377a29cff1 · outbound

This paper cites Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.956920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:7ecb049e25fcd791444d3e3eec17963de0e2a9eae27bcc6d6ee1cb6cd95a1e94

Observation 2a223616-955e-4695-beef-6ed93ad0b5d8 · outbound

This paper cites Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.647551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:0ca37efee1c300288e6777521ae117c49c73c22ad1b7bb266fd3df8f3b2f2b45

Observation 341c8c77-a45e-43c5-b927-0d000c51a526 · outbound

This paper cites An empirical analysis of compute-optimal large language model training.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate An empirical analysis of compute-optimal large language model training

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.649109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:3b4962082d2bb75d1d77cd76c902e9628bbdeec3a54b87818cf7b6936777be0c

Observation a25c7cef-17ab-492a-8c89-fa20506e0592 · outbound

This paper cites MiniCPM: Unveiling the potential of small language models with scalable training strategies.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate MiniCPM: Unveiling the potential of small language models with scalable training strategies

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.658846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:dea5eb3708c529fd549879f2c41765ea58e4ada139e99cec660b850928b3dd44

Observation 105ca246-4592-4698-8265-beb37beba02a · outbound

This paper cites Hyperparameter Transfer with Mixture-of-Expert Layers.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Hyperparameter Transfer with Mixture-of-Expert Layers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:04:27.931887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:2fe374a5fb2e199abe88891eb553eb20b5d72bb953f5b653ba1d18888a41c4ef

Observation 5df69093-255e-45e3-90fa-e2d1e9061e4e · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Muon: An optimizer for hidden layers in neural networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.664825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:fe1e5e8b2c34b2c14e8e3ef1e3c27717a92cd2e24a1f5790b49bc239f4f3e633

Observation 08d9609b-8193-47fc-b997-1d026b4fcc94 · outbound

This paper cites Why warmup the learning rate? underlying mechanisms and improvements.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Why warmup the learning rate? underlying mechanisms and improvements

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.638192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:017bcfb2d528dc5108adcb9d216861774bca306d8925fe4ee454f0942e82f2c8

Observation c3c2da91-5055-4caf-adbc-6747afaba377 · outbound

This paper cites Universal sharpness dynamics in neural network training: Fixed point analysis, edge of stability, and route to chaos.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Universal sharpness dynamics in neural network training: Fixed point analysis, edge of stability, and route to chaos

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.641688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:ea188928b3f1207d48f0a585a341d952bc5f4575514cd7c32b8acf6034a38c95

Observation 1f7aefa1-9268-4b7b-8ec0-b949486be19b · outbound

This paper cites Scaling Laws for Neural Language Models.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Scaling Laws for Neural Language Models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:03:57.953743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:d365d567218690b02f57fb6362c8e43d7d1abb905805c97afeeb47e200388f75

Observation f9d77f6e-8539-4e61-83d1-6ba760fdc530 · outbound

This paper cites Symmetry in language statistics shapes the geometry of model representations.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Symmetry in language statistics shapes the geometry of model representations

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T01:17:13.480674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:2b985bf117cb4bbe0587df1c2f222304764d03ef390a76daa20ade03301179ad

Observation 36b6ff84-5963-4b3d-b089-7dfdddbee848 · outbound

This paper cites nanoGPT: The simplest, fastest repository for training/finetuning medium-sized gpts.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate nanoGPT: The simplest, fastest repository for training/finetuning medium-sized gpts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.666566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:3eb9b73ad0cd6cc6d55d3d386a4852b057382398af6c92ebb8c7e52a68318807

Observation 4dd91632-60e5-431d-a328-4785cbf89c4f · outbound

This paper cites Weight decay may matter more than µp for learning rate transfer in practice.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Weight decay may matter more than µp for learning rate transfer in practice

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.636273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:a548aeea6429ed87c7da5f5cd02d5239222e099a5eed8f37c4cdbd58615d3db6

Observation 68afa20d-3f98-4b2e-9507-7f8edb63c538 · outbound

This paper cites Cifar-100 (canadian institute for advanced research).

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Cifar-100 (canadian institute for advanced research)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.631949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:7ecadda1ade4dd5020cc0a922d4e6fd31cf1a5f6e6210b39279f1f7c375db367

Observation 931597dd-a8b5-4142-8954-839468e0a7bc · outbound

This paper cites Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.973859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:94f7d3340ce5d059a758b0a2645f13f1d02a98739f3200ada30fb8d7f310151e

Observation aa9a5e87-a9a4-4286-a65d-e1b5177d18c6 · outbound

This paper cites Adaptive optimization in the $\infty$-width limit.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Adaptive optimization in the $\infty$-width limit

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.684429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:7a74343b30e4f1160ab575b3d6dce2d58e929a6ad07ea9b94222d6f21a38355c

Observation 40254e0d-da91-4b13-b354-1d44a5b8f043 · outbound

This paper cites The Llama 3 Herd of Models.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate The Llama 3 Herd of Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:03:57.976267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:e9c5f4b3de9903e842f0ae92003371d51c4cc8f85b603fb251a8d3f5609d0b52

Observation e86398ba-6468-4586-9b1a-3c1329e20761 · outbound

This paper cites Decoupled weight decay regularization.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Decoupled weight decay regularization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.630063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:7f079db1f8a0b6019583468ec77676187ee58fd07c1389f74d7d5ca70e858f51

Observation 4e04130c-406b-4290-ada1-1fdc75036a62 · outbound

This paper cites µ-parametrization for mixture of experts.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate µ-parametrization for mixture of experts

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.965196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:e99b87952897d362fe9398094590bca968b6277fd4b659c05de9735ff8bd0003

Observation fd7c740f-8eb3-4287-9736-607561d38a74 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Progress measures for grokking via mechanistic interpretability

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.625687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:664f81b46647076f45b189aa82212ad4da065fd760d00f02ea10e88be64a2ef9

Observation 81cf0dfb-a06a-45a5-a40c-d8187920c6b4 · outbound

This paper cites Super consistency of neural network landscapes and learning rate transfer.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Super consistency of neural network landscapes and learning rate transfer

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.627557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:0dd768f696716ab6e3c952d4cc6e9a2ce3320080d562a078d49a6ffea0dcc889

Observation f3969d39-71f5-4c4a-946f-e67e31b4288c · outbound

This paper cites 2 OLMo 2 Furious.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate 2 OLMo 2 Furious

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:03:57.978819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:d92dfc7e790cac81e792e6b9d24574e226ff58aba5970926cfb56266bf31d222

Observation b71bc578-9523-445a-9ada-6d9114cfa116 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate The fineweb datasets: Decanting the web for the finest text data at scale

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.634492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:721edf19c9263eb11a9fc2a27552f6d07c42c87521c29169e0de7b6591ff626a

Observation 2591375f-72b6-434c-b7b8-ded6f5501d98 · outbound

This paper cites Resolving discrepan- cies in compute-optimal scaling of language models.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Resolving discrepan- cies in compute-optimal scaling of language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.645280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:7dbb1abbf1d4f3445e7943ec0d6a9a178db1c8dc53525f1b578d1ed944053fbe

Observation f090bda2-61a0-4e8b-a103-c053707c5cbf · outbound

This paper cites Hyperparameter transfer enables consistent gains of matrix-preconditioned optimizers across scales.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Hyperparameter transfer enables consistent gains of matrix-preconditioned optimizers across scales

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.942472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:96cc3a3b0f4e82241f8755788519a4c0a72e3cb29f8b94280e8b4692a8b5e7f1

Observation 8d1391f0-43b8-4764-9e95-337cef06a1a9 · outbound

This paper cites Qwen2.5 Technical Report.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Qwen2.5 Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:03:57.981696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:4a0ee2fe9be7c12ff6cc6c6eecfc4396523f05c2bfb1c4808082851830c7d51d

Observation 12cab608-8552-497d-8cdc-825ae6c0ae81 · outbound

This paper cites Language models are unsupervised multitask learners.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Language models are unsupervised multitask learners

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.691495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:ab8a5f2ad848c17c7bc3e2570c73d36a2fe974bedf070f06e2be97ebbc2d79a1

Observation 70b76eaa-2f5f-4d4a-ae97-b91841feaf8b · outbound

This paper cites Roberts, Sho Yaida, and Boris Hanin.Frontmatter, page i–iv.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Roberts, Sho Yaida, and Boris Hanin.Frontmatter, page i–iv

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.693583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:90878b8883a90d1343adeb2194536b06b09c1323bcb67ed9fd30722f0222e29f

Observation 437a4816-6dc1-453d-894b-da3670683c40 · outbound

This paper cites On the infinite width limit of neural networks with a standard parameterization.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate On the infinite width limit of neural networks with a standard parameterization

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.984792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:f61204054ed6376c1208cc71452e6e3971778097342f5489d4cdb630a0b785ba

Observation c5221330-e3ca-49df-afc5-15af3e80a49a · outbound

This paper cites (how) can transformers predict pseudo-random numbers? InForty-second International Conference on Machine Learning.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate (how) can transformers predict pseudo-random numbers? InForty-second International Conference on Machine Learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.685637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:164d15587cc94d993befe490e15f0ffc141b491d5fa362eaaf8fc4f3ef70eb98

Observation ae3590ca-3896-46be-9f10-74e6a61f9fd2 · outbound

This paper cites On feature learning in structured state space models.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate On feature learning in structured state space models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.688013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:5ce8839222d0a761b0b71f035d6ca955f2de5697e0a8fef8e3bed1d5d27248a8

Observation 01963f8e-bf7f-4d72-98e7-4afd9d0a0ac6 · outbound

This paper cites an unresolved cited work.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-21T05:03:58.689757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:3aee123b929b83d40a5369166f3611c7fe2567d78ae2d01abab5c59b6eb52540

Observation 9080fe0e-3163-4552-ad8d-c1cf76545eda · outbound

This paper cites Attention is all you need.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Attention is all you need

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.695277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:8905ecdc56a78b5c1ac79fc9bc5afb7a87f872eb6b0ff3c56fc2ccd3097f0148

Observation bcc20965-99b7-4629-a9b1-c38ce91e6e7b · outbound

This paper cites Meta-Principled Family of Hyperparameter Scaling Strategies.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Meta-Principled Family of Hyperparameter Scaling Strategies

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.970923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:6794ce818a724985517c2623c0b8eb2e502b6816c71d79cf1e97f2bfd6d50813

Observation a8438573-213e-4618-8cbf-e6e9936bf45a · outbound

This paper cites an unresolved cited work.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-21T05:03:58.681913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:6c5a80c9eddce2dbed36339291c6795e52ba0dd3033cf200aebec3ec503850cb

Observation 9ab1a6ab-dfb7-43b5-801b-e8df0e228b85 · outbound

This paper cites Tuning large neural networks via zero-shot hyperparameter transfer.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Tuning large neural networks via zero-shot hyperparameter transfer

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.679500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:7e89ea929cbeb4510b1ba7103a5245c48a4639126b119d32e14983a85592d63b

Observation 61b9a4db-fcf8-4e77-b4bc-1e8260a5bd0d · outbound

This paper cites Tensor programs VI: Feature learning in infinite depth neural networks.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Tensor programs VI: Feature learning in infinite depth neural networks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:03:58.623087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:29a8e820f407baf741b5aed2dad4077ad448adc9504f382a9ac513ee0581eb66

Pith citing papers

No inbound Pith citation observations are available.