Pith. sign in

Paper Citation Record · LEDGER

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

As of 15 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2506.15025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15025 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:54:42.614806Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T17:31:32.533941Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:34:57.593618Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3f8b846-6e14-4517-be05-f687fba975d8 · outbound

This paper cites u-$\mu$P: The Unit-Scaled Maximal Update Parametrization.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size u-$\mu$P: The Unit-Scaled Maximal Update Parametrization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.444889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.444889Z digest=sha256:093c59140215a6bc92c6a3bea85777d49c1ea0fc5f281f95a4b094117c3ddc8f

Observation b4fc21cb-ea38-4af4-afab-63a2f63c7b5e · outbound

This paper cites Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.450340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.450340Z digest=sha256:a1425bb12faf00e7976768e6b95af040aa59bae7a069a252041b759cef6537de

Observation 4e9517dd-5fdd-4a8e-91e5-5d50efc79d17 · outbound

This paper cites On Lazy Training in Differentiable Programming.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On Lazy Training in Differentiable Programming

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.454722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.454722Z digest=sha256:022a0453d65255b2af16107a6186b9ad6cf10215efeca38fe3b6cbcfdb30ac2e

Observation 2e495ec2-61aa-4a99-b5c0-c7a317109251 · outbound

This paper cites Infinite- width limit of deep linear neural networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Infinite- width limit of deep linear neural networks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.459894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.459894Z digest=sha256:8fa4fc067279d6e09ac10d6f1191383b4876192f3a3f12989ee6088fc77a9718

Observation 52002095-c746-4854-844c-d90658e74767 · outbound

This paper cites Alemi, Roman Novak, Peter J.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Alemi, Roman Novak, Peter J

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.129312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.464142Z digest=sha256:e920f2f9383f1a99f9c32a1324fe4cf20a3c4431e5922bc0b7323e5ee85e25e6

Observation 82f197d9-4eb3-444a-83c2-647fe2f7f6c1 · outbound

This paper cites On the infinite-depth limit of finite-width neural networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On the infinite-depth limit of finite-width neural networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.115696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.473690Z digest=sha256:0c04c99dc6b93fc0447a348c4e0d8029954ff3c8a2d0ca2e8fc36233759a0d65

Observation 87e1e4de-8cd6-4997-a736-fe3820039dec · outbound

This paper cites On the impact of the ac- tivation function on deep neural networks training.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On the impact of the ac- tivation function on deep neural networks training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.102018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.477648Z digest=sha256:e9eb661d6df34d0e99d5090edf16e71b21be3619d4116a2c88151e10eeb7138b

Observation 2f285087-ba6c-4b17-946f-12894b8c479c · outbound

This paper cites Stable resnet.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Stable resnet

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.087750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.482191Z digest=sha256:19e679a957cddbb9c7fc941f0f12c70c112eb3576f743dc48218dada2cc076cd

Observation b5cf8226-42e4-4f65-8ba2-6adf96c288c4 · outbound

This paper cites Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.485996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.485996Z digest=sha256:a335b0469e5800a6aa7dc9b8ced7e8e590bef4d219ee49f70b30207d04fcfbf9

Observation fd4661cb-4347-492a-aca5-915c0d87ec2a · outbound

This paper cites Neural Tangent Kernel: Convergence and Generalization in Neural Networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Neural Tangent Kernel: Convergence and Generalization in Neural Networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.490672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.490672Z digest=sha256:ce911247ef4fdfd062bb08aef989e1affc5e98f634ef34d12f305b094d25988c

Observation 590b4620-eaa3-47ac-be14-15a5c3c5747e · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks,.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Muon: An optimizer for hidden layers in neural networks,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.070446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.495390Z digest=sha256:744bde1b80605cb80943249bfd0e1d2536967c7634d2ff7bacde580d72614cd6

Observation 0f979c3d-4478-4ecb-953e-6ca8318b1ec3 · outbound

This paper cites Kingma and Jimmy Ba.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Kingma and Jimmy Ba

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.503330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.503330Z digest=sha256:8c1139ea6a2fcbf92e2e5f19bcd0ee1d1fb11c6077a567c94a5546b8194447a0

Observation d8abe900-b9de-495f-bf1e-7662454d1e7f · outbound

This paper cites an unresolved cited work.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:54:43.056873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.499441Z digest=sha256:cd0ae6272bbb6f0fc2b266fbe989c7171982d89d80b1cfbe2b8f7f0990330852

Observation 3075e622-a266-4e46-b146-b163ba8bfff7 · outbound

This paper cites The llama 3 herd of models, 2024.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size The llama 3 herd of models, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.516190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.516190Z digest=sha256:fcf588ab4c7e781da7354662d3692ff3928e22e5e782c6e67a3e5fa006d8255f

Observation 230ffc16-b8eb-421f-8d85-49e6a96e7ec7 · outbound

This paper cites Pointer sen- tinel mixture models, 2016.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Pointer sen- tinel mixture models, 2016

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.028154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.521875Z digest=sha256:c6d794066de633f1114c43456cdc10e415623e80007400976d2e506cbf177186

Observation 4976277b-3b6d-46c9-ada3-c88d92d87a5f · outbound

This paper cites An Empirical Study of $\mu$P Learning Rate Transfer.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size An Empirical Study of $\mu$P Learning Rate Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.511780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.511780Z digest=sha256:fb1121e828391ecabebc0d58cebf2d737b41ba1006c581e65bf5322618e6e598

Observation a6df0e87-6b3a-459e-9930-6d91c199d3ae · outbound

This paper cites Poole, S.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Poole, S

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.015329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.530857Z digest=sha256:8d43dee3ba430c52a533de5f864c30b4adb08f6a4897f9fe07bde17bfccadefd

Observation 07b7ee97-7df3-4b56-9441-1ac76142e1c4 · outbound

This paper cites Schoenholz, J.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Schoenholz, J

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.001690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.535008Z digest=sha256:049a848fd3fa8998e58d60aa75b2cbb40c6955a1eb7da85374913b7725cbf91d

Observation 19eda179-91e6-419b-b004-cbde3f0c47a5 · outbound

This paper cites an unresolved cited work.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.526095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.526095Z digest=sha256:753f7ef6e0fb0a7a9234bdb1767d519844843de6c99788893acf5f145fc543c7

Observation 60dcbd05-82f5-42d8-827a-7668986fbb75 · outbound

This paper cites The Falcon Series of Open Language Models.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size The Falcon Series of Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.544080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.544080Z digest=sha256:761717476ba4a09a56072ac765d75f73ae192f559b5bac468dfa1f37a389e893

Observation 2d3fc3b8-693c-4e3e-9bc2-7b06c8d81e88 · outbound

This paper cites Gemma 3 technical report, 2025.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Gemma 3 technical report, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.548770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.548770Z digest=sha256:62286f2e8467bd2ce532377d3801c4aebbdf87785864edd256b4236023ef6070

Observation f736ab14-e453-4cf0-811d-2a28c79e8fe8 · outbound

This paper cites Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.539112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.539112Z digest=sha256:1f077d648e26dd79c88ad446a0bd17ca7478d4079a39976e7fa37eb7cac2a1fd

Observation ea63cfd8-d429-48e1-af50-f8099316e3ec · outbound

This paper cites Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.557483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.557483Z digest=sha256:5671fdb82fd62145a32372ca6f986f0d841e20c5f7c89d7ff6400476f874dfd3

Observation 9469a6ce-f8dc-40e0-b5ca-7656daaba731 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.562346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.562346Z digest=sha256:9fe57535b9c0fcab10328ba84c4265132c21ee2a2743266b4da45a1e833a2e37

Observation ec43db71-939e-494a-b99c-f35d63d9c7a3 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.552969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.552969Z digest=sha256:f3f56fcb5918c7f40f47c771f3b10e3f5f1b4f0cc33e99edd7cfd09b4a814725

Observation 1d57d817-22eb-4cb8-a62c-a0f2eb0de283 · outbound

This paper cites How does critical batch size scale in pre-training?,.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size How does critical batch size scale in pre-training?,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.981917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.571451Z digest=sha256:2525e7cbac27a2773a2f49b4e1cd48e8966115f62f8a21ece7ed4b22664107b7

Observation 23af0345-ba61-4677-a09e-a194ec2b363b · outbound

This paper cites Selected Studies of the Principle of Relative Frequency in Language.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Selected Studies of the Principle of Relative Frequency in Language

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.969342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.579612Z digest=sha256:04ab641ecdde58dc94efef0fe113821ab138ac9e3b4f80ef114b7bc835660749

Observation 301056cd-5a09-4409-b131-1c21484cc905 · outbound

This paper cites Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.567261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.567261Z digest=sha256:5b108bc9e066ba3f2bdc23c4028d85df965f852e7906880c5fe8b6ead669d224

Observation 69665084-8ddd-457a-aa2c-ee73b03ebf63 · outbound

This paper cites Because Wj∼N (0,Im) and Mj =⟨v,Wj⟩, for anyi∈ [m] the pair (Mj,Wji) is jointly Gaussian with correlationρ := vi ∥v∥.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Because Wj∼N (0,Im) and Mj =⟨v,Wj⟩, for anyi∈ [m] the pair (Mj,Wji) is jointly Gaussian with correlationρ := vi ∥v∥

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.956995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.585134Z digest=sha256:1c94e296613744c8537f4b5d4ada146c5985d727ea8d918c33638f0ad3915200

Observation f171a23a-e471-4f34-8e9d-ba48be49c6c9 · outbound

This paper cites Still conditional on v, sign(Mj) is±1 with equal probability, independent of the magnitude of Wj.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Still conditional on v, sign(Mj) is±1 with equal probability, independent of the magnitude of Wj

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.943251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.590226Z digest=sha256:9e17180e7000d918e16fb82c34db1e183d5398a97dd7cd58af573bbbd796992e

Observation a4358c98-7dd9-49d2-b588-4b01e82d6cf0 · outbound

This paper cites Because the Yj’s are conditionally independent, Ev Cov(X|v) = Ev h dX j=1 Cov(Yj|v) i =d Im− Ev µ(v)µ(v)⊤ | {z } = 2 πmIm = d− 2d πm Im.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Because the Yj’s are conditionally independent, Ev Cov(X|v) = Ev h dX j=1 Cov(Yj|v) i =d Im− Ev µ(v)µ(v)⊤ | {z } = 2 πmIm = d− 2d πm Im

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.930270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.597097Z digest=sha256:5b1d35d15381ef56dbd3cf3bb686a749c69f45ed740de2554e849a38f0f5d147

Observation 31f911b0-7376-4e10-a45a-49d7fdd94672 · outbound

This paper cites From Step1, E[X|v] =dµ (v), so Covv E[X|v] =d2 Covv µ(v) =d2 2 π Covv v ∥v∥ = 2d2 πmIm.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size From Step1, E[X|v] =dµ (v), so Covv E[X|v] =d2 Covv µ(v) =d2 2 π Covv v ∥v∥ = 2d2 πmIm

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.917670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.601434Z digest=sha256:7d52c75a66116448fa9d4a0eb129bc862f9686270af5b81436f7419c98fba6a8

Observation 99f7d6eb-a9a8-4b1c-8f78-015dd947a521 · outbound

This paper cites Conditioning on E, each entry of E⊤M a centered Gaussian variable, hence E[S(E⊤M)|E] = 0.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Conditioning on E, each entry of E⊤M a centered Gaussian variable, hence E[S(E⊤M)|E] = 0

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.903671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.606351Z digest=sha256:2bbe55d882e7574d694dc30e3833477b7d7978806ef81c179cf912c8cc66cfc6

Observation e1760d3f-2b56-4eaa-be4c-e3aca666982c · outbound

This paper cites Fix a column index k.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Fix a column index k

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.888341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.610552Z digest=sha256:7c139a51bcc694d34313f94d776005f437aaaf1bc51c1cf81bd599f3aee027c9

Observation ef577664-d8da-4277-9f3e-194bb3efff91 · outbound

This paper cites Different columns ofM (differentk) are independent, so Cov(X) is diagonal and each coordinate variance is the same as above.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Different columns ofM (differentk) are independent, so Cov(X) is diagonal and each coordinate variance is the same as above

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.871604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:54:42.614806Z digest=sha256:83b32dcec3e9361a0d7d34fe8239f2864a7c31bb015146a33691fb5c21d024d4

Observation 52994ea5-ed58-4778-9f43-0c7fd4e17cd0 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Adam: A Method for Stochastic Optimization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.507645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.507645Z digest=sha256:69a2dcde633301c8cbb66295ecdb38da98bb3fa02c29cb4ee64b43dd08001810

Observation 3c8dd48f-a720-4c59-839c-43c18097175c · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Exponents Across Parameterizations and Optimizers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.469149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.469149Z digest=sha256:690a8f49d0c12485025b04a348450924eec9e1dc9caaea4ef4ab2cf8d3f56ccc

Observation 08d43654-9d72-4310-8e60-28d09a1f0868 · outbound

This paper cites How Does Critical Batch Size Scale in Pre-training?.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size How Does Critical Batch Size Scale in Pre-training?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.575268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.575268Z digest=sha256:9fd672fe81497290d9985b18c0221f43cb159e682c54a0484b0cef5e2ec46544

Pith citing papers

Observation b12fee4f-c0fe-4629-9cfd-5c377a29cff1 · inbound

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate cites this paper.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.956920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:cf48ac7607b75cd8de1a100c3cde4a245d7c523576de3519ef0f31bebe91775d

Observation 00fd0d3b-1820-4c62-ad19-72b38dc51853 · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:44:42.593237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T07:44:16.677054Z digest=sha256:3cce61327fd98813caf5a04a498c91a33cd352cc51fc0833088a5f28b70fdc1b

Observation bc462715-4679-4150-bef0-3cee690c782f · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.595450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:73e4b82adce915b145bc5f9f8ce6c641b9b0643e5b0baba5d13545a5aa5ba683