Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:54:42.614806Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2506.15025.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:54:42.614806Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T17:31:32.533941Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T17:34:57.593618Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b3f8b846-6e14-4517-be05-f687fba975d8 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size u-$\mu$P: The Unit-Scaled Maximal Update Parametrization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4fc21cb-ea38-4af4-afab-63a2f63c7b5e · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9517dd-5fdd-4a8e-91e5-5d50efc79d17 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On Lazy Training in Differentiable Programming
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e495ec2-61aa-4a99-b5c0-c7a317109251 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Infinite- width limit of deep linear neural networks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52002095-c746-4854-844c-d90658e74767 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Alemi, Roman Novak, Peter J
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 82f197d9-4eb3-444a-83c2-647fe2f7f6c1 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On the infinite-depth limit of finite-width neural networks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 87e1e4de-8cd6-4997-a736-fe3820039dec · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On the impact of the ac- tivation function on deep neural networks training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2f285087-ba6c-4b17-946f-12894b8c479c · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Stable resnet
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b5cf8226-42e4-4f65-8ba2-6adf96c288c4 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd4661cb-4347-492a-aca5-915c0d87ec2a · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590b4620-eaa3-47ac-be14-15a5c3c5747e · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Muon: An optimizer for hidden layers in neural networks,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0f979c3d-4478-4ecb-953e-6ca8318b1ec3 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Kingma and Jimmy Ba
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8abe900-b9de-495f-bf1e-7662454d1e7f · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3075e622-a266-4e46-b146-b163ba8bfff7 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size The llama 3 herd of models, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 230ffc16-b8eb-421f-8d85-49e6a96e7ec7 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Pointer sen- tinel mixture models, 2016
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4976277b-3b6d-46c9-ada3-c88d92d87a5f · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size An Empirical Study of $\mu$P Learning Rate Transfer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6df0e87-6b3a-459e-9930-6d91c199d3ae · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Poole, S
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 07b7ee97-7df3-4b56-9441-1ac76142e1c4 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Schoenholz, J
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 19eda179-91e6-419b-b004-cbde3f0c47a5 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60dcbd05-82f5-42d8-827a-7668986fbb75 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size The Falcon Series of Open Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3fc3b8-693c-4e3e-9bc2-7b06c8d81e88 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Gemma 3 technical report, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f736ab14-e453-4cf0-811d-2a28c79e8fe8 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea63cfd8-d429-48e1-af50-f8099316e3ec · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9469a6ce-f8dc-40e0-b5ca-7656daaba731 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec43db71-939e-494a-b99c-f35d63d9c7a3 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d57d817-22eb-4cb8-a62c-a0f2eb0de283 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size How does critical batch size scale in pre-training?,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 23af0345-ba61-4677-a09e-a194ec2b363b · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Selected Studies of the Principle of Relative Frequency in Language
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 301056cd-5a09-4409-b131-1c21484cc905 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69665084-8ddd-457a-aa2c-ee73b03ebf63 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Because Wj∼N (0,Im) and Mj =⟨v,Wj⟩, for anyi∈ [m] the pair (Mj,Wji) is jointly Gaussian with correlationρ := vi ∥v∥
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f171a23a-e471-4f34-8e9d-ba48be49c6c9 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Still conditional on v, sign(Mj) is±1 with equal probability, independent of the magnitude of Wj
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a4358c98-7dd9-49d2-b588-4b01e82d6cf0 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Because the Yj’s are conditionally independent, Ev Cov(X|v) = Ev h dX j=1 Cov(Yj|v) i =d Im− Ev µ(v)µ(v)⊤ | {z } = 2 πmIm = d− 2d πm Im
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 31f911b0-7376-4e10-a45a-49d7fdd94672 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size From Step1, E[X|v] =dµ (v), so Covv E[X|v] =d2 Covv µ(v) =d2 2 π Covv v ∥v∥ = 2d2 πmIm
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 99f7d6eb-a9a8-4b1c-8f78-015dd947a521 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Conditioning on E, each entry of E⊤M a centered Gaussian variable, hence E[S(E⊤M)|E] = 0
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e1760d3f-2b56-4eaa-be4c-e3aca666982c · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Fix a column index k
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ef577664-d8da-4277-9f3e-194bb3efff91 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Different columns ofM (differentk) are independent, so Cov(X) is diagonal and each coordinate variance is the same as above
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 52994ea5-ed58-4778-9f43-0c7fd4e17cd0 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Adam: A Method for Stochastic Optimization
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c8dd48f-a720-4c59-839c-43c18097175c · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Exponents Across Parameterizations and Optimizers
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08d43654-9d72-4310-8e60-28d09a1f0868 · outbound
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size How Does Critical Batch Size Scale in Pre-training?
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b12fee4f-c0fe-4629-9cfd-5c377a29cff1 · inbound
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 00fd0d3b-1820-4c62-ad19-72b38dc51853 · inbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bc462715-4679-4150-bef0-3cee690c782f · inbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.