Pith. sign in

Paper Citation Record · LEDGER

Why Do We Need Warm-up? A Theoretical Perspective

As of 7 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 2 inbound Pith citation observations for arXiv:2510.03164.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.03164 v2

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:39:03.472216Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T20:13:36.440752Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T20:15:04.536047Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved76
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 728eded1-397e-4d7f-8e45-3cc1e04f527b · outbound

This paper cites plainlm: Language model pretraining in pytorch.

Why Do We Need Warm-up? A Theoretical Perspective plainlm: Language model pretraining in pytorch

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:52.774743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:52.774743Z digest=sha256:fc5babfd6ae76186e68e003d5aaf6ccc4522614a8e702712757575eaa75ef2f4

Observation 62a51332-e1c6-4389-b70f-fc7e06fd687a · outbound

This paper cites vision: Vision model pretraining in pytorch.

Why Do We Need Warm-up? A Theoretical Perspective vision: Vision model pretraining in pytorch

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:52.939248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:52.939248Z digest=sha256:1181d9c4c5cfa13248b28e020c1252df8c8a5c63776d4172bf72325ba152a371

Observation 431fc2e4-8aa5-48a6-ad74-3cd69f9691da · outbound

This paper cites Benefits of learning rate annealing for tuning-robustness in stochastic optimization.

Why Do We Need Warm-up? A Theoretical Perspective Benefits of learning rate annealing for tuning-robustness in stochastic optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:53.057902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:53.057902Z digest=sha256:f735aabf3fdfea2d44dc5f3841626efa507415afa9da759764e745f5cf206eba

Observation 149fb9d5-ab9c-4174-9e23-7fbde6b68020 · outbound

This paper cites Layer Normalization.

Why Do We Need Warm-up? A Theoretical Perspective Layer Normalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:53.264531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:53.264531Z digest=sha256:45a5846c8014978b6dd88c81915ad255904854579fcce62fbdca7de7d94583d1

Observation a889a417-0937-40e6-a32c-1016652dc328 · outbound

This paper cites A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets.

Why Do We Need Warm-up? A Theoretical Perspective A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:53.411038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:53.411038Z digest=sha256:d575cd095abd4b56ed40dec5a6b2ff3aba19ea5196785a6c2fb05a8d2b3239e8

Observation e74688c1-4fe4-4611-9377-81ff0ab82c92 · outbound

This paper cites On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach.

Why Do We Need Warm-up? A Theoretical Perspective On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:53.610360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:53.610360Z digest=sha256:ae73ebf77c56f16a1b7a7db99b8c7f292ab83efb9c0612d7f9ed0c86c6e0679d

Observation 57c806d0-96f4-4a92-99f6-aedbce155fcc · outbound

This paper cites Why Gradients Rapidly Increase Near the End of Training.

Why Do We Need Warm-up? A Theoretical Perspective Why Gradients Rapidly Increase Near the End of Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:53.777630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:53.777630Z digest=sha256:9f1227ff483ebd59f2aadc34b306a45c2bbaf6f4437e2c514fa50f91d083fb3e

Observation e1bde069-6bd4-47e7-9f02-85ea680d3b84 · outbound

This paper cites Optimal Linear Decay Learning Rate Schedules and Further Refinements.

Why Do We Need Warm-up? A Theoretical Perspective Optimal Linear Decay Learning Rate Schedules and Further Refinements

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:53.947945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:53.947945Z digest=sha256:9de83b4960c2fec7cf6a1cd328a0d269f4fda18af54a19f6ffd5568555478fbc

Observation 184c32fa-1ef5-4c17-a20f-3776cb854a4e · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Why Do We Need Warm-up? A Theoretical Perspective An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:54.083987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:54.083987Z digest=sha256:db996d2f11ae6a0cef36b25c3156c8e99537691399a366fc761aa5c44780d84f

Observation a5ed8a55-fe3a-47d2-8dbc-4f922557ea2f · outbound

This paper cites Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler.

Why Do We Need Warm-up? A Theoretical Perspective Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:54.214964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:54.214964Z digest=sha256:8f8b983edf17835c6790e625939cc0aaf401829084e3ff4842dee56e249c80d3

Observation 70b39ca0-925c-4eb8-8faa-2bdc87902b32 · outbound

This paper cites Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced.

Why Do We Need Warm-up? A Theoretical Perspective Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:54.348591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:54.348591Z digest=sha256:47c30adc4d3461501a6f58f3ee3918c698742c55f48777f332095e9e15729165

Observation 7d1f73df-10a9-4d79-b227-7833a31e70a4 · outbound

This paper cites Beyond uniform smoothness: A stopped analysis of adaptive sgd.

Why Do We Need Warm-up? A Theoretical Perspective Beyond uniform smoothness: A stopped analysis of adaptive sgd

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:54.532210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:54.532210Z digest=sha256:3b1d71a05ed100403da40a6c1f7148092773bcff3811de7c61682b4cbb98947a

Observation 9e70ac17-cdb1-4ac4-b931-0c8ea7a2d36d · outbound

This paper cites Accelerated stochastic optimization methods under quasar-convexity.

Why Do We Need Warm-up? A Theoretical Perspective Accelerated stochastic optimization methods under quasar-convexity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:54.667724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:54.667724Z digest=sha256:8eb7926b9c0ab610a7ef790c28855dfd04445b9c45616144e8ffa3f249e8d696

Observation 491f06aa-614d-407b-b324-f5fd86136d8c · outbound

This paper cites Convergence of Clipped SGD on Convex $(L_0,L_1)$-Smooth Functions.

Why Do We Need Warm-up? A Theoretical Perspective Convergence of Clipped SGD on Convex $(L_0,L_1)$-Smooth Functions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:54.874128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:54.874128Z digest=sha256:ca833377810bd6ca43bc4899821a5beb0044e2082b55ffaa2d60b1958fac18b0

Observation f3373c31-9333-445c-982e-05503cf50e41 · outbound

This paper cites A Loss Curvature Perspective on Training Instability in Deep Learning.

Why Do We Need Warm-up? A Theoretical Perspective A Loss Curvature Perspective on Training Instability in Deep Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:55.058339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:55.058339Z digest=sha256:29a23e719e3f161e071e2f26ce8a3f33e8c77542d989e21a2f8c99ec08e84e6b

Observation 72115b05-738f-407a-ae44-a5cd7453e73a · outbound

This paper cites Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity.

Why Do We Need Warm-up? A Theoretical Perspective Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:55.256221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:55.256221Z digest=sha256:6156adeff7e3978439e7d77c09f54737032448d54b5de1867a59c90a9eac43a8

Observation eee57c98-a2a0-4b37-b089-570b82feb3a2 · outbound

This paper cites A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation.

Why Do We Need Warm-up? A Theoretical Perspective A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:55.371093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:55.371093Z digest=sha256:9351a15c7be8490c1529b9d2629595d1a58924286dc370c4fdfa2e45358a02bf

Observation a24f7417-8014-4bcd-af6c-ac5c3a40258a · outbound

This paper cites Sgd for structured nonconvex functions: Learning rates, minibatching and interpolation.

Why Do We Need Warm-up? A Theoretical Perspective Sgd for structured nonconvex functions: Learning rates, minibatching and interpolation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:55.510663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:55.510663Z digest=sha256:dfba7e05a3508945970478362eba6c219c3a8a5663a1e0017e8002ff73d6f6e7

Observation cdc5061f-6659-4cd0-8063-56f34f4ec4bc · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Why Do We Need Warm-up? A Theoretical Perspective Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:55.605650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:55.605650Z digest=sha256:65aeb07a8a205636e88bca6c09288c6cb186e2bf698a3cd0b4ab48344f561140

Observation f633f650-1203-4683-aafd-08491a57fc17 · outbound

This paper cites No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths.

Why Do We Need Warm-up? A Theoretical Perspective No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:55.710545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:55.710545Z digest=sha256:91ee2545789de32058a9b81e883e4eb8432d72deacef742e506231c3804a63a4

Observation 590dcdf8-1c8f-4e78-a4f7-8312681dabf5 · outbound

This paper cites Scaling laws and compute-optimal training beyond fixed training durations.

Why Do We Need Warm-up? A Theoretical Perspective Scaling laws and compute-optimal training beyond fixed training durations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:55.886801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:55.886801Z digest=sha256:94f546060d08c8f7a32edfcc6701648e13307939fbb2922d215d14c172ace601

Observation eec30366-a3cc-4b14-a855-8cd81cc3c4a4 · outbound

This paper cites Gradient descent learns linear dynamical systems.

Why Do We Need Warm-up? A Theoretical Perspective Gradient descent learns linear dynamical systems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.010775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:56.010775Z digest=sha256:3930e5e526d6f65f2c2081bf49a292c613f555e63e76903c030a04ab0522877d

Observation 9da8f203-1a1f-49b6-9256-473c352cb250 · outbound

This paper cites Deep residual learning for image recognition.

Why Do We Need Warm-up? A Theoretical Perspective Deep residual learning for image recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.108598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:56.108598Z digest=sha256:ab40b98c374f3f3d4282fb6c07b578befd222b6ae23eee84fe60c2979db7f2dc

Observation b75d76bc-65a8-42f5-b00a-7d9f6a864020 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Why Do We Need Warm-up? A Theoretical Perspective Gaussian Error Linear Units (GELUs)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.263519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:56.263519Z digest=sha256:8a44edc26a6a227ccbe3d100a544754ef1418fc38ef28a73f08819efc21c3151

Observation 383d99fb-5cdd-4ed3-a254-cb89e9b486d4 · outbound

This paper cites Near-optimal methods for minimizing star-convex functions and beyond.

Why Do We Need Warm-up? A Theoretical Perspective Near-optimal methods for minimizing star-convex functions and beyond

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.359484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:56.359484Z digest=sha256:74475e29533346d21b314edf40ff21310f8746ef8b8da926dea7a31f7fe194ef

Observation d9a7c06d-9155-4ed8-87d9-c3264157ab62 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Why Do We Need Warm-up? A Theoretical Perspective Training Compute-Optimal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.463192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:56.463192Z digest=sha256:1dafae8022fc8b38521a229bf9b0193b44f586e4b4ba1d91f0e4ad37da962cd0

Observation 6269db09-27c3-495f-9c95-86dc0e42ccac · outbound

This paper cites An empirical analysis of compute-optimal large language model training.

Why Do We Need Warm-up? A Theoretical Perspective An empirical analysis of compute-optimal large language model training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.629540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:56.629540Z digest=sha256:23517030ecb4ab37be932c66d4e241480e6f40ea710c0a8b60c539b0ef14d799

Observation 4fb71426-0151-48e3-ba08-f7b28326a805 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Why Do We Need Warm-up? A Theoretical Perspective MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.790434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:56.790434Z digest=sha256:88a0a37292727a1bfce01f16ddcfc9869901dd0ed37cbcb4ab898a39064f4cff

Observation 036ce580-74e4-433c-a011-634ce01a186c · outbound

This paper cites Improving transformer optimization through better initialization.

Why Do We Need Warm-up? A Theoretical Perspective Improving transformer optimization through better initialization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.930852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:56.930852Z digest=sha256:5db5214802383517aad9c439096f63d25fdaaa5fdb47c446e946e8fd1351d95b

Observation 00671aad-5bd1-43cb-95ec-42f36a89434a · outbound

This paper cites Loss landscape characterization of neural networks without over-parametrization.

Why Do We Need Warm-up? A Theoretical Perspective Loss landscape characterization of neural networks without over-parametrization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.081181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:57.081181Z digest=sha256:7afb45e8299fe6acd7f3f512a9bfbfb642e524f8c74e7cc7560d0993060e4d5c

Observation 3d011ce7-8726-40bc-91bb-7a331b5e3f86 · outbound

This paper cites Why warmup the learning rate? underlying mechanisms and improvements.

Why Do We Need Warm-up? A Theoretical Perspective Why warmup the learning rate? underlying mechanisms and improvements

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.244841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:57.244841Z digest=sha256:1611c92a40651d857da865d7616f98e8381a8bd8d975db475aa6a03df7c4bee9

Observation aef0ebda-5f64-4efa-aed9-2b6e38e4f6ad · outbound

This paper cites Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition.

Why Do We Need Warm-up? A Theoretical Perspective Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.374428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:57.374428Z digest=sha256:07440d13d858fbdc6b397a1d011d78461f33d126fd4d5341f0394652f543183d

Observation 2b36a278-8f3b-4563-8896-c93683abe061 · outbound

This paper cites an unresolved cited work.

Why Do We Need Warm-up? A Theoretical Perspective Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.507142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:57.507142Z digest=sha256:00747b43a551d1aa92ff3895fae91275f5e1e67011d35cfa07b617b842feaf78

Observation 99514f4e-22b3-4f30-957e-afc94d8b474f · outbound

This paper cites Deep learning without poor local minima.

Why Do We Need Warm-up? A Theoretical Perspective Deep learning without poor local minima

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.609783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:57.609783Z digest=sha256:2f6c7f289990c29a8bb8d9ab9191cdacef1e25825c13c01aef86e9522c3b077d

Observation 116cd242-7c43-4402-9b22-ea8c50783d42 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Why Do We Need Warm-up? A Theoretical Perspective Adam: A Method for Stochastic Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.721743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:57.721743Z digest=sha256:b1bdc6aac9f052aa0b8ff37527626981d97ab2de185ded38280f611551f6edae

Observation 706da65f-d409-42cf-bd5b-c7c9ea81f64f · outbound

This paper cites An alternative view: When does sgd escape local minima? In International conference on machine learning, 2018.

Why Do We Need Warm-up? A Theoretical Perspective An alternative view: When does sgd escape local minima? In International conference on machine learning, 2018

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.819361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:57.819361Z digest=sha256:b1b515a03bca58a1ba5af97903a7de3534706c988889f9ebf804de3401513a04

Observation b1228a4f-78df-4eb7-9f23-4784716f46cf · outbound

This paper cites Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis.

Why Do We Need Warm-up? A Theoretical Perspective Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.947144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:57.947144Z digest=sha256:cfc98722e396ff4b4805022f3c750812af3012ac2919137dbd4105160712b642

Observation a0c02409-bb40-46c1-b44c-3037fcfe5985 · outbound

This paper cites Analyzing & reducing the need for learning rate warmup in gpt training.

Why Do We Need Warm-up? A Theoretical Perspective Analyzing & reducing the need for learning rate warmup in gpt training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.133345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.133345Z digest=sha256:95d0f0d04f89cd289c8a39571a3190ca8308fe2c9f4521af4a30d57c21f12620

Observation b53decd7-bd22-4517-b052-90244089927c · outbound

This paper cites Convex and non-convex optimization under generalized smoothness.

Why Do We Need Warm-up? A Theoretical Perspective Convex and non-convex optimization under generalized smoothness

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.310721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.310721Z digest=sha256:94cdd4e35de9b9b3929a4402aee24da9a93ce3b8b08378719dd417e3aa0a5b24

Observation 719634b7-1aa9-4edb-895a-08254fc9d44c · outbound

This paper cites Loss landscapes and optimization in over-parameterized non-linear systems and neural networks.

Why Do We Need Warm-up? A Theoretical Perspective Loss landscapes and optimization in over-parameterized non-linear systems and neural networks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.449756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.449756Z digest=sha256:2bfe92cc1a9812b72e5d3a2f0a1eec504a3c4b136aaf5588116bb8b6976cd54c

Observation 21efe338-a1bc-4817-b4e5-3e35f32d45ad · outbound

This paper cites Aiming towards the minimizers: fast convergence of sgd for overparametrized problems.

Why Do We Need Warm-up? A Theoretical Perspective Aiming towards the minimizers: fast convergence of sgd for overparametrized problems

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.650623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.650623Z digest=sha256:ac2798f700a80078f32a13364526ab2b97877407a10c9405af2e2baeb0e9beb3

Observation 4a969c6c-cba2-4130-8c93-7308ff18d135 · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

Why Do We Need Warm-up? A Theoretical Perspective On the Variance of the Adaptive Learning Rate and Beyond

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.819676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.819676Z digest=sha256:c4f82e8eeedaae7a0bfc8e8d348218d6ef56919e930240402cd8caf0da3631d2

Observation 89caaa2e-7bee-452a-9079-db1bb129da9e · outbound

This paper cites Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence.

Why Do We Need Warm-up? A Theoretical Perspective Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.891916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.891916Z digest=sha256:7188a8127121f0db023d7b8b4b442cbadc49e96a5bb7ae70e0fc0c5716e51724

Observation 648c8223-3bce-474f-ad61-e981523edfb9 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Why Do We Need Warm-up? A Theoretical Perspective SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.029439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.029439Z digest=sha256:b5e55c1280d796a49f71858ee3c012b4dd67b8da6a6847cb381c46c6308ac16b

Observation 88d91fec-cb0e-47cd-adb3-84a865af0838 · outbound

This paper cites Decoupled Weight Decay Regularization.

Why Do We Need Warm-up? A Theoretical Perspective Decoupled Weight Decay Regularization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.193626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.193626Z digest=sha256:46f4f5bc346f6884e8f0740b0233c4c625630ad899e77476c4fefdf271f02ae5

Observation a1bac448-bfbd-4ba4-b4da-9731b4c89095 · outbound

This paper cites The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning.

Why Do We Need Warm-up? A Theoretical Perspective The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.249809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.249809Z digest=sha256:2a7b15d09453ff15693def126a193642491aacd5b2787432b9576b99302a70cb

Observation ffa64622-a0a7-40a3-a49a-eeea6c54b6bd · outbound

This paper cites Matrix differential calculus with applications to simple, hadamard, and kronecker products.

Why Do We Need Warm-up? A Theoretical Perspective Matrix differential calculus with applications to simple, hadamard, and kronecker products

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.350720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.350720Z digest=sha256:84ac56004b002446949e93472658ccd618a402a74546c8a9609206caf1d33635

Observation ed497575-5314-43e2-91da-05d9b6c4f757 · outbound

This paper cites An Empirical Model of Large-Batch Training.

Why Do We Need Warm-up? A Theoretical Perspective An Empirical Model of Large-Batch Training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.426040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.426040Z digest=sha256:0bc5e8abf12c5a88a2420fc57d50c126d0f21a10143738aeb0e3232997a0587f

Observation dcbb8440-442b-4c34-9222-c176b5e1dabb · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

Why Do We Need Warm-up? A Theoretical Perspective The fineweb datasets: Decanting the web for the finest text data at scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.495969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.495969Z digest=sha256:56a928ca7ad3494b2265418285e3f3cdd9ab3bb0343c9a33c83298d9c1416cb5

Observation 75c04dd8-88c7-448a-bd6c-7b29f0a42f71 · outbound

This paper cites Gradient methods for minimizing functionals.

Why Do We Need Warm-up? A Theoretical Perspective Gradient methods for minimizing functionals

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.578629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.578629Z digest=sha256:5a3686160f434d2a3b2a94391c1c77ace438728283408c2da5045e529cd3b7df

Observation 5b7c22af-6cb4-40cd-b65c-e90f2ac141be · outbound

This paper cites Language models are unsupervised multitask learners.

Why Do We Need Warm-up? A Theoretical Perspective Language models are unsupervised multitask learners

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.674564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.674564Z digest=sha256:c679a7c9ac7142481077501dd19acdee5ec96bc6eb065ddae4aa7d046696eda7

Observation 95a62f5f-478d-46fb-89b6-07d1f2248832 · outbound

This paper cites Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs).

Why Do We Need Warm-up? A Theoretical Perspective Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.776219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.776219Z digest=sha256:2403d1fa6575930fd6e47404b90797da4f30f6974fca00e2e524017a72301f3a

Observation 05b0d5fe-f2e7-4478-b1b0-2521860b2175 · outbound

This paper cites Stepping on the edge: Curvature aware learning rate tuners.

Why Do We Need Warm-up? A Theoretical Perspective Stepping on the edge: Curvature aware learning rate tuners

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.901620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.901620Z digest=sha256:46bf698b065c4682b94678263cb00665c360862bf3490f57a6121063aa7b71b4

Observation ce80e345-2c2a-46c4-b820-4fbd67f88085 · outbound

This paper cites The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training.

Why Do We Need Warm-up? A Theoretical Perspective The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.978009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.978009Z digest=sha256:ad51e87059f2d4d1a8ea17da3570ace53dffb3ebafab616c686cf013085142b0

Observation e558cf7f-ddfc-4c8a-a555-a88298d98e2f · outbound

This paper cites GLU Variants Improve Transformer.

Why Do We Need Warm-up? A Theoretical Perspective GLU Variants Improve Transformer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.127984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:00.127984Z digest=sha256:1e25d72b6e873cd34d4dd87e97a7db3d94c461aea150f4bcda87cde1632a1a3d

Observation 7d5a15b6-fd2b-4aba-8e62-141cba87251f · outbound

This paper cites On the generalization benefit of noise in stochastic gradient descent.

Why Do We Need Warm-up? A Theoretical Perspective On the generalization benefit of noise in stochastic gradient descent

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.284259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:00.284259Z digest=sha256:13b18ac573564466c4c3c09774c7ea18c4bc40bf60a6eeb765cf040d2c342556

Observation 2ba07a3e-b380-47db-96df-167002843834 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Why Do We Need Warm-up? A Theoretical Perspective Roformer: Enhanced transformer with rotary position embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.425116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:00.425116Z digest=sha256:56590da41b911e147039ca48ff736bc92513e0045f9dc8aec98d10f132345630

Observation 217d4747-d2a8-4904-a697-80d9c57895c1 · outbound

This paper cites On the importance of initialization and momentum in deep learning.

Why Do We Need Warm-up? A Theoretical Perspective On the importance of initialization and momentum in deep learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.578854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:00.578854Z digest=sha256:a067d1dc4570b5de1f564ab3ea010b1cf24df368e1a9ea62aba30c6a18f07fff

Observation 4b7c4293-5cfe-4b91-af61-bbeaacf0670a · outbound

This paper cites Fast convergence in learning two-layer neural networks with separable data.

Why Do We Need Warm-up? A Theoretical Perspective Fast convergence in learning two-layer neural networks with separable data

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.714356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:00.714356Z digest=sha256:2fdbd9a40380063f2dda5aa56fcc2f69c3c39c663ef44e4b224c231a941a4120

Observation adeb38bf-a005-40ec-8257-9ae372c03ed9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Why Do We Need Warm-up? A Theoretical Perspective Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.861463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:00.861463Z digest=sha256:ed808a0931437eafd960e40eeb48a09eb24b750825ba28f2848a7f6d6c2f5597

Observation 84f68236-ec75-4826-868b-f8c0866214bc · outbound

This paper cites Empirical tests of optimization assumptions in deep learning.

Why Do We Need Warm-up? A Theoretical Perspective Empirical tests of optimization assumptions in deep learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:01.030025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:01.030025Z digest=sha256:cd5f467888b40fe6b8e1bf1fbc9fa24f95de44eff17478fd0b4560609fdb2fb2

Observation a80581cf-2114-4a12-88fb-cf5e581d546b · outbound

This paper cites Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods.

Why Do We Need Warm-up? A Theoretical Perspective Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:01.193581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:01.193581Z digest=sha256:a777a1caa20ef592b9b7398932ec69e95dbc0e6299c70f70b4e23d34f6e1375c

Observation a1650146-ef7d-4966-b190-4e8aa5da6398 · outbound

This paper cites Attention is all you need.

Why Do We Need Warm-up? A Theoretical Perspective Attention is all you need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:01.314115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:01.314115Z digest=sha256:cfc275494694b163f1722b6a207805f00f6e4bb3b8a173519e7156f7a1f03174

Observation 5f300618-eb15-4eee-8bcf-2c4bf1534160 · outbound

This paper cites Convergence of adagrad for non-convex objectives: Simple proofs and relaxed assumptions.

Why Do We Need Warm-up? A Theoretical Perspective Convergence of adagrad for non-convex objectives: Simple proofs and relaxed assumptions

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:01.594610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:01.594610Z digest=sha256:40d9cc6a1ec496255fbe4fa6e1575145cdebc9787fb29d830d1e98136fa6a9a0

Observation bc431704-a053-41e8-993a-0db5dd316058 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Why Do We Need Warm-up? A Theoretical Perspective Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:01.733820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:01.733820Z digest=sha256:7402cd86ce6fd724d646dd113d35d6774c8c0f2b1dae5518614b21ef5267e178

Observation 125f1764-fac9-4207-85eb-4803f87867a0 · outbound

This paper cites Small-scale proxies for large-scale Transformer training instabilities.

Why Do We Need Warm-up? A Theoretical Perspective Small-scale proxies for large-scale Transformer training instabilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:01.892794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:01.892794Z digest=sha256:6b466e32e50b68f6cda237b0f7bd279769b820a4a623f50b98fdf1695b0902b7

Observation a98b8919-b85a-4f15-abff-ca0412f2a3df · outbound

This paper cites On the overlooked pitfalls of weight decay and how to mitigate them: A gradient-norm perspective.

Why Do We Need Warm-up? A Theoretical Perspective On the overlooked pitfalls of weight decay and how to mitigate them: A gradient-norm perspective

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:02.101607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:02.101607Z digest=sha256:35a7d0228cdb380f01fb56687e5cf066b40176c98671805d8e7a393443479584

Observation 9abf9a17-3192-4b9a-abf4-93183a4fca72 · outbound

This paper cites On layer normalization in the transformer architecture.

Why Do We Need Warm-up? A Theoretical Perspective On layer normalization in the transformer architecture

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:02.305117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:02.305117Z digest=sha256:264f22f2b0e3efae899570695237c99e661658659d7e0ddfc53530d830fc3dea

Observation 1e455ef0-d364-4d76-a416-edcf5f517509 · outbound

This paper cites Dive into deep learning.

Why Do We Need Warm-up? A Theoretical Perspective Dive into deep learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:02.459232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:02.459232Z digest=sha256:6db8a80358eea85acc0ca1e3d5d3eb4069b4376a6922d3eda84ae39b93014ffe

Observation a2fe7e7d-00db-4cad-840f-f698d4c32a39 · outbound

This paper cites Root mean square layer normalization.

Why Do We Need Warm-up? A Theoretical Perspective Root mean square layer normalization

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:02.605680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:02.605680Z digest=sha256:55d3ea416a6807f78a312d40f902aa166ba6ecc76b42219c79bc8be7b0da4e32

Observation dcdfc5bf-6ca6-49af-94f2-6e3622e0ceda · outbound

This paper cites Improved analysis of clipping algorithms for non-convex optimization.

Why Do We Need Warm-up? A Theoretical Perspective Improved analysis of clipping algorithms for non-convex optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:02.697714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:02.697714Z digest=sha256:aff5fc3209adf63325f7c096575107280feb066359d2eeb14d591a87277498b0

Observation b8c02199-f48d-4b06-a70f-52f3feac5a49 · outbound

This paper cites Why gradient clipping accelerates training: A theoretical justification for adaptivity.

Why Do We Need Warm-up? A Theoretical Perspective Why gradient clipping accelerates training: A theoretical justification for adaptivity

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:02.842307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:02.842307Z digest=sha256:cefb9fde1efb749f575adf041002c5e8ea46aaf95519946ef8d4cc1e2ca0ffc1

Observation d90af45d-98dc-4e8a-9d6f-b03a9e6d6e45 · outbound

This paper cites On the convergence and improvement of stochastic normalized gradient descent.

Why Do We Need Warm-up? A Theoretical Perspective On the convergence and improvement of stochastic normalized gradient descent

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:02.990431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:02.990431Z digest=sha256:03b51a8d5b40268d0a30c0fc539feab81f861be6c4c1807f6660e6890ac4a8eb

Observation ba8c39b4-7f2b-47c8-aa2b-6a096121e6db · outbound

This paper cites @esa (Ref.

Why Do We Need Warm-up? A Theoretical Perspective @esa (Ref

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:03.122139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:03.122139Z digest=sha256:ce11f5a3c25822709e1c21456b037c14b1204bc658c62458fca804d1010560dd

Observation 28aafa4e-8ded-42f3-a3f4-0ad38030173d · outbound

This paper cites an unresolved cited work.

Why Do We Need Warm-up? A Theoretical Perspective Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:03.316608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:03.316608Z digest=sha256:3f4b48c69de0dc82517d704eca9297980833c3b974b1be56cef4e737d6f27f35

Observation c689fca2-05c8-4d5c-bb73-37bdf0933cbd · outbound

This paper cites an unresolved cited work.

Why Do We Need Warm-up? A Theoretical Perspective Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:03.472216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:39:03.472216Z digest=sha256:8228fefb5c0049285c661a1f6fc6411b7df17b7489ed0168c5e41f00d0863843

Pith citing papers

Observation 91e5ed69-2867-42ed-bbba-e29abcf80be0 · inbound

Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning with Real-Time Feedback cites this paper.

Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning with Real-Time Feedback Why Do We Need Warm-up? A Theoretical Perspective

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:16:10.224205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:15:57.763495Z digest=sha256:0bdfc8980db39e560dbfe4afb411a46e2a574d03f0fb499d0641ec9371e3acf1

Observation c2ce85ae-8831-4397-8cca-f6e072d2affe · inbound

Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness cites this paper.

Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness Why Do We Need Warm-up? A Theoretical Perspective

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T20:15:04.538172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T20:13:36.440752Z digest=sha256:6160664b2607fe4a5b1627bc33f968b633afd640781004b0ffe7bc1a77058d58