Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T12:00:49.127471Z
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2605.06654.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T12:00:49.127471Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T21:13:39.201865Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T11:56:55.992114Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c2207734-b6f6-42fc-a872-15b9fd029f84 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less arXiv preprint arXiv:2512.16928 , year=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1a26534a-2a18-4038-9238-482c3f75d6d6 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less The Geometry of Sign Gradient Descent
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7f285ec0-7c51-4989-94a2-5e67b5bff8f7 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Old Optimizer, New Norm: An Anthology
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ad4e9f03-6c4e-478a-826b-f9b537b6122c · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less LoRA Learns Less and Forgets Less
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2c37b115-b48a-4f6b-83e9-cfa5e1166495 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Why Gradients Rapidly Increase Near the End of Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9aa17456-7f4c-4915-8828-f9adb80ba8cb · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7d23d21a-41e1-4567-924e-045438fda472 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Gaussian Error Linear Units (GELUs)
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 75507299-eb8a-4078-83fb-a22a249ceb3b · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Measuring Forgetting of Memorized Training Examples
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7ee54a71-d2e5-47f8-929b-96ac55cfdafa · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 08cf0d11-a3bd-41b6-be38-ff7898efc49b · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Kimi K2.5: Visual Agentic Intelligence
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 36fc57b1-e1fc-445b-9dbc-effb24ae7b67 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Adam: A Method for Stochastic Optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f777e85e-7854-414a-a13e-ad3cc895b8ce · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4ac87229-543a-4618-82ed-a76d92920cdc · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Muon is Scalable for LLM Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 20218c8c-a913-4081-90e6-74a40eeb76b6 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less AdaGrad under Anisotropic Smoothness
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 96edba34-090c-4e5c-b12d-8427fd30a166 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4f098732-d3b8-4da8-b2ed-62c4bafda5c2 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Decoupled Weight Decay Regularization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f4527207-05ea-4dd8-9507-941d1a833de2 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1e85877e-44bc-4d15-a1a8-5749c9a97cc2 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Unbiased gradient low-rank projection.arXiv preprint arXiv:2510.17802
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8de20694-9b62-4668-9ea0-79d4fddf3e94 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Training Deep Learning Models with Norm-Constrained LMOs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f1f02a4f-4515-4026-a6d2-bd867ff84826 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less icarl: Incre- mental classifier and representation learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7a950eec-14aa-4e47-ab6c-b8982a1f4b85 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less On the Convergence of Adam and Beyond
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e69b531a-65e8-493c-9425-606bceeda5fb · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less (How) Learning Rates Regulate Catastrophic Overtraining
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c241f457-713a-41f8-9d3e-7a495da4d34b · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Benchmarking Optimizers for Large Language Model Pretraining
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fd97bb7d-23c4-4df2-b005-67030e28bb84 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less GLU Variants Improve Transformer
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ee1c7cdd-d492-4dec-a458-30197356ce85 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Lora vs full fine-tuning: An illusion of equivalence
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b97e7a15-a4d7-4248-9177-2ad6489cbcff · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Overtrained Language Models Are Harder to Fine-Tune
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ab44dce8-7b8b-44cb-902e-4aa2bd1a0bc3 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Less Regret via Online Conditioning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1e3575ad-ab31-4b6d-8858-e9664cfd7ddf · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less ArXiv Preprint: 2511.00674 , Year =
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d8f36dc5-1a7a-4d95-a955-201b4894fdc2 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 29e05265-12eb-40cc-95fd-de75bd8b3692 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less SOAP: Improving and Stabilizing Shampoo using Adam
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6a34c718-50c3-41dc-a453-89bed6c500bc · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Muon outperforms adam in tail-end associative memory learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8347914b-4c3a-4e8e-aefe-e7d261dbc576 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Magicoder: Empowering Code Generation with OSS-Instruct
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fbb0ff62-ebc5-4ca2-8e0d-a4c89742f877 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Fantastic Pretraining Optimizers and Where to Find Them
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 804a2766-31e4-49f3-92f5-fb3a064bf02b · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Structured Preconditioners in Adaptive Optimization: A Unified Analysis
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ccac7eb6-8554-4f1d-ba41-a2562d30ea4f · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Controlled llm training on spectral sphere
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c4701b1e-aca3-42f1-8db0-da32010a60c0 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less On the width scaling of neural optimizers under matrix operator norms i: Row/column normalization and hyperparameter transfer.arXiv preprint arXiv:2603.09952
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 300a9d68-6a8e-4151-aa35-e275248e36ae · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Qwen3 Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 982ace06-fbed-4f02-b303-9cc6e2c4cdfa · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less A Spectral Condition for Feature Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d01d0395-cad3-4f9b-8c96-e98959813614 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 06f3d469-ea31-425a-b541-7c6f1cd48023 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f9e4fb2e-0950-4211-b854-5b50d1359a24 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 538e25c1-2441-4600-8506-41ad000ad01b · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7998c599-731f-4125-89ab-d025e16f3398 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less ADADELTA: An Adaptive Learning Rate Method
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a21fa863-90da-41e5-9bfa-ccd8459b71d7 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2b360052-3f06-406c-91f1-0c3317a26c5a · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Understanding deep learning requires rethinking generalization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b7dff74f-a846-48eb-8e76-8b915e9acffc · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Why Transformers Need Adam: A Hessian Perspective
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9d9eb6cb-2630-4c13-a323-20ebf4406cca · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 34dcb554-30af-42cc-b424-357df4da4d06 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 54569520-ba6a-4626-91be-97757a02b5b6 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Here, LoRA rank 64 shows a greater forgetting compared to rank 256, mainly because rank 256 diverges for lr=5e-4, and a smaller learning rate leads to less forgetting
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 581d67f9-75c1-41ff-b6eb-d653bb622d9a · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2469ed4d-f60e-4a5b-8514-9601ba82e64b · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less B.2 More Detailed Activation Plots In Figures 10 and 11, we present the average activation sparsity of detailed modules and specific ac- tivation splits
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 57ef95fa-30bb-4711-b3c2-176177269826 · outbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less 23 •Ifα∈(α 1,∞], based on Lemma 2 and Assumption 4, it holds that ∥∆W∥ α1,β∗ ≤ ∥∆W∥ α,β∗ andE h ∥x∥2 α1 i = Θ E h ∥x∥2 α i
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0dcd802d-c6f2-4675-a64f-0c12f2d4c537 · inbound
Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 259292d8-88cf-4435-a378-1f4e34a0f944 · inbound
When Does Muon Help Agentic Reinforcement Learning? Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.