Pith. sign in

Paper Citation Record · LEDGER

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

As of 3 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2605.06654.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.06654 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T12:00:49.127471Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:13:39.201865Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T11:56:55.992114Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact41
  • verified fuzzy4
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2207734-b6f6-42fc-a872-15b9fd029f84 · outbound

This paper cites arXiv preprint arXiv:2512.16928 , year=.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less arXiv preprint arXiv:2512.16928 , year=

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:08.621624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:5e56a8ac5f6daa49b61bed23aca706cf9e75f1c29cd069a6330c838d181a31ba

Observation 1a26534a-2a18-4038-9238-482c3f75d6d6 · outbound

This paper cites The Geometry of Sign Gradient Descent.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less The Geometry of Sign Gradient Descent

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.613799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:ec02c87bf9cc5b4ce1b74969e1160f7f9c1b308fa8933ce1858fed39e2b590fe

Observation 7f285ec0-7c51-4989-94a2-5e67b5bff8f7 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Old Optimizer, New Norm: An Anthology

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:27:53.041705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:c0ef45cbf08667318b6a5212d82c54f44964c40d74bb1866165e110736c47a63

Observation ad4e9f03-6c4e-478a-826b-f9b537b6122c · outbound

This paper cites LoRA Learns Less and Forgets Less.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less LoRA Learns Less and Forgets Less

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:08.668232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:1ed4983f8c444420c1b49e26a38bb9880b5d14e5747d35aded14ed86430a1482

Observation 2c37b115-b48a-4f6b-83e9-cfa5e1166495 · outbound

This paper cites Why Gradients Rapidly Increase Near the End of Training.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Why Gradients Rapidly Increase Near the End of Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.465260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:2de5ad19f8eaa7ee24d41fc7b52f0e2dbabe59d9d8bde7ee4b42aeaf1b585e4d

Observation 9aa17456-7f4c-4915-8828-f9adb80ba8cb · outbound

This paper cites The Llama 3 Herd of Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less The Llama 3 Herd of Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.588850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:5ccc579951d75629c7e7fe52630bdd4446be963faa943f6554a888cf615fcead

Observation 7d23d21a-41e1-4567-924e-045438fda472 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Gaussian Error Linear Units (GELUs)

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.454674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:1c308970c8997efc29de80a3216c59e865e55faab54a4420c7471045e9c21786

Observation 75507299-eb8a-4078-83fb-a22a249ceb3b · outbound

This paper cites Measuring Forgetting of Memorized Training Examples.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Measuring Forgetting of Memorized Training Examples

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.593021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:595d59336ea992c85f16e8f5801ec4c4ffef435692cb439d005cf287666b9c97

Observation 7ee54a71-d2e5-47f8-929b-96ac55cfdafa · outbound

This paper cites Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.606390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:cae53d1e39b5d42fa8f40696701224af2affdd561e90a7a60548417f19e86bdf

Observation 08cf0d11-a3bd-41b6-be38-ff7898efc49b · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Kimi K2.5: Visual Agentic Intelligence

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.580519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:4230c63ed4fabb7a2c8ade862e074b257ea3b87546a7d102ef04575fc3b8e30c

Observation 36fc57b1-e1fc-445b-9dbc-effb24ae7b67 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Adam: A Method for Stochastic Optimization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.584831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:046c466efc2fc347996fc8dce8d51fb955e1c9acc690b4b4ef6969235926b4fb

Observation f777e85e-7854-414a-a13e-ad3cc895b8ce · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.565537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:7ae1c4775b1715c422e3ede682cf3de4024b850dd82c965e0d61f862fd72ebc6

Observation 4ac87229-543a-4618-82ed-a76d92920cdc · outbound

This paper cites Muon is Scalable for LLM Training.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Muon is Scalable for LLM Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:e897c2fe7b3b4d856b28d7ed7eae387aadce27ab1652c0d85be657c7b0d133bf

Observation 20218c8c-a913-4081-90e6-74a40eeb76b6 · outbound

This paper cites AdaGrad under Anisotropic Smoothness.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less AdaGrad under Anisotropic Smoothness

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.577041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:5ea7b774ba5afdaf957d5203e3d15a90188d1611cb26090c85bb91a4ec382dda

Observation 96edba34-090c-4e5c-b12d-8427fd30a166 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.573133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:1c0944b7740dfdfce7b2955c05c2755ecc035f2282281aa23ed4dd00c9ea7548

Observation 4f098732-d3b8-4da8-b2ed-62c4bafda5c2 · outbound

This paper cites Decoupled Weight Decay Regularization.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Decoupled Weight Decay Regularization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.600508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:d1eb9a1f2c97e8990280a3ba4da7cd5197d38b42781ac19572fd2311f76c1906

Observation f4527207-05ea-4dd8-9507-941d1a833de2 · outbound

This paper cites Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.634575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:3a8d8df754146132f03c19b5a6a9d2d851ff1c6982c3f9b7e7524fe591030b9f

Observation 1e85877e-44bc-4d15-a1a8-5749c9a97cc2 · outbound

This paper cites Unbiased gradient low-rank projection.arXiv preprint arXiv:2510.17802.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Unbiased gradient low-rank projection.arXiv preprint arXiv:2510.17802

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.558293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:706c2aa5a251d902d62d5378861c0aa62875f97b8152a0d3805768e16076f543

Observation 8de20694-9b62-4668-9ea0-79d4fddf3e94 · outbound

This paper cites Training Deep Learning Models with Norm-Constrained LMOs.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Training Deep Learning Models with Norm-Constrained LMOs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:22:37.653978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:10f3d3f3ed9edf71fb83a4070faaa8609acef1c9212ab7f4d408302f63b3757e

Observation f1f02a4f-4515-4026-a6d2-bd867ff84826 · outbound

This paper cites icarl: Incre- mental classifier and representation learning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less icarl: Incre- mental classifier and representation learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:47:25.028824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:0ddc159e0c829083e7a7cd601e2f12c37b67dd646ebdd04aa197bd5d68b7a47f

Observation 7a950eec-14aa-4e47-ab6c-b8982a1f4b85 · outbound

This paper cites On the Convergence of Adam and Beyond.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less On the Convergence of Adam and Beyond

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.480366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:c02fad674ebacb3adf3c910eedfd11b0c248791b3d41445781b8da3a49b5e32e

Observation e69b531a-65e8-493c-9425-606bceeda5fb · outbound

This paper cites (How) Learning Rates Regulate Catastrophic Overtraining.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less (How) Learning Rates Regulate Catastrophic Overtraining

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.516618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:dc6bb9b2a79f55469907d6d843358c4802ac6e63a435c51f5a8c664508dea4e4

Observation c241f457-713a-41f8-9d3e-7a495da4d34b · outbound

This paper cites Benchmarking Optimizers for Large Language Model Pretraining.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Benchmarking Optimizers for Large Language Model Pretraining

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.543311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:fed11f78e796762f010336b824848621e525ddb78eb560c4bf622a454c99161a

Observation fd97bb7d-23c4-4df2-b005-67030e28bb84 · outbound

This paper cites GLU Variants Improve Transformer.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less GLU Variants Improve Transformer

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.625847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:dc1adb76b6aab820f1a6861af098cd76df981a416171f2c42c55b62ecb829c3e

Observation ee1c7cdd-d492-4dec-a458-30197356ce85 · outbound

This paper cites Lora vs full fine-tuning: An illusion of equivalence.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Lora vs full fine-tuning: An illusion of equivalence

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.547227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:a96aba812b7b20ed47800bd761da00841c34f9c7a48b1038d5c63830b4edd5eb

Observation b97e7a15-a4d7-4248-9177-2ad6489cbcff · outbound

This paper cites Overtrained Language Models Are Harder to Fine-Tune.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Overtrained Language Models Are Harder to Fine-Tune

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.539517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:6585e6c9eb6585d8819bbfc634915d122bfcdcbc383fb4a3fbe7f88024756297

Observation ab44dce8-7b8b-44cb-902e-4aa2bd1a0bc3 · outbound

This paper cites Less Regret via Online Conditioning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Less Regret via Online Conditioning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.508615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:7d35660c999ecaddb8704da7cd40f02f7df6216c5b6903e77eacd5bb6dde5f78

Observation 1e3575ad-ab31-4b6d-8858-e9664cfd7ddf · outbound

This paper cites ArXiv Preprint: 2511.00674 , Year =.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less ArXiv Preprint: 2511.00674 , Year =

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:08.610339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:9dd1537428740a67d00f29701fbedb78b7da569a0080fa0a9faaf6d98a391c80

Observation d8f36dc5-1a7a-4d95-a955-201b4894fdc2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.512185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:2b15991bb232dd27c448589b40110ed8ef6bd4bcf42f776de2ee5f0032c6132c

Observation 29e05265-12eb-40cc-95fd-de75bd8b3692 · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less SOAP: Improving and Stabilizing Shampoo using Adam

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:02:28.631723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:6922e6d0d39a2e3560fefdc4c0dfdecea146f7188b7dc67ae530eba40c527445

Observation 6a34c718-50c3-41dc-a453-89bed6c500bc · outbound

This paper cites Muon outperforms adam in tail-end associative memory learning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Muon outperforms adam in tail-end associative memory learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.554861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:5a20cbbfcf0f8965edf24d76c5102d51c3f0e672d4505623c0f1b24fd835f436

Observation 8347914b-4c3a-4e8e-aefe-e7d261dbc576 · outbound

This paper cites Magicoder: Empowering Code Generation with OSS-Instruct.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Magicoder: Empowering Code Generation with OSS-Instruct

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.655977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:c26ec4d3a02b49dfd8974eb2c8471632d40ebe58b0deb0359e9cb676f2b97856

Observation fbb0ff62-ebc5-4ca2-8e0d-a4c89742f877 · outbound

This paper cites Fantastic Pretraining Optimizers and Where to Find Them.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Fantastic Pretraining Optimizers and Where to Find Them

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.470520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:115d98c2cad1345498af84be617dbccfdac555132656438e7c60b6820ae0f72a

Observation 804a2766-31e4-49f3-92f5-fb3a064bf02b · outbound

This paper cites Structured Preconditioners in Adaptive Optimization: A Unified Analysis.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Structured Preconditioners in Adaptive Optimization: A Unified Analysis

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.660083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:b7810c07eb73d57e0474e5822a1b4bec0e333753332c43ccf03eece988268716

Observation ccac7eb6-8554-4f1d-ba41-a2562d30ea4f · outbound

This paper cites Controlled llm training on spectral sphere.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Controlled llm training on spectral sphere

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.531367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:7e942794bbb80cbc939394570eddc78dfda700a0d58a8e915255dc2e3bc5b467

Observation c4701b1e-aca3-42f1-8db0-da32010a60c0 · outbound

This paper cites On the width scaling of neural optimizers under matrix operator norms i: Row/column normalization and hyperparameter transfer.arXiv preprint arXiv:2603.09952.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less On the width scaling of neural optimizers under matrix operator norms i: Row/column normalization and hyperparameter transfer.arXiv preprint arXiv:2603.09952

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.475919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:3d53a1fe03c0d2f14c15c1e032fd34fd63cbdd81813fc00c8bc45425dc3e4f3c

Observation 300a9d68-6a8e-4151-aa35-e275248e36ae · outbound

This paper cites Qwen3 Technical Report.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Qwen3 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.629870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:ded5904eaa01932ac401a40310444d6edee25bb833bba6558a5662e7f2328c21

Observation 982ace06-fbed-4f02-b303-9cc6e2c4cdfa · outbound

This paper cites A Spectral Condition for Feature Learning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less A Spectral Condition for Feature Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.596926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:920098d06cc4d6638acce1641a846b9a5dd3912493bf81e5850d563265ca8287

Observation d01d0395-cad3-4f9b-8c96-e98959813614 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:39:00.104164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:fa00880a4eeb08c9747bb21895a91df92916bf9af672fd7b34c6a871b8c0be0b

Observation 06f3d469-ea31-425a-b541-7c6f1cd48023 · outbound

This paper cites StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.562081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:57668715505b4e1bfa36dd4774fc6586536c5cb7ac9bee9acc6e23effce167cb

Observation f9e4fb2e-0950-4211-b854-5b50d1359a24 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:07:53.977132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:c3e3f7694ce4dd1e8304034a88750827df8b854a34ba06af761d601b557c4453

Observation 538e25c1-2441-4600-8506-41ad000ad01b · outbound

This paper cites MARS: Unleashing the Power of Variance Reduction for Training Large Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.551157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:7232fd7a068b42366cc4483f565c97ef5bafd12c4c5940ef60b9857598f80665

Observation 7998c599-731f-4125-89ab-d025e16f3398 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less ADADELTA: An Adaptive Learning Rate Method

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.535605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:071cfd193daa059d316096f9365e32847e3e8a832e543835bd2f52cd92cb1e3d

Observation a21fa863-90da-41e5-9bfa-ccd8459b71d7 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.642431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:2680336f00e5e2e9d31b68919f8d92e50d697b35fde6e3085ad708d718d7ee81

Observation 2b360052-3f06-406c-91f1-0c3317a26c5a · outbound

This paper cites Understanding deep learning requires rethinking generalization.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Understanding deep learning requires rethinking generalization

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:56:40.372790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:f2a4e09a66c294b32f97b0dc4f124d2312dadcf90cdd56e5b24e299652812506

Observation b7dff74f-a846-48eb-8e76-8b915e9acffc · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Why Transformers Need Adam: A Hessian Perspective

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:08.489778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:bbd685332deb6907fc46095b6872d94464d9e12cdf428f9c9b9b24e3aac9a3e5

Observation 9d9eb6cb-2630-4c13-a323-20ebf4406cca · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T19:26:08.526331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:0cd7f613cb58dccaa209ee59000517ce9789b9464d73619e14fb701afdf3d4f3

Observation 34dcb554-30af-42cc-b424-357df4da4d06 · outbound

This paper cites an unresolved cited work.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:47:25.031255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:527872c1c455a2ed8efb0c399cb5277742c9c8697cf64760bca9f9802ff80182

Observation 54569520-ba6a-4626-91be-97757a02b5b6 · outbound

This paper cites Here, LoRA rank 64 shows a greater forgetting compared to rank 256, mainly because rank 256 diverges for lr=5e-4, and a smaller learning rate leads to less forgetting.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Here, LoRA rank 64 shows a greater forgetting compared to rank 256, mainly because rank 256 diverges for lr=5e-4, and a smaller learning rate leads to less forgetting

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:47:25.025446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:00b1a90a11f97a51a3450977155139cfae0e0362c22f5b631b0e92dd6ab1f4e2

Observation 581d67f9-75c1-41ff-b6eb-d653bb622d9a · outbound

This paper cites an unresolved cited work.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:47:25.019542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:d5ebdbf674f28816987a3a5e126e8c48e55b7a95321f5e6df4163601af73855a

Observation 2469ed4d-f60e-4a5b-8514-9601ba82e64b · outbound

This paper cites B.2 More Detailed Activation Plots In Figures 10 and 11, we present the average activation sparsity of detailed modules and specific ac- tivation splits.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less B.2 More Detailed Activation Plots In Figures 10 and 11, we present the average activation sparsity of detailed modules and specific ac- tivation splits

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:47:25.016966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:c51402ab705b10ef2eea0999c6aa27b536020c5f6e90a0561dce054075574ee2

Observation 57ef95fa-30bb-4711-b3c2-176177269826 · outbound

This paper cites 23 •Ifα∈(α 1,∞], based on Lemma 2 and Assumption 4, it holds that ∥∆W∥ α1,β∗ ≤ ∥∆W∥ α,β∗ andE h ∥x∥2 α1 i = Θ E h ∥x∥2 α i.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less 23 •Ifα∈(α 1,∞], based on Lemma 2 and Assumption 4, it holds that ∥∆W∥ α1,β∗ ≤ ∥∆W∥ α,β∗ andE h ∥x∥2 α1 i = Θ E h ∥x∥2 α i

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:47:25.022703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:db08fa21ba1b90a854b6963113f8c9315e669e503b19d538818f8cdd7e69b030

Pith citing papers

Observation 0dcd802d-c6f2-4675-a64f-0c12f2d4c537 · inbound

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss cites this paper.

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

Reference 185

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:56:55.993275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T02:35:39.845487Z digest=sha256:e51c79b1bc9dfa8672ea647abdd6f09f6542473482bdb066b6f9631c3e90389c

Observation 259292d8-88cf-4435-a378-1f4e34a0f944 · inbound

When Does Muon Help Agentic Reinforcement Learning? cites this paper.

When Does Muon Help Agentic Reinforcement Learning? Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:13:39.201865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:13:39.201865Z digest=sha256:6f3839095cb3e5a8b3f90bc98b4175c3a793f52ebc3f0250d4e4e7287d330548