Pith. sign in

Paper Citation Record · LEDGER

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam

As of 18 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2507.06464.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06464 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:10:16.631122Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e944ecc4-1240-4dbb-8a4b-37159c5297a8 · outbound

This paper cites Adam: A method for stochastic optimization,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam: A method for stochastic optimization,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.201901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.336936Z digest=sha256:decf7a14feccc5ae3836db4b8992503abf19421c7903c6096608f457c85feb0f

Observation e6e02f0c-d61c-4d6f-80e1-ea04028cd716 · outbound

This paper cites Attention is all you need,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Attention is all you need,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.191786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.455074Z digest=sha256:0641610100d85771fbd355a7a0d1e25843d704bbf7a07184764b85060149deb3

Observation 130b3ed3-5f82-4fe6-8a9f-e2350eb1faee · outbound

This paper cites Language mod- els are few-shot learners,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language mod- els are few-shot learners,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.458605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.458605Z digest=sha256:4ddef490e61695f211fa51d9abf52717c0b3522a55fd5331dbd19d104d69d9e1

Observation 1a5ec753-fa43-4e58-b12d-202711d646bf · outbound

This paper cites Palm: Scal- ing language modeling with pathways,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Palm: Scal- ing language modeling with pathways,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.462151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.462151Z digest=sha256:c0d0f85cc4877da2ee6922a6ebe7a0dd7f1197431296df45d07b27e54de40f57

Observation 7091e28a-58fc-431a-b9f6-edc9e96e4b74 · outbound

This paper cites The Llama 3 Herd of Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.465335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.465335Z digest=sha256:2ba79671609dd9a2720324c1cd88b51cf03c6232f211c44122b49781f19b8476

Observation a7f18002-e37c-421a-998c-93439602f0dd · outbound

This paper cites DeepSeek-V3 Technical Report.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.468721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.468721Z digest=sha256:7dd13ddf25a49c5a04e99d2d14a2c47736639baf8b33b8b8b95e1346321e2f3e

Observation c344c854-4c34-4e60-8ee5-4d7fa5be9619 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Learning transferable visual models from natural language supervision,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.471944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.471944Z digest=sha256:f124659d307708e15d418ecf474d774997189eb02422d7f04cc04bc9c37eefd6

Observation 99a15b1d-518d-42bd-9632-e3469bc87f4b · outbound

This paper cites Segment Anything.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Segment Anything

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.474845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.474845Z digest=sha256:e4240a8fd0fe4f20028faa10be69c010964b6e1831cb62e92e8cf77f70ad494a

Observation 8c5b1c58-700a-4826-b81c-372942e56358 · outbound

This paper cites A convnet for the 2020s,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A convnet for the 2020s,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.478341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.478341Z digest=sha256:b90ffe0fe8e71a7b254610889eaeb592f93de0b092d3feea2a13c4d557e48527

Observation 7ea6ca0d-f1a1-4b7c-acfd-602228e86852 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoen- coders,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convnext v2: Co-designing and scaling convnets with masked autoen- coders,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.481830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.481830Z digest=sha256:3de9d023b6d3fc8bceccb0daf496593d8e0de849674321470d8cc3d69ac8102a

Observation c1825759-aff8-4476-ad47-2cf31d09d429 · outbound

This paper cites Imagenet classification with deep convolutional neural networks,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Imagenet classification with deep convolutional neural networks,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.485548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.485548Z digest=sha256:591d08439813cdb01a0a8c8bd3f9c034a36d51c1a2a6c3165f29a0f67b2eaebd

Observation c49b5738-5629-45bf-9d39-f9530cbe0292 · outbound

This paper cites Deep residual learning for image recognition,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Deep residual learning for image recognition,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.488397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.488397Z digest=sha256:8df4bd53d9cfd24a897417882f97cd523bafc41d02c11f2ab07ec51f5635d113

Observation 95fdae10-305a-4e29-92b0-50fbe1e9c141 · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Symbolic Discovery of Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.490853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.490853Z digest=sha256:7449e25180903705c798ae2462b486a8fff412b4f235eb0a1d764653a134fb89

Observation 387290f7-6e61-4fbb-9549-3f7e6b456b20 · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.493852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.493852Z digest=sha256:aca94dc260a92907c8af0cf26ba0b5f757862c8fc8ad33e48c6fa667ffd3bc33

Observation 64c549bf-2a0d-45a2-921e-0411c3fa38b6 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adaptive subgradient methods for online learning and stochastic optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.497361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.497361Z digest=sha256:99afc5feb408cb43373250aaa31641570a3032ef4d14f2fd45c8acac5efc635c

Observation 9034aef6-356b-4b25-bf4d-b3a7e75faeb4 · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.137025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.499783Z digest=sha256:f02127a3c296cacc857f5765bd03272ed65d64dc82f1e36fe438ae4630727722

Observation 8ba9cd53-5f3c-4235-a11e-ac7c73ec22b0 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam ADADELTA: An Adaptive Learning Rate Method

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.502866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.502866Z digest=sha256:06be6f7febeb142501c5f2abdc1ba3cdb5403d5f787a18071348fca429cd6bec

Observation 8fad584c-c762-4759-acd9-db97b754c220 · outbound

This paper cites Incorporating nesterov momentum into adam,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Incorporating nesterov momentum into adam,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.124513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.506673Z digest=sha256:cd838bc9967be2b26211277254183031e3b51ebd18aabf10201e303a0390f020

Observation b424d17a-bdd7-4d6d-84f5-938ee379f839 · outbound

This paper cites On the convergence of adam and beyond,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the convergence of adam and beyond,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.509896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.509896Z digest=sha256:47884cb83ebb2d6cd0f5e77f19790931e46bfd7724753d50dadcb400d7b602fb

Observation 6bf20935-3052-4582-a149-eba073699fa4 · outbound

This paper cites Decoupled Weight Decay Regularization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Decoupled Weight Decay Regularization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.513195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.513195Z digest=sha256:fd314a82faa6e0f6b139496ce7947f2b361cc336215541060ac93ea35ce4c1b5

Observation 882ac5b3-d99a-4380-afc0-42040410dad4 · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.109013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.516431Z digest=sha256:8f75c47db4a3cdabbfa95e4b0913058ab6be31c075d97884e6675205124f3fb3

Observation f6cc7300-be44-4848-8327-9827f0c8109d · outbound

This paper cites Adafactor: Adaptive learning rates with sub- linear memory cost,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adafactor: Adaptive learning rates with sub- linear memory cost,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.097992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.518984Z digest=sha256:aef9477841501c71395b8c2c535f438c6c645c23f739364054b4faa2abae0923

Observation d39444e8-0734-4f72-98ad-07c53c256862 · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.088592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.521637Z digest=sha256:796e12d7a7b1250fb495d0c19f68ebd1420d1cf206034cd43a4c79c8c4fc69b1

Observation ebea6fa2-4ce9-47dd-be8e-5ba659eea44f · outbound

This paper cites Signsgd: Compressed optimisation for non-convex problems,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signsgd: Compressed optimisation for non-convex problems,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.078021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.524327Z digest=sha256:e962d6936076fe9fb0d8f26a0f4fc5ed0695adfd5d26caad1030bedef1121382

Observation e926c48a-d74f-4e96-8149-97437b6e9db3 · outbound

This paper cites Momentum ensures convergence of SIGNSGD under weaker assumptions,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Momentum ensures convergence of SIGNSGD under weaker assumptions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.067344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.526948Z digest=sha256:2d85ff4b4fc4e7a1432f51c40ab278d45922705ad023d69555c1a0eca6bd3b01

Observation 8a4e6230-10b8-41d4-9d9c-120d5a4c958b · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.530074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.530074Z digest=sha256:f7d8a77214e1649d3477eb25e9c40276587047ed08cf3ed3d6a7a48fbb20cef6

Observation 82b4b94f-b472-4f8f-addb-70900f8590bd · outbound

This paper cites A direct adaptive method for faster backpropagation learning: The rprop algorithm,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A direct adaptive method for faster backpropagation learning: The rprop algorithm,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.055958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.533409Z digest=sha256:8e5d26d6e4f1802d8c07d503dabbd96e1c85e1a3a81cd17e3d9c085722e5bddf

Observation 8824a63e-45a6-459f-97fd-1ec59887e172 · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.045610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.535835Z digest=sha256:4aaee56781731ad0019b7067cf70119ec499bbc2549f16ca2e1a3faf3b802aff

Observation afb2ed8b-b8f9-47ed-bb3d-4140897acc52 · outbound

This paper cites Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.538230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.538230Z digest=sha256:20a71e2ef735906d234780abbee91398d74f4de96c6a982074a07d7ea24aeb0b

Observation 6036972d-a558-4688-963c-b919c3331931 · outbound

This paper cites Dissecting Adam: The sign, magnitude and variance of stochastic gradients,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Dissecting Adam: The sign, magnitude and variance of stochastic gradients,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.035995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.541329Z digest=sha256:7f13bd6490d37cfdeb1092c5845a161e3332c85c90a67a940338f09d28e285c6

Observation 03e8a653-b280-4c7a-9d79-158e223e8596 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.543726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.543726Z digest=sha256:120daaedc969498c3a3719b362b51f59d5397c080360999b5b1439e57264188a

Observation 0b7b1394-3226-48cb-a8e0-b8b9b09a1f9a · outbound

This paper cites A method of solving a convex programming problem with convergence rate o\bigl(kˆ2\bigr),.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A method of solving a convex programming problem with convergence rate o\bigl(kˆ2\bigr),

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.025018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.546651Z digest=sha256:e8c0e21d8c25256cb6d7798769c8f45b12e65bfe533a8a9e9f0a5af4000f3314

Observation 9d6fc65c-1ed3-4412-add0-8e36cd225a20 · outbound

This paper cites Springer Science & Business Media, 2013, vol.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Springer Science & Business Media, 2013, vol

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.015378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.549317Z digest=sha256:bb3483016b6a87494e1ea671dfba82129d4f5c63e470b0f80085929912651373

Observation 1f9894cf-5e9c-4c60-9bb3-92b9b905aecf · outbound

This paper cites Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.003645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.551946Z digest=sha256:6acbc0d673f0d6973225449d5c43154ee3a7390e3523e2df3455e4aca03e88ce

Observation 17c3485f-f796-4d31-97c8-1e15f18f969d · outbound

This paper cites Win: Weight-decay-integrated nesterov acceleration for adaptive gradient algorithms,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Win: Weight-decay-integrated nesterov acceleration for adaptive gradient algorithms,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.991917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.555136Z digest=sha256:707a278ed5af77d9a0a83909639948af4ab6b5b588dd92cc58cb7e2b02eb74f6

Observation 7e6bf4ef-121e-415e-a5f8-4efda35ca0fd · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam GLM-130B: An Open Bilingual Pre-trained Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.558707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.558707Z digest=sha256:7a4450588dbb1dd6fb428005529e30ab768dc6ffea64d65c7ce36df94846727b

Observation fef1eb67-c197-4ee5-84c5-067f88733fc4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.561798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.561798Z digest=sha256:f7934e96914caeee4c259c54217dca543ec554a6b1b00ff906f686d35993dd74

Observation 4389efc2-d313-4583-8363-821f45cddab5 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Baichuan 2: Open Large-scale Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.564494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.564494Z digest=sha256:332ec6293ad8190808591409329d770340abe69bbf3e3f9b86d03aa9d6d2db91

Observation a65fb9e5-510a-45de-9c7b-e09761f73eb6 · outbound

This paper cites What Language Model to Train if You Have One Million GPU Hours?.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam What Language Model to Train if You Have One Million GPU Hours?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.567629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.567629Z digest=sha256:648d69de9b4ee6573ba418b0a89ce5e55fb80418808e33bcdd14bf190229c6c1

Observation 6c436ba2-b562-42f2-af4e-208e2df607d3 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Theory on Adam Instability in Large-Scale Machine Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.570768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.570768Z digest=sha256:ff991c6c03d3325012158454a1ca9ebd0760778d32eb084d45d93f75efcdef05

Observation 4a9a43cc-6d72-4ab3-a56a-d57d4b7aaa05 · outbound

This paper cites A Mean Field Theory of Batch Normalization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Mean Field Theory of Batch Normalization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:10:16.753519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.573719Z digest=sha256:64e5469679fcc914c6fc7cf7715e4d6b57f826aa063ee6e39e5d0d0664aa8512

Observation 81412f7d-d3c9-443f-b776-78d06309f9d7 · outbound

This paper cites Understanding the Difficulty of Training Transformers.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Understanding the Difficulty of Training Transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.577022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.577022Z digest=sha256:46d51b53053ca97f6c3168766ab609cc8f0874eb50b55c7e81f5b35b9780a5f7

Observation ae162317-639b-4c10-8aa0-87babd01ece8 · outbound

This paper cites On layer normalization in the trans- former architecture,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On layer normalization in the trans- former architecture,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.981559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.580127Z digest=sha256:48700b237fc05fa1cc843c0473422dc027bff6862cb54c301ed6cf3abd15fe6d

Observation 2652fda8-ca7c-4f19-9ce7-dd05b222d562 · outbound

This paper cites The lipschitz constant of self- attention,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The lipschitz constant of self- attention,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.969465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.583966Z digest=sha256:74f278c91ad36252626df6eb4a3ca123b4ca2782709dd5d13ef3e8d872036e1a

Observation 8dfbbdde-7c7e-47bd-8fd7-2af0221cade8 · outbound

This paper cites LipsFormer: Introducing Lipschitz Continuity to Vision Transformers.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LipsFormer: Introducing Lipschitz Continuity to Vision Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.586759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.586759Z digest=sha256:ba63026b18028f33b9f72247ccd552ef72754fe14206851d4677ca147bdeb96d

Observation 8c0381e7-1410-476e-9265-b2dbab2a4363 · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.955532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.589251Z digest=sha256:b1827fbbde9be13fdd37c572a56a9eabedc87757a0c22a506c8968a17c4bc775

Observation 218ef676-cdc5-4fff-9afb-a88f3cad0e79 · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.591929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.591929Z digest=sha256:85e9b917399ab90b745d0f65394f0cdc2e3a2ee37c67f146290998dda9cc052d

Observation b1702c3e-1070-4189-83a0-ceba8119b40e · outbound

This paper cites Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.595651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.595651Z digest=sha256:eacbfe030e0eb266a63f7da7070b30d19cfd40e3ef2e01e2739cfe854609095c

Observation ee2e7d3e-cdc5-4a94-90b3-ffc1a0050369 · outbound

This paper cites Loss Spike in Training Neural Networks.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Loss Spike in Training Neural Networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.598782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.598782Z digest=sha256:de592b6e0f2d8c2bc5a721fbb8913ad77a81da0d7e663afe2b04dfbe7aced2c5

Observation 05f8d309-aa58-4cf9-ab76-89e4d31a7e70 · outbound

This paper cites Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.945621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.601925Z digest=sha256:852171d9b9f024e732b74cff609a564abca7c781a2575ce33994a65bb8144b47

Observation 0665f40e-517a-40fa-9b6b-57ba587a69f1 · outbound

This paper cites On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.605575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.605575Z digest=sha256:bfbe42b0069f1dea3eefcd305c6ef36180784cf244e7902026c6615cc96ab7dc

Observation aa8ccb70-02cc-4385-9382-1b1296c7d6a0 · outbound

This paper cites A Simple Convergence Proof of Adam and Adagrad.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Simple Convergence Proof of Adam and Adagrad

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.608689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.608689Z digest=sha256:3b870f19e69a7acd5a1418ca744507218841d1341f7d31981188d27ce4a1f1df

Observation 2da32d6b-006e-4c34-9b9f-d5b321245e07 · outbound

This paper cites Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.611740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.611740Z digest=sha256:54ca78564c0f2db5bb42a614479ba5b7df561899bb7b62c4f9d5c0233cb21421

Observation c793896d-c0f0-4540-b1f4-ba65969027a9 · outbound

This paper cites Adam can converge without any modification on update rules,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam can converge without any modification on update rules,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.935092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.614982Z digest=sha256:c2375b1628bf90d0391b5a472c63ad32ee5a671bd2fa075e1a27e986468ae97f

Observation edb696f8-338f-4603-bc9b-65aa87994a2a · outbound

This paper cites Convergence of adam under relaxed assumptions,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence of adam under relaxed assumptions,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.925114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.617374Z digest=sha256:db552f34be32a1be64b32bdf0f6344094e16099bb9b4f819cabb04c84f003890

Observation aef5e537-b67a-4191-baf8-44b8ee68e314 · outbound

This paper cites On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.619880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.619880Z digest=sha256:03c7e29ce4a4b516f2d10360173587370c624913a1f087aca51dc545694b2407

Observation d412bbce-b75e-4338-85a1-8d95ec0c4527 · outbound

This paper cites Lower bounds for non-convex stochastic optimization,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lower bounds for non-convex stochastic optimization,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.915141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.622696Z digest=sha256:bd89b54b13dbc4ce245503cf69e6bee3b47343740f81387f79efd46da51f69f9

Observation 8aa29741-fb40-45e8-9b6f-6eb49c4d653a · outbound

This paper cites A stochastic approximation method,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A stochastic approximation method,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.625262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.625262Z digest=sha256:c8d4f39a6e562d2893f5fbf908476d1a3ed371eb4b5a316c39a4bfa647f6e043

Observation 37318d86-763a-4520-a46e-43cf4fa59af7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.628081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.628081Z digest=sha256:078e2dd178461348795120a7a9bff76dc9e2e158fd9f7782b1704cc7ed4c41e4

Observation ef273616-c7fa-4b26-957c-cf3618c69d17 · outbound

This paper cites Language models are unsupervised multitask learners,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language models are unsupervised multitask learners,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.898796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:10:16.631122Z digest=sha256:e8c6be34f762b765710fe12e8ea166225b0c19bfd4733fa9dfc452b387a9628a

Pith citing papers

No inbound Pith citation observations are available.