Pith. sign in

Paper Citation Record · LEDGER

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam

As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2507.06464.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06464 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:10:16.631122Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e944ecc4-1240-4dbb-8a4b-37159c5297a8 · outbound

This paper cites Adam: A method for stochastic optimization,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam: A method for stochastic optimization,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.201901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.336936Z digest=sha256:f9c07b30176505435d7f8ec8caadb42c44b80b9e2b1b95092c0bf48593ff45d8

Observation e6e02f0c-d61c-4d6f-80e1-ea04028cd716 · outbound

This paper cites Attention is all you need,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Attention is all you need,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.191786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.455074Z digest=sha256:79a91efda540b3aa4a2f8c1883aac71305f01bf20ed83487b53acfe818b9e0a3

Observation 130b3ed3-5f82-4fe6-8a9f-e2350eb1faee · outbound

This paper cites Language mod- els are few-shot learners,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language mod- els are few-shot learners,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.458605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.458605Z digest=sha256:6c51a117d084699bab00a88ceb85c803b13011af8a0a92de6764608da6b92c74

Observation 1a5ec753-fa43-4e58-b12d-202711d646bf · outbound

This paper cites Palm: Scal- ing language modeling with pathways,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Palm: Scal- ing language modeling with pathways,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.462151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.462151Z digest=sha256:a1336839c169cd93007bb32bc53b274869140ec0684e0c078c4649a839338c3f

Observation 7091e28a-58fc-431a-b9f6-edc9e96e4b74 · outbound

This paper cites The Llama 3 Herd of Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.465335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.465335Z digest=sha256:fce9d77f59bd0ef6831641655d8de4ab791f235bc0bb634d57a0bdc234fe9b3b

Observation a7f18002-e37c-421a-998c-93439602f0dd · outbound

This paper cites DeepSeek-V3 Technical Report.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.468721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.468721Z digest=sha256:52a95b7bed5794ce51e9b22e7c507380a87e440258b58254f390782b3f830a09

Observation c344c854-4c34-4e60-8ee5-4d7fa5be9619 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Learning transferable visual models from natural language supervision,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.471944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.471944Z digest=sha256:05de7fcbca00a951e06078baac431cb3ae7fbcb4f6a8282e3bd76cdb162b6690

Observation 99a15b1d-518d-42bd-9632-e3469bc87f4b · outbound

This paper cites Segment Anything.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Segment Anything

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.474845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.474845Z digest=sha256:e63c7edc3af38c6847b43dae880fd5329ab16cd9e081dbd252d9e0d4f7388201

Observation 8c5b1c58-700a-4826-b81c-372942e56358 · outbound

This paper cites A convnet for the 2020s,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A convnet for the 2020s,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.478341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.478341Z digest=sha256:6fa11cdbe342977c7c24ed974d33ed8c0c52be80f55393738f623d7c12279947

Observation 7ea6ca0d-f1a1-4b7c-acfd-602228e86852 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoen- coders,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convnext v2: Co-designing and scaling convnets with masked autoen- coders,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.481830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.481830Z digest=sha256:56d6f968de7254ebb2e382e9dd0e799ba6d35ecb78852bee0174ecdf29b5ca15

Observation c1825759-aff8-4476-ad47-2cf31d09d429 · outbound

This paper cites Imagenet classification with deep convolutional neural networks,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Imagenet classification with deep convolutional neural networks,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.485548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.485548Z digest=sha256:dd0ebf0ab0180207c736eccfb8695406fda5701aa66782ecdb50ded67703d51b

Observation c49b5738-5629-45bf-9d39-f9530cbe0292 · outbound

This paper cites Deep residual learning for image recognition,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Deep residual learning for image recognition,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.488397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.488397Z digest=sha256:d4d7e26261e28e1f41b8e831125af9a2c1988532bf128d30cfe147ef9a424e18

Observation 95fdae10-305a-4e29-92b0-50fbe1e9c141 · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Symbolic Discovery of Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.490853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.490853Z digest=sha256:420eea6d02afb569592671602d7bed262074f126c090bfb184129bfad96208ae

Observation 387290f7-6e61-4fbb-9549-3f7e6b456b20 · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.493852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.493852Z digest=sha256:35f01b51831a1e6c7a2442d8cd710a96f5203e20a78d24cc43c35249c40d7919

Observation 64c549bf-2a0d-45a2-921e-0411c3fa38b6 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adaptive subgradient methods for online learning and stochastic optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.497361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.497361Z digest=sha256:fa22a4185f2ed7bbfe11b1a957d4e10ed2a08a489e7f06dd1bebeca0d1303aea

Observation 9034aef6-356b-4b25-bf4d-b3a7e75faeb4 · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.137025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.499783Z digest=sha256:4ed440daf5fbf9468a8fa558268e143bb521c7dde4667b3f03193e10af902f2d

Observation 8ba9cd53-5f3c-4235-a11e-ac7c73ec22b0 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam ADADELTA: An Adaptive Learning Rate Method

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.502866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.502866Z digest=sha256:a5fa94a3e4f816f539a2319808fe1c0d3536486c2c157463ed1d9574b0b471ac

Observation 8fad584c-c762-4759-acd9-db97b754c220 · outbound

This paper cites Incorporating nesterov momentum into adam,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Incorporating nesterov momentum into adam,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.124513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.506673Z digest=sha256:eceb7a209f3847ffe527e65bac43188f99e915e6a51a84d78218ac955bdb2f65

Observation b424d17a-bdd7-4d6d-84f5-938ee379f839 · outbound

This paper cites On the convergence of adam and beyond,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the convergence of adam and beyond,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.509896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.509896Z digest=sha256:ce63859a0f1ab7fb6ad11c6261d85cf59a030861bae2790de265ec00ec488f25

Observation 6bf20935-3052-4582-a149-eba073699fa4 · outbound

This paper cites Decoupled Weight Decay Regularization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Decoupled Weight Decay Regularization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.513195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.513195Z digest=sha256:52f314a6823ff106a66cca9dc315206dbe392dc458692046892dd47be794e16e

Observation 882ac5b3-d99a-4380-afc0-42040410dad4 · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.109013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.516431Z digest=sha256:94702e0910e67055de4ac30b2bb30a2684c947b4f224fd50a8a88624b181d5cf

Observation f6cc7300-be44-4848-8327-9827f0c8109d · outbound

This paper cites Adafactor: Adaptive learning rates with sub- linear memory cost,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adafactor: Adaptive learning rates with sub- linear memory cost,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.097992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.518984Z digest=sha256:68e4b0fc088ca1e4ed259b7098f28a457a6c383e1c8b93b0e48820378c6f8e96

Observation d39444e8-0734-4f72-98ad-07c53c256862 · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.088592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.521637Z digest=sha256:1ddfeb3e827acdcd70163daefbb79891666dbb3623208c9320c6fafc507c1bde

Observation ebea6fa2-4ce9-47dd-be8e-5ba659eea44f · outbound

This paper cites Signsgd: Compressed optimisation for non-convex problems,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signsgd: Compressed optimisation for non-convex problems,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.078021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.524327Z digest=sha256:05a3c388ad81040801d0b7127d975bd17a9dc1bb1b2b2e93896c42daf799ee5c

Observation e926c48a-d74f-4e96-8149-97437b6e9db3 · outbound

This paper cites Momentum ensures convergence of SIGNSGD under weaker assumptions,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Momentum ensures convergence of SIGNSGD under weaker assumptions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.067344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.526948Z digest=sha256:617806a10ab148f136e1dc44dffa0f5b985d6378d6a9d7fc77b3dc7a36daf67c

Observation 8a4e6230-10b8-41d4-9d9c-120d5a4c958b · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.530074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.530074Z digest=sha256:4f8e7f3575b86006dad1ad68062b63b4103e06aaf469fc142de1faf8e55db53b

Observation 82b4b94f-b472-4f8f-addb-70900f8590bd · outbound

This paper cites A direct adaptive method for faster backpropagation learning: The rprop algorithm,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A direct adaptive method for faster backpropagation learning: The rprop algorithm,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.055958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.533409Z digest=sha256:5832cf54644510d6c7b43b340e4cab53c107d6f856169485fc6fc57675f34131

Observation 8824a63e-45a6-459f-97fd-1ec59887e172 · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.045610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.535835Z digest=sha256:bf6c5533d8984f35a3ad86cb9024bec964f566e0cfe3156c6c16372a15a2c848

Observation afb2ed8b-b8f9-47ed-bb3d-4140897acc52 · outbound

This paper cites Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.538230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.538230Z digest=sha256:3205dbe9a87cebf2c7510a0650d283821fba359cc12a0b4a783d89b2ccfe3910

Observation 6036972d-a558-4688-963c-b919c3331931 · outbound

This paper cites Dissecting Adam: The sign, magnitude and variance of stochastic gradients,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Dissecting Adam: The sign, magnitude and variance of stochastic gradients,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.035995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.541329Z digest=sha256:52505f2989ffc240f94f148637651897729b459e1c5ccfcb3f2db9724157ac69

Observation 03e8a653-b280-4c7a-9d79-158e223e8596 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.543726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.543726Z digest=sha256:472c012c4e4c8b70464c43fc1d342d72bc70a94849078a7025c8b6e92d8bfd5c

Observation 0b7b1394-3226-48cb-a8e0-b8b9b09a1f9a · outbound

This paper cites A method of solving a convex programming problem with convergence rate o\bigl(kˆ2\bigr),.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A method of solving a convex programming problem with convergence rate o\bigl(kˆ2\bigr),

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.025018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.546651Z digest=sha256:316029572b473a2774866230a5e374f68e8a04bf3a260c9ffc660d292bac606d

Observation 9d6fc65c-1ed3-4412-add0-8e36cd225a20 · outbound

This paper cites Springer Science & Business Media, 2013, vol.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Springer Science & Business Media, 2013, vol

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.015378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.549317Z digest=sha256:03b2673551b832d392e2a904f58a037f1dc0eba83704bbebb3c901adf2beed54

Observation 1f9894cf-5e9c-4c60-9bb3-92b9b905aecf · outbound

This paper cites Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.003645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.551946Z digest=sha256:d4654f3a5654a3ea1eb4ce446f1db140f1d916ddb2c19ddb3a517809c3b02546

Observation 17c3485f-f796-4d31-97c8-1e15f18f969d · outbound

This paper cites Win: Weight-decay-integrated nesterov acceleration for adaptive gradient algorithms,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Win: Weight-decay-integrated nesterov acceleration for adaptive gradient algorithms,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.991917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.555136Z digest=sha256:6dbdf521bb7f7f21e6e2a67a0b25b57af78ad47a23f1fec700e8d375e3113923

Observation 7e6bf4ef-121e-415e-a5f8-4efda35ca0fd · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam GLM-130B: An Open Bilingual Pre-trained Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.558707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.558707Z digest=sha256:e442f0395c2f02c0503ebc473ed616773dca6a27c564092598929d1eee229c1e

Observation fef1eb67-c197-4ee5-84c5-067f88733fc4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.561798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.561798Z digest=sha256:49f261ce09e8180892ee7ad1ca858e2f9a15bf12ad192da96cfca075b4644675

Observation 4389efc2-d313-4583-8363-821f45cddab5 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Baichuan 2: Open Large-scale Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.564494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.564494Z digest=sha256:008184d65b45c1e8a74e18182fe4b159df1eb40d76423d64c4bedf38eeda635b

Observation a65fb9e5-510a-45de-9c7b-e09761f73eb6 · outbound

This paper cites What Language Model to Train if You Have One Million GPU Hours?.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam What Language Model to Train if You Have One Million GPU Hours?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.567629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.567629Z digest=sha256:a5fb4169f3df108e84a86eaf3d9cf83336f83da8f46139909021850a1c541e8d

Observation 6c436ba2-b562-42f2-af4e-208e2df607d3 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Theory on Adam Instability in Large-Scale Machine Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.570768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.570768Z digest=sha256:3e37fc4a6df1358edf5a46b461cdbf1f6df39820e8bd1f574add4a062829f0e8

Observation 4a9a43cc-6d72-4ab3-a56a-d57d4b7aaa05 · outbound

This paper cites A Mean Field Theory of Batch Normalization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Mean Field Theory of Batch Normalization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:10:16.753519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.573719Z digest=sha256:6007ef870205b84305363ee416dd370de3679ce4ef8722107b4d96878600566c

Observation 81412f7d-d3c9-443f-b776-78d06309f9d7 · outbound

This paper cites Understanding the Difficulty of Training Transformers.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Understanding the Difficulty of Training Transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.577022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.577022Z digest=sha256:625a29b391f091448c7e9155818975d07752e2269eba0e9a65ce8acc6c89fcf5

Observation ae162317-639b-4c10-8aa0-87babd01ece8 · outbound

This paper cites On layer normalization in the trans- former architecture,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On layer normalization in the trans- former architecture,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.981559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.580127Z digest=sha256:bf274a8b8919e6a8afbc7205a2d3c17f5e730b6c3c37602f5575240b1d81849b

Observation 2652fda8-ca7c-4f19-9ce7-dd05b222d562 · outbound

This paper cites The lipschitz constant of self- attention,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The lipschitz constant of self- attention,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.969465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.583966Z digest=sha256:da87a959fca23c2e9a274b24c793cf0e1ae0a18909ed608a28750bacba3c59a4

Observation 8dfbbdde-7c7e-47bd-8fd7-2af0221cade8 · outbound

This paper cites LipsFormer: Introducing Lipschitz Continuity to Vision Transformers.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LipsFormer: Introducing Lipschitz Continuity to Vision Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.586759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.586759Z digest=sha256:c889e451d65f0e26c918ce5852a9ebd6085612eec326921182a4413b3414376a

Observation 8c0381e7-1410-476e-9265-b2dbab2a4363 · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.955532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.589251Z digest=sha256:d54c44c242e34514747ce1a8b488b67b92d7822c711df595e4405dc9977451cf

Observation 218ef676-cdc5-4fff-9afb-a88f3cad0e79 · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.591929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.591929Z digest=sha256:8a102cb6791089c1f2bbd951b53396776bee1f9d2389450555406ffeb451c1f6

Observation b1702c3e-1070-4189-83a0-ceba8119b40e · outbound

This paper cites Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.595651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.595651Z digest=sha256:a35f49bcb056bea1204bf34cf1e69b761d1c84e276d88bbadf678b994e200c9c

Observation ee2e7d3e-cdc5-4a94-90b3-ffc1a0050369 · outbound

This paper cites Loss Spike in Training Neural Networks.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Loss Spike in Training Neural Networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.598782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.598782Z digest=sha256:79511b4f8668eb5260e31dd87cdf106b4a2874fadf2b7cc8844087c6c8f37e9c

Observation 05f8d309-aa58-4cf9-ab76-89e4d31a7e70 · outbound

This paper cites Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.945621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.601925Z digest=sha256:81264271f68f22c91fd1aa781a0a0c4968e1a3e9ef8f0b321d2213500e6a8cc4

Observation 0665f40e-517a-40fa-9b6b-57ba587a69f1 · outbound

This paper cites On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.605575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.605575Z digest=sha256:71bd2ef870fd06e84583774837e8a3cc70c07ca6d5e90e9d035ba7c183d25f54

Observation aa8ccb70-02cc-4385-9382-1b1296c7d6a0 · outbound

This paper cites A Simple Convergence Proof of Adam and Adagrad.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Simple Convergence Proof of Adam and Adagrad

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.608689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.608689Z digest=sha256:a5baf050f4168b5a97acd51b2cea0518350e397a08bbc8adad37361e4d170f75

Observation 2da32d6b-006e-4c34-9b9f-d5b321245e07 · outbound

This paper cites Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.611740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.611740Z digest=sha256:02f5ffadeab13ddcb6094d426df35e1fd58d68000b2b784a83e7195b471f5836

Observation c793896d-c0f0-4540-b1f4-ba65969027a9 · outbound

This paper cites Adam can converge without any modification on update rules,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam can converge without any modification on update rules,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.935092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.614982Z digest=sha256:035162c8520c6f1eb85921c7f61f57ec34e5e65e28ac0a65941121b08d7b3058

Observation edb696f8-338f-4603-bc9b-65aa87994a2a · outbound

This paper cites Convergence of adam under relaxed assumptions,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence of adam under relaxed assumptions,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.925114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.617374Z digest=sha256:5f22e0054dcbe7da7542fd983b277316e50aeb7a5260c8ca65bb5c8e7493e3bc

Observation aef5e537-b67a-4191-baf8-44b8ee68e314 · outbound

This paper cites On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.619880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.619880Z digest=sha256:2c686b8b077d76909dfdd6721d9136d038bf91524818e554093392bebda25abe

Observation d412bbce-b75e-4338-85a1-8d95ec0c4527 · outbound

This paper cites Lower bounds for non-convex stochastic optimization,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lower bounds for non-convex stochastic optimization,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.915141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.622696Z digest=sha256:88879356f038c98822fcea443f6d8b1a4b403baead28cfcd7a21a2ed553ac2d7

Observation 8aa29741-fb40-45e8-9b6f-6eb49c4d653a · outbound

This paper cites A stochastic approximation method,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A stochastic approximation method,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.625262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.625262Z digest=sha256:d13952d20ac9492b0827a24e7fa6b192f055de12411f3b293b4ee9ce5f864b69

Observation 37318d86-763a-4520-a46e-43cf4fa59af7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.628081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.628081Z digest=sha256:ad18a91c12dd67513b6e11b4fae6a851e778ed1de34b9bf9da8cd665b10cf19a

Observation ef273616-c7fa-4b26-957c-cf3618c69d17 · outbound

This paper cites Language models are unsupervised multitask learners,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language models are unsupervised multitask learners,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.898796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.631122Z digest=sha256:aac32390e8e79cb79b76db190f61909ece0faf6cc1e147484e3d09258f01bcaf

Pith citing papers

No inbound Pith citation observations are available.