Pith. sign in

Paper Citation Record · LEDGER

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam

As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2507.06464.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06464 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:10:16.631122Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e944ecc4-1240-4dbb-8a4b-37159c5297a8 · outbound

This paper cites Adam: A method for stochastic optimization,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam: A method for stochastic optimization,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.201901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.336936Z digest=sha256:02acbf8ec83c3f1d9bb84602bff8813cb1f34863034962c6295cd46aa582d42d

Observation e6e02f0c-d61c-4d6f-80e1-ea04028cd716 · outbound

This paper cites Attention is all you need,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Attention is all you need,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.191786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.455074Z digest=sha256:742c096d5640f9016cafad812b845545e8344c6ee3fb54a8d8248859cfeb9726

Observation 130b3ed3-5f82-4fe6-8a9f-e2350eb1faee · outbound

This paper cites Language mod- els are few-shot learners,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language mod- els are few-shot learners,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.458605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.458605Z digest=sha256:959cee2d612d510e73297d60cfeeb83fdb7f45d7d38ca27b0da331ac4df775bf

Observation 1a5ec753-fa43-4e58-b12d-202711d646bf · outbound

This paper cites Palm: Scal- ing language modeling with pathways,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Palm: Scal- ing language modeling with pathways,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.462151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.462151Z digest=sha256:d98f406fa256fa161690ae5d87b8d2a86f85b5dbb2d797e98e366d84d419d02e

Observation 7091e28a-58fc-431a-b9f6-edc9e96e4b74 · outbound

This paper cites The Llama 3 Herd of Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.465335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.465335Z digest=sha256:8a6146778f14427902182157e331080e3a1a0ff37add30185574895a13592754

Observation a7f18002-e37c-421a-998c-93439602f0dd · outbound

This paper cites DeepSeek-V3 Technical Report.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.468721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.468721Z digest=sha256:37acbfbda81f7a126c89913362c72a833294ef346ca45f53ce1567d5a859c3d0

Observation c344c854-4c34-4e60-8ee5-4d7fa5be9619 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Learning transferable visual models from natural language supervision,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.471944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.471944Z digest=sha256:e8f6714706d80061b4f95780a9714b03512b5bbdd5e4ed2781989b5f11b02a4e

Observation 99a15b1d-518d-42bd-9632-e3469bc87f4b · outbound

This paper cites Segment Anything.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Segment Anything

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.474845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.474845Z digest=sha256:d345efaff2de16151419d25637fcae1ec075fa2ed412ba6c93692b8df3cf5cd3

Observation 8c5b1c58-700a-4826-b81c-372942e56358 · outbound

This paper cites A convnet for the 2020s,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A convnet for the 2020s,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.478341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.478341Z digest=sha256:ff969d0862dc09cf20dd369c500c7688235f3bd42fbf96a3168c4e0686464519

Observation 7ea6ca0d-f1a1-4b7c-acfd-602228e86852 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoen- coders,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convnext v2: Co-designing and scaling convnets with masked autoen- coders,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.481830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.481830Z digest=sha256:9ee84db85a77c138afd2afc2c4ec52a5bc185f04b6d51409534cb82795cd4887

Observation c1825759-aff8-4476-ad47-2cf31d09d429 · outbound

This paper cites Imagenet classification with deep convolutional neural networks,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Imagenet classification with deep convolutional neural networks,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.485548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.485548Z digest=sha256:e7ae6f442e6e8389abfe18795089ff1b09c70d47064e36d03831097e8f35261e

Observation c49b5738-5629-45bf-9d39-f9530cbe0292 · outbound

This paper cites Deep residual learning for image recognition,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Deep residual learning for image recognition,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.488397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.488397Z digest=sha256:24cce8f67471b5983390ab43b810cb8a69aa374405528a92a06770909d2933c3

Observation 95fdae10-305a-4e29-92b0-50fbe1e9c141 · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Symbolic Discovery of Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.490853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.490853Z digest=sha256:5267e331149ed822077d3048fb4e6e1517babac28fe84dd9158eab7768a55e6c

Observation 387290f7-6e61-4fbb-9549-3f7e6b456b20 · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.493852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.493852Z digest=sha256:7b485ab852eff7b45abf50f525e7590852ece1c6279a57448d519a450a9c1240

Observation 64c549bf-2a0d-45a2-921e-0411c3fa38b6 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adaptive subgradient methods for online learning and stochastic optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.497361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.497361Z digest=sha256:8dc7c0887762507f4e5cbd9c2bfd900066f8ecc6c42192c6170a95c812f117b0

Observation 9034aef6-356b-4b25-bf4d-b3a7e75faeb4 · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.137025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.499783Z digest=sha256:c93d580b1dfb8e80c5e6f8241aa68dfeaf393f132bd01ca19327ee68da5c8da6

Observation 8ba9cd53-5f3c-4235-a11e-ac7c73ec22b0 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam ADADELTA: An Adaptive Learning Rate Method

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.502866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.502866Z digest=sha256:8fe1a0ed0d8b27a98d1604f33684a1f9547ea86ef54f6b3b1d9a9c3278b595fb

Observation 8fad584c-c762-4759-acd9-db97b754c220 · outbound

This paper cites Incorporating nesterov momentum into adam,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Incorporating nesterov momentum into adam,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.124513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.506673Z digest=sha256:c8dc9f661e2a64c97aa8277e9bcd589dc60065b100fe3bc99f2dcc7727ed1c0c

Observation b424d17a-bdd7-4d6d-84f5-938ee379f839 · outbound

This paper cites On the convergence of adam and beyond,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the convergence of adam and beyond,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.509896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.509896Z digest=sha256:7270b3af7a8a15ea837030e3ca7d0a25fc488aa824ae11e9e336843c7bc414d2

Observation 6bf20935-3052-4582-a149-eba073699fa4 · outbound

This paper cites Decoupled Weight Decay Regularization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Decoupled Weight Decay Regularization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.513195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.513195Z digest=sha256:cca8bd5b7894fbb4815ffad2f4933ecc51082bc04423acd90754413f2ce0ba43

Observation 882ac5b3-d99a-4380-afc0-42040410dad4 · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.109013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.516431Z digest=sha256:2f9cadb0afd837ea7e92a74d72980b945398006be5b51155151a8f61ad08a480

Observation f6cc7300-be44-4848-8327-9827f0c8109d · outbound

This paper cites Adafactor: Adaptive learning rates with sub- linear memory cost,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adafactor: Adaptive learning rates with sub- linear memory cost,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.097992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.518984Z digest=sha256:78e8812e3906f9db7f6924aec6c3f06851cc584fd72f3783783963d235836ff1

Observation d39444e8-0734-4f72-98ad-07c53c256862 · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.088592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.521637Z digest=sha256:a25dacc3ab1e35f5505064e8f14e66dfc91664eb5cbed7fa9635492905f19ff0

Observation ebea6fa2-4ce9-47dd-be8e-5ba659eea44f · outbound

This paper cites Signsgd: Compressed optimisation for non-convex problems,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signsgd: Compressed optimisation for non-convex problems,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.078021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.524327Z digest=sha256:44f4d2f47fefbb4cd0b45531a98386dec1cc52b2f18d12b107ed8dd13b55e487

Observation e926c48a-d74f-4e96-8149-97437b6e9db3 · outbound

This paper cites Momentum ensures convergence of SIGNSGD under weaker assumptions,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Momentum ensures convergence of SIGNSGD under weaker assumptions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.067344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.526948Z digest=sha256:7144f735ed39cab6e886349b0022fb72f727d2dd834bc5758dfd148421744f6e

Observation 8a4e6230-10b8-41d4-9d9c-120d5a4c958b · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.530074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.530074Z digest=sha256:efaf5edda02e0df799784d1c8c2159ac8f6a0ea4462ab48e228a9320a494de95

Observation 82b4b94f-b472-4f8f-addb-70900f8590bd · outbound

This paper cites A direct adaptive method for faster backpropagation learning: The rprop algorithm,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A direct adaptive method for faster backpropagation learning: The rprop algorithm,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.055958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.533409Z digest=sha256:731bb546b5e040a3d4d6ff3ba18bf97f0232182ce07923a87df90b6835fb56d3

Observation 8824a63e-45a6-459f-97fd-1ec59887e172 · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.045610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.535835Z digest=sha256:50aac22c3336cb41c1f7ab6c0c60b66396dee4f6d3aa8a1869b12a7de34bd367

Observation afb2ed8b-b8f9-47ed-bb3d-4140897acc52 · outbound

This paper cites Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.538230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.538230Z digest=sha256:2d10c63bfb630413e2f9ad23a7ab02127ba23252052eff29926c641d10ea1b71

Observation 6036972d-a558-4688-963c-b919c3331931 · outbound

This paper cites Dissecting Adam: The sign, magnitude and variance of stochastic gradients,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Dissecting Adam: The sign, magnitude and variance of stochastic gradients,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.035995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.541329Z digest=sha256:c834c1c1870fa5b6065ed31fb2dd915b7eb088b2cefe2c0e5f6c13959958be4e

Observation 03e8a653-b280-4c7a-9d79-158e223e8596 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.543726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.543726Z digest=sha256:9bbf8d7ee2948563897b5461915baaafb89605e31a33a0ce64b434f3c22ebeac

Observation 0b7b1394-3226-48cb-a8e0-b8b9b09a1f9a · outbound

This paper cites A method of solving a convex programming problem with convergence rate o\bigl(kˆ2\bigr),.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A method of solving a convex programming problem with convergence rate o\bigl(kˆ2\bigr),

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.025018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.546651Z digest=sha256:a1a0d54b5bf9f33c35e63b99f62f2036680c7078f5d1c18f50d1892cfbcd5994

Observation 9d6fc65c-1ed3-4412-add0-8e36cd225a20 · outbound

This paper cites Springer Science & Business Media, 2013, vol.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Springer Science & Business Media, 2013, vol

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.015378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.549317Z digest=sha256:fb91bab8a985e8a29cfab9ddb67601a06c6eb700f29e83a00b00f78d3a233a0e

Observation 1f9894cf-5e9c-4c60-9bb3-92b9b905aecf · outbound

This paper cites Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.003645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.551946Z digest=sha256:e4c812d5bfe24c5049d2b80999b1b96f9c9748844e8a7a5300f304112df06300

Observation 17c3485f-f796-4d31-97c8-1e15f18f969d · outbound

This paper cites Win: Weight-decay-integrated nesterov acceleration for adaptive gradient algorithms,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Win: Weight-decay-integrated nesterov acceleration for adaptive gradient algorithms,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.991917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.555136Z digest=sha256:6e9d1e6a22f3575616e9bc96615ad216074b27052c3e2024fd8b6ff5bc7955cb

Observation 7e6bf4ef-121e-415e-a5f8-4efda35ca0fd · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam GLM-130B: An Open Bilingual Pre-trained Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.558707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.558707Z digest=sha256:8ae056f32ee112e8c412c17e1284a4ff717e400abb98e702f8a0e69250f55c17

Observation fef1eb67-c197-4ee5-84c5-067f88733fc4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.561798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.561798Z digest=sha256:3db6cb56ed8525cf6144321e13bb99e0f21bd076f61ff55924f919989878408b

Observation 4389efc2-d313-4583-8363-821f45cddab5 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Baichuan 2: Open Large-scale Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.564494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.564494Z digest=sha256:26fa9400da5c7b7146b672cf2454fb435612166b7fa1588b62643e3253366e47

Observation a65fb9e5-510a-45de-9c7b-e09761f73eb6 · outbound

This paper cites What Language Model to Train if You Have One Million GPU Hours?.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam What Language Model to Train if You Have One Million GPU Hours?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.567629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.567629Z digest=sha256:9cf89d2dabcbf3500271657203e67728b3b40c4ce590eb97cfafbc029f2aaa18

Observation 6c436ba2-b562-42f2-af4e-208e2df607d3 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Theory on Adam Instability in Large-Scale Machine Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.570768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.570768Z digest=sha256:4d0eb4ab7e82b433685f90e69b8e662005a047ca0aa299b3aba67431dfb25e5b

Observation 4a9a43cc-6d72-4ab3-a56a-d57d4b7aaa05 · outbound

This paper cites A Mean Field Theory of Batch Normalization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Mean Field Theory of Batch Normalization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:10:16.753519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.573719Z digest=sha256:b35d9a2d6fe292488586f4f830aa431275f89d2f522333e9c59c59193c862936

Observation 81412f7d-d3c9-443f-b776-78d06309f9d7 · outbound

This paper cites Understanding the Difficulty of Training Transformers.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Understanding the Difficulty of Training Transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.577022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.577022Z digest=sha256:d5950f4232388a9d14346392b7705416d803e9cf3ee8f7d73b23feaec145673f

Observation ae162317-639b-4c10-8aa0-87babd01ece8 · outbound

This paper cites On layer normalization in the trans- former architecture,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On layer normalization in the trans- former architecture,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.981559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.580127Z digest=sha256:248a1f8daf4c17169407d0d2f4f507328894cdc78bc96627f98cf7135605d0a5

Observation 2652fda8-ca7c-4f19-9ce7-dd05b222d562 · outbound

This paper cites The lipschitz constant of self- attention,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The lipschitz constant of self- attention,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.969465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.583966Z digest=sha256:d8da611b75f813c889ea3d8ee294889adf0fca2da924fbe00b347b7623ccab5b

Observation 8dfbbdde-7c7e-47bd-8fd7-2af0221cade8 · outbound

This paper cites LipsFormer: Introducing Lipschitz Continuity to Vision Transformers.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LipsFormer: Introducing Lipschitz Continuity to Vision Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.586759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.586759Z digest=sha256:7d341214ae0f97abc88d096d8c2f75353afea0a7608b2081dfff45f50261dfc7

Observation 8c0381e7-1410-476e-9265-b2dbab2a4363 · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.955532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.589251Z digest=sha256:0d91e9c05f7425e0226596f15212e0b8aeed6c1f250fc5e2b814b071b4df4faf

Observation 218ef676-cdc5-4fff-9afb-a88f3cad0e79 · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.591929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.591929Z digest=sha256:2397d10195f91ef94f25e0970f7fb7ddc66c5e92a4c6fb123a3f8046e6fa3d95

Observation b1702c3e-1070-4189-83a0-ceba8119b40e · outbound

This paper cites Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.595651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.595651Z digest=sha256:33d0c17a6363ac2bf7885ef72ae9c859af957e5907448ca96f9a6e0e44dea0d0

Observation ee2e7d3e-cdc5-4a94-90b3-ffc1a0050369 · outbound

This paper cites Loss Spike in Training Neural Networks.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Loss Spike in Training Neural Networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.598782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.598782Z digest=sha256:ad39d3cd420f98064c5909e9f4704b2aaa561d9bd09765687f1bdbb73ffa2ebe

Observation 05f8d309-aa58-4cf9-ab76-89e4d31a7e70 · outbound

This paper cites Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.945621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.601925Z digest=sha256:7425ca9cfd6adac228bfb8f79ff65c19d596a30958ba953d2093f8af9370b963

Observation 0665f40e-517a-40fa-9b6b-57ba587a69f1 · outbound

This paper cites On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.605575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.605575Z digest=sha256:d324ae873f6cd0624dcf14a75b41ce0acc0d426fd0a54b8bb98d1e4f65e4b6c5

Observation aa8ccb70-02cc-4385-9382-1b1296c7d6a0 · outbound

This paper cites A Simple Convergence Proof of Adam and Adagrad.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Simple Convergence Proof of Adam and Adagrad

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.608689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.608689Z digest=sha256:c9bc4031696120fbb13ca4004ccc27d431d41ac2c95ca6beb98cf27155d2f95a

Observation 2da32d6b-006e-4c34-9b9f-d5b321245e07 · outbound

This paper cites Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.611740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.611740Z digest=sha256:4ace9288785b97cd8e29778fa9562179f63a80dc9c5f0ed7a9fec79046bd53f4

Observation c793896d-c0f0-4540-b1f4-ba65969027a9 · outbound

This paper cites Adam can converge without any modification on update rules,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam can converge without any modification on update rules,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.935092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.614982Z digest=sha256:f106d5deb653caf1971013d609a655d9261c211da22182b2e6a65a4682baa076

Observation edb696f8-338f-4603-bc9b-65aa87994a2a · outbound

This paper cites Convergence of adam under relaxed assumptions,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence of adam under relaxed assumptions,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.925114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.617374Z digest=sha256:3275c771f4c3edbd94117505ad7203a109db94f8ec98fb256b29eec68c2ac20a

Observation aef5e537-b67a-4191-baf8-44b8ee68e314 · outbound

This paper cites On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.619880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.619880Z digest=sha256:0a27d277f91866bac05e13f41e38794b9e2acfa0c788673bde22670c9e58be33

Observation d412bbce-b75e-4338-85a1-8d95ec0c4527 · outbound

This paper cites Lower bounds for non-convex stochastic optimization,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lower bounds for non-convex stochastic optimization,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.915141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.622696Z digest=sha256:b4c9b9a124050352b6428365d5dd6bd19c88ba794a777cd386af018836d71a9b

Observation 8aa29741-fb40-45e8-9b6f-6eb49c4d653a · outbound

This paper cites A stochastic approximation method,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A stochastic approximation method,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.625262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.625262Z digest=sha256:0d690b99a8fc27f6a5472e48b5179ee89c3e327b5824e5ddf67ad05ece63fef8

Observation 37318d86-763a-4520-a46e-43cf4fa59af7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.628081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.628081Z digest=sha256:152e5aae7c2ec6504214967ef1aa8cd98a811780ab9fee6fdad4aff17ef790c8

Observation ef273616-c7fa-4b26-957c-cf3618c69d17 · outbound

This paper cites Language models are unsupervised multitask learners,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language models are unsupervised multitask learners,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.898796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:10:16.631122Z digest=sha256:7c6fcf3a7efd720c3aa0001f4c3ac4afd0147240e5a373d01956a10e89652c99

Pith citing papers

No inbound Pith citation observations are available.