Pith. sign in

Paper Citation Record · LEDGER

Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2305.14342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.14342 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:18:01.465897Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T17:36:25.374536Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6704f74b-430b-4ac6-aabf-bd407257fa08 · inbound

MARS: Unleashing the Power of Variance Reduction for Training Large Models cites this paper.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.838625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.838625Z digest=sha256:32f0ac352d59dd90cb82c84dae192c0edb57668e6192fb5906ba527461323881

Observation d60d0ddb-4d26-45c6-a4b3-b4e4c5826b3a · inbound

HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization cites this paper.

HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:30:54.969095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:30:54.969095Z digest=sha256:38336dd4842f7a3d7ba321b66a04031a053b38521e98bb8258eb82b928d376c8

Observation aed8cdd0-072e-4cdc-8d38-e525c030185c · inbound

An Enhanced Levenberg--Marquardt Method via Gram Reduction cites this paper.

An Enhanced Levenberg--Marquardt Method via Gram Reduction Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:57:13.148559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:57:13.148559Z digest=sha256:8d2ec621d145f43ca752e6c3dd0e4f139b67bd53f96087f4b0ea014cbdb9d7d3

Observation 816c38e3-096c-4bc3-8bd0-0af2dfd509f5 · inbound

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training cites this paper.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.964729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.964729Z digest=sha256:28456ecc6d907ba7aa788ac5f901f7897c332042adf031c00c1000ebb138cdcc

Observation d72f25a4-55e2-4c77-a360-1ff99fd6cf31 · inbound

Open Problems in Machine Unlearning for AI Safety cites this paper.

Open Problems in Machine Unlearning for AI Safety Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:10.338748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:10.338748Z digest=sha256:d0d73937861e9ed7bc7a22c61fa8f9bc8906f39cb9759bbfc296443569420fee

Observation f13b28c7-d635-4820-bafe-ead489f3cdae · inbound

Physics of Skill Learning cites this paper.

Physics of Skill Learning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:20:51.014910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:20:51.014910Z digest=sha256:ea84d5c5ff4bbdfe9cb05fae5f5a29763256ab852274084973d235d421c3e983

Observation 6ce8f0f0-0954-4e5e-90e8-879eb8ae34cf · inbound

Explicit Eigenvalue Regularization Improves Sharpness-Aware Minimization cites this paper.

Explicit Eigenvalue Regularization Improves Sharpness-Aware Minimization Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:03:52.681306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:03:52.681306Z digest=sha256:c04ad9c6f227bcb98ccdd090194d214b68b870f8af6a3c3287a137852fe92a26

Observation d1ff73d5-fed1-4f72-bc32-107140871d21 · inbound

Spectral-factorized Positive-definite Curvature Learning for NN Training cites this paper.

Spectral-factorized Positive-definite Curvature Learning for NN Training Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T16:15:07.263627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:15:07.263627Z digest=sha256:647fccf379b39c940846b771429ad99aa20c2b895aae763b18f2a98580a4a874

Observation 7e82254b-2dae-417a-a908-44bf23178cdc · inbound

Improving Adaptive Moment Optimization via Preconditioner Diagonalization cites this paper.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.278139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.278139Z digest=sha256:c305fdd4fec83ee6e0507315ec1311f57842a4c1e71afd879fb4454e6cc2f899

Observation 28103d01-4a27-4def-af13-4ab16277a96d · inbound

AlphaGrad: Non-Linear Gradient Normalization Optimizer cites this paper.

AlphaGrad: Non-Linear Gradient Normalization Optimizer Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:18:01.465897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:18:01.465897Z digest=sha256:c8215adb6d193ea9da3b8a3b9ed150fdc7302ed6ceee4d190cfa7a4889a1e496

Observation 53b5ac0a-c32a-4cdf-9c2f-5000aebe686b · inbound

HessFormer: Hessians at Foundation Scale cites this paper.

HessFormer: Hessians at Foundation Scale Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:26.092494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:26.092494Z digest=sha256:72c53ad210fdcc9cf5351d88bdc2971642c938d46056d88ace4ca60d903bf007

Observation f92b09ec-4ed6-42b5-a8d8-cd983770081c · inbound

A Physics-Inspired Optimizer: Velocity Regularized Adam cites this paper.

A Physics-Inspired Optimizer: Velocity Regularized Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T14:34:54.274212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T14:34:37.468180Z digest=sha256:ca92fed379593e07a70b477cb381b7dcfc59220b591074a4f21598e0c064ce3f

Observation 79d2afd9-892c-4856-80e5-74ea0ec09f49 · inbound

A Physics-Inspired Optimizer: Velocity Regularized Adam cites this paper.

A Physics-Inspired Optimizer: Velocity Regularized Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-15T20:24:29.576710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:24:29.576710Z digest=sha256:1a21b70455885c194880377f90b0ed1ce2f57c9379a9259d2c448c6b7c5aca72

Observation 9d5a6cc2-5fce-499b-b2a9-d380d7eddbc3 · inbound

AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training cites this paper.

AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:13.476967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:09:13.476967Z digest=sha256:3453d36eb53355226f090a7483426c45c8e402c9e3872e8b234b2c35d0355964

Observation 41d1d060-2d07-4b60-82c3-1d07c83b3c6b · inbound

Accelerated Training of Federated Learning via Second-Order Methods cites this paper.

Accelerated Training of Federated Learning via Second-Order Methods Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:35.556742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:35.556742Z digest=sha256:69135a633e0716c1b03828865ba5aa909451a79134185dfcc9aa23af4c8a6764

Observation f061df55-1140-45c8-8105-1c0ed0ec6369 · inbound

Stochastic Diagonal Estimation Based on Matrix Quadratic Form Oracles cites this paper.

Stochastic Diagonal Estimation Based on Matrix Quadratic Form Oracles Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:44:33.269049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:44:33.269049Z digest=sha256:4986eb1138aabe4dceba4e7886af71e80ea0bf01f5072340bd4c0b2883b6b34f

Observation d489ca37-7038-47a4-bbe0-bc28e55b7f10 · inbound

Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track cites this paper.

Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:17.783769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:13:17.783769Z digest=sha256:a604a455cfe6d71132c48dc9c7f26eb0df9690c242af1c45de82b89cc71ff069

Observation 8a4e6230-10b8-41d4-9d9c-120d5a4c958b · inbound

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam cites this paper.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.530074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.530074Z digest=sha256:f7d8a77214e1649d3477eb25e9c40276587047ed08cf3ed3d6a7a48fbb20cef6

Observation 8ff146d4-94b4-4b36-9368-70d78bbf45b9 · inbound

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD cites this paper.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.829273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.829273Z digest=sha256:b2edc6748f23681452b675b16035d4c06d47d60390a04a02134ba17a79917a9d

Observation bd57ed43-8e23-4aee-b56d-253bba041ffb · inbound

Sketched Gaussian Mechanism for Private Federated Learning cites this paper.

Sketched Gaussian Mechanism for Private Federated Learning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T21:12:50.922764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:12:50.922764Z digest=sha256:7f1947d13de83bff8cad4ea2d840547178190ebd500f281924132e20b1c5b886

Observation 6d51b4e2-2428-4ea0-835e-65e80319d235 · inbound

On the Convergence of Muon and Beyond cites this paper.

On the Convergence of Muon and Beyond Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:56:34.183303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T15:56:30.602824Z digest=sha256:75f6a4f0a35c9a34a5bf954a798d4654ee396c65bf8238f0d2e2c6e7df7a9373

Observation fa4a4d9d-b91c-4d2b-a413-ad83348928b1 · inbound

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure cites this paper.

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:52:41.131143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T14:51:30.312509Z digest=sha256:f424635e5a9fcb1772438d08cd4491a0a5dd65654b36aa2e4bad1582ccd36fb7

Observation 524fef99-d863-46aa-929b-efe21266c15b · inbound

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning cites this paper.

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:46:16.991047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T10:44:53.516653Z digest=sha256:1d8902344f2bcefe872ac5e2133b977efe03b9b8e30186b1bdb8fd7308fc76da

Observation 4162e015-673a-42ae-bd0b-59b21bd62987 · inbound

GraphMind: Theorem Selection and Conclusion Generation Framework with Dynamic GNN for LLM Reasoning cites this paper.

GraphMind: Theorem Selection and Conclusion Generation Framework with Dynamic GNN for LLM Reasoning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:34:17.957013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T18:31:46.510611Z digest=sha256:c9eb5b337eda0e6b70c1e8d41777bedac203cc9b7cb13c63a48633f75fc93106

Observation c1581cf0-e851-440d-acb1-cb73c7511f13 · inbound

Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for Safe Reinforcement Learning in Optimal Control cites this paper.

Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for Safe Reinforcement Learning in Optimal Control Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:38:03.167730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T15:38:00.110683Z digest=sha256:6076ea1841d91e0ad1e5acdb7ed7adc470a734c27b0bbf0766e0d2a83c3ffd54

Observation 20291dcc-a6e8-4039-9249-45414dc50937 · inbound

Depth, Not Data: An Analysis of Hessian Spectral Bifurcation cites this paper.

Depth, Not Data: An Analysis of Hessian Spectral Bifurcation Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:54:11.462030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T13:53:29.750448Z digest=sha256:c90665c5d3594ad2c763a740ba132e88c4b800bdd342e3b92ae32f6b8a50cb68

Observation 7cb9d17c-2b5d-4eec-bab4-aafb3d5eabbc · inbound

Depth, Not Data: An Analysis of Hessian Spectral Bifurcation cites this paper.

Depth, Not Data: An Analysis of Hessian Spectral Bifurcation Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T06:08:34.918720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:08:34.918720Z digest=sha256:1d59e8ffb4c7d1064bb3c0e07d68525518f6757d9fa9b36adf20d2b08c821698

Observation f65aeb92-3dfa-4e91-9a24-53946232c73b · inbound

HTMuon: Improving Muon via Heavy-Tailed Spectral Correction cites this paper.

HTMuon: Improving Muon via Heavy-Tailed Spectral Correction Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:55:26.310837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T06:50:29.893126Z digest=sha256:b1e549bab042f0cccff7c74d8c4d750ab95670429f5609fe7d70bebb8d83f132

Observation abd7b83e-755f-4e06-9d4e-7809f7361549 · inbound

FLeX: Fourier-based Low-rank EXpansion for multilingual transfer cites this paper.

FLeX: Fourier-based Low-rank EXpansion for multilingual transfer Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:25:53.403860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T19:50:04.707303Z digest=sha256:b07c27b57fe84b9355c513acaa5a0f439ce28086d628429560e1470c6608a98e

Observation 1e31ab96-ce91-475e-8898-20a5a0dda58d · inbound

STQuant: Spatio-Temporal Adaptive Framework for Optimizer Quantization in Large Multimodal Model Training cites this paper.

STQuant: Spatio-Temporal Adaptive Framework for Optimizer Quantization in Large Multimodal Model Training Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:50:57.041214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:49:40.758234Z digest=sha256:43bcb10480831b25ddd2b5d74e8bcea94553aaed11b108908a3c5895fe2f7da2

Observation 9c0dcc7e-ec11-4da9-9feb-eea5c9b0be8a · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.357258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:568b9a90b9c7206b324f78463952f5c68c4117c63a7866a3900e4cf45ccc0ddf

Observation ddde0999-e6c4-457f-a4b9-87ec3efecd8e · inbound

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments cites this paper.

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:07.780240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T19:50:50.653184Z digest=sha256:4298738247c35e9c438d5089fadae9ae0b410b7d097d45c2dde352d4526d420e

Observation d0f2b442-edfd-42e7-a38a-27a19a23f2c2 · inbound

Anon: Extrapolating Adaptivity Beyond SGD and Adam cites this paper.

Anon: Extrapolating Adaptivity Beyond SGD and Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:09.762686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T16:25:36.073417Z digest=sha256:b9c8669f0d0b1cd28a95aca867ce1342090424bd9597dee9553031509efbc7e9

Observation 0d7a88ee-e398-474d-910f-ddad5766b1cb · inbound

Fast Gauss-Newton for Multiclass Cross-Entropy cites this paper.

Fast Gauss-Newton for Multiclass Cross-Entropy Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:09.639399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T14:01:26.202636Z digest=sha256:63edc124cb86257053b6e88be6517cf2a7a3fd8061b0941918b4605366199160

Observation b4d6b4c9-d03d-450e-aa4d-4cd04cb06175 · inbound

When Descent Is Too Stable: Event-Triggered Hamiltonian Learning to Optimize cites this paper.

When Descent Is Too Stable: Event-Triggered Hamiltonian Learning to Optimize Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:15:58.135974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T01:52:21.146199Z digest=sha256:2cb77045aeb8b826b33570efe9707b4c87ae85d838054658a36b188105f49e6a

Observation 38f7f36c-4495-4e6c-9046-1dc5bab39699 · inbound

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling cites this paper.

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:55.545930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T02:47:08.766380Z digest=sha256:0aa561e673e88419c5a7a2362ef6ce3492bd8cff071c12184a0214bc81912257

Observation 768681cb-4d85-4433-b1de-9f55c8b3656a · inbound

Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered cites this paper.

Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T21:23:44.499777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T21:19:55.074853Z digest=sha256:706342857835a22ead9c5d13ed214df17cbaadc21c9683165fee891452e56c22

Observation a70a80e3-06e8-4f18-8e64-1e9c65c9463f · inbound

TBP-mHC: full expressivity for manifold-constrained hyper connections through transportation polytopes cites this paper.

TBP-mHC: full expressivity for manifold-constrained hyper connections through transportation polytopes Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.417380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-22T10:00:07.226252Z digest=sha256:8ea5404f82e71a682e946980a1f1d4d0a697a01980d343fe28fc3b95d65ba111

Observation e278e906-69ee-4b96-b4e4-31171763cf2e · inbound

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics cites this paper.

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:11:17.272744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-22T08:06:52.309619Z digest=sha256:fd5903bfb73dc63a21305c469691d10c92a3a437aa4a03777260193905e050c7

Observation dd4ffc87-30e8-45ae-ae74-045f5c499fc3 · inbound

MAdam: Metric-Aware Multi-Objective Adam cites this paper.

MAdam: Metric-Aware Multi-Objective Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:28.457859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T11:21:53.393381Z digest=sha256:1c3c26baab1ad067ab7e35c4475513feb23328a1a43a446a6a7f68cb60a9e544

Observation 0f5bdaee-8582-4180-a923-3c8dd4fa77e6 · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:06:44.868441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:a1a7119f2ef7c15890fee370a0935b6a5fc7200d7100a509cd56a9a4b6385e3a

Observation 23065f76-49d6-41ff-ae8e-b74256f84141 · inbound

Nonreversible Gauge Fields in Fokker--Planck Dynamics: Supersymmetric Hamiltonians and Learned Finite Forces cites this paper.

Nonreversible Gauge Fields in Fokker--Planck Dynamics: Supersymmetric Hamiltonians and Learned Finite Forces Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.266798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T22:42:16.293180Z digest=sha256:47c5550276a692d13fc549906441258c298c2f4124fe1c68d2f1c1b45284b39c

Observation dd75d8df-e762-4841-adeb-7804f04027d1 · inbound

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning cites this paper.

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:41.033120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T05:03:06.472027Z digest=sha256:9fed1bacd722a57ba9ac28f6a1a5f2cfe1d8cfd0144d2f0c71eea5b2b5c9c809

Observation 639c8233-e634-4a3b-9ac5-2cc2e2935678 · inbound

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models cites this paper.

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.420186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T00:09:54.353782Z digest=sha256:0d4d211449ab8a52f611fdd789cd2227eb5a65d7588892b3a72df6e3ee255f67

Observation 65aa25fd-9971-43bd-aad9-911f003c6835 · inbound

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models cites this paper.

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:25:41.284177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T06:36:46.924135Z digest=sha256:bbcc109f3de6e1fb26630fbc5663b8e1db2788f95f540d57206a6e79c32afbc4

Observation a5c9ca12-1df1-4865-8c30-f836d80cec43 · inbound

Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules cites this paper.

Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:21.498373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-30T07:31:11.050545Z digest=sha256:9ded400a7a9ef20419f733dc9cf6d602962f67509e2c02394a33e19a166f4242

Observation ad27b670-2eea-4251-82d4-ef71158d43d9 · inbound

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions cites this paper.

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-07-09T17:36:25.376052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-09T17:31:09.270342Z digest=sha256:a6647b9456be8a914b495459b388ea506a26537266885ffc346277cd171203dc

Observation 72064397-e298-4e98-bb53-26beab5072c7 · inbound

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions cites this paper.

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 163

Resolution
unresolved
no resolver link, observed 2026-07-14T15:50:41.329016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:50:41.329016Z digest=sha256:23129d86655df2af4b7749ebf42229a2e8dcad8854d8dd05889f89ed0afa003d

Observation a0823259-c6e0-43b6-be8a-72bef1918ab1 · inbound

OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining cites this paper.

OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T12:30:37.114754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:30:37.114754Z digest=sha256:2f6ee0f0f3f4f7c97430ca66e1d7146c0e8e5e099c5c9d7f7d4ded7322797711