Pith. sign in

Paper Citation Record · LEDGER

Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2305.14342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.14342 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:08:34.918720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T17:36:25.374536Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f92b09ec-4ed6-42b5-a8d8-cd983770081c · inbound

A Physics-Inspired Optimizer: Velocity Regularized Adam cites this paper.

A Physics-Inspired Optimizer: Velocity Regularized Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T14:34:54.274212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T14:34:37.468180Z digest=sha256:afc6f2a3186610a276bb44963d49bd2448fad3b7af04cadc73742d9cbbb98f47

Observation 6d51b4e2-2428-4ea0-835e-65e80319d235 · inbound

On the Convergence of Muon and Beyond cites this paper.

On the Convergence of Muon and Beyond Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:56:34.183303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T15:56:30.602824Z digest=sha256:57e7e6324c9c5bf2d64585cb6a9cc45500fb93360c643d23d60c87ebf296f785

Observation fa4a4d9d-b91c-4d2b-a413-ad83348928b1 · inbound

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure cites this paper.

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:52:41.131143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T14:51:30.312509Z digest=sha256:cf11dffc76253ab80f1ce5bfa01810789e24dd2e0a196adc60b7e8accbd6f81f

Observation 524fef99-d863-46aa-929b-efe21266c15b · inbound

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning cites this paper.

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:46:16.991047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:44:53.516653Z digest=sha256:c4583bc345465f0b907af86d9cb69c556e017e8d6dbd94e28b23c6ce7f11e28c

Observation 4162e015-673a-42ae-bd0b-59b21bd62987 · inbound

GraphMind: Theorem Selection and Conclusion Generation Framework with Dynamic GNN for LLM Reasoning cites this paper.

GraphMind: Theorem Selection and Conclusion Generation Framework with Dynamic GNN for LLM Reasoning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:34:17.957013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T18:31:46.510611Z digest=sha256:02e529f4d9c4ea5088ab4d3f4d11d2e6fe993352f072caac0820c73665afe9ea

Observation c1581cf0-e851-440d-acb1-cb73c7511f13 · inbound

Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for Safe Reinforcement Learning in Optimal Control cites this paper.

Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for Safe Reinforcement Learning in Optimal Control Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:38:03.167730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:38:00.110683Z digest=sha256:0a2ff01a3b46555be800c49b768eee51bd610320eab04c64f1df7d33ccee8d5a

Observation 20291dcc-a6e8-4039-9249-45414dc50937 · inbound

Depth, Not Data: An Analysis of Hessian Spectral Bifurcation cites this paper.

Depth, Not Data: An Analysis of Hessian Spectral Bifurcation Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:54:11.462030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:53:29.750448Z digest=sha256:2ac760aa46f2bc44f80eadc82b24f6211ba01b87edb3e5d190263cdc3877fd45

Observation 7cb9d17c-2b5d-4eec-bab4-aafb3d5eabbc · inbound

Depth, Not Data: An Analysis of Hessian Spectral Bifurcation cites this paper.

Depth, Not Data: An Analysis of Hessian Spectral Bifurcation Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T06:08:34.918720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:08:34.918720Z digest=sha256:1e22dbd00553cb3e5672fefbd7723cf53a6529ac13c3a5de2fde2c856b64dcd4

Observation f65aeb92-3dfa-4e91-9a24-53946232c73b · inbound

HTMuon: Improving Muon via Heavy-Tailed Spectral Correction cites this paper.

HTMuon: Improving Muon via Heavy-Tailed Spectral Correction Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:55:26.310837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T06:50:29.893126Z digest=sha256:1a291610103dfc7740510b4f041b43b8cda39c2e58369592c30a887202ae7029

Observation abd7b83e-755f-4e06-9d4e-7809f7361549 · inbound

FLeX: Fourier-based Low-rank EXpansion for multilingual transfer cites this paper.

FLeX: Fourier-based Low-rank EXpansion for multilingual transfer Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:25:53.403860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:50:04.707303Z digest=sha256:a7491a15c1e0c17d06ae950be3f08448a83fdf88c4dca8d78e8bb66aad97fa4f

Observation 1e31ab96-ce91-475e-8898-20a5a0dda58d · inbound

STQuant: Spatio-Temporal Adaptive Framework for Optimizer Quantization in Large Multimodal Model Training cites this paper.

STQuant: Spatio-Temporal Adaptive Framework for Optimizer Quantization in Large Multimodal Model Training Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:50:57.041214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:49:40.758234Z digest=sha256:16014996613e5a7d3c86d6cd271d4041a4a6b3e09ecb99e9672218db9723347a

Observation 9c0dcc7e-ec11-4da9-9feb-eea5c9b0be8a · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.357258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:9d964f33c2a73ca527060090903287ac9f110bc30b510eec609b4e02571da805

Observation ddde0999-e6c4-457f-a4b9-87ec3efecd8e · inbound

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments cites this paper.

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:07.780240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T19:50:50.653184Z digest=sha256:f1f7b7cee9ea121d0be885699ecb19951741ae1d0a0372cb9e8eeec6cba27f01

Observation d0f2b442-edfd-42e7-a38a-27a19a23f2c2 · inbound

Anon: Extrapolating Adaptivity Beyond SGD and Adam cites this paper.

Anon: Extrapolating Adaptivity Beyond SGD and Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:09.762686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T16:25:36.073417Z digest=sha256:4f6e99427c6b70b45098c87f9d6054951aedad54041f0d1f12be879a79ee4b02

Observation 0d7a88ee-e398-474d-910f-ddad5766b1cb · inbound

Fast Gauss-Newton for Multiclass Cross-Entropy cites this paper.

Fast Gauss-Newton for Multiclass Cross-Entropy Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:09.639399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:01:26.202636Z digest=sha256:16ff9de5a1c24ce422ce1bdaa30e403cb7a0dad50bdde4b65fce7425f70fc733

Observation b4d6b4c9-d03d-450e-aa4d-4cd04cb06175 · inbound

When Descent Is Too Stable: Event-Triggered Hamiltonian Learning to Optimize cites this paper.

When Descent Is Too Stable: Event-Triggered Hamiltonian Learning to Optimize Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:15:58.135974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:52:21.146199Z digest=sha256:6eb54de1010151947ef0ae6b34576c2e68505638b1d6648292f134b97da66089

Observation 38f7f36c-4495-4e6c-9046-1dc5bab39699 · inbound

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling cites this paper.

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:55.545930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:47:08.766380Z digest=sha256:d85663ea432f4b545960b7822b178b330946958d20f627e44a68db2d71d1cf86

Observation 768681cb-4d85-4433-b1de-9f55c8b3656a · inbound

Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered cites this paper.

Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T21:23:44.499777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T21:19:55.074853Z digest=sha256:37f90ba7db5757f1fbba972c457e0ccb592f2966f59d49c4cd83be9c57153201

Observation a70a80e3-06e8-4f18-8e64-1e9c65c9463f · inbound

TBP-mHC: full expressivity for manifold-constrained hyper connections through transportation polytopes cites this paper.

TBP-mHC: full expressivity for manifold-constrained hyper connections through transportation polytopes Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.417380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:07.226252Z digest=sha256:c143ae686e103212df87a553d4cbde96adbab420d6629adde692394a8d509b26

Observation e278e906-69ee-4b96-b4e4-31171763cf2e · inbound

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics cites this paper.

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:11:17.272744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T08:06:52.309619Z digest=sha256:de12ed2b22d42c7b88ca1387332365778d34a2591936400799275e301c7d8284

Observation dd4ffc87-30e8-45ae-ae74-045f5c499fc3 · inbound

MAdam: Metric-Aware Multi-Objective Adam cites this paper.

MAdam: Metric-Aware Multi-Objective Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:28.457859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:21:53.393381Z digest=sha256:1f0d50d20fd2c1cfebd10afd4eeaf35b054f4f2005b5c518ff1cf0eb2831cdc7

Observation 0f5bdaee-8582-4180-a923-3c8dd4fa77e6 · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:06:44.868441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:f44ebf39ddb155a722d7cb51bf35c91c489dac6a3e5d0bf27fdbd05e9dd177df

Observation 23065f76-49d6-41ff-ae8e-b74256f84141 · inbound

Nonreversible Gauge Fields in Fokker--Planck Dynamics: Supersymmetric Hamiltonians and Learned Finite Forces cites this paper.

Nonreversible Gauge Fields in Fokker--Planck Dynamics: Supersymmetric Hamiltonians and Learned Finite Forces Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.266798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T22:42:16.293180Z digest=sha256:1574e2a143a09715c81fe0032f7915a1ee1b736437fd280ae0a46d9e9c1d136b

Observation dd75d8df-e762-4841-adeb-7804f04027d1 · inbound

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning cites this paper.

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:41.033120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T05:03:06.472027Z digest=sha256:cc221cebd911126728e1ee00ee8756984e16c76d2d718fddbdc5ac61cb2d24b1

Observation 639c8233-e634-4a3b-9ac5-2cc2e2935678 · inbound

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models cites this paper.

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.420186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T00:09:54.353782Z digest=sha256:8dd9b274479be90fbb2f8ed7c4aacd4afd718b3132956d74caca95fb19c0af45

Observation 65aa25fd-9971-43bd-aad9-911f003c6835 · inbound

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models cites this paper.

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:25:41.284177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T06:36:46.924135Z digest=sha256:eae51e02c25a8b7d514a83c55b6f13d03d75e3b0205e3f412ab757c8e2501006

Observation a5c9ca12-1df1-4865-8c30-f836d80cec43 · inbound

Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules cites this paper.

Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:21.498373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T07:31:11.050545Z digest=sha256:eaea6e7a839b19131a93a88b285871a85f11d57d3226dd51c81faa6380092721

Observation ad27b670-2eea-4251-82d4-ef71158d43d9 · inbound

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions cites this paper.

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-07-09T17:36:25.376052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-09T17:31:09.270342Z digest=sha256:7dfbba468578e21a093ff876b88c6905410f5d93706f80578d03f6888a9be1ff

Observation 72064397-e298-4e98-bb53-26beab5072c7 · inbound

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions cites this paper.

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 163

Resolution
unresolved
no resolver link, observed 2026-07-14T15:50:41.329016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:50:41.329016Z digest=sha256:193eaf8420f31291d5d1a81086c544c115f9505ace95f661f916470ac192b46d

Observation a0823259-c6e0-43b6-be8a-72bef1918ab1 · inbound

OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining cites this paper.

OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T12:30:37.114754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:30:37.114754Z digest=sha256:1ee73a73cdcbf6017e2ec6552461b87e296379adda98bee57a047d68a3c2a91d