Pith. sign in

Paper Citation Record · LEDGER

MARS: Unleashing the Power of Variance Reduction for Training Large Models

As of 23 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 20 inbound Pith citation observations for arXiv:2411.10438.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10438 v4

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:43:02.115828Z

measured 115 of 115 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:20:51.021889Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:56.218978Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 32587236-7371-4bc6-8cae-c4c9317b1f20 · outbound

This paper cites and Yuan, Y.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Yuan, Y

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.585508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.585508Z digest=sha256:c0e2e3291297736708e0dc2c4661a422faa4a66fcd23451b67021967fcd4dd3a

Observation c23494a3-6110-4f42-b1fb-ef0b3e02ddac · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Scalable Second Order Optimization for Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.591712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.591712Z digest=sha256:70278e094e6b7db25842569100e97759a7811b413b47d6d7c4c29db7ac937e38

Observation 489e2f04-193e-4726-9e4f-70c67a967204 · outbound

This paper cites C., Foster, D.

MARS: Unleashing the Power of Variance Reduction for Training Large Models C., Foster, D

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.597515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.597515Z digest=sha256:766e896ff0ec92f0dcc0454b9a3f801495eb1431cd2b5159e37febb5df2191fe

Observation a3cade9e-05c8-4b36-a45c-aa930683b285 · outbound

This paper cites and Glynn, P.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Glynn, P

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.603039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.603039Z digest=sha256:ddf9c988d9f041a94827c4b6bddc6b4ce0dc733963dd03bf92a89a452a2d3c86

Observation 75163b56-d8a4-4c74-b8e2-2ff5c75ab8cd · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Old Optimizer, New Norm: An Anthology

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.608492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.608492Z digest=sha256:7ff2988997dc449709ca88229ff00be96ab52255ed82a32f46f00881f60ba2be

Observation 34e7797b-eb8b-4859-a7d8-4b4e61f5733d · outbound

This paper cites L., Gao, J., and Choi, Y.

MARS: Unleashing the Power of Variance Reduction for Training Large Models L., Gao, J., and Choi, Y

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.615584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.615584Z digest=sha256:9a37552822be527474bafd844336ea17c333cf143184f9ea89b4f00acb711825

Observation 0af577b4-de4b-447b-a9c0-132775e4e360 · outbound

This paper cites Language Models are Few-Shot Learners.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.622016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.622016Z digest=sha256:80300c5e383058fbd6bcbe5ba59350d991696b18e98305ef89cb878daa8c1e9d

Observation 18cd5d9d-fd4e-4d79-9728-7799a3146168 · outbound

This paper cites Stochastic spectral descent for restricted boltzmann machines.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic spectral descent for restricted boltzmann machines

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.629196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.629196Z digest=sha256:e99a8bae3acde6672413255c054ad62d63e36c9921aae226c40b7bf5f5496fea

Observation f271a367-fd5f-4c90-be6b-be2c2a5d4c4f · outbound

This paper cites Stochastic spectral descent for discrete graphical models.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic spectral descent for discrete graphical models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.635158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.635158Z digest=sha256:79be33fafb729b5bce71b0635085a8140986f6f469b4a70781c9fb8df06bc35e

Observation 6a327634-7f86-4bdd-abcb-41771812372d · outbound

This paper cites Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.642149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.642149Z digest=sha256:aaf04e50cbd5e709831804724db47eeb65335ac6fe51baf6ef23c38a63f182d4

Observation 5a69e144-f8f0-4d1e-a106-01c139073854 · outbound

This paper cites Symbolic discovery of optimization algorithms.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Symbolic discovery of optimization algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.647714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.647714Z digest=sha256:b576e6d9f618293d6caf5a8522d1620306e5b4d4ceec2b2e291aa99eb066f302

Observation de0da958-9185-4240-ab0b-e12a0a20c986 · outbound

This paper cites W., Sutton, C., Gehrmann, S., et al.

MARS: Unleashing the Power of Variance Reduction for Training Large Models W., Sutton, C., Gehrmann, S., et al

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.652838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.652838Z digest=sha256:c2592a4a8302a27f6978bdccd6b4e5bb4201d3a8657d630fba5b0702771b0a73

Observation 4dfdbe6d-757e-4d09-aded-a15b2248b3ce · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.657716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.657716Z digest=sha256:0915964bf8cae8f4c26dd45df581df5fd06c30144f6d0e1c2a4212fb806d897a

Observation c6986463-4542-4b2e-9862-06881e5d84fe · outbound

This paper cites and Orabona, F.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Orabona, F

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.662494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.662494Z digest=sha256:7930cb6a7256d9720350296bac01ba3ac94b1eafc771cafa874482977c7bf647

Observation 11856636-9353-4a77-a17b-83886cbbd8a4 · outbound

This paper cites and Bottou, L.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Bottou, L

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.667415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.667415Z digest=sha256:425e9f47fe13293c89b8d4b82fc232b6f6da99d93a33f701a0c50e554b1750df

Observation 383cc93d-6e19-4562-b8f7-bf25784f1e2f · outbound

This paper cites Saga: A fast incremental gradient method with support for non-strongly convex composite objectives.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Saga: A fast incremental gradient method with support for non-strongly convex composite objectives

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.672447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.672447Z digest=sha256:c6785c526530334e604dd2ffc79a1e67c0b6b749b019a92494b800036e533997

Observation 473f2f71-6b58-4f5c-baaf-687600880774 · outbound

This paper cites The Road Less Scheduled.

MARS: Unleashing the Power of Variance Reduction for Training Large Models The Road Less Scheduled

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.677433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.677433Z digest=sha256:5f6cf138c4f2f6794ec8fd27dd395c61f1c4eb10538a03b570b24035bce34d07

Observation 311715d3-4e4b-4421-a965-f5f1ccda8f22 · outbound

This paper cites Stochastic variance-reduced newton: Accelerating finite-sum minimization with large batches.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic variance-reduced newton: Accelerating finite-sum minimization with large batches

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.683546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.683546Z digest=sha256:8e3317bfd060aaef3f1a5d714332d863b1c1c03aee0eb8f5b013b20b9f5c81fc

Observation a9ce96ff-9c55-4f14-b1e6-4d48be2963c5 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

MARS: Unleashing the Power of Variance Reduction for Training Large Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.688005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.688005Z digest=sha256:a2fdba29d4044bdbac99e21fa85797d2b469b0b27911b295f496b5c6eb1aee93

Observation c084cd7b-690e-4278-8d2c-dd4f5fcfbf75 · outbound

This paper cites The Llama 3 Herd of Models.

MARS: Unleashing the Power of Variance Reduction for Training Large Models The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.692307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.692307Z digest=sha256:97bbc3dcf146aaf7e7f33617037cbca26d7490ce7122214fffeb0cdf5d9277eb

Observation 82f877d8-2988-4dd1-9e2e-4e596f507325 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Adaptive subgradient methods for online learning and stochastic optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.696550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.696550Z digest=sha256:0d185f77ce3f563720af7f6cc1e5bae5308e5ca20fc8cdc774d87b3b2ed2599a

Observation e7450097-b671-4ddc-96e0-bb465215a037 · outbound

This paper cites J., Lin, Z., and Zhang, T.

MARS: Unleashing the Power of Variance Reduction for Training Large Models J., Lin, Z., and Zhang, T

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.701454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.701454Z digest=sha256:be881f643fa9215452b0943edd0d2bdfc2ce3a4491d84bc1af0c3eb3f1956eef

Observation d2867a9b-db15-4931-a3b8-dbaa05ed9173 · outbound

This paper cites Promise: Preconditioned stochastic optimization methods by incorporating scalable curvature estimates.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Promise: Preconditioned stochastic optimization methods by incorporating scalable curvature estimates

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.977171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.706512Z digest=sha256:7176fcd73d7368018f4be221a34c6da77f355d04b65e711f67b1181681c925fe

Observation 28f611b6-5a1a-40d5-8130-e98ebebb5e23 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

MARS: Unleashing the Power of Variance Reduction for Training Large Models A framework for few-shot language model evaluation, 07 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.711300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.711300Z digest=sha256:212a4f9c6264233ad00a320eed15b37d38a6ced7931ddb4d6da96fa443d51f22

Observation 332c3eeb-a786-460f-91d7-5586729a2c81 · outbound

This paper cites and Lan, G.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Lan, G

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.715934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.715934Z digest=sha256:831873ccc2e786a8690ad11e900749549b5d68e1274bb6aa95be9f03eff80cfa

Observation 0a9c5b38-f97a-4ac3-ad29-1b5a50e3f11e · outbound

This paper cites Openwebtext corpus.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Openwebtext corpus

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.720780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.720780Z digest=sha256:5ad71fb55da23ca30b74882f026506f68e5c9d2550d5f4d886ca6feec2782a87

Observation f0090491-0933-4f80-86d8-14c98d7ad1e8 · outbound

This paper cites and Graves, A.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Graves, A

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.725724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.725724Z digest=sha256:6cb32b49f6e7193ce2e31ee3685a007759e10653c2650277436526d026dabe6a

Observation c99be9de-a6d6-4e29-ab8a-075002b83bb4 · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Shampoo: Preconditioned stochastic tensor optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.730667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.730667Z digest=sha256:718366294317709d2271fd3d25595404e45220aa79a5efeac15a65acc5f0d996

Observation 837873a5-4868-487b-9676-6051d1199ade · outbound

This paper cites Beyond convexity: Stochastic quasi-convex optimization.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Beyond convexity: Stochastic quasi-convex optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.735999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.735999Z digest=sha256:5f9ebb35eb68573be2e6ca9dae6a0341db45bf198e6a0f682e41728da30fe05f

Observation 0eb16ae7-421a-4793-818f-b9012f330f17 · outbound

This paper cites Deep residual learning for image recognition.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Deep residual learning for image recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.740817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.740817Z digest=sha256:0a1c822c0582a6e3b3668b81f6d48b801b5d7abebcc2eb64afefab44ab2a0d9a

Observation 365be778-5806-4500-81d9-2650b84c20ee · outbound

This paper cites Measuring massive multitask language understanding.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Measuring massive multitask language understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.863789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.745796Z digest=sha256:4dbd28210d2bd8ed1f63bfcf75c545cb2366ca89eda3ed32eb3270834bdf7fa0

Observation 7c0c1a4f-42c9-4135-80fe-d10eab8c40bc · outbound

This paper cites an unresolved cited work.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:43:03.843349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.750649Z digest=sha256:d13296dcc66862d45983a63c4d03e11ba39b93def9d7bbed3aa629d1f41c7365

Observation a44fc610-5fee-424c-903f-eeff8f848111 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

MARS: Unleashing the Power of Variance Reduction for Training Large Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.756402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.756402Z digest=sha256:1d706ba66327a50529a564d762e12888890465fa4eadfeefd073008f120d716f

Observation f860c1b5-000b-43e3-90d6-ad060a880bd7 · outbound

This paper cites Super-adam: faster and universal framework of adaptive gradients.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Super-adam: faster and universal framework of adaptive gradients

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.815290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.762055Z digest=sha256:f7933a658dd81c8a5ad5b8542a359c51815dacd03611467b6191a4e2ce32836e

Observation 095ffd0c-484f-441a-958f-0ccf4ac6f5a4 · outbound

This paper cites and Zhang, T.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Zhang, T

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.766979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.766979Z digest=sha256:648838a1713854419b5fe9a7e04b395ce35f5eef3f729e42ca66c1e493c66659

Observation e6ac791e-5db5-4467-b101-9f19dc00ae43 · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Muon: An optimizer for hidden layers in neural networks, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.771826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.771826Z digest=sha256:2b1059ff3528b1188e97b50c6e5c7ff3d8066e492fc157b02fcdff59389c9d90

Observation 22458ebd-08da-48ea-9dec-10abb76e7a98 · outbound

This paper cites an unresolved cited work.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:43:03.761712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.776602Z digest=sha256:26c92b3847101621afabc4ffeb21efcc286abbb0a742aaf3bc0ce9bbfeb9b4a7

Observation d99fb412-a9cf-421e-93ec-20e084cb9f06 · outbound

This paper cites an unresolved cited work.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:43:03.735491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.781457Z digest=sha256:bcf9230a35cd10496f93e3d4e1d676ce5bdfb2aac164a584f9d4f3bfc3b5999b

Observation 7db134d0-4a1a-44a0-8ce9-914866c1da56 · outbound

This paper cites T., and Cevher, V.

MARS: Unleashing the Power of Variance Reduction for Training Large Models T., and Cevher, V

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.699534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.786181Z digest=sha256:f874e1c75f51163fb3b7078f322ae1529d6494812558f93608fe3aa1f90fdef4

Observation e0978da1-f323-430f-9617-c1a31cc3e799 · outbound

This paper cites an unresolved cited work.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.791022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.791022Z digest=sha256:84f70c22b5f64efa6c6aad678d5202c412c9b04c251df738fe198729c2c8dbd7

Observation 579e6a5e-ef91-48af-ae00-7b31d2e79891 · outbound

This paper cites Learning multiple layers of features from tiny images.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Learning multiple layers of features from tiny images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.795890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.795890Z digest=sha256:9c826c43b67ddb373afffce0fb56c3161cf62dac0b3a3a8be7be26b105b24038

Observation cce1d6db-0167-4dcb-b47f-59bb14ca9253 · outbound

This paper cites On the computation of the matrix k-th root.

MARS: Unleashing the Power of Variance Reduction for Training Large Models On the computation of the matrix k-th root

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.646413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.801647Z digest=sha256:69766a74e4999b9d8523d2cdc19bf42532852a1e4fa8c9de2d5bc054dbfe67bd

Observation 6b6aa78d-cce2-4b06-9429-5f6a6637dc43 · outbound

This paper cites S., Moeller, T.

MARS: Unleashing the Power of Variance Reduction for Training Large Models S., Moeller, T

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.621762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.807152Z digest=sha256:a2165385a0fc5b193ba1fdd07a193caa2021c199679745d45cc47701d204f472

Observation 564e5370-2bc4-40a4-bf8c-75b95e12b00b · outbound

This paper cites Gradient-based learning applied to document recognition.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Gradient-based learning applied to document recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.813004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.813004Z digest=sha256:22bfab8957ebda6c19321fc1f59f2e59e27b02ec92b39e790dd0302f3c43d245

Observation 29b4ed67-9353-458f-8251-93feac329703 · outbound

This paper cites Storm+: Fully adaptive sgd with recursive momentum for nonconvex optimization.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Storm+: Fully adaptive sgd with recursive momentum for nonconvex optimization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.583059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.817989Z digest=sha256:abdcfee727fc67b76217210dfd738710f8e0114de24f4dc3628dd34d79e28c42

Observation 49389a72-448f-4e61-b719-5bd1f2b428dc · outbound

This paper cites Smoothness and Adaptivity in Nonlinear Optimization for Machine Learning Applications.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Smoothness and Adaptivity in Nonlinear Optimization for Machine Learning Applications

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.562677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.823101Z digest=sha256:dc3ea23ff2cdda2c44c434739966cfabe01a6315d855c2adcb828fd5249efdca

Observation 28f3b14e-136d-4d1c-a3ae-94871d9d4cae · outbound

This paper cites Convergence of adam under relaxed assumptions.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Convergence of adam under relaxed assumptions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.541573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.828261Z digest=sha256:8200085808a8bc5d121e8c47d7d55673f2cf0e1f256241463c18043f28368d41

Observation 476a60c7-8422-4d96-af61-7dabfc07b7ce · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

MARS: Unleashing the Power of Variance Reduction for Training Large Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.833269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.833269Z digest=sha256:ab8d056ce61b1e31b79a00972ac5230763e0327bc20b0d007db5031bbadcccb1

Observation 6704f74b-430b-4ac6-aabf-bd407257fa08 · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.838625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.838625Z digest=sha256:b6aacc83575e64420307fd8a37f049fd209eae191b729faa8f9a70540d770b1b

Observation 929c6321-4e29-4615-9ed1-0bff2eb4f95e · outbound

This paper cites Adam$^+$: A Stochastic Method with Adaptive Variance Reduction.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Adam$^+$: A Stochastic Method with Adaptive Variance Reduction

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.844215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.844215Z digest=sha256:f6849e6514de27ddb25d2df4163890df52b2f77e0d3c760c455e34de922cd884

Observation 98bfc228-49f1-4c1d-90ea-7fcf22fbdc3f · outbound

This paper cites and Hutter, F.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Hutter, F

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.849372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.849372Z digest=sha256:e55bebdbdef8fe431c0a27c0b715a8055944a0a427239f2d845f96fc84898fe6

Observation 7feaa9b5-925b-4996-a482-6bb4a927d258 · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Fineweb-edu: the finest collection of educational content, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.498210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.855256Z digest=sha256:741b0375ae7cfa8e60f8c2889bc9680f201fb1762cabc14dad855a410c8f7a7d

Observation 0d21d3e8-2f8b-4f1a-b690-c0ab786f0e98 · outbound

This paper cites and Grosse, R.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Grosse, R

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.860686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.860686Z digest=sha256:4eaa313ae34dd7e38c5939c7ca99842e81d6b429f38a7f4922d138723590bafb

Observation c7297947-5172-401e-93a6-a3807bcf4d5b · outbound

This paper cites and Tropp, J.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Tropp, J

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.453737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.866570Z digest=sha256:e57168d2103811de98a2160fb9e3e97f4ec70f97dc5cfec068e61813bef6a1a6

Observation c3cd4b70-9543-4b57-8d1f-e0488227881d · outbound

This paper cites An Empirical Model of Large-Batch Training.

MARS: Unleashing the Power of Variance Reduction for Training Large Models An Empirical Model of Large-Batch Training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.872324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.872324Z digest=sha256:aa5e2a8bf069b52a406735ace4c4e709877fc02ad86114cc4b5134b27a238ab8

Observation b2c25452-1b46-4637-9aff-884ce3bff0e5 · outbound

This paper cites Adaptive Bound Optimization for Online Convex Optimization.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Adaptive Bound Optimization for Online Convex Optimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.879308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.879308Z digest=sha256:b51673d0ea958717d39868d9eec131d4ce3ceeb5e3d13f0298e200bfe28e2b59

Observation 68ea52d1-0b27-4be9-be39-b1dd82794446 · outbound

This paper cites Can a suit of armor conduct electricity? A new dataset for open book question answering.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Can a suit of armor conduct electricity? A new dataset for open book question answering

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.885057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.885057Z digest=sha256:352d121c37d4db7072ed3c79bdba95bda6e2669c0341644e1ac2c26ae8243238

Observation af18d8a0-4378-440b-9257-b35ebf83d8c2 · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

MARS: Unleashing the Power of Variance Reduction for Training Large Models A New Perspective on Shampoo's Preconditioner

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.891022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.891022Z digest=sha256:0960d4e9bd6fd219db7965e56175621dcead25602d4505bcdd7911f92f8af752

Observation 85cfd310-4490-42d3-a135-aa6de65968f7 · outbound

This paper cites A method for solving the convex programming problem with convergence rate o(1/k^2).

MARS: Unleashing the Power of Variance Reduction for Training Large Models A method for solving the convex programming problem with convergence rate o(1/k^2)

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.898101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.898101Z digest=sha256:31eadc0621df9484f515fb350926a1bfe53d5cb7abb2a1a6bec10856b2bb6a36

Observation 2845d882-48fc-4566-8ce3-ed4e2923103f · outbound

This paper cites Introductory lectures on convex optimization: A basic course, volume 87.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Introductory lectures on convex optimization: A basic course, volume 87

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.410000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.903254Z digest=sha256:b0a5167168c2a1be7a4ea63de3082f46f79eaba7293aad57e19bb99901368e5a

Observation 4ff75ab9-714f-4a67-b7f3-503bb0df5637 · outbound

This paper cites M., Liu, J., Scheinberg, K., and Tak \'a c , M.

MARS: Unleashing the Power of Variance Reduction for Training Large Models M., Liu, J., Scheinberg, K., and Tak \'a c , M

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.379053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.908303Z digest=sha256:aa184e150a9b45c16d63117eb0507071b9651a79db816ea52e2495b1f5d3fb2f

Observation 3f2a9ec6-1ccb-468c-bb19-2d2e95021808 · outbound

This paper cites Stochastic Recursive Gradient Algorithm for Nonconvex Optimization.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic Recursive Gradient Algorithm for Nonconvex Optimization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.913923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.913923Z digest=sha256:89be31257d4885c99b9cce530727373c615de18e16cdc3444116ad35109fb8e4

Observation 16aaf30a-2629-4f06-990b-25a4de6a3f11 · outbound

This paper cites Language models are unsupervised multitask learners.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Language models are unsupervised multitask learners

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.919275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.919275Z digest=sha256:f229c00f6ac88aee4234e22b4618f4e8a9af275e2c9c6d843330a62c5640c714

Observation 28c67352-ff2d-457b-b4a0-f5a51d15074e · outbound

This paper cites J., Hefny, A., Sra, S., Poczos, B., and Smola, A.

MARS: Unleashing the Power of Variance Reduction for Training Large Models J., Hefny, A., Sra, S., Poczos, B., and Smola, A

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.332688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.924796Z digest=sha256:91f1db9ddf02d25750e9f192568a7c7314079ec9b5c37490471d25613939f721

Observation 64e703b0-6003-4ae7-ac02-cbd442a56063 · outbound

This paper cites On the Convergence of Adam and Beyond.

MARS: Unleashing the Power of Variance Reduction for Training Large Models On the Convergence of Adam and Beyond

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.936920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.936920Z digest=sha256:665953f7c4741126a1bfd6d7d5e9c8b8bdd9a937ce0217ccb2615bd55c8999c5

Observation 43854652-eacd-442b-9a6d-4ecb8e5021ab · outbound

This paper cites and Braun, H.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Braun, H

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.306454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.943511Z digest=sha256:55a1d58703679ac22e2181b39feae63b09e32b4a496b2abe9e0dccd0e4dcfd37

Observation eea572f5-4d30-4393-b5c5-794629c95ca1 · outbound

This paper cites A stochastic gradient method with an exponential convergence \_rate for finite training sets.

MARS: Unleashing the Power of Variance Reduction for Training Large Models A stochastic gradient method with an exponential convergence \_rate for finite training sets

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.950051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.950051Z digest=sha256:6f3d22eab8d99881f42e8d268890e68122a1318da03b558b75feaf2f8aae71ef

Observation 7709864d-3762-4722-a6f1-245e4196a9ce · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

MARS: Unleashing the Power of Variance Reduction for Training Large Models L., Bhagavatula, C., and Choi, Y

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.262329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.955568Z digest=sha256:cd14f635c01d7d47491136e5cf47b150589566a94d14991cd46d893dcea542bc

Observation ed306893-a7be-47ee-b441-72d8d69ab381 · outbound

This paper cites English Conversational Telephone Speech Recognition by Humans and Machines.

MARS: Unleashing the Power of Variance Reduction for Training Large Models English Conversational Telephone Speech Recognition by Humans and Machines

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:43:02.356237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.960751Z digest=sha256:5e9fd9bc85b2b7837436d1c5fd947bad8ddac8ee205ae28bc4c5b85bb37b2bd0

Observation b4caae8a-a9b1-4331-a710-0b162f46e2bc · outbound

This paper cites and Smola, A.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Smola, A

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.966520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.966520Z digest=sha256:6c84114dfc741bb8c2e5d279d80526b15affdf8559f05f8ebcd10e60d36f2927

Observation 85de6cc9-7d30-49f6-baaa-a23a13a6ab46 · outbound

This paper cites Iterative berechnung der reziproken matrix.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Iterative berechnung der reziproken matrix

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.209999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.972196Z digest=sha256:4404a964949da96bf7b0d96f6aff4c33054e15dec57514183ad9848aa5e8bbe7

Observation 8bbaeb57-1417-41df-9d90-d1be52929e5c · outbound

This paper cites and Zhang, T.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Zhang, T

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.173456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.977890Z digest=sha256:0881bc0f7993b3fb2fce0a9271be37055f4f3925df5d3a77442431ea661bb2d2

Observation c262343c-682c-4473-8583-183ee134a835 · outbound

This paper cites and Stern, M.

MARS: Unleashing the Power of Variance Reduction for Training Large Models and Stern, M

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.983245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.983245Z digest=sha256:7b8c537323f6d13e025e3ea6f49f7bcaac2818fef9d2d14f802a441debccc8d7

Observation 840b886c-4ca4-4976-9d25-026b41b288c8 · outbound

This paper cites A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale.

MARS: Unleashing the Power of Variance Reduction for Training Large Models A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.988790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.988790Z digest=sha256:c992258b9f9bc3a94d365894815123ab43b004ff1037e7137672009bc2741d4d

Observation 7d974669-e5d9-4404-b931-c8068638f8c8 · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Dropout: a simple way to prevent neural networks from overfitting

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.993564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.993564Z digest=sha256:18fb0521941cdb2052640dd60a855101446b07d54679ca722ebb54cbab02c62c

Observation cfa80dc0-fdc0-4942-8de4-99571633fcc4 · outbound

This paper cites Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.107413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:01.999285Z digest=sha256:06e6e1802de788318cafc5f99e08c5dbbe97a2ee4cf4ec01b3b89ecc32437c9c

Observation 0024312a-7550-4bd0-ada2-61fad6a32166 · outbound

This paper cites Attention is all you need.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Attention is all you need

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.005034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.005034Z digest=sha256:fa56ce21f3a5896cc79c16102c52e0e30f6e2b04f5d089ba0265cb98317ac96d

Observation 4a752e3d-2adc-4526-a48a-4cfed46f2a60 · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

MARS: Unleashing the Power of Variance Reduction for Training Large Models SOAP: Improving and Stabilizing Shampoo using Adam

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.010064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.010064Z digest=sha256:509e5230336e645b8e91b5aaaebd9a3ecff1f3cdf3cd9674d4e29358d80c0096

Observation 012af07a-9e4d-4b3d-836a-77b90ec463da · outbound

This paper cites Spiderboost and momentum: Faster variance reduction algorithms.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Spiderboost and momentum: Faster variance reduction algorithms

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.075581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:02.016660Z digest=sha256:33b5c33590239bdb4081f0c10ea876918237ba54a42ce1625f8da7025ba39ae8

Observation b3dc8bbf-c6f3-4315-8dad-6e039b3dea6a · outbound

This paper cites Adagrad stepsizes: Sharp convergence over nonconvex landscapes.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Adagrad stepsizes: Sharp convergence over nonconvex landscapes

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.023654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.023654Z digest=sha256:92cd8d6b5e537b2b8429c097992307caaa6becd72e62ae69c9f82e5ba9db7e85

Observation be36d960-c952-4b93-986c-5cdb9aa09499 · outbound

This paper cites Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.032612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.032612Z digest=sha256:a73893bf461deb0ab5045aabb9721663cfa2a80351ad101fa6d15da2ee26e7cf

Observation bf3a7227-59be-4ac9-a24e-864cf24dc490 · outbound

This paper cites Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.039030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.039030Z digest=sha256:23497083267dfdd083d3e3f36b757262a6b17aea19c8aa55ca99855251717634

Observation 08c21778-335f-47a5-9962-26cecaa098da · outbound

This paper cites A Coefficient Makes SVRG Effective.

MARS: Unleashing the Power of Variance Reduction for Training Large Models A Coefficient Makes SVRG Effective

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:43:02.285588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:02.044225Z digest=sha256:1c921bfe3e6e1654d52fe51ea457de947dd9122ce2a6fc1f02d4fd4e03895199

Observation a3c723bf-65ea-48d1-8cee-0424688e03d4 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.050371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.050371Z digest=sha256:a6a03311777e2f84e74973f5a79e26c9cd397e3eecc011bb131621dfbf24280c

Observation 5d6dde87-279b-4b4a-b5b7-79bfaf0dd268 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

MARS: Unleashing the Power of Variance Reduction for Training Large Models ADADELTA: An Adaptive Learning Rate Method

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.057300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.057300Z digest=sha256:f1d9883d3da1bba8a8d05a2084267856f52ea77d1e1b99ef6b8ffd5577e4d0d9

Observation 11b4d049-9e91-4b71-bc3f-4cd504d7f7fc · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Korhonen, A., Traum, D.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Hellaswag: Can a machine really finish your sentence? In Korhonen, A., Traum, D

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.063883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.063883Z digest=sha256:7ff325444f376255f7ec4c7e59060d9fd48273452e31d8e2c554e6d2d7da1cd5

Observation 2fd7335a-5e17-4e65-8923-dc475745f55a · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

MARS: Unleashing the Power of Variance Reduction for Training Large Models OPT: Open Pre-trained Transformer Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.069466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.069466Z digest=sha256:fc0783862f42547100bbbfbda5880310a185428b4549872a2012772e3bd7e0b0

Observation 107f6011-4904-4dc7-b21d-236b708babfd · outbound

This paper cites Adam can converge without any modification on update rules.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Adam can converge without any modification on update rules

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:03.004859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:02.074897Z digest=sha256:f7f6babcdafc788fcc6495765e944f2a73845a542ca0a62bb8a2f08901b831b0

Observation c62ad04f-5834-4868-a3e1-459c08c6a12a · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Deconstructing What Makes a Good Optimizer for Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.080704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.080704Z digest=sha256:bb83de3ab6ef27604a1b539577499a947201b41396315794a0bec672dd75dbda

Observation 5e1daf55-ba58-4897-a71d-419a408daae3 · outbound

This paper cites Stochastic nested variance reduction for nonconvex optimization.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic nested variance reduction for nonconvex optimization

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:02.984590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:02.086996Z digest=sha256:3a5b63020be0ee120d86ed3391959a338686df9e8a3ebbdc5fc3d4cfbb7725fd

Observation 54a450a1-df44-4fc5-b952-58b06cdd5cb6 · outbound

This paper cites Towards understanding convergence and generalization of adamw.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Towards understanding convergence and generalization of adamw

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:43:02.960616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T19:43:02.093463Z digest=sha256:f0db94abfce61ca919754593e6fca491ba6525ca2d54786befd0f17eda14453c

Observation 5465fbce-eb1c-4d54-a499-e15f24becaf1 · outbound

This paper cites C., Dvornek, N., Papademetris, X., and Duncan, J.

MARS: Unleashing the Power of Variance Reduction for Training Large Models C., Dvornek, N., Papademetris, X., and Duncan, J

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.099002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.099002Z digest=sha256:f1589ad544a59f06e56afafcfaf344497e6b7bc985c1b5f067484a7616544ffc

Observation ad9cbca0-2e2f-4589-ae04-316cdc5154a9 · outbound

This paper cites @esa (Ref.

MARS: Unleashing the Power of Variance Reduction for Training Large Models @esa (Ref

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.104287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.104287Z digest=sha256:d2db49e4d7ace679aef25c0fb2da97e62cd12aaf6c3c157bfa625c6ced34a0ae

Observation 209059a6-ecd1-410f-926f-bbbd5bc28d06 · outbound

This paper cites an unresolved cited work.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.110478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.110478Z digest=sha256:be715335d7b124774aacb9bd3c23e5f97591fb4e9ee31586f452fbf272f27c07

Observation bafc3be5-4fca-4f01-9a61-c7cbb5b25d94 · outbound

This paper cites an unresolved cited work.

MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:02.115828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:02.115828Z digest=sha256:9e7faa5e891b4c063d043deff6dd30b56243558271903b933d09507aa238f859

Pith citing papers

Observation 82111d87-d0d8-4a25-a1e8-c4c8250ce563 · inbound

Physics of Skill Learning cites this paper.

Physics of Skill Learning MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:20:51.021889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:20:51.021889Z digest=sha256:5d5bd084e7a53b5de3395955a0b87df76b0536c5a9099286adb814ab7110a247

Observation b5c46eee-9c29-44dd-865b-18abd9153a49 · inbound

Training Deep Learning Models with Norm-Constrained LMOs cites this paper.

Training Deep Learning Models with Norm-Constrained LMOs MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:22:37.163437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T21:22:36.870292Z digest=sha256:bc89bc5466c5cf09003703cd07fa04d32b9330edb05e42dca0ffe4110da8668b

Observation 073a7acf-1bb7-471c-953e-9a2431184c4a · inbound

Continuous Cardiac Arrest Prediction in ICU using PPG Foundation Model cites this paper.

Continuous Cardiac Arrest Prediction in ICU using PPG Foundation Model MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T04:32:13.501247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:32:13.501247Z digest=sha256:2ae2f7a7854d3161bdcd73f37000e94bb3ae224e2e0599ed9c648403d6f8c1f4

Observation a2b38aa2-0c1f-4d77-81d3-f81653fca415 · inbound

Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise cites this paper.

Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:37.720082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:37.720082Z digest=sha256:b400c13ee11b7cfb41a55316fa83a685ff9d51dfa95f21276eebf996d16e1d7e

Observation a11c4c39-2f8d-4a92-a779-1caf95de60bb · inbound

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD cites this paper.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.857286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.857286Z digest=sha256:d5dfe436b566c1c8908e36507de9af7e072824d8f6c69d2b5f77f907b0b79fe2

Observation 96305ccd-7657-4bf1-b514-55b65d9d72a7 · inbound

On the Convergence of Muon and Beyond cites this paper.

On the Convergence of Muon and Beyond MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:56:34.038686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T15:56:30.602824Z digest=sha256:d326ec9afdd7bfa7ea4d498d97371cac40527b5f0ea415cf7a7da86c0eac661d

Observation 837ee271-41b3-4b54-aa56-2611a1f51e82 · inbound

Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization cites this paper.

Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:46:12.512191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T09:44:53.004293Z digest=sha256:62080dd412d8740dba09256d6cdc490c6683f1e361ccf9d6d3f0577b47be7cab

Observation 9d406b4a-3007-4676-bb2c-6be5b7626b11 · inbound

Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum cites this paper.

Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.532381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T08:12:57.430291Z digest=sha256:c06a19ebeac97c2acf5fc8a11506ccb0bd3f368754f9a437ec5fa8bb1d23315c

Observation a767a7f1-1c97-42be-904b-2253a4f48c5c · inbound

HTMuon: Improving Muon via Heavy-Tailed Spectral Correction cites this paper.

HTMuon: Improving Muon via Heavy-Tailed Spectral Correction MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:55:26.336736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T06:50:29.893126Z digest=sha256:8eedf5bb47c436fa1b1cfe48984dc6922e2639b3de4b0d08293b77913aad4f33

Observation 8f4fb87c-5a45-480f-a4f1-3e547fba4e02 · inbound

CLion: Efficient Cautious Lion Optimizer with Enhanced Generalization cites this paper.

CLion: Efficient Cautious Lion Optimizer with Enhanced Generalization MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:55:20.585950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T11:53:28.411087Z digest=sha256:a581cf81ea7a13aa8548b0e99e48dad2003e74fa889db0b7b746aebf0f7a960f

Observation 020babd5-7115-47a9-b824-8019276a58f3 · inbound

Personalized Federated Learning for Gradient Alignment cites this paper.

Personalized Federated Learning for Gradient Alignment MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:50:26.268022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T19:32:49.261153Z digest=sha256:93109147e165546187419f5d8195bc4a0b31de3512921980d7ea85580bb6e071

Observation 538e25c1-2441-4600-8506-41ad000ad01b · inbound

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less cites this paper.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.551157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:e493633cff9214fa4d973fea8a51f285901ef3ae6ce531fc6f5e1ab18cfa5841

Observation 1c50ea8b-05e4-4edd-a635-6d520343531d · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 165

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:38:11.033711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T09:34:45.186929Z digest=sha256:5d4dacf6e26c71a42fdef0df29c3a8aa65cb76063a7a9e26d0c6d886af0c074c

Observation e4e9ed29-c096-40da-a023-b20670e9989e · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.292382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T18:42:01.854481Z digest=sha256:569d37b343805fc9e187073bdd4fb6a2cc5fffa46caa6e67b015aa773ba8a034

Observation 8c1a22ae-c71f-4321-af06-e8f23ac2b9cf · inbound

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics cites this paper.

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.303623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-22T08:06:52.309619Z digest=sha256:6ba855bcd22c25106935b7fe5d7c2fee44e6989943362503a9baa48597bfcd92

Observation b4b54b21-6276-4641-8006-bcd108a796c8 · inbound

MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic Optimization cites this paper.

MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic Optimization MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.220618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:31:53.550342Z digest=sha256:3b06534a9b364ad31297ab3304cbd2dae32c9224254d548f25c2c15d8b36fe88

Observation 72e002c3-65e5-4a39-875c-59bb0765fa72 · inbound

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining cites this paper.

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:18.839572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T06:41:55.732230Z digest=sha256:2e8e0ceb89e7367475e37c73bcfd1f8d8b27ba645e26d0f8dc0c0e6c2b6fab52

Observation 92f28a8e-159e-4574-89fd-27f8a1ea7be4 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 128

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:c28f22ccb3db90833a70e5104e5b03fe1220b86bb8b3d08cd3310cf8ad9dcfd4

Observation bd0f53ac-6b1c-47ec-96bf-825f239333de · inbound

Muse: Representation Geometry of Muon Beyond Normalized Momentum cites this paper.

Muse: Representation Geometry of Muon Beyond Normalized Momentum MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:58:58.664681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:58:58.664681Z digest=sha256:27a380246927b63fe383a685a698865f2f3c360f7fec26565eadced3650feae4

Observation 0f32997d-8304-4d9a-9fa0-e4c2397f32b3 · inbound

OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining cites this paper.

OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:30:38.237293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:30:38.237293Z digest=sha256:979e77911519e92411805b2f4ad4789a2ba6d9126b27f010c96eaa7507039b65