Pith. sign in

Paper Citation Record · LEDGER

Memory-Efficient 4-bit Preconditioned Stochastic Optimization

As of 12 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2412.10663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10663 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:50:24.396837Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3baabe3b-ea9d-4129-8036-9a93820a5ae4 · outbound

This paper cites Disentangling Adaptive Gradient Methods from Learning Rates.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Disentangling Adaptive Gradient Methods from Learning Rates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.038813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.038813Z digest=sha256:1510af93af1072595634d5e6478c374a34df3e9d3ceb05c2301e985afc8e647f

Observation d1f9b074-f827-40ff-a6c8-66c8a94d6cc3 · outbound

This paper cites Qsgd: Communication-efficient sgd via gra- dient quantization and encoding.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Qsgd: Communication-efficient sgd via gra- dient quantization and encoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.681430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.044650Z digest=sha256:667897e18d5825bab74b3ea29aacd7e2318aae80f4e15ad945ebb57d2e8443b0

Observation ef0786a3-b958-43c2-87fa-15b60fbfcef7 · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Scalable Second Order Optimization for Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.049744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.049744Z digest=sha256:2ce3c6df00c2d0cf8c93a8bc6d881e8341e126b7d4821bd7fb2f73558bd03ec0

Observation 54b4f2df-400d-4138-a2fa-08bb085dd44d · outbound

This paper cites Stochas- tic approximations and differential inclusions.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Stochas- tic approximations and differential inclusions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.660354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.055011Z digest=sha256:482c47c0461a10dd49cec468c5606ab3666f6ab432e68b5169963bed4733d9ab

Observation 2f77c16d-6a36-455e-a837-ce2365108578 · outbound

This paper cites Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.643518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.059985Z digest=sha256:e09c4875526a15f92b4b30d3a89f2144fefc340a2628eedb0f8cf622e33b8361

Observation 5e71a2e1-7644-48ae-ac6a-f08c8ff71905 · outbound

This paper cites Stochastic approximation: a dynamical sys- tems viewpoint.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Stochastic approximation: a dynamical sys- tems viewpoint

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.618021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.065354Z digest=sha256:748bea45765e9d9dc175f905122520a00d574f99cb2c0f2518741e5218a02864

Observation a2e85daa-e140-4eaa-afc7-71be80c143eb · outbound

This paper cites Language Models are Few-Shot Learners.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.070170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.070170Z digest=sha256:f738744c107c1f93f8e1423e058bc9da714544c59852cdefd99f9f46f297bb8d

Observation 0c3beb52-263d-4181-97d1-8e84ea942ce5 · outbound

This paper cites Lower bounds for finding stationary points ii: first- order methods.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Lower bounds for finding stationary points ii: first- order methods

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.596438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.075803Z digest=sha256:036382252d888dfb8e3a23209e1be5cd8e57905c667415f150a09af500e6cde2

Observation 219db309-1fd4-4e00-86d8-e886a1485330 · outbound

This paper cites Optimization and nonsmooth analysis.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Optimization and nonsmooth analysis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.571868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.080337Z digest=sha256:19cf3a7347a8a863adb53285e638201c83347cfdda01e68c4d30e33f872935e1

Observation 044afc28-ef1d-4aa3-a969-8aab1f5ce6d8 · outbound

This paper cites AutoAugment: Learning Augmentation Policies from Data.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization AutoAugment: Learning Augmentation Policies from Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.085329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.085329Z digest=sha256:0642bb69b7990393c5e539dcb29ed32d1abd2f8d66bf6a3ed3be810c5a090045

Observation 4c7b912a-605b-4cc0-9da7-d29d372c7f5b · outbound

This paper cites Randaugment: Practical automated data augmen- tation with a reduced search space.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Randaugment: Practical automated data augmen- tation with a reduced search space

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.551165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.091242Z digest=sha256:2ebc7918d7bcb2c82320154e7b7d96b3864d97208bef609d984256e99f0d85cb

Observation 6d53f9a6-1289-4727-bab1-ffa30862597d · outbound

This paper cites Pathological sub- gradient dynamics.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Pathological sub- gradient dynamics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.525584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.095983Z digest=sha256:7ddda14bb406224bd7e54c2041571b1d96c7d5b94815b8d33de67ee83e214e9a

Observation 8103e450-7622-49ed-a284-260752fa330a · outbound

This paper cites Stochastic subgradient method converges on tame functions.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Stochastic subgradient method converges on tame functions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.502286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.102733Z digest=sha256:470a65ea44f7bf5786a5a58a256742eb2339b47dfaabfb8b71ad8eadc542e86c

Observation 74e2a6d3-c45b-44a7-9f0a-62fc56fe814f · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Imagenet: A large-scale hierarchical image database

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.110579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.110579Z digest=sha256:ca45abe3fd4e9b12ec1e2234495d39faaaa523934480f611a362c92683ce1736

Observation c3a1db67-6fec-43e1-bcb1-1f138d582f8d · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization 8-bit Optimizers via Block-wise Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.116174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.116174Z digest=sha256:2b766c5e832824b2d906eaf8a6f9d4782d686cced499eb40432d9cc7e0d9859c

Observation cb132225-4516-47d1-9830-d914a8a3fa9c · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.123963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.123963Z digest=sha256:31ebc0a40e697636015d23a5922b8b51be29c7b33df77565392b880419cc658e

Observation f36936a6-d005-4d6f-8c47-47a1a693c816 · outbound

This paper cites Adaptive sub- gradient methods for online learning and stochastic opti- mization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Adaptive sub- gradient methods for online learning and stochastic opti- mization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.456438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.131062Z digest=sha256:32607201e72c03d4886afd6281c08b8b5ef4c74f6e65437c4370774f595b39ca

Observation 83b72dc2-c9a3-4805-a166-c81408a46899 · outbound

This paper cites Stochastic methods for com- posite and weakly convex optimization problems.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Stochastic methods for com- posite and weakly convex optimization problems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.437387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.137144Z digest=sha256:3f3aaf482fae83b5d6c23eba84f97bb516b14c3b494920bf001b7eeee3037d81

Observation 83881c3b-dbff-440b-b03b-4a1640ce0314 · outbound

This paper cites A survey of quan- tization methods for efficient neural network inference.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A survey of quan- tization methods for efficient neural network inference

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.415832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.142495Z digest=sha256:f8289f011c29d2847b5c3c46a5bebe9a8e3e882f7143064f403f95b0247b4cde

Observation a85be752-776d-4b3b-9302-cf9ca20aa6e7 · outbound

This paper cites Practi- cal quasi-newton methods for training deep neural networks.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Practi- cal quasi-newton methods for training deep neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.396650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.147386Z digest=sha256:e1354bf57678d21ce5e4c5877971da9244c871e460392be1469b6d1f1fc1d9d6

Observation 4f6e5407-8681-490f-b325-faeedf0713f0 · outbound

This paper cites A schur–newton method for the matrixp th root and its inverse.SIAM Journal on Matrix Analysis and Applications , 28(3):788–804, 2006.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A schur–newton method for the matrixp th root and its inverse.SIAM Journal on Matrix Analysis and Applications , 28(3):788–804, 2006

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.378818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.152394Z digest=sha256:519e589e9d4ac47b5664ebd23aeefdae7166941f864ddc5c5d2f39305d9c7efc

Observation 6e4e80f8-2677-4a8e-9401-40c5f8321e56 · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Shampoo: Preconditioned stochastic tensor optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.362689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.157728Z digest=sha256:42fc6171422918638d067cebdad3403e6814b474f5c8fe2f2dd010c2f9a18976

Observation 5240ece3-6e13-42db-a26f-e792f39bf890 · outbound

This paper cites Deep residual learning for image recognition.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Deep residual learning for image recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.345719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.162986Z digest=sha256:461237b7136fab7e1a493a26cf4d27b9be1fd3b13869b887df6c1789531176ff

Observation 0e2e08b5-2303-43f5-aa24-323fde4246d3 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Training Compute-Optimal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.168423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.168423Z digest=sha256:ce698e89b3d5b3d3f8d8ce117410bf6ccc183e48dc11cbf0e163f373b2f815de

Observation d15b6993-2368-44b3-aed2-a014844042e7 · outbound

This paper cites Scaling Laws for Neural Language Models.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Scaling Laws for Neural Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.174002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.174002Z digest=sha256:a078089907f02554d3dd7091c69a2a1086d61a5ad6199fe4f1203486acb6f4d8

Observation 284003b4-e4c2-43ec-a958-e14673caa2df · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.179731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.179731Z digest=sha256:5d795bdee25d2628662eb7244abe8f4a1b63158d85038cb1111d1d5e1462c51b

Observation 508fc494-7374-4f12-8b47-ed211b1650cc · outbound

This paper cites Learning multiple layers of features from tiny images.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Learning multiple layers of features from tiny images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.185289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.185289Z digest=sha256:4fbbdcb3a8b05f0331b0807a9a01a98856e4b826ebfa91d0ce30a20e588ee88e

Observation 4abdb69c-71f3-47c5-b718-afe0724d1c5e · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Imagenet classification with deep convolutional neural net- works

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.190866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.190866Z digest=sha256:71eec19cc0db47032c5ea6f6a4eb48000b53449f941b247b8593bb6d50e93a60

Observation c4002d35-b1a8-414e-9d1b-0928c4dbfaa7 · outbound

This paper cites Tiny imagenet visual recognition challenge.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Tiny imagenet visual recognition challenge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.196055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.196055Z digest=sha256:85b774b009c41990af1f5e0a8646963fd7a12f5d17de58f9bef92ca2f2fd0651

Observation 3e485a9b-809e-4cde-8b23-87d1fe47ac09 · outbound

This paper cites Biobert: a pre-trained biomedical language representation model for biomedical text mining.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Biobert: a pre-trained biomedical language representation model for biomedical text mining

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.284735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.200658Z digest=sha256:85f1f19a47b8fcad0b67b15b269131a35287e1c5e0ce78e28b5695a84dd161c4

Observation 6c4c0482-d15c-4101-8629-360cf7d432ea · outbound

This paper cites Vision Transformer for Small-Size Datasets.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Vision Transformer for Small-Size Datasets

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.205713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.205713Z digest=sha256:ff2ab4ced0b002551978f73448633f4e60f641fcbf7c571247588ac39d4243f0

Observation 71d89b49-da1a-47fb-9fe4-ca8648433408 · outbound

This paper cites Memory efficient optimizers with 4-bit states.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Memory efficient optimizers with 4-bit states

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.268482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.210780Z digest=sha256:8da744e391ddc896fe281bb98363f6d70829b3d398ed8f1c6e2072bdfb090ecc

Observation bd284cb7-e48f-44f4-91c2-705c9cac06ff · outbound

This paper cites ReLoRA: High-Rank Training Through Low-Rank Updates.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization ReLoRA: High-Rank Training Through Low-Rank Updates

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.215832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.215832Z digest=sha256:72d43b65de560ee2a904699163c3231145aa60a87780ee23095ab287e9346373

Observation 1229df2f-866f-4300-8857-8c7fff14e9b3 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.251290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.221264Z digest=sha256:eafac6472cc2ccd01afafa06ea079353dc392bf2795de02ecb4ef1aa3ffbfcaa

Observation d549d709-24d2-44c8-9b15-3d8a98ff84fc · outbound

This paper cites Decoupled weight de- cay regularization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Decoupled weight de- cay regularization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.235801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.227137Z digest=sha256:c1cc1cb1f63c18980e3753c50bfc741eddedb70c2e95e679765c1a146282134a

Observation 530cdb4b-f617-4bc3-a572-5ddbbaa0ceb3 · outbound

This paper cites Optimizing neural net- works with kronecker-factored approximate curvature.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Optimizing neural net- works with kronecker-factored approximate curvature

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.217693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.232474Z digest=sha256:175fa194b2e94deb0279ab327d3f348e11f44da2c2a4194872053d3bcba4887a

Observation c6ddb572-8caa-4109-8cd5-b596d7cece6a · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A New Perspective on Shampoo's Preconditioner

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.237399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.237399Z digest=sha256:65127d4f636b6d65b27ffd043a1cd29417c126246c20a0679b2e8cb16a175c3f

Observation 491079d2-2698-4a27-9360-c1acb16f9854 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.243153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.243153Z digest=sha256:db6baa568ccf557c13b071ae5b919e18925d7befcfaa19682e956e7c95952960

Observation 8a9e99e1-4006-4404-9ccb-b7571a4def38 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.248004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.248004Z digest=sha256:0741949521304421efe801f7dab963366bb79d71dbc194026e04445464536bf4

Observation a77dfca3-3598-4f5e-927d-6fbfa004e624 · outbound

This paper cites Ef21: A new, simpler, theoretically better, and practically faster error feedback.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Ef21: A new, simpler, theoretically better, and practically faster error feedback

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.171369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.252774Z digest=sha256:f4e3cbda304763ccbf544429647e7a643b98004987fb1ee4ca934d0d1101dbff

Observation 82b856b9-e98d-4d56-afe7-03f7d4924448 · outbound

This paper cites A stochastic approxima- tion method.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A stochastic approxima- tion method

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.151379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.257589Z digest=sha256:54f8a43f9f2b0e5623c05230ebb3f39c4e88d4dabdd94134895824086299f24b

Observation 9c3e612e-aed1-405f-99e7-9ff27088862d · outbound

This paper cites 1-bit stochastic gradient descent and its application to data- parallel distributed training of speech dnns.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization 1-bit stochastic gradient descent and its application to data- parallel distributed training of speech dnns

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.134140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.262743Z digest=sha256:3703ffbb778f86b55ea9bd9f8871d19113393947f877a9bc7fc6418b43a3d2bd

Observation 4da20aa6-8ab7-4c1e-8e21-8c47e36fb61a · outbound

This paper cites A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.267978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.267978Z digest=sha256:682595f1dc51372a72f1dd0974ec805b2edcab54005d606cfe30e52160b763ac

Observation ff676b04-675c-4a94-8bdb-b603b6613c71 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.272997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.272997Z digest=sha256:25e8de6b845a4ffae0cd67afe4d49d96273f361878dcd7d163504a385ff29f3f

Observation 94ae2857-c901-4213-8b4a-5a185fbe7f38 · outbound

This paper cites On the importance of initialization and momentum in deep learning.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization On the importance of initialization and momentum in deep learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.113071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.277637Z digest=sha256:eace1e2008ad42935648c29affd8bc7f7d01a094ede481f32076fce4a97d3de1

Observation 11f7b997-8f6e-4413-8fd2-6457c0d2fdb8 · outbound

This paper cites 1-bit adam: Communication efficient large- scale training with adam’s convergence speed.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization 1-bit adam: Communication efficient large- scale training with adam’s convergence speed

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.092633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.282445Z digest=sha256:e11edfd7a6e26de3e8a3a95bed93eadef8443860d2950e5eb5c81d48d720a9ca

Observation d4ea303b-9cd2-427d-8ae7-3aa68cf20ed3 · outbound

This paper cites Lecture 6.5- rmsprop: Divide the gradient by a running average of its re- cent magnitude.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Lecture 6.5- rmsprop: Divide the gradient by a running average of its re- cent magnitude

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.070539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.287286Z digest=sha256:eb9a4eb1a55a14ceab7e741a388bc69349e6d8f467b8734985123ffd85046854

Observation 83951d8a-9aa4-41c5-b9a4-294e56bef14b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.293143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.293143Z digest=sha256:e4a16e56af1891346da2a349ee532dd9bf91fec6e26c5b6f000f9411402534c1

Observation 9a9287f3-272b-40e3-a161-a985cca76413 · outbound

This paper cites Powersgd: Practical low-rank gradient compression for dis- tributed optimization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Powersgd: Practical low-rank gradient compression for dis- tributed optimization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.045328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.298708Z digest=sha256:1f10911f403add3e26860db6ebb2d671a0f7f3b9fcfe5a131ec7cb92b570d52e

Observation bd7348a7-9db4-4a20-bbd2-d939d6120adf · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization SOAP: Improving and Stabilizing Shampoo using Adam

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.303608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.303608Z digest=sha256:9ba417f43ed16e84576312e7c91cfa422f81f81db586a650c0da126c0139bd1b

Observation 2d1e864b-319e-4c3e-a327-d32a25ef9dbe · outbound

This paper cites 4-bit Shampoo for Memory-Efficient Network Training.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization 4-bit Shampoo for Memory-Efficient Network Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.308787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.308787Z digest=sha256:a3fa5eeac2bdc4710f8c9f10339e0534a0c1cd515bd77f1030b5dfc55cf80da7

Observation 2c9988b1-a79c-4bd4-b47f-251750f56fd6 · outbound

This paper cites Terngrad: Ternary gradients to reduce communication in distributed deep learning.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Terngrad: Ternary gradients to reduce communication in distributed deep learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.025285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.313567Z digest=sha256:ee8ad9ac4bc0ccafbd9bd716ad5bfd5d7dbcf8c7343bfbe8f066e6c888d1aaff

Observation 0d779a2b-6e52-42c8-98de-36510fe533fe · outbound

This paper cites ResNet strikes back: An improved training procedure in timm.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization ResNet strikes back: An improved training procedure in timm

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.318482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.318482Z digest=sha256:a27a128576117c53758569ece2d0b518ad6f79450f5b2e82e1646fb4712c9236

Observation 837dac85-0c35-4d77-8720-c24c9d328772 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization BloombergGPT: A Large Language Model for Finance

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.323992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.323992Z digest=sha256:5505cdb59aa876cd4f9521f0a08207d9f199f6dcbf244e9dc3069686f6612743

Observation 9f2669cf-3312-4c83-8134-f4730221d530 · outbound

This paper cites Crystal graph convolu- tional neural networks for an accurate and interpretable pre- diction of material properties.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Crystal graph convolu- tional neural networks for an accurate and interpretable pre- diction of material properties

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.007259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.330016Z digest=sha256:b2dc92f171de567ca82e2e22590ab64569ec4d2a769491603e8ae55b856355dc

Observation b532440f-c87c-43b1-9ed8-218200bdd7f0 · outbound

This paper cites LoCo: Low-Bit Communication Adaptor for Large-scale Model Training.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization LoCo: Low-Bit Communication Adaptor for Large-scale Model Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.335113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.335113Z digest=sha256:82c23e2091688af68c4bc0eca5d7b6dd13f5c96e4e8827471d9ba8ceac84da40

Observation 0514f0bb-b1c6-4db2-a352-774fcb21dfc7 · outbound

This paper cites Adan: Adaptive nesterov momentum algo- rithm for faster optimizing deep models.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Adan: Adaptive nesterov momentum algo- rithm for faster optimizing deep models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.990483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.341648Z digest=sha256:ebc9da899db491baa6c11532700d7560bbbe97d4f795c81a1f1ebfed8ef0e8d4

Observation 6c659bcd-1664-45b7-a297-0384a93ec693 · outbound

This paper cites Zeroquant: Ef- ficient and affordable post-training quantization for large- scale transformers.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Zeroquant: Ef- ficient and affordable post-training quantization for large- scale transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.347114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.347114Z digest=sha256:ca9df71892a942171ea2b86f6ddb31969006577b8b8372d1eca09c5583fc1a0c

Observation e45e018e-1712-48c0-b86c-642f11f8746d · outbound

This paper cites A general regret bound of preconditioned gradient method for dnn training.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A general regret bound of preconditioned gradient method for dnn training

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.960505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.352290Z digest=sha256:4fa7578aff8a3e7e80623b278740420ee6ab259cd27eed443e5fd5645cf0e5ec

Observation 103ed47c-d57b-4d5c-a7ec-531d6af374ad · outbound

This paper cites Cutmix: Regu- larization strategy to train strong classifiers with localizable features.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Cutmix: Regu- larization strategy to train strong classifiers with localizable features

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.943136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.358057Z digest=sha256:385116fc605da984b327106ca18a69b7d0b65311bcd58a7d511c498b77fa7746

Observation 167cede4-e469-48ae-89f0-4b21c3efc5be · outbound

This paper cites mixup: Beyond empirical risk minimiza- tion.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization mixup: Beyond empirical risk minimiza- tion

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.926154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.363316Z digest=sha256:ce72ba86c92b43d08a83ae2514fc421b2b33e6e05d462a6a50eff09fada41a7a

Observation f5e70d96-6696-41f9-abb7-21b535705d11 · outbound

This paper cites Why are adaptive methods good for attention mod- els? Advances in Neural Information Processing Systems , 33:15383–15393, 2020.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Why are adaptive methods good for attention mod- els? Advances in Neural Information Processing Systems , 33:15383–15393, 2020

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.908110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.368363Z digest=sha256:afe8b1f1deb1319bf302203d6782eaf6bde243b6d2e53012add34f4844576f95

Observation 9948481d-ec79-41c1-98c9-912879ad5dea · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.373990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.373990Z digest=sha256:e08631ddb7ff055ef4550d3a9298a9599194ff99dc0059c07963289d90fc96cf

Observation 3f7081ec-f753-40ea-8b75-b1bdf44ee7bb · outbound

This paper cites Random erasing data augmentation.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Random erasing data augmentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.890889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.379527Z digest=sha256:f5db586f6cd6e8a9680bccaa57d5653613cbbcf0b934694290d3fa968b01cbd1

Observation e543c5ac-48f1-47ae-ba25-6244b883ca8a · outbound

This paper cites Definition B.2.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Definition B.2

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.871998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.385998Z digest=sha256:4645869754d0801350c5d5e6e623a92036c9bf5518b663885824e8760ef9a8b9

Observation 2c018be2-19f1-4e67-9b6c-2a4afb0248bc · outbound

This paper cites an unresolved cited work.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:50:24.853790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.391447Z digest=sha256:9783c73f1fa77455828f7064c233e521827a40a926412fb6d1c05fd057e49537

Observation cacbd0aa-ca71-4ff6-9f9e-6952fac1849c · outbound

This paper cites Here NMi is the normal space of Mi.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Here NMi is the normal space of Mi

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T15:50:24.834907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:50:24.396837Z digest=sha256:cda482585043efc2e4545f1e131417f6a138c2cf8e4f2846a13b0d74ab370c86

Pith citing papers

No inbound Pith citation observations are available.