Pith. sign in

Paper Citation Record · LEDGER

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2507.17501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17501 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:54:58.891090Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce36854f-edba-4a91-9615-b9841d518ab7 · outbound

This paper cites Layer Normalization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.796795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.796795Z digest=sha256:cecbc8bd2a56772b14e653a848e4b11a471eead63525a6f770bda9f895e18c17

Observation d219215c-3b81-437b-81be-c36abb8427cd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Adam: A Method for Stochastic Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.818406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.818406Z digest=sha256:81b3472048f03fecbeb770231a05ebeaf37672f446101d81382a8151ee97404b

Observation 05b68beb-efea-4b68-b80d-19d4b1057d2d · outbound

This paper cites DeepSeek-V3 Technical Report.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.823828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.823828Z digest=sha256:bd4e1c47813c3bbb73517bdef61b258c18f82073208afa5af70354be84643896

Observation fc2c3735-121c-44aa-80ff-0ef1832a9784 · outbound

This paper cites A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.844812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.844812Z digest=sha256:b149ab64d6f5a5e83c3710ed7914f75b7e51dc5947762f4b21f2c66bff5f4cda

Observation c3f65aae-289e-4c31-be1a-4d27a400a888 · outbound

This paper cites Large Batch Training of Convolutional Networks.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Large Batch Training of Convolutional Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.851371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.851371Z digest=sha256:e9fb2cbb251f3329ab57668550b74c7cb9075e22eac5dca28511594fad514f21

Observation a4a68fd6-09ed-4b1f-b3d8-372136ec03ff · outbound

This paper cites Transformers without Normalization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Transformers without Normalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.862382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.862382Z digest=sha256:23acdb76f2363d39d3f93bb8bef8f79e0c20aa6996ac01e9a3345d0403b0c896

Observation ca7a4ae5-d782-4073-8b6c-93d8e67b6134 · outbound

This paper cites In a backpropagation, since we have obtained ∂L ∂vec(Y ), we would like to further analyze ∂L ∂vec(Wq) , ∂L ∂vec(Wk) , ∂L ∂vec(Wv).

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD In a backpropagation, since we have obtained ∂L ∂vec(Y ), we would like to further analyze ∂L ∂vec(Wq) , ∂L ∂vec(Wk) , ∂L ∂vec(Wv)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.233831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.868032Z digest=sha256:fe6c6a8185b39a604060d6dbfb3c0f10ad4f69be579e13bb17d9fb55f03650c1

Observation 55947400-400a-4948-8d24-66b547135587 · outbound

This paper cites Défossez et al.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Défossez et al

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.160647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.891090Z digest=sha256:15790f0221522a1572952d9bfd35594df859c499a357e1822e981ac64b17331b

Observation 38da0786-31b5-44cb-a816-8b63bc9fd89e · outbound

This paper cites This method augments the gradient direction with a fraction of the update vector from the previous step, allowing faster convergence and helping escape local minima.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD This method augments the gradient direction with a fraction of the update vector from the previous step, allowing faster convergence and helping escape local minima

Reference 1983

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.180836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.884769Z digest=sha256:710ce2370a4e8b63c1b523708d41808e8c28cfbfcd6b50a9f3b3f794f1627a3c

Observation e20bf3e9-e019-4bb4-8551-e98612ef886b · outbound

This paper cites Qwen Technical Report.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Qwen Technical Report

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.834431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.834431Z digest=sha256:971d77ebec78a0f9cc7e87072ea99a326af2df00b09967351123f7327462183e

Observation a11c4c39-2f8d-4a92-a779-1caf95de60bb · outbound

This paper cites MARS: Unleashing the Power of Variance Reduction for Training Large Models.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.857286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.857286Z digest=sha256:5cd904c19b953365b0d80ff277bc78cbb0a75217b0181c53215466de5f6f1b77

Observation e4d077cf-903a-497f-9a1a-68817dae4d15 · outbound

This paper cites Language models are few-shot learners.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Language models are few-shot learners

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.802656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.802656Z digest=sha256:af87371a9fd53360e73241e5e83905febd14f7aa3c484de57e91844abaf91d56

Observation bfe13c81-6ebe-4925-ab8b-32859137c785 · outbound

This paper cites All language models were trained on OpenWebText, using GPT-2 tokenizer.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD All language models were trained on OpenWebText, using GPT-2 tokenizer

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.215255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.873604Z digest=sha256:c99fb80baab021912108ce623c641421c35d9fea3e0cbfcf5158dbf4effb7da7

Observation c3ff864a-a53e-40a7-ae5c-7b7fbb974fe1 · outbound

This paper cites The Llama 3 Herd of Models.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD The Llama 3 Herd of Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.808030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.808030Z digest=sha256:9404a8b8f1341521efd5009248fb16b5c32ae218f450c0d44fb2bf4c2cf0e3da

Observation 271477db-e557-4ed9-9925-3ced02bd0778 · outbound

This paper cites Query-key normalization for transformers.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Query-key normalization for transformers

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.253015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.812922Z digest=sha256:972066107e14bd8a9e54dd1d2622c5ffe107202e6565667c3d86ff1cd0f71fc3

Observation 2937fc9b-9bca-464a-9646-ea17b6a75406 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.840074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.840074Z digest=sha256:c55b647440a6507105d601ee5c124c1327e3cd8586c1a7e560288c0c73bfa78e

Observation 8ff146d4-94b4-4b36-9368-70d78bbf45b9 · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.829273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.829273Z digest=sha256:aaf4348ba3ace41aea5f37723dbdd9af8be7a9d4e4a16c5ddd41eeb81797745d

Observation a612d0c9-61dc-4944-b71d-eba64cdf5b11 · outbound

This paper cites an unresolved cited work.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Unresolved cited work

Reference 2242

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:54:59.197679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.879365Z digest=sha256:6a193e71cd5aac32f0f7ea2d4a37770ebb3ff85a5828f5625c097ada13b4d041

Pith citing papers

No inbound Pith citation observations are available.