Pith. sign in

Paper Citation Record · LEDGER

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2507.17501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17501 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:54:58.891090Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce36854f-edba-4a91-9615-b9841d518ab7 · outbound

This paper cites Layer Normalization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.796795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.796795Z digest=sha256:43795ce9e63accf4019d09ae8618a87875208de5c67cb01405669a06a279ceb7

Observation d219215c-3b81-437b-81be-c36abb8427cd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Adam: A Method for Stochastic Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.818406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.818406Z digest=sha256:a62b05d40e6a67faa62c146fd4e073b6a3df91ac70649887c109fbc8056d6b05

Observation 05b68beb-efea-4b68-b80d-19d4b1057d2d · outbound

This paper cites DeepSeek-V3 Technical Report.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.823828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.823828Z digest=sha256:cc3ba518acfb0c78325291b4a1edd144a1786f75641432a1fea31321302c3db9

Observation fc2c3735-121c-44aa-80ff-0ef1832a9784 · outbound

This paper cites A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.844812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.844812Z digest=sha256:e081716373ef3b49abb655c3b6361d9eabb8c0681a95b7fcdd542333a1175174

Observation c3f65aae-289e-4c31-be1a-4d27a400a888 · outbound

This paper cites Large Batch Training of Convolutional Networks.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Large Batch Training of Convolutional Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.851371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.851371Z digest=sha256:2d932aeba33b666508839f2ef1f9750bc255e47a2dee17ca491759e502205918

Observation a4a68fd6-09ed-4b1f-b3d8-372136ec03ff · outbound

This paper cites Transformers without Normalization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Transformers without Normalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.862382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.862382Z digest=sha256:50ba15e9d4b12cb1bab1e5c0d72d2ec31a3cae8b2b4bc931d8201e2b0cab74d0

Observation ca7a4ae5-d782-4073-8b6c-93d8e67b6134 · outbound

This paper cites In a backpropagation, since we have obtained ∂L ∂vec(Y ), we would like to further analyze ∂L ∂vec(Wq) , ∂L ∂vec(Wk) , ∂L ∂vec(Wv).

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD In a backpropagation, since we have obtained ∂L ∂vec(Y ), we would like to further analyze ∂L ∂vec(Wq) , ∂L ∂vec(Wk) , ∂L ∂vec(Wv)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.233831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.868032Z digest=sha256:a33e8df76a61bdf9f1aad10f7ee96392789af31e9ba2225dbd1d8b40427462bc

Observation 55947400-400a-4948-8d24-66b547135587 · outbound

This paper cites Défossez et al.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Défossez et al

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.160647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.891090Z digest=sha256:96f4b14e4553be6e6fc5b1f3e70eb0074b1bc6d1fde67654047b1f663b5e7e43

Observation 38da0786-31b5-44cb-a816-8b63bc9fd89e · outbound

This paper cites This method augments the gradient direction with a fraction of the update vector from the previous step, allowing faster convergence and helping escape local minima.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD This method augments the gradient direction with a fraction of the update vector from the previous step, allowing faster convergence and helping escape local minima

Reference 1983

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.180836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.884769Z digest=sha256:8adcbfead07eca1988089c0a8be28419c03909c897b49f71184b7a1d14bd9b45

Observation e20bf3e9-e019-4bb4-8551-e98612ef886b · outbound

This paper cites Qwen Technical Report.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Qwen Technical Report

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.834431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.834431Z digest=sha256:f4501c40c16d07c7626aea646769051ecf84ca123693006bb90b173d7982584e

Observation a11c4c39-2f8d-4a92-a779-1caf95de60bb · outbound

This paper cites MARS: Unleashing the Power of Variance Reduction for Training Large Models.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.857286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.857286Z digest=sha256:d438d9de111a41d5cd4ab4c140402f338d09a51842f00979362807a035851423

Observation e4d077cf-903a-497f-9a1a-68817dae4d15 · outbound

This paper cites Language models are few-shot learners.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Language models are few-shot learners

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.802656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.802656Z digest=sha256:1fda68d84da3709037a4007a8b45342ad94541909453fa7a412d4e2382985056

Observation bfe13c81-6ebe-4925-ab8b-32859137c785 · outbound

This paper cites All language models were trained on OpenWebText, using GPT-2 tokenizer.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD All language models were trained on OpenWebText, using GPT-2 tokenizer

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.215255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.873604Z digest=sha256:62b6ec8582d1c908d22d526959d56c0312d42e399c0e52a0802272cf04f255a7

Observation c3ff864a-a53e-40a7-ae5c-7b7fbb974fe1 · outbound

This paper cites The Llama 3 Herd of Models.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD The Llama 3 Herd of Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.808030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.808030Z digest=sha256:3c672d9e853a4cb8f2fc170f8b7cab631920c788a8a1c56af069a91221dd5f31

Observation 271477db-e557-4ed9-9925-3ced02bd0778 · outbound

This paper cites Query-key normalization for transformers.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Query-key normalization for transformers

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.253015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.812922Z digest=sha256:ce3500442560de1b8b0086f1218500565a764d19a4828164085a8a8ba389f69a

Observation 2937fc9b-9bca-464a-9646-ea17b6a75406 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.840074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.840074Z digest=sha256:4efdfafa91b9825df79eb234be9df6ca0faf4029c56f325c325e29b83b818047

Observation 8ff146d4-94b4-4b36-9368-70d78bbf45b9 · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.829273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.829273Z digest=sha256:380a4a718d9a7d356acb217a27bf9cfb58f995dc2434cbbaad34fb18c3669f1d

Observation a612d0c9-61dc-4944-b71d-eba64cdf5b11 · outbound

This paper cites an unresolved cited work.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Unresolved cited work

Reference 2242

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:54:59.197679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:54:58.879365Z digest=sha256:89eaca06af60ff8bdc8697db94dc8ea205967356fc24a7ff07f82081cbec2722

Pith citing papers

No inbound Pith citation observations are available.