Pith. sign in

Paper Citation Record · LEDGER

nGPT: Normalized Transformer with Representation Learning on the Hypersphere

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2410.01131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.01131 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:54:27.436646Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:30:07.817010Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 14f872da-23c2-4972-a125-50234aaaeedd · inbound

Transformer models are gauge invariant: A mathematical connection between AI and particle physics cites this paper.

Transformer models are gauge invariant: A mathematical connection between AI and particle physics nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T12:14:53.901488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:14:53.901488Z digest=sha256:701da1c425ed6039adbe383a8914297405ce78a8df70a9bb5c6e6da173c83c8b

Observation 387a7cb2-6a81-4870-b072-6f18762890be · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.961082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.961082Z digest=sha256:e08e432aa90a909522771e53c92ca7f5f79f100a23e62f2cdf693e19a54bb020

Observation 88573fc3-ad48-4eaa-b307-dfd6c924cc3e · inbound

Normalized Matching Transformer cites this paper.

Normalized Matching Transformer nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:05:11.573058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T22:03:16.161201Z digest=sha256:ffbf396928a645a382682d6e6ab2e819f070c01bd395a023126bdc1ba049ca18

Observation 7b47b369-0bb6-4d15-bd65-c14e732def0f · inbound

Superposition Yields Robust Neural Scaling cites this paper.

Superposition Yields Robust Neural Scaling nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:36:25.534812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T14:38:44.789822Z digest=sha256:d310655e68dc2843814935ee623ca54119c8c290f62fe3255305ed319dabb153

Observation 3d9ed69f-484d-48ca-8a46-892eb45a7214 · inbound

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation cites this paper.

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:01:46.189690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T18:56:48.722344Z digest=sha256:ca8d2adb248070f1b13bef21692fbe3d8508d527e7a394bbe940a1ba1fc70e26

Observation d37b01ab-8058-4db7-8d15-a58eb8b1d551 · inbound

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation cites this paper.

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T10:25:17.146093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:25:17.146093Z digest=sha256:401fdad8d094816e4c3f43af6150679e4777fdc42dc49a7d1342f58ad31066a1

Observation 1d683da5-7593-4f5f-913d-30ab77be134f · inbound

ENSAM: an efficient foundation model for interactive segmentation of 3D medical images cites this paper.

ENSAM: an efficient foundation model for interactive segmentation of 3D medical images nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:54:27.436646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:54:27.436646Z digest=sha256:18247ea449c9ce865cfef0929c4e53c99c85d23498690c44031401e5ddd828bb

Observation 31d74673-2c9c-49e5-bac5-b2bbd047cde4 · inbound

Universal One-third Time Scaling in Learning Peaked Distributions cites this paper.

Universal One-third Time Scaling in Learning Peaked Distributions nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T05:01:12.799568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:01:12.799568Z digest=sha256:32f9868e700a39e30f4f0bc01c3e59c9671af70efbd4e0d64270ec318a5355fa

Observation d37d4404-41ed-4d4c-afe2-b8aaa8b02984 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:49.967839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T20:04:56.512544Z digest=sha256:55c095adacc3ca07e1f83f09f4b413371e9acb986221b978c05fbae29ba61274

Observation 1995a4cf-7eb9-4beb-9620-08691b5fd5f0 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:12:41.324396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T17:08:31.770889Z digest=sha256:60fca66a2ce21b2ec1cba07f59d59014a6bd85ec223196f6d53b42b1480607a7

Observation 116c1d4c-9209-4a51-ae74-c921b6830765 · inbound

When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer cites this paper.

When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:10.441727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T08:29:58.495518Z digest=sha256:77795343b48a1c7c8ec78ff90674b538997a7b7f18cbdddb1195bb9dd0101276

Observation e5ba9dbe-d9d5-43d1-b2eb-7cb27ed5d92b · inbound

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning cites this paper.

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.092346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-09T20:05:21.069489Z digest=sha256:849ede1eb3c862fd3c93b41231e3fc4ea44b5e53138242bd772a2738e646335c

Observation 2f6df2fc-544a-4582-b3a9-75b4789c61fe · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:08.766353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:d32b2d68f3b79ceb5d7e982f2afe2d6c233a6f168dc7c8e2d496aaa8fcd4dd6f

Observation 740f07ba-42f1-4ad7-97f0-13eb8ed12c3a · inbound

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity cites this paper.

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:08.524404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:11:04.146711Z digest=sha256:d188ef5214be3135d9aaa2b04c1dcb333f9f707107f6dc6ac16204c3a95abfb0

Observation b627a3f8-9cea-4adc-9bcf-1d5ab827f2cf · inbound

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives cites this paper.

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:37:19.228142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:34:48.195468Z digest=sha256:eb02119672ee97f545016f81c79c966d2bd6d994920981f941d3bab1ec8abf43

Observation 7db5a631-d2b0-4884-94ad-538826e3c8e2 · inbound

Chem-GMNet: A Sphere-Native Geometric Transformer for Molecular Property Prediction cites this paper.

Chem-GMNet: A Sphere-Native Geometric Transformer for Molecular Property Prediction nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:57:53.678529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T19:55:19.362468Z digest=sha256:43ff9f897194c2c9ea2c34fb622efa5dbb924308abaabe7fd041932bc3a5fa23

Observation 65f73d7f-438a-4fc4-9137-c19d6dcc5240 · inbound

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders cites this paper.

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:15:29.656032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T07:15:16.674714Z digest=sha256:c19fd1467793a566c361b8504504425101e0e51aa50c36863ca3e3bdd005b061

Observation 1cb962d3-3aad-45f7-a761-3b2b78183328 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.818744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:94d32466c5fa6717d9a21035c43d5fdc08db91ec4906bb1e828e7292b3a5e1a1

Observation f2e2411a-d06e-4433-add0-6b0d47495558 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:09.204824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:09.204824Z digest=sha256:da5dc5f52b2c2b63e5c2fc937565c5d627206978d5647820aff45d51f4a6dfdd

Observation b4aa417d-9913-4a99-86ef-2adb63cb354e · inbound

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control cites this paper.

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 240

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:46.590081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:46.590081Z digest=sha256:9dbf81b410910be444ee0c0978d623e88c1d2e765a0a427a15d9f5b82d763303