Pith. sign in

Paper Citation Record · LEDGER

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2506.22049.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22049 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:21:11.749664Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1e56cb1-b782-4045-8322-3306d5d4fc46 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.855557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:09.405154Z digest=sha256:0cb9538fd1644ed4379603ae29f4ef848cc3c94d90134ccc417ed9feccf963f8

Observation f6ae2540-0da1-4afd-8286-7f3a109c9cc2 · outbound

This paper cites The Llama 3 Herd of Models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.452244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.452244Z digest=sha256:0d42eb3b43bc4b647b9433544f9d37ae548ff4c686d77e5cc16a57b138316d49

Observation 0bd11535-26f4-403f-a4b6-860fed6d6cc7 · outbound

This paper cites Qwen2 Technical Report.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Qwen2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.516987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.516987Z digest=sha256:8e112e116568511e6a6249f43dd23ed116b80418c5c2f22db450341817c2d2a3

Observation b4bb5521-5c02-4c3a-b75c-b1b989834350 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.605367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.605367Z digest=sha256:aaeb68c74fef3d98712e770136ebdc53fde5a41a867562e79f27a8ca04b9fd19

Observation caca3065-fec6-45b6-a839-e549e05e583c · outbound

This paper cites Adaptive input representations for neural language modeling.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Adaptive input representations for neural language modeling

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.683118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:09.660591Z digest=sha256:c06e317a6f61d35908b2a23870e5d3657a7fa41de52401da817885f19c9552d7

Observation e77e122a-9f1d-40f1-9926-4ca44946bdcc · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.708524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.708524Z digest=sha256:8ca707931f8c03d2b8e9669731834500c8d08e0a166b587dc4e6cbd1132ff8e3

Observation bc0dcd1f-e364-4b0a-bc50-bd37b63d1aa5 · outbound

This paper cites Learning deep transformer models for machine translation.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Learning deep transformer models for machine translation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.545995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:09.764755Z digest=sha256:73fd733887e81777b04a54d5857fc2f437364d0d551e9b45211f07eb123b7470

Observation b457ab1c-8055-469a-b204-a5afb9ca33d8 · outbound

This paper cites Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.387759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:09.865572Z digest=sha256:ed934d330efb0d8c3ea2835d815729ecb2579cdb45b27969fbcdf580824c5e1e

Observation b716bc9a-db6d-4c36-bcd7-c917ef91e4cc · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The Unreasonable Ineffectiveness of the Deeper Layers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.909236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.909236Z digest=sha256:6c0870df7ab0ecf094185941cd4b99a4ea2594fc793f6e264233b48c33229375

Observation bbda0a57-d6e7-4e8b-84bd-8262e9067d5a · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.026345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.026345Z digest=sha256:4633fd112b658c0d027f5f664e177f5c1da66579deb2215a3a1d7e76a7ac6ad9

Observation 6d9c51fc-52b9-45eb-8887-1129ca86a587 · outbound

This paper cites Layer Normalization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Layer Normalization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.068870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.068870Z digest=sha256:e5ae41fd88812d94f6a425cf04dd13178f83815817cfceeb27d70583d14c3f3f

Observation a7ead766-3b43-466f-b27b-331c1a18ecf0 · outbound

This paper cites Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.135364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.135364Z digest=sha256:57fb099ee5157fc64f80efc25a323f05f4936deb0a0e8809942130c96445dfb4

Observation 765c349a-a682-4772-b025-c7490566712c · outbound

This paper cites The curse of depth in large language models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The curse of depth in large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.187845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.187845Z digest=sha256:c595c6d5d545c6c3a9b135847b4cf658c67d466cadc81134789e1e424e50025d

Observation b71d0a8e-4860-43bd-b0ab-c7d4c60477ef · outbound

This paper cites Deepnet: Scaling transformers to 1,000 layers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Deepnet: Scaling transformers to 1,000 layers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.284486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:10.234776Z digest=sha256:808305c6ac80e77400b23858f7027ab1c35503fe042b45c7871f3011045943bc

Observation 19ca331f-9c38-4f44-bd14-3e0f5396d1dc · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Cogview: Mastering text-to-image generation via transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.176152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:10.268765Z digest=sha256:421dbce202b1ae125d14526d24aa10940b9213bc6390e24fd5ef5cb71ff294af

Observation c04172f7-369d-4e02-a47b-9ef09d6682a4 · outbound

This paper cites Group normalization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Group normalization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.091570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:10.300584Z digest=sha256:e2b725609add76956e32d626cbcdc543ff9251d6f7a2888d27410afd5780e298

Observation a22f81b1-924a-4493-a895-affa81442b78 · outbound

This paper cites Root mean square layer normalization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Root mean square layer normalization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.952220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:10.352396Z digest=sha256:c5b85c3d77d269094ac52d4c86f1db548f91411172cccfd1254007c062b78dcb

Observation cce29769-20ca-4659-be46-5d740a038946 · outbound

This paper cites Understanding and Improving Layer Normalization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Understanding and Improving Layer Normalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.438131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.438131Z digest=sha256:caa490d456417eb8620aaaf7335fbaa1632ea934c86a4eba123f682ac4c4f9fa

Observation 0b006660-0b36-4b59-949b-c47f4450d93a · outbound

This paper cites Transformers without normalization, 2025.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Transformers without normalization, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.515415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.515415Z digest=sha256:f66bab253135b844fa663e4213f7df47e181da3be6cf62bece3752e581876e21

Observation 214245db-f125-4b6d-b055-c67a9c129e3f · outbound

This paper cites On layer normalization in the transformer architecture.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling On layer normalization in the transformer architecture

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.865768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:10.573822Z digest=sha256:23629e7c9bd2e6ac2c5c51e60d0da4b66a74a535d3edc8abed835d12cbc3f5ac

Observation accca324-9de0-47f8-b790-2b45b3eed285 · outbound

This paper cites B2t connection: Serving stability and performance in deep transformers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling B2t connection: Serving stability and performance in deep transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.733130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:10.615392Z digest=sha256:6d5030629d0f295255792f57207bd9cfb6506a5111f2177017d3ad6575210f42

Observation 5ed40db9-0ac8-49d2-9536-daa3fb9bbafe · outbound

This paper cites Peri-LN: Revisiting Normalization Layer in the Transformer Architecture.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.661789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.661789Z digest=sha256:55cb1d56b9d9232570f0a293898f9afc7942b512dfb360b01978841e7c35bfa8

Observation 99d9f61f-7345-4669-8584-0e8a34e7dd7c · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.741333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.741333Z digest=sha256:9d1e75733429f710e95919515a1de925bd2c4da06c3a5aecdb780f246b880df8

Observation b3be30c0-e5f2-4ecc-b23e-b4811567bf73 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.769395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.769395Z digest=sha256:f1f96831a4a3ce02b0a346d81920a90320aeee23a6730e8bc15f2d691bea5a65

Observation 6d7f4837-08e2-4a40-b736-2a68eb5ec173 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.825974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.825974Z digest=sha256:fd5cfb717d731f16b5786f45e6581eb1ff0a7f04b7aa1c85d669cd298a73def9

Observation f6768497-b4eb-4e73-be14-ddcad43e1eb5 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.900239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.900239Z digest=sha256:dabc894684fdb9e2b0f587035f0c8b31b7a916a686f00da722ee126d8e45c0df

Observation 90c34a68-cf92-43ca-b45b-4a76e4068468 · outbound

This paper cites an unresolved cited work.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.954572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.954572Z digest=sha256:dcb5662c71b15915decca025dfb924df55f84e2aeabb2a2b417b43af4a037584

Observation ed815e55-47d4-48ac-8f61-9bb187ea90f2 · outbound

This paper cites Relora: High- rank training through low-rank updates.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Relora: High- rank training through low-rank updates

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.595864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:11.014246Z digest=sha256:20aca900eae9669dc3e8ed6aff8acddb726bddb17260b617120ff8570fab9354

Observation 8a8621e3-d424-4be8-a038-6fa8409f93d5 · outbound

This paper cites Galore: Memory-efficient llm training by gradient low-rank projection.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Galore: Memory-efficient llm training by gradient low-rank projection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.528482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:11.043414Z digest=sha256:793f951c0c63f18a4966e8b435eb53284f267a9db6cafa93174cca159596d502

Observation 6472a69a-7795-4c60-86c3-14ae6c78ed73 · outbound

This paper cites Transformer-xl: Attentive language models beyond a fixed-length context.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Transformer-xl: Attentive language models beyond a fixed-length context

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.357730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:11.118726Z digest=sha256:950786199bdb93ff598ba3f5bb2ca48801b03db0878c7d5f991774555877433d

Observation 740f9556-949b-4e22-a5cd-91292529a5ab · outbound

This paper cites Sandwich batch normal- ization: A drop-in replacement for feature distribution heterogeneity.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Sandwich batch normal- ization: A drop-in replacement for feature distribution heterogeneity

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.196929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:11.168935Z digest=sha256:429d5fb018c3389d71f3fb74cd4834a3f594429c9f167f14e01bc9cbc6a1e5a1

Observation bc3ca830-ae9d-4668-b417-79cd9809d6bc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.222328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.222328Z digest=sha256:69e0aa806c859f62ee43faca1ae580a20ca4aca2326b22e042760bb8a12a38cb

Observation 5ac6186a-dee7-43ad-b92c-55d3030089bc · outbound

This paper cites GLU Variants Improve Transformer.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling GLU Variants Improve Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.265526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.265526Z digest=sha256:3a3880dc0db952c0bae4d967b5f26027a896e24114cf2298bd35f9b77af8ed29

Observation 992f5f56-b36b-417e-bdae-16dc89ca8156 · outbound

This paper cites Spike No More: Stabilizing the Pre-training of Large Language Models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.325852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.325852Z digest=sha256:5317dd6cb1b7163ae5f36f7b9d74e860cd859db5ef26643ec84dc634762814b2

Observation bfb1a337-b77a-4518-b7f5-97fef0e6e41c · outbound

This paper cites Adam: A method for stochastic optimization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Adam: A method for stochastic optimization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.057360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:11.363870Z digest=sha256:48e4d2aecf12d770dd49cb9d677f956f6ee885662108bfed59e92bdeb49b7984

Observation 182ed576-34bf-4d7f-95f1-448bcf0a792b · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.394382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.394382Z digest=sha256:033091dce444c048f5a0b865d6dcc694eafc35eaa0f0c7b9491fc1eebd5a4d21

Observation 3d83aa1b-b6b4-4f13-9356-d47e3b1b86a0 · outbound

This paper cites Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:12.830440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:11.433172Z digest=sha256:4c30fdb249afdb13f1c4927c2d151c88817eb3a1bbc5fb71a9bd0002e8cf53c3

Observation 7be4cea1-d926-40e8-83b3-286c529a8def · outbound

This paper cites Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:12.697867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:11.518549Z digest=sha256:3a4c7b273ca3da997f79eb9bec9dea01ce8acd7acffb8f15f78416294d02b365

Observation 8d26608b-4133-44f8-9cb5-2b2c2635f1f1 · outbound

This paper cites Lisa: Layerwise importance sampling for memory-efficient large language model fine-tuning.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Lisa: Layerwise importance sampling for memory-efficient large language model fine-tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:12.512251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:11.584659Z digest=sha256:fc404fc6affe0144f779bf0b13c49b387b2330aa8ceaaa118150eace2bf0b649

Observation fb50ee85-7b04-441f-8740-cd7134cf8679 · outbound

This paper cites The language model evaluation harness, 07 2024.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The language model evaluation harness, 07 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.650939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.650939Z digest=sha256:d26243a84b149f47cb1a0058dd00871ea911ce25a42585f2113f76a2814baba5

Observation f67d85ab-c70e-4247-817e-8c9c74823dcb · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.722895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.722895Z digest=sha256:324c4fa57bc60d56f7863b879bb61d084f0a4bef2109c0d155e46e0f5c1b4d9f

Observation bf37f784-a868-4bb5-9853-f70e99755612 · outbound

This paper cites an unresolved cited work.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:21:12.211536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:21:11.749664Z digest=sha256:d72c53297fb1e2400a45d386b758336340f70d48db006fef4a2cd269a9d829eb

Pith citing papers

No inbound Pith citation observations are available.