Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:21:11.749664Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2506.22049.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:21:11.749664Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a1e56cb1-b782-4045-8322-3306d5d4fc46 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6ae2540-0da1-4afd-8286-7f3a109c9cc2 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The Llama 3 Herd of Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd11535-26f4-403f-a4b6-860fed6d6cc7 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Qwen2 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4bb5521-5c02-4c3a-b75c-b1b989834350 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caca3065-fec6-45b6-a839-e549e05e583c · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Adaptive input representations for neural language modeling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e77e122a-9f1d-40f1-9926-4ca44946bdcc · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Generating Long Sequences with Sparse Transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc0dcd1f-e364-4b0a-bc50-bd37b63d1aa5 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Learning deep transformer models for machine translation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b457ab1c-8055-469a-b204-a5afb9ca33d8 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b716bc9a-db6d-4c36-bcd7-c917ef91e4cc · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The Unreasonable Ineffectiveness of the Deeper Layers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbda0a57-d6e7-4e8b-84bd-8262e9067d5a · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d9c51fc-52b9-45eb-8887-1129ca86a587 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Layer Normalization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ead766-3b43-466f-b27b-331c1a18ecf0 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765c349a-a682-4772-b025-c7490566712c · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The curse of depth in large language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b71d0a8e-4860-43bd-b0ab-c7d4c60477ef · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Deepnet: Scaling transformers to 1,000 layers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19ca331f-9c38-4f44-bd14-3e0f5396d1dc · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Cogview: Mastering text-to-image generation via transformers
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c04172f7-369d-4e02-a47b-9ef09d6682a4 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Group normalization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a22f81b1-924a-4493-a895-affa81442b78 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Root mean square layer normalization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cce29769-20ca-4659-be46-5d740a038946 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Understanding and Improving Layer Normalization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b006660-0b36-4b59-949b-c47f4450d93a · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Transformers without normalization, 2025
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 214245db-f125-4b6d-b055-c67a9c129e3f · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling On layer normalization in the transformer architecture
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation accca324-9de0-47f8-b790-2b45b3eed285 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling B2t connection: Serving stability and performance in deep transformers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ed40db9-0ac8-49d2-9536-daa3fb9bbafe · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d9f61f-7345-4669-8584-0e8a34e7dd7c · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3be30c0-e5f2-4ecc-b23e-b4811567bf73 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7f4837-08e2-4a40-b736-2a68eb5ec173 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6768497-b4eb-4e73-be14-ddcad43e1eb5 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c34a68-cf92-43ca-b45b-4a76e4068468 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed815e55-47d4-48ac-8f61-9bb187ea90f2 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Relora: High- rank training through low-rank updates
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a8621e3-d424-4be8-a038-6fa8409f93d5 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Galore: Memory-efficient llm training by gradient low-rank projection
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6472a69a-7795-4c60-86c3-14ae6c78ed73 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Transformer-xl: Attentive language models beyond a fixed-length context
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 740f9556-949b-4e22-a5cd-91292529a5ab · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Sandwich batch normal- ization: A drop-in replacement for feature distribution heterogeneity
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc3ca830-ae9d-4668-b417-79cd9809d6bc · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling LLaMA: Open and Efficient Foundation Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ac6186a-dee7-43ad-b92c-55d3030089bc · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling GLU Variants Improve Transformer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 992f5f56-b36b-417e-bdae-16dc89ca8156 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Spike No More: Stabilizing the Pre-training of Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb1a337-b77a-4518-b7f5-97fef0e6e41c · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Adam: A method for stochastic optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 182ed576-34bf-4d7f-95f1-448bcf0a792b · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d83aa1b-b6b4-4f13-9356-d47e3b1b86a0 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7be4cea1-d926-40e8-83b3-286c529a8def · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d26608b-4133-44f8-9cb5-2b2c2635f1f1 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Lisa: Layerwise importance sampling for memory-efficient large language model fine-tuning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb50ee85-7b04-441f-8740-cd7134cf8679 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The language model evaluation harness, 07 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f67d85ab-c70e-4247-817e-8c9c74823dcb · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Llama 2: Open foundation and fine-tuned chat models, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf37f784-a868-4bb5-9853-f70e99755612 · outbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.