Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:17:00.114314Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2501.12486.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:17:00.114314Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T20:44:48.441119Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-08T20:44:48.658319Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ce9ef419-260e-4283-aa1b-0a5107b71a07 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9528b09-50c2-4c18-91d7-a832ce1f7951 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Scaling to very very large corpora for natural language disambiguation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7156d400-e406-4431-9fac-e755958e10b8 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Data scaling laws in NMT : The effect of noise and architecture
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 52aaac97-a2f7-4630-a828-ac7cfe0e7a9e · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8346fa33-7a86-49a0-ad3b-2f9f332f5561 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Language models are few-shot learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 28a8bd68-561d-49a5-b373-78f702202d9d · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Maxtext: A framework for training large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4f251a5c-7200-4f73-8e9a-db254ccdf6c7 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Rigging the lottery: Making all tickets winners
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ae5c608-9276-4e89-b23f-d3b75a19057b · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws The lottery ticket hypothesis: Finding sparse, trainable neural networks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a653bcc9-d401-4d34-8e32-f7c4e9736e36 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Roy, and Michael Carbin
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7bad8e8b-5cc5-4aa9-98c5-df1b08d312d1 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ac1616-ed28-4232-b787-54cc007102b0 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Scaling laws for sparsely-connected foundation models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4eca3aca-c646-42d9-ae2c-0888e1a9bcaf · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws The State of Sparsity in Deep Neural Networks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eef64b4-c40c-4770-bed9-e7b5ffcb7273 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Scaling laws for neural machine translation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f13d3e0f-28a6-48fa-8fe0-ad918ca095b7 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws A bit of progress in language modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 03a3b10c-41b7-4079-a90d-12520cad9d84 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Data and parameter scaling laws for neural machine translation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 09d58f16-439b-430a-a611-9e0aebde9244 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d230d64-ca76-4445-826c-98b3d6806e51 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws OLM o: Accelerating the science of language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ab8b59cf-004c-4efe-b35c-535996ab21d0 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a9b99fb7-c20a-4bef-abc8-b0aaba0784b7 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Hassibi, D.G
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 24ffe59a-d148-4b01-a98c-3f17e98f40c8 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Channel pruning for accelerating very deep neural networks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 86361b72-1253-44c1-bb09-e59a79836d7a · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Rae, and Laurent Sifre
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0a45dba8-44d0-4069-906f-4ebc9ce1e0ac · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Roy, Jonathan Frankle, and Gintare Karolina Dziugaite
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6eede097-3a48-49eb-a575-518a720a226a · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Scaling Laws for Neural Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b082d7c6-63d5-435f-a286-a1a3a199f23c · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Accurate neural network pruning requires rethinking sparse optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 10c1ef59-edad-41e5-bc5d-410c336c6aaa · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Optimal brain damage
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation be89fc12-e283-473f-99f5-e9b0f974ecf4 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Liu and Jorge Nocedal
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d9cacea0-5b4f-4c4b-88b3-4f4ff7b30cb2 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Step: learning n:m structured sparsity masks from scratch with precondition
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bfd94445-578b-44a4-b1fd-91aec21b718c · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Deep double descent: Where bigger models and more data hurt
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 71129094-cad7-4543-b03e-c1f061da146c · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Efficient Stagewise Pretraining via Progressive Subnetworks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07ce358e-d6a6-4c51-8bb7-a13241a96faf · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Larsen, Jonathan Frankle, Surya Ganguli, and Gintare Karolina Dziugaite
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7bc78c98-6da9-4506-8516-46f7dabf842c · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws AC / DC : Alternating compressed/decompressed training of deep neural networks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a636485-d803-4ceb-9350-1f3db6eea43c · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 42a17963-76f5-4f1c-ad46-b8f66109493c · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Comparing rewinding and fine-tuning in neural network pruning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 52213121-9f54-4082-8198-b3768486833c · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws On the predictability of pruning across scales
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4d5e24ad-e14b-4e7e-a2f3-2e1693b67378 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Creating sparse gpt-3 models with iterative pruning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0d106158-7be4-414a-a0e6-f4d78520ad2f · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Beyond chinchilla-optimal: Accounting for inference in language model scaling laws
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd0b05eb-44a7-4faa-ba56-b4e036b379c4 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9097ced0-efcd-4c0d-9a84-1cfe77d6aaf4 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws A simple and effective pruning approach for large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9532bd64-8d83-489f-8a03-ec3b80573d42 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef670600-49db-4bcd-a463-d569c7db7785 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1302555c-d9c4-4f5b-99a4-ab07679ad2a4 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Step: Staged parameter-efficient pre-training for large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05875338-c573-416e-8b4b-bc885dba8c09 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Masked structural growth for 2x faster language model pre-training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5b8abaec-5155-4d34-a2b1-d90e3f19736b · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws To prune, or not to prune: exploring the efficacy of pruning for model compression
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bb11f40-05a0-45f6-9b66-10bbc0a1aece · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws write newline
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f655db2a-d1c5-43e6-aadb-7b6bcab2a6c4 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws @esa (Ref
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1045d6bf-01f8-4215-b817-2a579dbdeb83 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1705959-45eb-49f7-9703-de5493a14c13 · outbound
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be29f8b8-dde7-43ba-8519-5dbca8c31da0 · inbound
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.