Pith. sign in

Paper Citation Record · LEDGER

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws

As of 11 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2501.12486.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12486 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:17:00.114314Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:44:48.441119Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T20:44:48.658319Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce9ef419-260e-4283-aa1b-0a5107b71a07 · outbound

This paper cites Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:59.750276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:16:59.750276Z digest=sha256:3fd29b99ed24abeb1e80744a11647c62e60e1404c9af862a1a3fe92a6b4e4227

Observation c9528b09-50c2-4c18-91d7-a832ce1f7951 · outbound

This paper cites Scaling to very very large corpora for natural language disambiguation.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Scaling to very very large corpora for natural language disambiguation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:01.248876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.758530Z digest=sha256:54c153ad4459c35d95ec1b1797581f3084cad504d4464c67d9290115960abf06

Observation 7156d400-e406-4431-9fac-e755958e10b8 · outbound

This paper cites Data scaling laws in NMT : The effect of noise and architecture.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Data scaling laws in NMT : The effect of noise and architecture

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:01.219679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.765340Z digest=sha256:9eb52965bdbd4cfc159d1e0afb1a07630bface922c63c0ec014b078774f4d667

Observation 52aaac97-a2f7-4630-a828-ac7cfe0e7a9e · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:59.772536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:16:59.772536Z digest=sha256:e3cab83d0d41a34d1c5818deb4ceaa7f53c5e909d93540536d083dae95210425

Observation 8346fa33-7a86-49a0-ad3b-2f9f332f5561 · outbound

This paper cites Language models are few-shot learners.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:01.183455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.779895Z digest=sha256:d9420c14ebe51f2950c4705f2509dfbd93766230647c198548247ed6cb578982

Observation 28a8bd68-561d-49a5-b373-78f702202d9d · outbound

This paper cites Maxtext: A framework for training large language models.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Maxtext: A framework for training large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:01.144782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.786655Z digest=sha256:3a371e59d85466407808a044e6fbfd98312b8cd4fdb7fbc52d5c847f97c9a462

Observation 4f251a5c-7200-4f73-8e9a-db254ccdf6c7 · outbound

This paper cites Rigging the lottery: Making all tickets winners.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Rigging the lottery: Making all tickets winners

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:01.118432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.793479Z digest=sha256:98576ad99483f03c5adc9126556eb4936e3b167548a7db2a2ea04caf8bca8bb5

Observation 8ae5c608-9276-4e89-b23f-d3b75a19057b · outbound

This paper cites The lottery ticket hypothesis: Finding sparse, trainable neural networks.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws The lottery ticket hypothesis: Finding sparse, trainable neural networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:59.802953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:16:59.802953Z digest=sha256:c5faa059e2d08c67fd8e157a961bef09826dc280a753591056abcba31c1fd6b9

Observation a653bcc9-d401-4d34-8e32-f7c4e9736e36 · outbound

This paper cites Roy, and Michael Carbin.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Roy, and Michael Carbin

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:01.074493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.809967Z digest=sha256:f4b9de3fb93b098428f01a81f482e29e7165f250f916fc0f85a0a07788acb4f8

Observation 7bad8e8b-5cc5-4aa9-98c5-df1b08d312d1 · outbound

This paper cites SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:59.816539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:16:59.816539Z digest=sha256:a429d7a4cd5b1accc7992e761b240aa09a29dac35def06ad98b34d6b5b4ebfdc

Observation a9ac1616-ed28-4232-b787-54cc007102b0 · outbound

This paper cites Scaling laws for sparsely-connected foundation models.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Scaling laws for sparsely-connected foundation models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:01.051629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.824374Z digest=sha256:0b34cde01e0a699b02e05589768a7f49ff50710841189268d5d636f644f46e5f

Observation 4eca3aca-c646-42d9-ae2c-0888e1a9bcaf · outbound

This paper cites The State of Sparsity in Deep Neural Networks.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws The State of Sparsity in Deep Neural Networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:59.831134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:16:59.831134Z digest=sha256:18183950b600207d6577f6dede9a48346c803522e8feb7d2182164accdb46e0e

Observation 5eef64b4-c40c-4770-bed9-e7b5ffcb7273 · outbound

This paper cites Scaling laws for neural machine translation.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Scaling laws for neural machine translation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:01.023635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.838292Z digest=sha256:c83b905341159482b248a2867bff1eeb15764d34058e88550f1bf26ef38dc9a7

Observation f13d3e0f-28a6-48fa-8fe0-ad918ca095b7 · outbound

This paper cites A bit of progress in language modeling.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws A bit of progress in language modeling

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-10T17:17:00.393001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.846701Z digest=sha256:cccf70a2e3b288b71cdd49d6b124b20d033a9cb1eab5f2d85e0e71e18b1a9955

Observation 03a3b10c-41b7-4079-a90d-12520cad9d84 · outbound

This paper cites Data and parameter scaling laws for neural machine translation.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Data and parameter scaling laws for neural machine translation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:01.002645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.854996Z digest=sha256:d0e6092b26d161781fc30a3e779fd86899668890bf46d9e3de470520ba9fe6a1

Observation 09d58f16-439b-430a-a611-9e0aebde9244 · outbound

This paper cites The Llama 3 Herd of Models.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:59.861295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:16:59.861295Z digest=sha256:43baf2d114fff15dedad1803413b91c6b422c1ee95bb27e1b64e8be89acc828e

Observation 0d230d64-ca76-4445-826c-98b3d6806e51 · outbound

This paper cites OLM o: Accelerating the science of language models.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws OLM o: Accelerating the science of language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.981103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.872152Z digest=sha256:7a1ff0204ca65c2c6976a922da379995b5dca3c84c70f9367e077469e6784337

Observation ab8b59cf-004c-4efe-b35c-535996ab21d0 · outbound

This paper cites an unresolved cited work.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T17:17:00.954867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.882300Z digest=sha256:cb9f3d1483dccaa6ba7e62354758d39c3c68aa53086e95a59e155b433dd38620

Observation a9b99fb7-c20a-4bef-abc8-b0aaba0784b7 · outbound

This paper cites Hassibi, D.G.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Hassibi, D.G

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.933243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.888823Z digest=sha256:ca959f991fb23e61ee725010d861fc67434b7c4833dd2bc088aa1b39b364deda

Observation 24ffe59a-d148-4b01-a98c-3f17e98f40c8 · outbound

This paper cites Channel pruning for accelerating very deep neural networks.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Channel pruning for accelerating very deep neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.909558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.895169Z digest=sha256:b20d7e1be7b4bb088edcfda7705367140e80f400d8146b179111f80138595126

Observation 86361b72-1253-44c1-bb09-e59a79836d7a · outbound

This paper cites Rae, and Laurent Sifre.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Rae, and Laurent Sifre

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.889938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.900702Z digest=sha256:65a07d63993498d9c9d97da8d66fe2ecc4ba262feb12f4d3d661d42acc0ea8d1

Observation 0a45dba8-44d0-4069-906f-4ebc9ce1e0ac · outbound

This paper cites Roy, Jonathan Frankle, and Gintare Karolina Dziugaite.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Roy, Jonathan Frankle, and Gintare Karolina Dziugaite

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.871425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.906464Z digest=sha256:bcebdcbb13aa14804d57706cd1c84e8b23adece4cb03a1343880e37eb60c8bab

Observation 6eede097-3a48-49eb-a575-518a720a226a · outbound

This paper cites Scaling Laws for Neural Language Models.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Scaling Laws for Neural Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:59.912276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:16:59.912276Z digest=sha256:afad20605cc315e04439cd72c81f1cc7b07a0ce49dd30563e40a8abe8d2a289b

Observation b082d7c6-63d5-435f-a286-a1a3a199f23c · outbound

This paper cites Accurate neural network pruning requires rethinking sparse optimization.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Accurate neural network pruning requires rethinking sparse optimization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.849992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.919879Z digest=sha256:f05575a55f83a4a854396ecd5ddcfdc238adcc4344adf25895e18a628dad9ea5

Observation 10c1ef59-edad-41e5-bc5d-410c336c6aaa · outbound

This paper cites Optimal brain damage.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Optimal brain damage

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.827869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.928811Z digest=sha256:c95c62b35c14e903c529125ce3f7256c4614132578e96d03039f68c4f8ba552e

Observation be89fc12-e283-473f-99f5-e9b0f974ecf4 · outbound

This paper cites Liu and Jorge Nocedal.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Liu and Jorge Nocedal

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.801977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.937495Z digest=sha256:cec8ceedce806ef914254d04d31ace09e0f189f3d3b15d5ebbd66cb55bad73dc

Observation d9cacea0-5b4f-4c4b-88b3-4f4ff7b30cb2 · outbound

This paper cites Step: learning n:m structured sparsity masks from scratch with precondition.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Step: learning n:m structured sparsity masks from scratch with precondition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.782868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.945271Z digest=sha256:cc65141ed65f5d525ecd068f58040d900c5b8581df43d54473d9f86b1e73bc2e

Observation bfd94445-578b-44a4-b1fd-91aec21b718c · outbound

This paper cites Deep double descent: Where bigger models and more data hurt.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Deep double descent: Where bigger models and more data hurt

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.762746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.952581Z digest=sha256:24b7bf2c1436cea333d4e606cf3967f3d7cbb0792a4c338ac04e0384ea1eeedf

Observation 71129094-cad7-4543-b03e-c1f061da146c · outbound

This paper cites Efficient Stagewise Pretraining via Progressive Subnetworks.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Efficient Stagewise Pretraining via Progressive Subnetworks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:59.963976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:16:59.963976Z digest=sha256:0e2b74d0f2224fcc86531331135b7d8fb9e67527228786854a34d34401cf2e90

Observation 07ce358e-d6a6-4c51-8bb7-a13241a96faf · outbound

This paper cites Larsen, Jonathan Frankle, Surya Ganguli, and Gintare Karolina Dziugaite.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Larsen, Jonathan Frankle, Surya Ganguli, and Gintare Karolina Dziugaite

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.742715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.973612Z digest=sha256:7f054d397a45d54f42b795cd4ee6d61bca6b579a7e6c2257f970ca32edf5a492

Observation 7bc78c98-6da9-4506-8516-46f7dabf842c · outbound

This paper cites AC / DC : Alternating compressed/decompressed training of deep neural networks.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws AC / DC : Alternating compressed/decompressed training of deep neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.721218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.980118Z digest=sha256:fcce46f424d7273e32ed12edeb037720c71ec24343e4a47fc91c9dc79abc8162

Observation 8a636485-d803-4ceb-9350-1f3db6eea43c · outbound

This paper cites an unresolved cited work.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T17:17:00.702200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.986686Z digest=sha256:00de957aeb1fef3738c83e3353cf72f75d539d420acfe0c1cb80a3dd34fba3d4

Observation 42a17963-76f5-4f1c-ad46-b8f66109493c · outbound

This paper cites Comparing rewinding and fine-tuning in neural network pruning.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Comparing rewinding and fine-tuning in neural network pruning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.683346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:16:59.995673Z digest=sha256:ec48c62403fd5b4afcb1222f8b4151cc734e3233016670321112bc4bb6c04959

Observation 52213121-9f54-4082-8198-b3768486833c · outbound

This paper cites On the predictability of pruning across scales.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws On the predictability of pruning across scales

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.664019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:17:00.002162Z digest=sha256:b5febe206522c9d8a4a116c982c2026be630796f14f1ca4244a8c6a1ecddcc91

Observation 4d5e24ad-e14b-4e7e-a2f3-2e1693b67378 · outbound

This paper cites Creating sparse gpt-3 models with iterative pruning.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Creating sparse gpt-3 models with iterative pruning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.643198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:17:00.012910Z digest=sha256:9b69bda60db0666bd4740ebd92fcc0aef8e5ebd8461d8568ab6b3fb9e2964a1b

Observation 0d106158-7be4-414a-a0e6-f4d78520ad2f · outbound

This paper cites Beyond chinchilla-optimal: Accounting for inference in language model scaling laws.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Beyond chinchilla-optimal: Accounting for inference in language model scaling laws

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.622880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:17:00.021734Z digest=sha256:d2e4a2dc9776cd1c4cef7b23fa46911a95011a9d3fe640b08e2e1886c96d1d7c

Observation bd0b05eb-44a7-4faa-ba56-b4e036b379c4 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T17:17:00.032815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:17:00.032815Z digest=sha256:475a95c6005958dff15b528f2452a0382c61ed4743771f45b706320ae89072c3

Observation 9097ced0-efcd-4c0d-9a84-1cfe77d6aaf4 · outbound

This paper cites A simple and effective pruning approach for large language models.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws A simple and effective pruning approach for large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.603837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:17:00.041648Z digest=sha256:275a0875e02118784b1424b5075f75b77a2b57cdb16c037ad6ccf7ce34d987e3

Observation 9532bd64-8d83-489f-8a03-ec3b80573d42 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:17:00.053314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:17:00.053314Z digest=sha256:99c1aed5dc4bdd7a404f61cffd891ff5c184cba50a049a0bfc084635a3d678cb

Observation ef670600-49db-4bcd-a463-d569c7db7785 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T17:17:00.060950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:17:00.060950Z digest=sha256:acbc568255a42ec7eab455d9f24b834613f3b390cdbb95261cb0712612f5d4f8

Observation 1302555c-d9c4-4f5b-99a4-ab07679ad2a4 · outbound

This paper cites Step: Staged parameter-efficient pre-training for large language models.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Step: Staged parameter-efficient pre-training for large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.584714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:17:00.068020Z digest=sha256:6f6e29220f54d98d8844807d72b45eeafb18f0a629a732d9f34b4f9a7b498861

Observation 05875338-c573-416e-8b4b-bc885dba8c09 · outbound

This paper cites Masked structural growth for 2x faster language model pre-training.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Masked structural growth for 2x faster language model pre-training

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:17:00.564110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T17:17:00.074557Z digest=sha256:2e5e7f07ea96e8d88faee8a6e54d0037b01f585ad1afe6aa4ffc96b192e6a73f

Observation 5b8abaec-5155-4d34-a2b1-d90e3f19736b · outbound

This paper cites To prune, or not to prune: exploring the efficacy of pruning for model compression.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws To prune, or not to prune: exploring the efficacy of pruning for model compression

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T17:17:00.080795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:17:00.080795Z digest=sha256:13449b526d5900bb28971127a70648b8ee58e689412b6e1d579c9f5e9462e1b2

Observation 8bb11f40-05a0-45f6-9b66-10bbc0a1aece · outbound

This paper cites write newline.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws write newline

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T17:17:00.090198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:17:00.090198Z digest=sha256:3e64bd256f082c770c2e62209419e44d861eaffb0dcb082a6d558625f8c3a65c

Observation f655db2a-d1c5-43e6-aadb-7b6bcab2a6c4 · outbound

This paper cites @esa (Ref.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws @esa (Ref

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T17:17:00.096909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:17:00.096909Z digest=sha256:ddee39c311f3dbc352fc4532fa2cd9824b5cf4b3f22d042d5f3b8c1db5f16b58

Observation 1045d6bf-01f8-4215-b817-2a579dbdeb83 · outbound

This paper cites an unresolved cited work.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T17:17:00.102574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:17:00.102574Z digest=sha256:6407e25bc778c6f9f4705846b89fdce66990d3b4a7a143faac090680e8211826

Observation a1705959-45eb-49f7-9703-de5493a14c13 · outbound

This paper cites an unresolved cited work.

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:17:00.114314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:17:00.114314Z digest=sha256:0d0851326c80408ccfe2e1d8382c2d31be423cb03a85a8911bf0830e200c0589

Pith citing papers

Observation be29f8b8-dde7-43ba-8519-5dbca8c31da0 · inbound

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations cites this paper.

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T20:44:48.662191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T20:44:48.441119Z digest=sha256:61b92020286d8b90bb1f2c40a9ecfc09804a973b84ee48b4ab57a40fabafa153