Pith. sign in

Paper Citation Record · LEDGER

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models

As of 18 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2505.12216.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12216 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:44:32.136821Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38e2a364-1366-44e6-9eb5-38f3223ad163 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:31.996022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:31.996022Z digest=sha256:7b0996c8dbe61c08acac00dec82e8518b60425bac91193e03a7f4d3c3cccc82d

Observation 9ce29c8e-504f-47db-af1b-a2b46962e63b · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.000215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.000215Z digest=sha256:fd99aa132dcbaf2fbbf9752727c2198255ea2db07442e3aaf465b5fa87105a23

Observation 7a2efdac-6b4f-4fd4-a41a-6e89c241eca1 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.106556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.106556Z digest=sha256:da5d63c23bfb95019f53dbf8a71f89772f023dfc78502872c1545d09bd336445

Observation 4b9f4fd5-1703-47c2-8ff1-acd17001c7a7 · outbound

This paper cites EvoPress: Accurate Dynamic Model Compression via Evolutionary Search.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models EvoPress: Accurate Dynamic Model Compression via Evolutionary Search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.115409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.115409Z digest=sha256:4fec632c347fed7ab2457e0fd8d69872a5a50e992669491ca353e5c5a793640f

Observation d6712259-3d05-4354-8759-50afda2488a3 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.123193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.123193Z digest=sha256:2ffcff116d3b3eb6e10b4e199057bf83564b612dda86980004851b152a643923

Observation 0ac7c410-f328-4c0c-b2d5-37322410a269 · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.127470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.127470Z digest=sha256:b627bac84dbc3db1f35544880a01551a275610c58f8aa93d97c1c0a806730900

Observation 5556c198-136c-4e34-ac83-1475d5b3fda7 · outbound

This paper cites Transactions of the Associa- tion for Computational Linguistics, 12:1556–1577.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Transactions of the Associa- tion for Computational Linguistics, 12:1556–1577

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:32.280711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:32.132173Z digest=sha256:fe22d16d457cd69ae842a6a8733205c8b8ce147c0e1506e99b6d550cd1121180

Observation b9fb8bea-e17b-4f0e-a736-526f62c90390 · outbound

This paper cites SparseGPT generates spar- sified blocks with varying sparsity levels across layers.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models SparseGPT generates spar- sified blocks with varying sparsity levels across layers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:32.268129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:32.136821Z digest=sha256:df968fdc78b443521acdf24c9e085cc4eaebbc215469c2e9679db75df9d23847

Observation e7ef7af7-77ce-4ae9-8a5c-52ce8f7785a0 · outbound

This paper cites In 15th In- ternational Conference on Scientific and Statistical Database Management, 2003., pages 141–150.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models In 15th In- ternational Conference on Scientific and Statistical Database Management, 2003., pages 141–150

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:32.292280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:32.119355Z digest=sha256:ec827dad3d38cb501aa0fd4a88e576613afbaa22e93c323e879679e1477ad378

Observation 3a0d2829-fc77-4fce-a0fd-2f7d1ddc7fc8 · outbound

This paper cites Pointer Sentinel Mixture Models.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Pointer Sentinel Mixture Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.101934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.101934Z digest=sha256:61fdc30b99daaae0373762cf8a9d25fadb92c1b56d271f081d07fd05f7d012d5

Observation 4bdc28a0-2ff6-4cb9-b7f0-fcd5f2a05e0c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:31.986746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:31.986746Z digest=sha256:df492d82340b43a63183441cefa3c522fcbbf222b3044b2e254fbfbf6ff691a8

Observation 33b7391d-9ae5-4c39-82ad-7f837963e7c7 · outbound

This paper cites Weight subcloning: direct initialization of transformers using larger pretrained ones.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Weight subcloning: direct initialization of transformers using larger pretrained ones

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.111122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.111122Z digest=sha256:8ddf1419be26236bb11ecaef6765d856f17ac6f671458cfa5bfcaf1a3e8a253e

Observation 45bf08ae-5f02-4f3c-a600-d612d4339a6c · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models The Unreasonable Ineffectiveness of the Deeper Layers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:31.991534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:31.991534Z digest=sha256:1e34fa71165503bd0a537d9bb7446545f02d99b08acc26e7201f40ddefa521b5

Pith citing papers

No inbound Pith citation observations are available.