Pith. sign in

Paper Citation Record · LEDGER

HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2502.12574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12574 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:23:08.655628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:15:19.863525Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 73b00083-ebe0-4ad6-b566-9c8e9e4d6279 · inbound

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference cites this paper.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.655628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.655628Z digest=sha256:b5779687915ed39611253f91971725b9772e0332e7f779d4693b12096b063ee1

Observation d906d757-0644-4a84-b9bb-0e41f73bdc2d · inbound

Beyond Grading Accuracy: Exploring Alignment of TAs and LLMs cites this paper.

Beyond Grading Accuracy: Exploring Alignment of TAs and LLMs HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T23:49:35.266340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:49:35.266340Z digest=sha256:dc327562d2887a46836c64c8597b99818eb6ac33a3620093d3e93385ca55945a

Observation 416a4a35-3e1e-4cfc-9481-b02eb2d554f5 · inbound

CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference cites this paper.

CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:53:02.178203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:52:25.148104Z digest=sha256:b35acd1159b2df0957d94a76eba1cb365f16129a9be93c29d3ca05e4a83bd734

Observation 92ff89c6-6fb5-45d0-af7a-8a8d5b296d4e · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.412506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:fe5cde5c258c0cfd9347e77760bb50ed1540650f9e403e06f703096e23984aeb

Observation c3615dea-185f-4c33-a72b-e2bac65e33f5 · inbound

Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning cites this paper.

Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:23.445841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T05:30:46.989747Z digest=sha256:d41c14bb731f375c09ab574c923e5e281b4b10f46e8d8e36614b22a9341dd51c

Observation ce3b0670-1ab6-40f6-a70b-f415b16f32d4 · inbound

CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference cites this paper.

CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:15:19.866746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:14:44.387842Z digest=sha256:cea07b5640877f705122e877370ecb5c27ae1cb7e3a66a26b1949cdb6d3aba64