Pith. sign in

Paper Citation Record · LEDGER

vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2405.04437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.04437 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T02:48:35.073858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.326797Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4cbac68f-45a7-42af-88ab-5209a864596e · inbound

eLLM: Elastic Memory Management Framework for Efficient LLM Serving cites this paper.

eLLM: Elastic Memory Management Framework for Efficient LLM Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.910745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:41:45.544399Z digest=sha256:0f83961ffb5085088acbf219cf686cb06d0024f167b6ed1f63531267a2b52cb3

Observation 0f85f268-14ee-492e-93a6-fd93f1c5731b · inbound

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving cites this paper.

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T22:20:51.838309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:20:51.838309Z digest=sha256:179f3fddfda2c8e2e9fb87b9ed40a0f7fcc8eb3394fc46b66f835903c7c34ff2

Observation deea181f-38ac-4696-886d-b92bb444cd29 · inbound

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations cites this paper.

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T09:09:32.932051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:09:32.932051Z digest=sha256:598f39e9ff533915e92b4d8bffcee29290a2f4de77a9320c7b68d11027c0fa33

Observation 9d0d39ea-3a51-4936-870d-f1591f511c78 · inbound

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference cites this paper.

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:52.993784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:45:26.555395Z digest=sha256:9f15ca583c5b9eec7be2b5165bbb414472bbc82f92da32c7dfbe44a131210be2

Observation 1421fa3b-f095-4da7-88e5-d953020acd85 · inbound

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving cites this paper.

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:27.063199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:57:30.104609Z digest=sha256:0c8df428d75d168e4960afeec7cfc5a412f62c01c44430849c077d98e7f060f8

Observation 85c56d4b-216e-499d-8bbe-7f370039632a · inbound

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving cites this paper.

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.450518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:05:44.256565Z digest=sha256:0010fbb515073ed5955ea3a7c133d9b7dae061a9bc6e87eee9e44772ebec91a6

Observation 861b87d2-2475-4f59-9a8c-87be60491b3a · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.541010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:859a4cf63a8e4ed07284dfd285751947c78a7982607ef37dfeab1b185b1761a3

Observation 420d3c7e-fe8b-4d58-91c7-782b6db6bb72 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 187

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:22.213295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:22.213295Z digest=sha256:02746f52f830c86dac7d1b83147f318dcf1f237ab3f1b53064ac82b2b99ff947

Observation 2ead7c60-3ad8-49b8-a333-4a79cadf0d32 · inbound

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving cites this paper.

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:29.923241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:39:47.477338Z digest=sha256:a3162b793ccd719a573d96c31672374d41b2a95235c1da8252d352fed34d04bb

Observation c6687955-f488-49d7-b25e-5c4ee1f5523a · inbound

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers cites this paper.

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T10:40:08.018921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:40:08.018921Z digest=sha256:fa0859f9e4d6eb5bdf98ef785a309df263255df3987f5bf29034f5b3744128a4

Observation 0778cff1-9044-4a61-a263-8fbf6d1c9011 · inbound

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers cites this paper.

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T02:48:35.073858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:48:35.073858Z digest=sha256:0f13b1e4b2dc5e64abea1213f35d91195992e63ff94e339a6fd3577ac3b3d535

Observation d400d8ad-2f2c-46c2-9112-85c7cd931a85 · inbound

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs cites this paper.

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:50.308008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:23:08.673299Z digest=sha256:af900a656181318288be5322c796ba4e6d9c6a175b35bacc8865c2e47fd5b80e

Observation 1b961289-45f5-427e-93d7-a5c03cd83641 · inbound

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs cites this paper.

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.419736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T21:04:43.874528Z digest=sha256:f4019e95d8ee061589954daa77b5ca74fe7869d59c32a96b97caeb86f6869586

Observation a56edc14-4e35-4800-815f-b329b3a2a82f · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.327959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:7a005fac69d8c8f5d440af509741047205e0b3bbd09c9bfc026adf9b7f6a0108