Pith. sign in

Paper Citation Record · LEDGER

Efficient Memory Management for Large Language Model Serving with PagedAttention

As of 5 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 100 inbound Pith citation observations for arXiv:2309.06180.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.06180 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T15:03:07.651839Z

measured 169 of 169 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 180 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T15:17:55.992637Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact16
  • verified fuzzy6
  • unresolved45
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 08dccdd5-98d0-49e9-8fc2-68424e18710b · outbound

This paper cites DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale.

Efficient Memory Management for Large Language Model Serving with PagedAttention DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:07.819111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:7226eb43b3f166f7d591d42f67ced93855f87ca0b964c46aff65c484b334e34d

Observation ea5d1a3a-06f2-4b50-92b3-4ff93d6f7d9b · outbound

This paper cites Layer Normalization.

Efficient Memory Management for Large Language Model Serving with PagedAttention Layer Normalization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.793926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:885f0321ff1ef951430ca2c6a30b7c0cc353c5bbde052702af3203557c673449

Observation 020a3db1-7f86-418e-a32c-f59844f53d99 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.961032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:0a2b58f1c6d38964e9c154815d5f8b1260b99a8aa60e533188c352b218563bbd

Observation a93fed71-3f98-4bbe-830a-fcc14605e2ba · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.964835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:90483c3707e29399f30dfaf62e7cb00cc4cf91a935eef44c6422789737ed910a

Observation dbca0eb2-521c-44bd-b4ae-1ac2f4b43420 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.968780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:96610212451838b6777634750d5fc8a74f7ad69d7ddbd7158d7ef6f742d5707c

Observation fa3293c6-1d77-49e6-addc-9ad3407bc999 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Efficient Memory Management for Large Language Model Serving with PagedAttention Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.724048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:c29236fac13b44fe6cab75fe1de8a34be76b98d0374ea246c4f375b131245abf

Observation 824e324e-c329-43f7-a910-2250e37ef738 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Efficient Memory Management for Large Language Model Serving with PagedAttention Training Deep Nets with Sublinear Memory Cost

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.738692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:c95942e5a57b43e52b07a5538a7b7736f07ed611f104310a6b8893936e8221c2

Observation a908b3f5-513c-4a51-bfce-6ff7685937dc · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Efficient Memory Management for Large Language Model Serving with PagedAttention Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:03:07.972494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:67cd3be37c92e0c6c605cf7b1d7502892af7e319f03d110c9cfcb0f0bba99819

Observation 7fc5054a-6c19-48f0-ba31-5b0a7cb2c84a · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Efficient Memory Management for Large Language Model Serving with PagedAttention PaLM: Scaling Language Modeling with Pathways

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.716855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:fbe411676b349e4e9fc948a6a124cad0c6f633d86d37a3939114c267f4bdbee7

Observation c5b99f57-bbde-4dcd-ba2a-b0c62cecee45 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.976033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:90f3a4ea1590fda593588f786112380cecf6d00fe1121e025d5aa323d1914404

Observation 3c795a00-0234-4a5e-9838-52853962a2c9 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.980069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:e55922b31f22e32bbc0d9c3f6737b81c3dbf5f9341b8f9b02cb3f74e5feb6367

Observation f60ccf6b-6748-4095-a8ea-2770d22cd497 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.983910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:4e0b3eb9fe7d7f843ad7d890e56bb4513910d9ee5dc21851c81135c7e5d354f5

Observation 5421875f-a657-44c1-bfaf-46a7ab549b3f · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.987914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:adc8cfcfdf2d80804b12da4f7a5d66427a1900d0bf9429361e8b504dda6c9451

Observation 8475f1f3-5c14-4de7-be9d-5744bad783c3 · outbound

This paper cites Advances in Neural Information Processing Systems 35 (2022), 16344–16359.

Efficient Memory Management for Large Language Model Serving with PagedAttention Advances in Neural Information Processing Systems 35 (2022), 16344–16359

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:03:07.991726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:680694cde69ef28a06486341f7c70b6d283abf018cb2e76fa65ce07e6bf4c2f3

Observation 22f35f2e-51c1-411f-8b07-44a03537b375 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.995879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:c42f4453cc27606a322cec69c5db5e99cf2dd30cb264b065203fc6b9423d6a15

Observation c0b92cc6-7b6f-4c43-ab9a-64c35c09c07d · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.999490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:25f2e5528b852c4d0ad90e46e02ba7454bd6438eca9f89caf60c1d8ae0f7abde

Observation ecbb2d0a-e0ea-41db-8bac-7d9947ce92af · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:08.003049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:b41b685860d89367c3f6bd9f1d1ec2e4d3405e00eae47c641b691e0ed86b1c44

Observation 7f0b4fcf-41f2-4301-87d0-aded4ffa0da1 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:08.006385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:84a49e2b01a933727c2e0bc993ee51d014e8a6cca1a6a75e7fbaff0367b13082

Observation 47d9e4d5-8cb7-4862-9c75-82ccc985a1c0 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:08.010300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:0593e95c703c8b894e0f47719d396232ad1fd31a9fc2d945b7161993262787e6

Observation 649cd9d0-3c15-4f92-8c73-69860ecf4dac · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:08.013932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:dcfc576cd4d03e8d904a7764f4649f49f2bc2f6aee2f24e9edeacb50662768cf

Observation 7171c841-29bd-4898-8258-adc70234a42c · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:08.017449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:d9cd2fa7f2e8a89ef43639a7ac09815829b3afc179b2bffef50af25c78cca6b5

Observation 41b994ad-0cee-4d6e-b3ac-4f54a8316759 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:08.021083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:5b467bf1dc452b71ac82d02a0808557d427df3d2c6c028ac27797c79ce7768b4

Observation 69f0cf92-1c2e-4a5d-8b51-f3002f882278 · outbound

This paper cites In 16th USENIX Symposium on Oper- ating Systems Design and Implementation (OSDI 22).

Efficient Memory Management for Large Language Model Serving with PagedAttention In 16th USENIX Symposium on Oper- ating Systems Design and Implementation (OSDI 22)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:03:07.824156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:4d13ffbee14e3e55507f47eeaa8dc3bb837be4f6fa04292cb79c2a161f32acaa

Observation 585c1e67-09ac-4976-8be7-93d98b0e0134 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.829337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:9eb4b370c4066de0b9e309a75ddace239a96b95491e5fddf7c45773d8967843a

Observation 6b4f3895-665d-4efb-8785-9c4633f6c278 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.833668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:6acb0e159faf14f3234fb6db60ecb981871ee9a95451f9ebbb0c2ddad14bd774

Observation 5db4699b-9297-40fb-b135-f11ec619aa4e · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.837719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:e218fed279ca07694f7e991be1036b2940949d6cb7e66282e73bd501cb7466da

Observation bf0d9868-991c-4e70-97a0-d86b8857242b · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.841979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:577ceaedf04ee9b11ba7dbd99170a87714d8792778ee29076f8c805aa359be90

Observation 4734b47f-db5b-4741-bfc8-23bc9e216ee7 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Efficient Memory Management for Large Language Model Serving with PagedAttention The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.753723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:854c6b5ab1b45cd28d8d4a6336570abcc8fcd8d8c1016e5f6d0273c3b48a42bf

Observation 090d8bc3-9a77-4681-8587-00d9448e823b · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Efficient Memory Management for Large Language Model Serving with PagedAttention Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.767558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:7aa6c7ebb3cef345582963c476d037837d24d0007ddcf5faf80ab09fbba25ed6

Observation 42994f68-102c-4f9c-a533-fa2ce75b0a5c · outbound

This paper cites AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving.

Efficient Memory Management for Large Language Model Serving with PagedAttention AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:07.780556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:af4f8898a8ab77050a50bfcb6eb48eb37a95949ea4d1f97a3e3ec6e33f2adbb9

Observation 6611d062-ea7b-4b1a-8e1c-a04ea637c734 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.846713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:3053d9851890c2f0fe819c3a637fc6aacc0040ba0fcf2247e31ce6a956ab77ab

Observation 7d738a63-1597-49db-8fff-96a626340e43 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.850427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:05746ad03d8984d7bd1accaec646d0c5c0245b8a8599511660696daa6129d5fe

Observation 20103cf5-3f6f-4484-8635-05f8794c2528 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.854142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:a7a40bd19c6849c3008ae787475bc91b33b8065fb4c4b2bdf54516563686dd6d

Observation ecff823f-3c41-4263-bcf4-36e0aadf93f2 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.858277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:90cd39982322db70a94cbb4c2182ac3f68b111b3ad337a29b7a276f35e58065c

Observation d242a37f-13c8-4bf9-9225-05daabbd8863 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.862706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:dda498a3717696a80c4a37ebcfeb627742c1870ea6686846cd59f9e5b7fbc105

Observation 6573c59c-e979-4709-aa4d-49841a90a965 · outbound

This paper cites TensorFlow-Serving: Flexible, High-Performance ML Serving.

Efficient Memory Management for Large Language Model Serving with PagedAttention TensorFlow-Serving: Flexible, High-Performance ML Serving

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:07.761225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:74102d259db5ec5e124eb3a2b6490892415221c1ee545452684c27de9c8a0d4c

Observation fe748ac2-8082-48f9-af31-e99b892a8716 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.867075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:11522b8c2da6a353acbca2e5e8defce8ea558afc6ca939624d979fc0f5e007f4

Observation 6970e59a-e3bd-45f9-a0e1-1ca53edf0330 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.870878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:554e7684988db4648c7d501a62140df9e9189fd68f80f7ea5110e7a967842c46

Observation 5f9a882f-fecc-46b0-880b-1058a20b94bc · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.874552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:4fd779dd0849c52ec144d3bb0f339d5cc16e4f7c04adfcfcd8b0664c48678b2a

Observation 30a47c21-5c3f-4249-bff7-a8cc348693f6 · outbound

This paper cites GPT-4 Technical Report.

Efficient Memory Management for Large Language Model Serving with PagedAttention GPT-4 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.708918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:899b3cf1d87f8cba65f12767969442721df6d1fd80726a243819929e8a781a74

Observation a1cc0d46-d309-403d-82ff-3d711eff3e3a · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.878270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:b0d1c0f38b37fd63c127bdb43acc81a5945e0da20e8647239b23884cc399659f

Observation fd199c6e-1b18-4516-8610-557303632b5a · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.882604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:0985b73dd84cacc520c5a9d2261d3fe9bd36ae75056b3bd96b60a8b62b87b640

Observation eefaaf41-bbe0-44e1-a186-ceee8ffa759b · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.886267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:284f988b5ee376efbc5e4ef4e6e39a2f45082e215c8b5023901266f3cc8233d4

Observation 659423ce-ded9-4c05-970b-b38ca297a6d8 · outbound

This paper cites Efficiently Scaling Transformer Inference.

Efficient Memory Management for Large Language Model Serving with PagedAttention Efficiently Scaling Transformer Inference

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:07.746609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:181e0473d259592e3984525f947c98832011121ee3deba290921c908462fa635

Observation 17795cef-1a39-41a4-b047-882332367102 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.890597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:775254eb23d57a79ce84b5b218912ef379e713f76331c3dc81295e399b625dc7

Observation 464d2b8d-cfa6-4cba-b873-82cf65375e79 · outbound

This paper cites In USENIX Annual Technical Conference.

Efficient Memory Management for Large Language Model Serving with PagedAttention In USENIX Annual Technical Conference

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:03:07.895483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:42696d8e60271ba5a5fc9daaed65f5cbe4baf66cdabe4344dfaa89f48fa2dcda

Observation 9e37e6ef-7bb2-4691-816c-8a6faf31ca45 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.899533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:2525af11d1ea5cd1b00eea33697efbe831cad871f85e4d6c355c1a3d9844ee42

Observation 27a18ba0-8337-42ad-8a2c-12987079e10f · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.903474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:c75287a091078125adf00b7bac0a03fbea6b48813adcc3c6356e207fc5441536

Observation 04f25914-1bb7-4542-b0fe-bc53824a95b6 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.907553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:feba1e5b4c6f2f873d2e93df903eed58ff40dcba001e9d079457e85f96596968

Observation b514fe48-8fb9-4ec3-9311-fb3b4b968682 · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

Efficient Memory Management for Large Language Model Serving with PagedAttention FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:07.800919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:c575696e41cfacd8154fa74368f77dbd7522fa2dd5c981536a11917996cc9f2c

Observation 3cb0556b-ffb3-46a6-a7b9-f38e2f5426be · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Efficient Memory Management for Large Language Model Serving with PagedAttention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.812810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:885d8d13f74edbdf8ef186ed2f83e23cfdf07f899dabde8e9396d4a204845f53

Observation 453495f7-8cb5-4a78-9245-7bd54665c69b · outbound

This paper cites OLLA: Optimizing the Lifetime and Location of Arrays to Reduce the Memory Usage of Neural Networks.

Efficient Memory Management for Large Language Model Serving with PagedAttention OLLA: Optimizing the Lifetime and Location of Arrays to Reduce the Memory Usage of Neural Networks

Reference 52

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T15:03:07.701241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:3903bdc2ba70174856cbd813c91ef7d2554df8e0793177bbf3f60f12ffc3278f

Observation 6272fdd7-76f8-4f80-8260-be9929539786 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.911343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:a93f47426c422fcad7b5b41455ce96b774d7210e13cd46d5f7bc37d6992e1990

Observation 18a53911-7127-4192-a747-a445c79dc7b9 · outbound

This paper cites Hashimoto.

Efficient Memory Management for Large Language Model Serving with PagedAttention Hashimoto

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:03:07.914879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:3c1feaede39beac5c46a0afa70b3e74b9da44e247ca487544c4912a0154759cc

Observation fbc8220c-365a-45f1-892c-eb1ab0919310 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.918408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:f12a9c72685b0f30427adc5a6949235c21d461fb04fc551d8b85cba1f9b2b24a

Observation fd745b65-5172-4485-b2fd-d57ee7c8f11e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Efficient Memory Management for Large Language Model Serving with PagedAttention LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.731206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:72c7be30cd64f8993ce9b151ffde68338a02918b5fb54cda44b6e5c12b87a56e

Observation df58395b-c622-403e-a2a9-7aa76c595b9c · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.922356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:ca4a7cad972e203ace8708f77ef3f25af4d9d3b278b4a90cfdc0a61feec4fa93

Observation e39f4097-4adb-442f-a6bb-c8bbc759b2d2 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.926067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:3c25a9045db3adebbb336574042238352fd68819a2a43de5a8d3153e04a5d732

Observation 75222c7e-88ef-4087-bdfc-ce459d8a17e0 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.929825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:94e44f87e74dfd792aa7cf0f3b297ae8b39174bfbe940bcd63075c92507d583c

Observation 06833f31-c947-4493-95c3-cd1f6d8c7e6b · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.933604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:69d751770c10578136b508753c1586f08f8e072a2bae8a251dd61a7eff9e5e4e

Observation 4e8f5a54-6b76-468d-8568-6e67e32a93a4 · outbound

This paper cites In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies: Industry Papers.

Efficient Memory Management for Large Language Model Serving with PagedAttention In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies: Industry Papers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:03:07.937977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:cfa0cdac541d962237d00b0807c2381f87ced47f04128e15833ddfa00128aa95

Observation 53f81328-4859-432f-9bf9-3a24ff5e5060 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Efficient Memory Management for Large Language Model Serving with PagedAttention Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:14:51.325303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:9380623b62110ad55d24b9b9179603f93e7e1a909c203dd4514ac9309ff4dba2

Observation c21b7371-c2d6-4eed-8a18-faf9cb7ef895 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.942106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:51b72ecb93809206816602f3477f8c329f80d16e2002587004ce9b56c643dc7c

Observation 0bfcdb19-46a3-40da-b3db-a74e46841850 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Efficient Memory Management for Large Language Model Serving with PagedAttention Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:21:29.028955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:5cc562a32812aef80ff42d4703dacc7c0e7295c786b2f4877a42f60e948014b1

Observation 7ff26025-541d-4a6a-90c8-18935974fb7a · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.945945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:068dad98b75b5e3f125d0bdc0e78fbfb680a5e8f98c6a16b7974ec6e7fe4fedd

Observation 84b51dc0-35c3-45c4-b7bc-e799058e8007 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.949746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:bcff96dffaef1871adf182f55f047126af8ca807fa945660daad1a697df59243

Observation 75626fe5-d11e-4f77-8cd3-5785faa1cee0 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Efficient Memory Management for Large Language Model Serving with PagedAttention OPT: Open Pre-trained Transformer Language Models

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.806825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:2f2c0d6549fae6fb270b4bce7eb2205989d42a8c18092c2842cb197deb9f66b3

Observation cfb028ce-306c-4577-a080-51df01dc9fe1 · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.953903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:2bad41161f4a8eef068ca5408be24e073fbb9f100c3b1ddf19e5c77bc6ce8ed4

Observation 7d18360e-836f-41ae-9d50-4b3b6f8916ca · outbound

This paper cites an unresolved cited work.

Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:03:07.957425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:ddcf68f8c87a9846f07ba59d1d54d61531e6d237bb9657086aafad0958443c05

Pith citing papers

Observation a09cee7c-e6fb-4a8d-8a7f-8e6db35e1ce7 · inbound

Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs cites this paper.

Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T11:11:21.608774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T11:11:21.460613Z digest=sha256:ff56e7c5a54a613dee990813d33b65a38084fa373ea26b28daf4f3d0514ee406

Observation 44a29f19-7905-4a89-a7ad-7673f5708cf5 · inbound

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection cites this paper.

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T14:15:10.907921Z digest=sha256:198fa999dc3f8ef128181604232da14214bd52f2c2fdea0410f1913697af2b12

Observation e90940b9-37a1-46b0-bbfb-bbd9e07b8577 · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:47:27.868094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:daa53e62a3673281cae7ff5e962a77527be34134e24c26d637b452f39e0c2eca

Observation fd045068-0ad5-4337-b55a-975fd7ff9702 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e1d6afd09662bb930f04502e416fe30f71e6fb20c53407a30a21c82ed5d799bb

Observation 84d3b738-2ec1-488c-a7a9-80e4e4f191b3 · inbound

Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies cites this paper.

Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:32:50.776285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:31:49.857895Z digest=sha256:faa629a066ac323ec004e656e2a068a74ba5213ed98f995fe2eb4b5aeb645aff

Observation 1a5eeee8-d283-4b6c-88b8-539ab86f7d08 · inbound

YAC: Bridging Natural Language and Interactive Visual Exploration with Generative AI for Biomedical Data Discovery cites this paper.

YAC: Bridging Natural Language and Interactive Visual Exploration with Generative AI for Biomedical Data Discovery Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T22:14:23.425786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T22:10:43.859139Z digest=sha256:b8bc9bc72ff6c2dc6212a612d2943d1e786eba89f6b3a2e1376b3227477ccf81

Observation 6e306f6b-cea3-4bfa-824d-2702ab97836a · inbound

Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services cites this paper.

Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T15:01:31.493359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T14:59:38.194894Z digest=sha256:354a93e83fc003208cf3acf762170e7d10025134dab7f90f5adb1a63c3c09bda

Observation 8352fb9f-2190-44d3-9802-aee7ad2f0fb8 · inbound

Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning cites this paper.

Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T14:41:29.206997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:2e98ee5d6212ea6c6baae72416d251229f824a0d2fb64aa5f5af22a548a85d4f

Observation 16ce4b77-a270-4a88-b99e-4a8ff69b7ace · inbound

SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs cites this paper.

SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T13:41:25.122310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T13:38:29.809045Z digest=sha256:a4f49e220d0b93133bb7f0b8c7d29210ca6228b7d824ec3c372a06465b8942c9

Observation fa27d06a-24dd-487a-a6f2-55b48744a6e2 · inbound

SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs cites this paper.

SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T15:17:55.992637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:17:55.992637Z digest=sha256:284cd8e86b995f71d70d5d0c6d0d3a592fda7fbcc55a31870f594b6722001f59

Observation 3f5828a3-4a97-4d7e-acc2-6b66459574f5 · inbound

OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models cites this paper.

OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 58

Resolution
malformed identifier
local_arxiv, observed 2026-05-18T10:11:14.059288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:09:08.307199Z digest=sha256:3b986d33fd737eb8b05de0a8fd1fc470b231878ed2d6097cb880cba04bbe2216

Observation b09c3e4c-bd94-49fa-a538-a7ac835e08a2 · inbound

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation cites this paper.

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T10:06:13.713378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:04:39.223895Z digest=sha256:f1107476f989e76666d433b4a4bda34433c4612f960b261b4abfb7d68858d86f

Observation 0fdd69aa-b0f5-4dbb-93be-cdb0a736e4fc · inbound

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit cites this paper.

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T09:11:09.985826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T09:06:25.049493Z digest=sha256:985d65f3f359060578e65dc0a903fb11a6d4bff9060bf497b55c28f4ab43f7e7

Observation f7b06a5f-0f9a-4511-bc40-5fe4f38a3c81 · inbound

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit cites this paper.

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:16:46.216400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:16:46.216400Z digest=sha256:9cf3a86d94dfe8267811d8dcbb530b31edc2631a151adc240383b23deb1c0ae6

Observation 91c3ed04-8613-45e9-bef5-9f8ce13c9ed3 · inbound

Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting cites this paper.

Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T00:32:58.992599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:32:58.992599Z digest=sha256:e53c1bf566840fe99c09c3b8757d92e84fcc7ea514316a6e50213b25338d08b4

Observation f76bc791-9b01-41e9-a8b3-5726bb1c4118 · inbound

SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators cites this paper.

SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T02:00:39.352517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T01:59:30.583027Z digest=sha256:01a3ab0c6280a47a46ba00c48baaac7dd77e4ceb23cd876e41f51102b194d9e4

Observation e12f028e-a27a-49da-8839-6559f4736ef6 · inbound

CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization cites this paper.

CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T23:30:25.092756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:30:25.092756Z digest=sha256:f51c0d1b29b3f9415253d5be800ab63b48f8ba36c276111071774499b8a13aeb

Observation 8a04c1cb-01ee-4d22-8017-2e79f859b5a6 · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:50:30.208417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T18:46:04.926179Z digest=sha256:9f57b183dc46084763936f8869224c82c5f55cac23144b00d620b48318452955

Observation 7bae5c0c-e4bc-491b-952a-6a0e910ed617 · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T23:21:39.362759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:21:39.362759Z digest=sha256:560be10bb3ed91389ea2e9e29cde6386a3bd9ca043d3a3f71b682d65945d305c

Observation df5a381d-ad51-42b3-9a13-c7d69b8bf89b · inbound

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation cites this paper.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:15:26.729709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:9646ada5056d37a6bafa56668d60da5448d306c47e0c69ae739fa40412264202

Observation 1ffd6fc1-8044-4e95-8afa-9db274f6991d · inbound

Parallel Decoder Transformer: Planner-Conditioned Latent Coordination for Model-Intrinsic Parallel Generation cites this paper.

Parallel Decoder Transformer: Planner-Conditioned Latent Coordination for Model-Intrinsic Parallel Generation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:22:04.070787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:22:04.070787Z digest=sha256:12ca3536158ecbfb3778ec3d6f8e0413a56ea68578e9f7a728aa4e6e17e2d6e4

Observation 21438ac9-8943-42c3-a3a1-ee107a5e15fb · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 111

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T01:40:42.554096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:707ee23e596b7dac737c680e3c4c6fbf58f44200a133d4aa88cdabdfc80fe46c

Observation dabb644d-35e5-417d-9992-56045faa4ed3 · inbound

EVE: A Generator-Verifier System for Generative Policies cites this paper.

EVE: A Generator-Verifier System for Generative Policies Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T14:10:35.629946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:10:35.629946Z digest=sha256:c8bea17e6be4ae45703f19ad3819fbc2b323a3032374e6fec479c9cc52d366d8

Observation 83f47a5b-8cb9-4c34-b86d-fe721771d66b · inbound

Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models cites this paper.

Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T13:19:47.766001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:19:47.766001Z digest=sha256:bcf2a6265ac9e12e2e6429faa99eb92458aef7b776d2918a96303c61f1564cc5

Observation e32ca7d8-4c0b-4a10-93a1-8148a62a2950 · inbound

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist cites this paper.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.908771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.908771Z digest=sha256:7a78398e61fced02c4eef07ce99d6576338a27e09383ef91379a2ff90cd2ace5

Observation 79106028-0630-4467-b74d-e580f9316f61 · inbound

WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching cites this paper.

WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:12:58.654290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:12:06.034679Z digest=sha256:97962c050e8ef8e3453610825aab4198a237f7335cafd7cda6baa8759e9e8e11

Observation 55908e28-d2b4-4fda-8dfa-a68a54216ec3 · inbound

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch cites this paper.

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:12:54.825297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T13:12:01.889341Z digest=sha256:32d701243a4123523cae2518c8381b93cd1ce2aa074984320e30cab154e1c0eb

Observation 0145f447-cc1b-4b41-8253-502b0bac7606 · inbound

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs cites this paper.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.354187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.354187Z digest=sha256:041d083d9a0dd1178eccb0502114dc3ae073e6047b216d420dc76449883a8023

Observation 59a51cf2-28e8-489c-860c-e93b170f367c · inbound

Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis cites this paper.

Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T08:03:24.576463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:03:24.576463Z digest=sha256:6b8f2ed5a15448a461d84b65674e024843b6f4955e9ab3833c71cf35e5b04a50

Observation 53a6b8c7-9528-42b9-8dcf-2db0fd1a2a66 · inbound

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference cites this paper.

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:07:29.603499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T07:05:56.380048Z digest=sha256:602590f0700762b7022f4780b867403cf80c219e004a6695665a9ed56204ce45

Observation f5e16b7c-324e-4620-9fcc-84330660ff6a · inbound

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System cites this paper.

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T23:32:12.337470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:32:12.337470Z digest=sha256:7e6c850ca279af6e4581d6259e11eb5bcb6348e5ce6805aa57f90325483aaec8

Observation 9c365dc1-ac4e-4c10-9549-d0a342db390b · inbound

Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows cites this paper.

Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 714

Resolution
unresolved
no resolver link, observed 2026-08-02T23:06:03.369714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:06:03.369714Z digest=sha256:2824e7c716ada03bb7083bae99f8cf24be27fff9174fee03ed780284626847eb

Observation fd911f9d-4ae1-4fbf-8312-e4656226fca4 · inbound

SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation cites this paper.

SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:46.493319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:46.493319Z digest=sha256:932402c259757a5d3078d7b948ab5ec3e2300601a698179c0594d44a33b866f4

Observation 52ac49f6-aba6-43bb-97c0-4fc05963c61f · inbound

MemFactory: Unified Inference & Training Framework for Agent Memory cites this paper.

MemFactory: Unified Inference & Training Framework for Agent Memory Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:03:28.661682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T00:02:14.821285Z digest=sha256:dcf8548622efead06f5e5b842f17acfe754da7cbbd8b222d1ff661c9534c7cac

Observation 00658988-ccf1-4cf2-8028-d7dca793c69e · inbound

Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning cites this paper.

Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:23:15.999195Z digest=sha256:fc5450f3c3373ef491f0bd0cc945fd8e606678d2cb78f415c5277c5657a0b6a9

Observation e4a8daeb-d2d5-4605-9713-562afeb2418b · inbound

Reduced-Mass Orbital AI Inference via Integrated Solar, Compute, and Radiator Panels cites this paper.

Reduced-Mass Orbital AI Inference via Integrated Solar, Compute, and Radiator Panels Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:07:40.881627Z digest=sha256:504eb86ba65efc9d889a159eee3155c0a6b87d6f3238ea9d38166b2e3205ec8f

Observation e52651ee-5d58-4c47-aa5c-48ab9dbc99c4 · inbound

Benchmarking Compound AI Applications for Hardware-Software Co-Design cites this paper.

Benchmarking Compound AI Applications for Hardware-Software Co-Design Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T16:20:09.892247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T16:17:08.045511Z digest=sha256:da97e3117f9044baffda4e9a30128d680d585672db08712b9e647e06e4fbfafe

Observation 6d71ee04-a1e3-418d-8710-12d1b98f8dc7 · inbound

SLM Finetuning for Natural Language to Domain Specific Code Generation in Production cites this paper.

SLM Finetuning for Natural Language to Domain Specific Code Generation in Production Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T16:38:30.184565Z digest=sha256:ba5a2a42b674e040f0e6c3d592fd8a8154c618058bb5e036bb12b04f5b01fc13

Observation 10a927ef-35f0-4dea-bde0-019c641a9532 · inbound

Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search cites this paper.

Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 35

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:55:04.835146Z digest=sha256:c706d51e4c89ddbff833fff17780221cdd6e91157fc21790c8915fe4043ec19b

Observation e61aa8b3-c2d6-468f-a5e1-0615aa80d029 · inbound

LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models cites this paper.

LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:06:03.504804Z digest=sha256:a1b82125c8850cfdd53b254b0795294002600a632a1a8d8efcd2bddd23330fb6

Observation a87505b2-e5e1-4561-9e1c-d1794343e3d4 · inbound

Local Rules for Directing the Emergence of Global Properties in Complex Structures cites this paper.

Local Rules for Directing the Emergence of Global Properties in Complex Structures Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T21:37:29.667707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:37:29.667707Z digest=sha256:2658998a4c6bf87283d8997a3eedd1728d4f83aabd6323d4611a9dd4e47c775f

Observation a964f57f-d5ad-4310-aeaa-d56d086dd49e · inbound

Hierarchical vs. Flat Iteration in Shared-Weight Transformers cites this paper.

Hierarchical vs. Flat Iteration in Shared-Weight Transformers Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T12:59:34.289283Z digest=sha256:856a8f37cf9b161ef00d504c9a1e95e439c8665a9c5f577ddc317bc25c109623

Observation bf9c7bb6-5a24-4be6-a233-0cba20a5f1e3 · inbound

LLM4Log: A Systematic Review of Large Language Model-based Log Analysis cites this paper.

LLM4Log: A Systematic Review of Large Language Model-based Log Analysis Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 133

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:25:19.087407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T08:21:04.998443Z digest=sha256:a4249a6f9ac508c043a1bd9d606eaa4bfaf79c88b328b76dbab168f0d6fadeb0

Observation f8510faa-f26c-49eb-9d44-bcd0b9754c9e · inbound

LLM4Log: A Systematic Review of Large Language Model-based Log Analysis cites this paper.

LLM4Log: A Systematic Review of Large Language Model-based Log Analysis Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 133

Resolution
verified exact
local_arxiv, observed 2026-05-21T10:04:06.618589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T10:00:03.036921Z digest=sha256:c31786f96053c9da2d3eb23bf6bfe52bd8ea98e187312b1bd9f48adb9301f1d8

Observation 49a2065e-77cc-4752-894b-5225250cbd7d · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:5d21b8f413749adb899ccb0c9939ff5e6beb6b7ce7e3f7013dbf0ec2ef12e7c3

Observation 4806eeca-cb61-45cc-9895-cf72e050fb07 · inbound

Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon cites this paper.

Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T07:18:26.936652Z digest=sha256:d5d7d0167031bb243dc75bbd435890aa3bd4cc8b79731547b1bd301f35c73763

Observation 969bc451-96bf-41e1-85d8-f8af4cfde9ca · inbound

Neural Garbage Collection: Learning to Forget while Learning to Reason cites this paper.

Neural Garbage Collection: Learning to Forget while Learning to Reason Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:57:12.923258Z digest=sha256:b5c7fb055af4f78b85c5ea5705e8fc7be04d855cdcf3958b6c5c8221d59434f5

Observation d3a847f2-9f36-4042-bdbf-bb6d3ec55701 · inbound

PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training cites this paper.

PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:42:16.848186Z digest=sha256:cd17417cc9ea56b724563b3ea4c8d189c6808fe6d81c7992f2d99dd6cd7ddbe1

Observation 147a6f06-0d99-4e6c-af08-006dcfa5ea24 · inbound

Secure On-Premise Deployment of Open-Weights Large Language Models in Radiology: An Isolation-First Architecture with Prospective Pilot Evaluation cites this paper.

Secure On-Premise Deployment of Open-Weights Large Language Models in Radiology: An Isolation-First Architecture with Prospective Pilot Evaluation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:23:22.366584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T00:22:23.136915Z digest=sha256:a826bd44ee387c3e7d8df77f42822173303adcce5ad38345dcacb93a2bb542de

Observation 56f30735-a94e-4f8b-83f8-bf582ad11a9b · inbound

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling cites this paper.

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:39:37.485602Z digest=sha256:7877708058dd3681af70c3d675f7ad459595b4d3c894e54e91ae058563ce69d1

Observation 7e876114-b1b1-4689-b641-77fa36e076c3 · inbound

CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration cites this paper.

CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T15:32:32.808799Z digest=sha256:95b3ff139d96fe49095acf3aa659bfda8dd5ac99b0b5c8a58b3510063a1ba65c

Observation 88eabef9-6c84-43c2-bf68-82201e3f6ab9 · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T08:30:35.264804Z digest=sha256:d7de5a5bcf2d1fb30c959daba859933c6222d3d5bce2e62a377015422fb5be21

Observation a5cc1126-612b-47d4-950e-78cad4bfa8b0 · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:55:34.739614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T08:54:41.845820Z digest=sha256:880eef23153808df68682491227345d36e044b9681cd5de422bd6a392ca60470

Observation e09ce1cb-c9df-48ce-bcec-b1abddffe5b7 · inbound

Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference cites this paper.

Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:23:21.456692Z digest=sha256:4877dd1efcd2f8534fae1f2282415af97aaa58e66262000754f88fbde153b6be

Observation c0185c61-96ba-4647-8b6e-0a23c9aab31a · inbound

StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k cites this paper.

StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T18:44:42.456111Z digest=sha256:e90981496ec813250ea525fbb3b93e61c9ec98fbb4cc104b1bd20b24aef8fc3f

Observation a75bdfc5-750c-4315-8524-cfeecea13224 · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:20485122c77ab21329d16cc69c2c22c39a7991523eb7776be762dba0dae37a49

Observation 6cb5c210-50d6-49c3-a800-a33efa8ae065 · inbound

How Does Chunking Affect Retrieval-Augmented Code Completion? A Controlled Empirical Study cites this paper.

How Does Chunking Affect Retrieval-Augmented Code Completion? A Controlled Empirical Study Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T17:26:49.023109Z digest=sha256:864f43a391e7c52c78bef8e627f026b12014eb3f1a96a8405158700cb1469d1a

Observation c992578f-207e-46e3-8e13-f0010e0a8f38 · inbound

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving cites this paper.

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:49:24.880528Z digest=sha256:1ba69edf7929d99b1097254010e96c03cec525dd3a6e3cb6dc80c1a5d75a847a

Observation 32b523ad-def9-4fe4-950d-d00400fcfe55 · inbound

KL for a KL: On-Policy Distillation with Control Variate Baseline cites this paper.

KL for a KL: On-Policy Distillation with Control Variate Baseline Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T03:31:07.462474Z digest=sha256:8bd6f9cdc4719b1150d80df28cb36f49a90ae81a76b6f5d4a10e7dbd71740a35

Observation c4817d65-5487-43c3-916f-1ff8151a4cdc · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:d41740917ebe2d2632671445a6a38ddb1bbc5a3e96ce6dc39d7f5ab92865ee0f

Observation 49752ddc-d2a5-441d-a809-c9c8cb5f4677 · inbound

Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents cites this paper.

Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:04:55.618417Z digest=sha256:ad5c201df9d393311b3041ee4e51fa679557b68c46d6894f164cfd2c0c4455d2

Observation c2d228a4-6dfa-4e57-9ad0-9d888623c70c · inbound

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes cites this paper.

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:53:36.091037Z digest=sha256:60999fcdf515b33fe05e18bb3726a094210dc93d200e109474659ce17abfdce8

Observation 91304a7c-a0e0-47fa-9aa1-2eab54439058 · inbound

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes cites this paper.

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T05:09:46.085858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T05:05:12.267070Z digest=sha256:f61037184b052af609c84b53efbad0a0fd9260286e85c035cef5ebcc8e757b2f

Observation b4764dc5-d6b6-4fad-808f-2a81a08cff98 · inbound

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference cites this paper.

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:30:58.729425Z digest=sha256:84d078874c46f871cec2dec0fc30b5276ca39d0eeb8d789e4ab54b2f27ae15eb

Observation 22f72be2-9076-4ace-9a53-8d280b79fc15 · inbound

An Executable Benchmarking Suite for Tool-Using Agents cites this paper.

An Executable Benchmarking Suite for Tool-Using Agents Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:52:06.217699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T01:08:46.020900Z digest=sha256:d1ff4e41c29f4d3197b483ce6cc842bb40ed1788913853873d9e7bb968e67e9e

Observation 43461c72-10b5-4330-91ed-4783fa6158c8 · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 55

Resolution
malformed identifier
local_arxiv, observed 2026-05-13T02:22:06.763384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:8ba39124d8470ba446760891088755720df4c293bee2d4cfa8ea446b104c71e4

Observation 2f701fc2-3453-4ce1-866f-64a16063b7a8 · inbound

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production cites this paper.

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:19.432534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:17:24.147248Z digest=sha256:255bd71156c6215d6868b63ee9d0ba141091184dff76b40829f645023a442235

Observation 98e77faf-e7b7-497a-ab03-9963f6cb71b2 · inbound

The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures cites this paper.

The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T04:52:17.201341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T04:48:46.006687Z digest=sha256:3b38feddd80e98f16593440fdea204d244c1fdc3e532cecd1afb5353a5575c83

Observation 69df9bfd-99df-488c-954f-f41289fc3d83 · inbound

NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding cites this paper.

NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T03:37:11.490744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T03:32:28.187814Z digest=sha256:14cef212febf0521f0a45691b59cd1a04e3d9c6f8358692a61ede1d5f269ee71

Observation 3fa28c05-8db8-4bc8-a1ad-319aee585b18 · inbound

OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning cites this paper.

OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:05:06.027944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T21:57:38.752094Z digest=sha256:6419d754bcb9c87273ee4276d5204814cd696e35b0c791b34085eb77afe8eec9

Observation c54eed09-b2f9-45a1-a00a-7ce456243ab3 · inbound

Generative Floor Plan Design with LLMs via Reinforcement Learning with Verifiable Rewards cites this paper.

Generative Floor Plan Design with LLMs via Reinforcement Learning with Verifiable Rewards Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T05:05:02.434663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T05:01:22.151862Z digest=sha256:9f4bc95754c09c6a68681edfb1a47fd5dcbfd3779109c5aa31f45af3eb2df5ed

Observation 96c141c2-fe94-4275-b0e6-076a59498710 · inbound

Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation cites this paper.

Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T20:55:04.034104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:52:32.567107Z digest=sha256:ece774c69fc86ffa39f9e09b44d25dcf9dd61b1579376463de88b5d0fcee0c17

Observation 2402af29-3a3a-43ca-968a-235d89c6b7e9 · inbound

MeMo: Memory as a Model cites this paper.

MeMo: Memory as a Model Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T03:19:43.362686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T03:17:23.202604Z digest=sha256:16249dcef4786ee7eeea782013d760a883642f16f4b7df6f3073245c9b20decd

Observation e84f747f-2ce1-48b9-87ea-656b6fdaa540 · inbound

MeMo: Memory as a Model cites this paper.

MeMo: Memory as a Model Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T08:39:53.593616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:36:31.022046Z digest=sha256:81ddf5c7fee6e89c2de53fa89a86dcd28ec72897dc0e26eae5fa0e1e26e09a0b

Observation e058c1c4-0739-4b7b-9e41-db8e0a35eb58 · inbound

Policy-Grounded Dynamic Facet Suggestions for Job Search cites this paper.

Policy-Grounded Dynamic Facet Suggestions for Job Search Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T21:37:47.818605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T21:36:48.182204Z digest=sha256:dae082d9a40f66679631df7236d3346ed0432e681982b9afd4d1031f8cafc678

Observation 3d04e3f3-a7b8-40a4-895a-56ecab78a467 · inbound

R2V Agent: Teaching SLMs When to Ask for Help cites this paper.

R2V Agent: Teaching SLMs When to Ask for Help Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:28:59.748806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:26:29.804427Z digest=sha256:531fe44a631a3096bf6f1c995750f1e52a94d04b53488344d4d2c8a25e5420f2

Observation 09ddf539-c9e4-441e-b92e-8413f0c31e30 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:13:21.193394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:11:58.106397Z digest=sha256:c9f769791e9bbd1e893d632de7c5e8a154d6a3caaf93b89995d644f2077faa74

Observation 7672d4f9-ed35-468c-9705-6137f79c72c3 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T08:59:55.801689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:55:31.298030Z digest=sha256:a147a3f3ec0221ff810b250114762e68deb270f578cdaed0238bffd91591236b

Observation b99c7520-b32f-4b93-a5ca-c30ad5988397 · inbound

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference cites this paper.

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T22:12:51.171962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T22:09:15.782104Z digest=sha256:51bc7f25dd75527f0f75ca03319c736e827e27c1f4038944ed7c7e93ce414c1d

Observation 72746384-2afe-4780-a416-fdbe9be3c996 · inbound

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction cites this paper.

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:13:15.959185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T12:13:14.295659Z digest=sha256:4c66a0ee00edbdaac2913539aaf4eee9ca3c7ed59ac2bcfff1690acca4ee5ddc

Observation 03917308-d875-440b-9684-266e62329fc9 · inbound

KadiAssistant: A conversational AI Agent for information retrieval in Kadi4Mat cites this paper.

KadiAssistant: A conversational AI Agent for information retrieval in Kadi4Mat Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T21:43:45.328808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T21:41:01.078882Z digest=sha256:96f6bdd853beede42ff51640b53ba14986310ece102e989a1b19c2648efe3edf

Observation 14b8c744-e225-4851-b42e-957d4467c921 · inbound

CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing cites this paper.

CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T07:08:07.070717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T07:06:01.621365Z digest=sha256:fd6df64f72c4bde2c69b9bd255ca393859b2f2a79def71ed24c31ca38e75c74a

Observation b011f7fa-9874-453b-8649-d5dc50d75c65 · inbound

LLM Agents Make Collective Belief Dynamics Programmable: Challenges and Research Directions cites this paper.

LLM Agents Make Collective Belief Dynamics Programmable: Challenges and Research Directions Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:38:04.623177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T04:37:01.368357Z digest=sha256:dc1a3f30c46f116365869ebe1fb7b43277b9f13087b22b788db009a1569dda08

Observation 2846e575-b3af-49bc-85bc-7c86f40ffed2 · inbound

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR cites this paper.

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T02:29:25.352053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T02:24:48.872065Z digest=sha256:a7f223896fcf38532212f102754c37891babe19d905600f13ddf6a6adc43801a

Observation ad17b4d6-3f03-4b2e-a6e9-6cd28c9f670e · inbound

Runtime-Certified Bounded-Error Quantized Attention cites this paper.

Runtime-Certified Bounded-Error Quantized Attention Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:49:41.119692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T05:45:42.295527Z digest=sha256:29f06bacf472da16b0700e4a9500ff01eb5f71e24ccaf0ae100c098696165ee8

Observation ccdd2310-86f0-4a47-a25e-e61b8922d932 · inbound

When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning cites this paper.

When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.608066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T09:25:37.960991Z digest=sha256:20ac08520d5fada10838945e6661c888ce4e79e48be864c6db24e94a34d32863

Observation 85dcd48c-f689-411c-b247-c166ab2f5352 · inbound

TO-Agents: A Multi-Agent AI Pipeline for Preference-Guided Topology Optimization cites this paper.

TO-Agents: A Multi-Agent AI Pipeline for Preference-Guided Topology Optimization Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:34:46.904666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T09:33:43.733883Z digest=sha256:44e2c636932358606476a69923b8a33f16cc6aa005769e59175b2169129e5aa0

Observation d5bdee5d-6b54-491d-9c74-b70c16387b58 · inbound

Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference cites this paper.

Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:31:14.086093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T07:28:35.020478Z digest=sha256:fd9d3662e47e1ca2c0d0ae75931fd1949b9f8cc36625987b53de93879731a14d

Observation 34c4f42d-43e7-4577-8b93-866d8b07cbc4 · inbound

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks cites this paper.

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:44:48.300700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T15:42:21.405913Z digest=sha256:2e6696698b933a8350ae608ed9344dd06f19122586a377f7004f72d153bb3beb

Observation bd290409-6fe2-431b-8ffb-8ac5e58023bb · inbound

Resident KV Claims: A Conformance Contract for Future Reuse under Active KV Pressure cites this paper.

Resident KV Claims: A Conformance Contract for Future Reuse under Active KV Pressure Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:24:45.205057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T14:18:35.209571Z digest=sha256:954767354486200793bd5bd6d6103f2bc6a444fb638530d8ab0fc0c5a0c40675

Observation 7d7ed762-0779-442a-b35b-a5737d33461a · inbound

DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving cites this paper.

DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T20:53:57.592169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T20:46:51.896596Z digest=sha256:6e572b25ac53e4feaf8e50569e746af76c67c2d535ead591dbea1aeb742bcd1f

Observation 0465413f-89c2-4459-a945-f5b3dacbc004 · inbound

The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models cites this paper.

The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T17:44:57.816321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:39:16.469049Z digest=sha256:2104760fbcb14c2510f56d51e43b907a76d1334cfb2a8751fea92b9a0250feb9

Observation e0df5a6f-056a-48b6-bf7a-342314744147 · inbound

ARBITER: Reasoning Trajectory Basins and Majority Vote Failures in Test-Time Sampling cites this paper.

ARBITER: Reasoning Trajectory Basins and Majority Vote Failures in Test-Time Sampling Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:23:59.947279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T22:22:02.802499Z digest=sha256:69314f07d8b40a87e1b82ffd2895c9e4137928d6d5ac56f51ee326fc07f5bcc0

Observation 8e49efcf-4d58-4f29-b07a-0610410dff25 · inbound

A Unified Structured Query Understanding Framework for Industrial Semantic Search cites this paper.

A Unified Structured Query Understanding Framework for Industrial Semantic Search Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:54:44.842416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T14:51:33.630394Z digest=sha256:026e26c4feaf17b945beaf7bcfc2410a6fe4bea0a230bf0416c0cb44badcf7ee

Observation 428c5f72-fa48-40ad-a416-05461800ea45 · inbound

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers cites this paper.

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:13:48.664439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T18:08:42.533804Z digest=sha256:42eab55a5db070757700571f3773086b3d31e4b22222685ea339f58bc6d5f5a1

Observation 2524a5f8-f395-47b8-8e88-026ff4d2f44c · inbound

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving cites this paper.

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:03:47.859616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T18:03:19.798168Z digest=sha256:f2d075abddbea96ba0c63e12fa8ba71cbf1c5f59a0569062b979cdf7756c9b4b

Observation 4fc5ae9e-3a64-42ef-95cb-bd2983a67c28 · inbound

MRMMIA: Membership Inference Attacks on Memory in Chat Agents cites this paper.

MRMMIA: Membership Inference Attacks on Memory in Chat Agents Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:13:27.042754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T12:04:53.511344Z digest=sha256:5601ee690e6ecd3a91017a6574402af1c6152b2b10ae14cea58091b95c845995

Observation 86cf3189-d1ac-40de-9ed0-67bd3350a318 · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:14.945933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T08:40:10.152344Z digest=sha256:4a20ce6bffd84ce33281e9f35733fe9597966c3f8f52e0d829ddf0de6d43d1c9

Observation eb8ab6dd-fd67-44ed-bb74-56c90c3eace4 · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T12:54:41.542034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:54:41.542034Z digest=sha256:56aa1e74b9522deda566030a7fcc2d6d8ae214dbfe63df69ccf4ad4b82f0bffa

Observation 88cc7254-93c5-475d-bf14-610287411264 · inbound

Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward cites this paper.

Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:32:44.005264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T22:30:00.735630Z digest=sha256:cf9b117cf6f5a407a6969bfaab4b3a442511236a64e6eb2c2595bd56c77455c8