Pith. sign in

Paper Citation Record · LEDGER

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

As of 16 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 4 inbound Pith citation observations for arXiv:2501.06709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06709 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:58:08.498820Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:45:32.847427Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved5
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 70838d38-4447-49df-a558-80af7a61cc6b · outbound

This paper cites Brown, B.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Brown, B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.653172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.213684Z digest=sha256:9573cdfb959a0d8a1827d6fd20f5f4c82d8fea41ce1a38a9c771eb5beeaf6061

Observation e9b3d3ab-052a-46ce-b479-6114c41f6f03 · outbound

This paper cites GPT-4 Technical Report.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.221119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.221119Z digest=sha256:ea46e19938dd18cdfc46d00f3368f841c8776d281922df3e42386ef1a11f0a29

Observation 2642f22a-d560-412d-a3b7-041a69b91997 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.227140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.227140Z digest=sha256:8eefafbe72b11307467cbe8ee7ea1862f31ae04188a702aac3fc5732b77503e6

Observation c8b27bae-ca69-4c32-aab8-045993c3a9e1 · outbound

This paper cites Characterization of large language model development in the datacenter,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Characterization of large language model development in the datacenter,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.623471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.234544Z digest=sha256:ca6b5ac5657bd6ff0697140c3b5cf2a496db86531a6dd16d26361a65fa685629

Observation a9b7d5f8-4f80-4113-9703-d45e9911c44a · outbound

This paper cites Orca: A distributed serving system for Transformer-Based generative models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Orca: A distributed serving system for Transformer-Based generative models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.600018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.240643Z digest=sha256:0d1502ccde59641c9faf31b8719ab09e4e026e2b00b55702fbc0f45f7a2af46a

Observation 5a53d823-1a65-4e4a-859f-72126ad12211 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Splitwise: Efficient generative llm inference using phase splitting,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.576977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.248315Z digest=sha256:8cc723063169bc1ad1ec858029f5ed8ce86b9605ad1e3cc34ef0b12841a57500

Observation d3401967-7f6b-4079-82e4-878eb18686a4 · outbound

This paper cites Optimus: Warming serverless ml inference via inter-function model transformation,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Optimus: Warming serverless ml inference via inter-function model transformation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.552481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.254850Z digest=sha256:d2e3acac810972f6f9088ff47f0a91daa8fe71f2a8f787688063a92f5496face

Observation 995dcd7f-b0f1-4a81-8ef3-f9f7b29ef31b · outbound

This paper cites Otas: An elastic transformer serving system via token adaptation,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Otas: An elastic transformer serving system via token adaptation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.528062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.262315Z digest=sha256:d3815bbd6fdfd1099e25013d28ec75f0785699b49851d7be56acd8ecee9f1c16

Observation c51ead8a-cd79-482f-9ad7-581802a112da · outbound

This paper cites Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.507764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.267931Z digest=sha256:c2f24b5fac1e0114b94654cbaa7fc27a78ead0a46ea4ede2d9acc0f9eed10b01

Observation 49e0f30b-0397-499f-9da2-81df2420cc4a · outbound

This paper cites Efficiently scaling transformer inference,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficiently scaling transformer inference,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.484082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.276900Z digest=sha256:e87debb334d97c53cdfb84329bc93751cfb3f74f571c581324a773ef4ee55c31

Observation b98744f7-ea30-4b0f-8c29-375784e0e37e · outbound

This paper cites LongloRA: Efficient fine-tuning of long-context large language models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management LongloRA: Efficient fine-tuning of long-context large language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.461417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.287438Z digest=sha256:dafae2b577052e8ce96269cdb08956a4cb53a68f32df54e5d0b252377cb096a0

Observation c4a6f511-e017-42b4-aa5f-014cac17659c · outbound

This paper cites Efficient memory management for large lan- guage model serving with pagedattention,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficient memory management for large lan- guage model serving with pagedattention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.433817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.293309Z digest=sha256:ca189621dd223d05da21a075dba72cb6745bd7db6351659d4a635978878733e6

Observation 36fe6834-c623-4a4a-b6ee-baf3aa39282c · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.408044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.298553Z digest=sha256:07f07f05a6909fc55f579b620a194303d6f05eaab394e70414f5f6dc0107596f

Observation d00bb195-ef32-4aaa-85a2-30ee35428d73 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.386345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.305038Z digest=sha256:bc51567e4900430d39c29450eb82cfe215c0ee0de9c84941f3866f63bde904cd

Observation f871c23b-05d6-4ad8-9056-f3344e591c07 · outbound

This paper cites Efficient streaming language models with attention sinks,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficient streaming language models with attention sinks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.359107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.312308Z digest=sha256:d7bbdf1042423ca8e24b5a272220744dc78baa15c7d9f719989f33d59083f829

Observation c370e43a-eb65-47c6-908d-b817162cd06d · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.339682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.317601Z digest=sha256:a805f47349506a5ecdfb1f07a9c7bebaeeac13984cab2409ec12aab70bc245f3

Observation f6791e28-2a64-42b7-901d-bd756b1bba73 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.323613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.323613Z digest=sha256:a9d50c8caab26337993ad2763244ec6940a5954e8cac6fabe364a77ee04777ac

Observation d7d8c44a-c3e6-421b-9e35-83e93800a937 · outbound

This paper cites Model tells you what to discard: Adaptive KV cache compression for LLMs,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Model tells you what to discard: Adaptive KV cache compression for LLMs,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.320919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.335736Z digest=sha256:fbd6dcb908042c40d0262f2105387c3be16f4251dcd5caa1d556b21765e1074f

Observation 92d5fa76-fd6b-45b5-8513-c5cddeeb8fac · outbound

This paper cites Flexgen: high-throughput generative inference of large language models with a single gpu,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Flexgen: high-throughput generative inference of large language models with a single gpu,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.289777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.342949Z digest=sha256:931d86f73a46f2e7a67ef9949f1a931c223f68a86311ebb45ec365d8e6db44a3

Observation 21a3dae1-09ea-4da9-9f75-686ecf9d61e7 · outbound

This paper cites Infinigen: Efficient generative in- ference of large language models with dynamic kv cache management,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Infinigen: Efficient generative in- ference of large language models with dynamic kv cache management,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.265498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.349106Z digest=sha256:0c01ef72e57d228dc3ef992706796e88e585033b45717e317e77719e76a11396

Observation 1668b1ff-fff7-482d-9975-177f6099fe7d · outbound

This paper cites Cost-Efficient large language model serving for multi-turn conversations with CachedAttention,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Cost-Efficient large language model serving for multi-turn conversations with CachedAttention,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.241225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.356264Z digest=sha256:680f56ef086350d832de8bde0ee5225767b0f8877affb2adb3d6851b3c6ea091

Observation 17a3a078-f1d3-4e2a-a5b9-b39cccd494d5 · outbound

This paper cites Deepspeed- inference: enabling efficient inference of transformer models at unprece- dented scale,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Deepspeed- inference: enabling efficient inference of transformer models at unprece- dented scale,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.218362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.363065Z digest=sha256:ef4cd0b57ae2a729ac3aeace13ad8f8d9bbde0f5f1f83b2e98880d1ed2842095

Observation e7f660de-3832-4453-8725-a33dbb518132 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Llumnix: Dynamic scheduling for large language model serving,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.193397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.369100Z digest=sha256:b5e9c25248bf7d0ce348d3a560c3f208dae763e2518e0d667419aed4e3212bee

Observation 9057a506-236a-438c-90e5-4d64452995a6 · outbound

This paper cites Serverlessllm: Low-latency serverless inference for large language models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Serverlessllm: Low-latency serverless inference for large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.173370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.374194Z digest=sha256:6258a9c38c50ebc4d55203b9c032ffa30886a0b4e7d34b5d628077a4c8eb4bb6

Observation 19df26f4-727a-4b8a-b54f-a3a7326ff70c · outbound

This paper cites Turbotransformers: an efficient gpu serving system for transformer models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Turbotransformers: an efficient gpu serving system for transformer models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.146756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.379805Z digest=sha256:5bdb7ea00c00a4af2bdfb8fb2e2a7d04d569d5d50c94d069984a253a8a52f884

Observation 7e83071a-3a68-4526-94d9-b8d3d42306be · outbound

This paper cites Taming throughput-latency tradeoff in llm inference with sarathi-serve,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Taming throughput-latency tradeoff in llm inference with sarathi-serve,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.122150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.385167Z digest=sha256:e8fcf5d28ce7455285d16d6fb69f7874c5580a180c29eb86a3808ab2859ff3e6

Observation 3cc9f462-3e53-4d19-876a-a09fa998e019 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.086283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.394284Z digest=sha256:6f32cf1c5eb0d35e901ffa6868296b4b25e147ff862bf1e1443d1dd9afacb9b1

Observation a9042608-ac4d-4cde-97b3-35f7cd5fc350 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.401041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.401041Z digest=sha256:445834e4dfa2553fd6863596211b8f68d837e959fced9ade5abbc610bae29094

Observation 3405f68f-6281-44eb-ba11-2d40878f4747 · outbound

This paper cites dLoRA: Dynam- ically orchestrating requests and adapters for LoRA LLM serving,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management dLoRA: Dynam- ically orchestrating requests and adapters for LoRA LLM serving,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.061030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.406766Z digest=sha256:ff8c06943e5d8bedb5e2bc826bed491160db7f6e10e2c18172598c7737e09418

Observation 7546fca3-7b42-40bd-9489-c5b6ea9e92ce · outbound

This paper cites LMSYS- chat-1m: A large-scale real-world LLM conversation dataset,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management LMSYS- chat-1m: A large-scale real-world LLM conversation dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.033618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.411772Z digest=sha256:bcadded8cf74fb3368fa36a3a8a127e34cd51726b914b729db132d0ffe800b2a

Observation 2f10efc9-75f1-4584-8455-1b3a8a0032fb · outbound

This paper cites Wildchat: 1m chatGPT interaction logs in the wild,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Wildchat: 1m chatGPT interaction logs in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.014848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.417714Z digest=sha256:1be4e6e2619a31123348aa30624377f42e0dfb71b8f3998dc7ba4c308608bc0b

Observation 207f5195-686a-466c-a0d1-ae5395559b42 · outbound

This paper cites Judging LLM-as-a-judge with MT-bench and chatbot arena,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Judging LLM-as-a-judge with MT-bench and chatbot arena,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.995713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.423055Z digest=sha256:6df64b3fb526547552488d380aed79aec19ef250fa59d18da3474fb38e767e38

Observation 7b768cf5-06c4-4aab-8dc5-4266b15924aa · outbound

This paper cites Koala: A dialogue model for academic research,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Koala: A dialogue model for academic research,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.429322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.429322Z digest=sha256:19467275c9ccf6167add96d7c1227c04076a54027068151230a5c7e21bb7af70

Observation 37831d52-d151-4bfd-b0d0-10ecd0a38cdc · outbound

This paper cites Response length perception and sequence scheduling: An LLM-empowered LLM inference pipeline,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Response length perception and sequence scheduling: An LLM-empowered LLM inference pipeline,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.965068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.435506Z digest=sha256:f531e2b58b72e802efc978d75030c9414a96313b150ce11b28468dbde3083496

Observation 53e40a07-cfdc-4642-8716-f2bc105a4cce · outbound

This paper cites an unresolved cited work.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Unresolved cited work

Reference 35

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T20:58:08.946688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.442068Z digest=sha256:e446557ee9cd6d5a83f79c1c16ca92a246732f98bc17f9bc7ad7e16112e8d885

Observation f145516f-1315-48cf-b63d-90fa4948dcbb · outbound

This paper cites Adaptive resource provi- sioning for the cloud using online bin packing,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Adaptive resource provi- sioning for the cloud using online bin packing,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.928833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.448959Z digest=sha256:9f300801b2397d15b9343c1d0bc8bf9ae91fb0a8efa8b97e824f5d671987997f

Observation 4abe209a-994f-41df-9d29-5bf5a9aac970 · outbound

This paper cites Efficient online strategies for renting servers in the cloud,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficient online strategies for renting servers in the cloud,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.908341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.454685Z digest=sha256:392094893c5f8b352939bfde0a040e7709e337d1bf63e0d90697a530a2b544c7

Observation 71d93163-586b-4e8f-8e20-fef44de7b11d · outbound

This paper cites Powernap: eliminating server idle power,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Powernap: eliminating server idle power,

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-10T20:58:08.460910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.460910Z digest=sha256:a8ddfc71d13900ff36d58c4f440d5ebd2d3da06957fea09e5d433a4e93ffff73

Observation 99c54bf0-4ef5-498b-9932-3f39f792bc10 · outbound

This paper cites Easy, fast, and cheap llm serving for everyone,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Easy, fast, and cheap llm serving for everyone,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.885991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.466106Z digest=sha256:4e17eca88c95e502b2ab213842ae9641660435d57c29cd79f4785d1a85e8a592

Observation e58c2e49-e9eb-416a-a8f6-336b99cdc3ec · outbound

This paper cites Ray: a unified framework for scaling ai and python applications,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Ray: a unified framework for scaling ai and python applications,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.859845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.474077Z digest=sha256:2f096b1c0ac7c73522165cddbf9a8fc0a9a4e784a69f76df863a2ddb2e890221

Observation 8ecb158a-0e6c-4f84-a2c1-6053038575de · outbound

This paper cites Gloo: Collective communications library with various primitives for multi-machine training,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Gloo: Collective communications library with various primitives for multi-machine training,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.810006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.486606Z digest=sha256:d7cf1938e1c67e34af962fd2e64a8701d3829b1157524aeea200b2af51b394be

Observation a36461d3-72c8-4209-b669-8f95285db950 · outbound

This paper cites Openai platform document,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Openai platform document,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.790533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.491510Z digest=sha256:ece0818c232c6888551f3234569fbbd820cbe55cbc018fe1c9820c6ba22a69e5

Observation 8f967afb-874f-4d4b-957e-8dee2255b2e4 · outbound

This paper cites Anthropic platform document,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Anthropic platform document,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.769470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.498820Z digest=sha256:60cf9165d7285d5561cb15871552e53e13fdd829d65e3a1f515304f12ee92178

Observation 4a2d838e-b267-4d3b-8d9e-a7e1500860a8 · outbound

This paper cites Available: https://github.com/ray-project/ray.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Available: https://github.com/ray-project/ray

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.836802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T20:58:08.479825Z digest=sha256:7891cd5c179fc80c60b0c1e443c9e6e40c6775dc2c233daf11069345b7d0d8fc

Pith citing papers

Observation 01ad7691-2c93-438b-9311-2acf41d527b1 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

Reference 260

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:49.523271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:49.523271Z digest=sha256:ad1e6910c657bfd6ba91d96e8c3956f694da70498d9745464d783c56c7eb6e6e

Observation 80459404-6387-4874-b2a6-aef638b43d33 · inbound

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning cites this paper.

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:45:32.847427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:45:32.847427Z digest=sha256:885c15accb84400f7adaa725d1416e37e74e8d28e80d682542b9427790f1b854

Observation 0b31f61f-9a83-4899-bc23-8de4f04438e7 · inbound

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache cites this paper.

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:46.874022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:16:49.292491Z digest=sha256:7d9140bb685c3e2107605afa7d004a718ee1795292c1a85ad145a92e627df49d

Observation 87c72b18-1423-4970-92e9-e9b572cdfcb1 · inbound

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs cites this paper.

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:55:56.631328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:55:56.631328Z digest=sha256:df204d64874d0747141dfb626a5ea9807f37717b119ebf7983c82865f62d0141