Pith. sign in

Paper Citation Record · LEDGER

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer

As of 23 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2607.28150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28150 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T16:21:19.946822Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved86
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a1b21e5-e5f4-4521-aeb6-f6de3eb6b26f · outbound

This paper cites Phi-4-reasoning Technical Report.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.065335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.065335Z digest=sha256:c81b1a77848573e05f78102b4a10417116971e5f92703301a54af4f1c2b93d63

Observation 36e1e90f-7664-46b6-9f7b-327600ed9d06 · outbound

This paper cites Gulavani, Alexey Tumanov, and Ramachandran Ramjee.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gulavani, Alexey Tumanov, and Ramachandran Ramjee

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.156641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.156641Z digest=sha256:e583517b9e4e32a8127d893806375eacd7db36ae42af8a35f889ffafacc95ebc

Observation 92ea6a80-45d0-4643-a3d6-1f1627dbc47a · outbound

This paper cites Infiniband architecture specification volume 1 release 1.8.https://www.infini bandta.org/ibta-specification, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Infiniband architecture specification volume 1 release 1.8.https://www.infini bandta.org/ibta-specification, Accessed: 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.216603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.216603Z digest=sha256:0a2dec5f71ccfc8e4143da9f078c2cb574b197100689d3c4ab9896a49c9637db

Observation b1bd9d1a-57e5-47ab-a13b-5e2095273bfe · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer LongBench: A bilingual, multitask benchmark for long context understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.259742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.259742Z digest=sha256:ba51c0e618939c90adf397cb312f4370907d5ed01d0a901788beadbe10fc48a4

Observation 90f88231-f561-42ce-b31d-034983fb7af3 · outbound

This paper cites TokenFlow: Responsive LLM text stream- ing serving under request burst via preemptive schedul- ing.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer TokenFlow: Responsive LLM text stream- ing serving under request burst via preemptive schedul- ing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.360134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.360134Z digest=sha256:233eb0acef260003c2b1c23a43f25ba389227bdffe17d1e39a41c05c1425e1c8

Observation acceb86b-4b7c-4af9-9190-945eabe05800 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.479848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.479848Z digest=sha256:c06fa75cb9fe552159f6c464d9e4f29f2d37bf73ff630bcc8bf7c53103f1d11e

Observation 20c871b3-a6b3-431d-bc2a-69a82736f6ea · outbound

This paper cites Retroinfer: A vector storage engine for scalable long-context LLM inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Retroinfer: A vector storage engine for scalable long-context LLM inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.584938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.584938Z digest=sha256:b41be14544bda3ceb8059d7b1bdc1c67375e3f259fe00a0cb4c01f17260055de

Observation 16178e22-2491-464b-88e5-19fd8f266fd0 · outbound

This paper cites Elastic GPU service instance families.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Elastic GPU service instance families

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.721313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.721313Z digest=sha256:b76dbdeace649d3133fd555f7f34247ba2e3d60f69cb0c6f3057a9ee76b9e004

Observation 18c0429d-1695-4288-acce-c52a2a710416 · outbound

This paper cites eRDMA.https://www.alibabacloud .com/help/en/ecs/user-guide/elastic-rdma-erdma, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer eRDMA.https://www.alibabacloud .com/help/en/ecs/user-guide/elastic-rdma-erdma, Accessed: 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.833863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.833863Z digest=sha256:80112f9bdb601681e8887ec53a0c42113ccb1a81ba5644c6f0cccc7ebf883939

Observation bd7a11f3-3dd1-4e26-8a27-dbf2bf135df9 · outbound

This paper cites A2 ultra machine types.https://docs.c loud.google.com/compute/docs/accelerator-optimiz ed-machines#a2-ultra-vms, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer A2 ultra machine types.https://docs.c loud.google.com/compute/docs/accelerator-optimiz ed-machines#a2-ultra-vms, Accessed: 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.959928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.959928Z digest=sha256:42c6e3ff727b6d2dc4855572ecd04f4e893bf8e7602c7edea98b98b1e2d314b8

Observation c50a4781-be92-40b3-a776-e8fc18801422 · outbound

This paper cites Computing instance.https://www.te ncentcloud.com/document/product/560/19701#GT4, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Computing instance.https://www.te ncentcloud.com/document/product/560/19701#GT4, Accessed: 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.064781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.064781Z digest=sha256:1ef0b9955b623f490db44631258beb25e2c1c825247527b0984de3346a1cf9aa

Observation 5257c786-1809-4850-8017-37a35a81bf17 · outbound

This paper cites NVIDIA dynamo platform.https: //developer.nvidia.com/dynamo, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer NVIDIA dynamo platform.https: //developer.nvidia.com/dynamo, Accessed: 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.135063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.135063Z digest=sha256:c239dc8b70f084d56634c957c277ad7861f438289c8bad4eb16e16feb00e2427

Observation b9706a4c-a753-493a-aaba-c22ae4175ac1 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.200993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.200993Z digest=sha256:7f7dfbea3a1ecac1360503f0dde78a58df439561417ad14da49ab54283517dd6

Observation 16a4b3a7-8577-4ef3-9b2a-abe8b57b95cf · outbound

This paper cites Deepseek-v3.2-exp: Boosting long- context efficiency with deepseek sparse attention.https: //github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/ main/DeepSeek_V3_2.pdf, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Deepseek-v3.2-exp: Boosting long- context efficiency with deepseek sparse attention.https: //github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/ main/DeepSeek_V3_2.pdf, Accessed: 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.262681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.262681Z digest=sha256:f446a8bc1c46436a48dc5af984bd02b8528a8846639bab55abbdb6d1f94c1462

Observation 7ef8f0b6-e997-48f8-89f3-4205d6870e1e · outbound

This paper cites DeepSeek-V3 Technical Report.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.321893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.321893Z digest=sha256:2c42159fbd4486202dd4e5fa9595576e0e202706aebf08b919db921df7bcd9be

Observation ab829f9a-1c01-4adf-9ac4-70e779731ecc · outbound

This paper cites Pre- fillOnly: An inference engine for prefill-only workloads in large language model applications.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Pre- fillOnly: An inference engine for prefill-only workloads in large language model applications

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.361409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.361409Z digest=sha256:36392d05b734b5a783063d62bfb70a7b7d972aa4d733679454e3c7a867444738

Observation 8eeae7a4-03c7-47c3-9a1c-e632c27c51a4 · outbound

This paper cites The design and operation of CloudLab.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer The design and operation of CloudLab

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.412761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.412761Z digest=sha256:dc78a14d9cd563553624303bb7ad8195d74d6c1adcef37966ee5dcb0c0ec6dd2

Observation 7f4ba4ca-34bf-4642-ae95-a4168b81eb4e · outbound

This paper cites Graham, Artem Y.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Graham, Artem Y

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.463354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.463354Z digest=sha256:33c3107d104bcedce0c2524daa98a54ccaf914595693c84e7a7dd2a933ebaa69

Observation 68f34fcd-5a60-45f1-81f3-51f9404a14e6 · outbound

This paper cites Cost-efficient large language model serving for multi-turn conversations with CachedAtten- tion.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Cost-efficient large language model serving for multi-turn conversations with CachedAtten- tion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.530574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.530574Z digest=sha256:81732cad2e550336fbb62ecb52ab32c48c266bc1ca7c423e24b4911653e477c1

Observation bdfec739-9496-4adb-8558-49c9238dc7a1 · outbound

This paper cites Fast state restoration in LLM serving with HCache.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Fast state restoration in LLM serving with HCache

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.606072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.606072Z digest=sha256:0b5296f86fea7b1474baafe1c1c8e45432e2a773899b29ee649fd077f7fd6f02

Observation 91e69d8b-16b4-4014-adb1-f4fafcaa5a53 · outbound

This paper cites Weaver: Efficient multi-llm serving with attention offloading.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Weaver: Efficient multi-llm serving with attention offloading

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.665888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.665888Z digest=sha256:9b3367b3196d85611e5b1531b8c09038221a78d6e22f0c5b667e9fa2033d0d87

Observation f8146844-e22f-4abb-8533-1663cabdc776 · outbound

This paper cites Hybrid multi- document summarization using pre-trained language models.Expert Syst.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Hybrid multi- document summarization using pre-trained language models.Expert Syst

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.762133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.762133Z digest=sha256:3e3ae973a30f60f350c306f056dc37ea5410d378d837cdbe7998f4aea24a3937

Observation c050fe89-1eab-40a2-805b-e11282e526ba · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.813892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.813892Z digest=sha256:036f6c9c03f4c68671f292c68a246033fa59c8978eac475bda4b53d2d5b946e3

Observation 219f1b3b-220d-4a35-92e8-f0d330fcdfa8 · outbound

This paper cites HATA: trainable and hardware-efficient hash-aware top-k at- tention for scalable large model inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer HATA: trainable and hardware-efficient hash-aware top-k at- tention for scalable large model inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.882652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.882652Z digest=sha256:0b230e16d1b21bb6d424ca69cbf7022bb9d8eaf0a2e9dcc5ef0dd16ad98a825d

Observation db2f6316-a3f1-415c-84ab-0a00d025d108 · outbound

This paper cites Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.049340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.049340Z digest=sha256:9754daa9518ee752a0d8f36975504d67270f8ce7f12e91e1c6e8eec4995feb60

Observation ec1bfb5d-b426-49a8-90d9-e39889c655cf · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.163297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.163297Z digest=sha256:5178bfa9a54afc52fabef5f81cfa17bc3a65371c41709842ae27619b8797366e

Observation 768defe2-4ffd-475a-85c2-35e87c35f7f0 · outbound

This paper cites OmniKV: Dynamic context selection for efficient long-context LLMs.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer OmniKV: Dynamic context selection for efficient long-context LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.336284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.336284Z digest=sha256:458ea7f767058f1d11b078c2ce64a9c9520ea51859e2ba00be105eb12a8b1ef6

Observation e3109b47-3e74-44ff-8ce5-7f54b51ffd9e · outbound

This paper cites WaferLLM: Large language model inference at wafer scale.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer WaferLLM: Large language model inference at wafer scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.452112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.452112Z digest=sha256:4807ad30e0197e16c9ea732d1b06d147f4bc442fb6b2671ff23598b235847ff1

Observation 3c1493b0-fd7f-4b3e-9cfd-d922e82bcefc · outbound

This paper cites Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.600544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.600544Z digest=sha256:f10ed8b1f511c32ba61fa9c2d95294070d88b93d7b8506fd5ed662f0b7413348

Observation a198e9ad-aeb4-4d46-9291-79305597b009 · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.757015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.757015Z digest=sha256:8754136c0a2dffbff5930e1ba7f67404bbd5707c8641e236adcc5999dc66a62c

Observation 88f9298f-b660-4f3d-8b42-61c58cd035e9 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.862617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.862617Z digest=sha256:1ba69f67ffb34a6b3ec17e45ff9cb7c68b196435f2afd62311b9407a630d0a1e

Observation d39c685a-19f3-414b-8a27-6b841a53c3f4 · outbound

This paper cites Efficient attentions for long document summarization.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Efficient attentions for long document summarization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.985128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.985128Z digest=sha256:8533ea64999754903b516abb65084c816f7e911015aeddb53a051d2d13d20806

Observation b40def8a-c8e6-42a3-a7f2-cc4873d9822f · outbound

This paper cites Kamath, Ramya Prabhu, Jayashree Mohan, Si- mon Peter, Ramachandran Ramjee, and Ashish Panwar.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Kamath, Ramya Prabhu, Jayashree Mohan, Si- mon Peter, Ramachandran Ramjee, and Ashish Panwar

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.132083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.132083Z digest=sha256:4af4178e7126dea85a43eb891c6ea4b719c7c2d523075540f248284c3a6819c8

Observation 8fa02f65-eb89-4e6b-bea6-7c9341a4a8a8 · outbound

This paper cites Oaken: Fast and efficient LLM serving with online-offline hybrid KV cache quantization.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Oaken: Fast and efficient LLM serving with online-offline hybrid KV cache quantization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.245943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.245943Z digest=sha256:d77791d526c6df0f8e2b651aa1e238c1c97b405f5ed4a0394f8da74147e50303

Observation f7ba4d0b-9423-4b9d-913c-c231645e266c · outbound

This paper cites Aqua: Network-accelerated memory offloading for llms in scale-up GPU domains.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Aqua: Network-accelerated memory offloading for llms in scale-up GPU domains

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.349109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.349109Z digest=sha256:fedfe6d3c1e29bf141145f3e47d8eba57ad1d320c2f17f276e0a6d2be64906bc

Observation 24cbc52d-330e-4989-9f4a-90e53d334753 · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Efficient memory management for large language model serving with PagedAttention

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.510728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.510728Z digest=sha256:cbcc94b02d2840cf3713c76dcf19d8c6e081589124b49061dcab39dbfcd5e71b

Observation 17372a46-8264-4d95-afef-b95fe357b59f · outbound

This paper cites InfiniGen: Efficient generative inference of large language models with dynamic KV cache management.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer InfiniGen: Efficient generative inference of large language models with dynamic KV cache management

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.620746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.620746Z digest=sha256:0218c5f44de090e3238db2637e3203bd5f9891543358c76a076452cbaea21eb0

Observation 480830a5-bd9f-437d-af28-cd77bd7bdc71 · outbound

This paper cites ClusterKV: Manipulating LLM KV cache in semantic space for recallable compression.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer ClusterKV: Manipulating LLM KV cache in semantic space for recallable compression

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.687846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.687846Z digest=sha256:ee46760f76b2ae9e6bcae71d68f57f16b15e4879ff1346086d394ecc53769c87

Observation 9168d745-635f-4f6e-8fa4-8e125894b1df · outbound

This paper cites Cachegen: KV cache compression and streaming for fast large language model serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Cachegen: KV cache compression and streaming for fast large language model serving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.781253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.781253Z digest=sha256:8bc2cdfed92d3a542f5ade977f75fb4404cee52fc29ff3dab636cbf6cd259aee

Observation 6d4893e6-b746-483c-bb52-f4cf6674b763 · outbound

This paper cites Helix: Serving large language models over heterogeneous GPUs and network via max-flow.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Helix: Serving large language models over heterogeneous GPUs and network via max-flow

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.839066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.839066Z digest=sha256:eaff94952aecb933138e3f752783720c03c33c2baa4e548ed0051905ca463f34

Observation 4a679fda-2331-4e5e-9665-2264f296e165 · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.924508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.924508Z digest=sha256:47b8213f45c1f1aadb50714b58ee09c7f9c43651d1a1bc9a5f4b820e368277a6

Observation d78d0d49-af54-4119-9d9a-f3f00887bd17 · outbound

This paper cites Heterogeneity-aware cluster scheduling policies for deep learning workloads.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Heterogeneity-aware cluster scheduling policies for deep learning workloads

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.042849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.042849Z digest=sha256:d4ff9d2bdc33ea1b94bea599865671143d03aa770e2096f95ee1ee8521bc3af6

Observation 71bf4129-a72d-4573-bac6-9efb850ce811 · outbound

This paper cites GPT-5 is here.https://openai.com/gpt-5, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer GPT-5 is here.https://openai.com/gpt-5, Accessed: 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.128417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.128417Z digest=sha256:5a9652c03c081a656b1b20786aab3924193c40c9859739fa83f131eff22cbb57

Observation b7f1b082-5cf6-48bb-9c06-8d658b5b644d · outbound

This paper cites InstAttention: In-storage attention offloading for cost-effective long-context LLM inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer InstAttention: In-storage attention offloading for cost-effective long-context LLM inference

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.241941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.241941Z digest=sha256:5a7b6843c109d44299646cc1a149656d937d7103a57bb6ec0e9115ab71e71c59

Observation abe3057a-892d-4051-bb43-c6f60a269d71 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Splitwise: Efficient generative LLM inference using phase splitting

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.318560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.318560Z digest=sha256:a2de0f557c20f9f2099d81cbd8b19d9ad992667a279e454fc42f4d275ae67eaf

Observation 27130887-cb0f-40a2-b12a-76381da4c0e3 · outbound

This paper cites vAttention: Dynamic memory management for serving llms with- out PagedAttention.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer vAttention: Dynamic memory management for serving llms with- out PagedAttention

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.451886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.451886Z digest=sha256:42b894aa953e709f7f94351292380eb6b7b6a17f80c8735a72df892b0ad1b80a

Observation b5e26cee-d669-45e7-882f-c3eb61e3177f · outbound

This paper cites Mooncake: Trading more storage for less computation - A kvcache-centric architecture for serving LLM chatbot.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Mooncake: Trading more storage for less computation - A kvcache-centric architecture for serving LLM chatbot

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.537189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.537189Z digest=sha256:c3976884f48a70a66d8752d499a9993d216b95d689b5a44f7558e89489e57e54

Observation d6ea45bf-cee7-42e9-9983-df5588d67c2d · outbound

This paper cites Breakfast of champions: to- wards zero-copy serialization with NIC scatter-gather.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Breakfast of champions: to- wards zero-copy serialization with NIC scatter-gather

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.641220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.641220Z digest=sha256:dbb97c4186e9b973b87a7ebfe8448295d5aac503975c13dc1e28afcea8b31913

Observation 0d950655-6fe7-4926-aac1-f3e6f96b0c3f · outbound

This paper cites Code Llama: Open Foundation Models for Code.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Code Llama: Open Foundation Models for Code

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.739282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.739282Z digest=sha256:aeafc2afbb5855397fb95702d7d3b0f785a097b53e578db30c1235430b891bc4

Observation 447a7dc3-5d2d-423e-bdfa-5ff0092eab9e · outbound

This paper cites Partner success with AWS.http s://aws.amazon.com/partners/success, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Partner success with AWS.http s://aws.amazon.com/partners/success, Accessed: 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.813479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.813479Z digest=sha256:ab5a83a5a756814bc1581da3c6676bdd622fe55fb8786d8f3cd94e55ca0b51d1

Observation ee8d85d9-7e40-4039-838e-17fa95eaebbb · outbound

This paper cites Recommended GPU instances.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Recommended GPU instances

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.899455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.899455Z digest=sha256:27d6363aaa0dd8599b8e8b6a3bed943594c67e806246b640894cb6ffbb7b0839

Observation c047fda4-07cc-4f71-b038-ff25ef0f7af2 · outbound

This paper cites Gonzalez, and Ion Stoica.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gonzalez, and Ion Stoica

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.985377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.985377Z digest=sha256:2a4a65af37476efff6ecfe73d6f75edec4356981f6711ee37cb6f0f020dc101c

Observation 9af42044-f506-4a87-9fec-2e54abad43cf · outbound

This paper cites FlexGen: High-throughput generative inference of large language models with a single GPU.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer FlexGen: High-throughput generative inference of large language models with a single GPU

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.076373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.076373Z digest=sha256:ca6b05a5b56d58c1919a1fe324858f639bb4e25b9593876ae4157ab04de6a366

Observation 707ee9f0-01b3-45d7-a105-d5db55c090ac · outbound

This paper cites Maguire Jr., and Dejan Kostic.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Maguire Jr., and Dejan Kostic

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.144547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.144547Z digest=sha256:8708191d87a159fbbf80d0ba67b54299b5d327a49b9d422a898f59c968b3d386

Observation dbc46c3f-e186-4485-85f6-8d25c5de5894 · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.265620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.265620Z digest=sha256:08598c37c2c79d3833a9e843a691bf8e4ba699716759f091ddbfe5dba3feaa26

Observation 89ed69c4-7aa6-4a59-bc1c-b0c462225e9b · outbound

This paper cites PowerInfer: Fast large language model serving with a consumer-grade GPU.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer PowerInfer: Fast large language model serving with a consumer-grade GPU

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.444363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.444363Z digest=sha256:101d25e233d92da5330c8989f5a4bb3a9e3aaf54c77fd5ab80e8e2d5c7a7b72a

Observation 51f82ed7-c4e8-4c3d-b5f3-aeb6c480360b · outbound

This paper cites Preble: Efficient distributed prompt scheduling for LLM serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Preble: Efficient distributed prompt scheduling for LLM serving

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.514460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.514460Z digest=sha256:1d545a6ab61290abe7f8f6af67131f5eb303689fdbf38cd6fd156f876bab2f80

Observation 5d28ecb4-9dd4-4959-a4b1-fb00d6b26d62 · outbound

This paper cites Déjàvu: Kv-cache stream- ing for fast, fault-tolerant generative LLM serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Déjàvu: Kv-cache stream- ing for fast, fault-tolerant generative LLM serving

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.609483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.609483Z digest=sha256:7a86b4c97f6f8e33e28c9d830ec44746883535f11c1a9f264d0fc02987f9fbb0

Observation 53cef9d2-a140-470a-a865-0969cd1daecd · outbound

This paper cites QUEST: query-aware sparsity for efficient long-context LLM inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer QUEST: query-aware sparsity for efficient long-context LLM inference

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.673473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.673473Z digest=sha256:a3d1c132aa9a2ed0448c023114a683f768c98852bea2a07f5df2ba147ee3b5f2

Observation 89b885dc-3922-4b61-ba37-07abbaaa4c40 · outbound

This paper cites Gemma 3 Technical Report.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gemma 3 Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.746835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.746835Z digest=sha256:4d6ca5022601eafe632affa133fb04a86cd94c9319f4e9f7d7ecd70e26fc66c5

Observation c0ec2726-551d-4284-97bd-3eac364d487d · outbound

This paper cites The Llama 3 Herd of Models.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer The Llama 3 Herd of Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.878772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.878772Z digest=sha256:12ae87dc66230984c71aa11d98ff9d139904e84f8e4b05678b9e2a04f20fb096

Observation 690116a3-8c48-4e8c-a601-00df14918396 · outbound

This paper cites Qwen3 Technical Report.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Qwen3 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.977097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.977097Z digest=sha256:54129feea9f568a446f658a5b0a5a31230c29e2198725ea61c84ddbbfa32d5fd

Observation b44aee88-2c5d-4a16-9047-addd2b89c0bd · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.061323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.061323Z digest=sha256:0b1e7bb6bd656ca1d65e0660db4deea8fc922dbf91597e3cee48e1a6900914c8

Observation 07710b8f-929d-4d72-b195-898d83b6990d · outbound

This paper cites KVCache cache in the wild: Characterizing and optimizing KVCache cache at a large cloud provider, 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer KVCache cache in the wild: Characterizing and optimizing KVCache cache at a large cloud provider, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.149540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.149540Z digest=sha256:57adcaa62d74606987cd17ccc27a9b8dab9eaea44cedc7c55785a8538a89b185

Observation 20d9eb3b-868a-40b7-b2fa-190b65c8cbee · outbound

This paper cites From prefix cache to fusion RAG cache: Accelerating LLM inference in retrieval-augmented generation.CoRR, abs/2601.12904, 2026.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer From prefix cache to fusion RAG cache: Accelerating LLM inference in retrieval-augmented generation.CoRR, abs/2601.12904, 2026

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.205939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.205939Z digest=sha256:7de453c54ecf6210b4033ee2fdabc87cb3fd6367c377297f42e49f66ba1ffd12

Observation 24c399bc-9622-4b97-9ed9-f5c1554eaaa7 · outbound

This paper cites Element-aware summarization with large language models: Expert-aligned evaluation and chain-of- thought method.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Element-aware summarization with large language models: Expert-aligned evaluation and chain-of- thought method

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.286738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.286738Z digest=sha256:39ade30e66b907db0575ceb4a3df8b8443b276608bdf4a01ba59bfc855a89f84

Observation 30eca377-2f8e-4662-b1aa-9842445cd7ed · outbound

This paper cites Phoenixos: Concurrent os-level GPU checkpoint and restore with validated speculation.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Phoenixos: Concurrent os-level GPU checkpoint and restore with validated speculation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.371796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.371796Z digest=sha256:990f3cf254eb73cc2bf040e8afb670fd19d88a4e083a3dc7f35df17fc79e6904

Observation faa59009-37cc-4db8-96e1-f6165a693b8d · outbound

This paper cites LoongServe: Efficiently serv- ing long-context large language models with elastic se- quence parallelism.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer LoongServe: Efficiently serv- ing long-context large language models with elastic se- quence parallelism

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.481672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.481672Z digest=sha256:4d42418b9402b9c5560b4cdcd80c01cb5fa9e4ff2c1fecb08079133acbafcf5f

Observation afb05c1d-4a7a-4e15-88a7-17e8ff8fff54 · outbound

This paper cites DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.554551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.554551Z digest=sha256:1d57910767165d065b1ed7e291db69914840597d433da9b9e4808e1e488afa74

Observation a29c98f4-00f5-4917-b2c9-7d948f907eca · outbound

This paper cites Efficient streaming language models with attention sinks.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Efficient streaming language models with attention sinks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.644504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.644504Z digest=sha256:52a14de5e01095d485b342a60a3ce464689209c4467aaa6168d3eb2f9bc9c78a

Observation 6dbbf075-439e-47e9-bfd3-d104192ffc0d · outbound

This paper cites CacheBlend: Fast large language model serving for RAG with cached knowledge fusion.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer CacheBlend: Fast large language model serving for RAG with cached knowledge fusion

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.715514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.715514Z digest=sha256:c524468e4f98868ca748a28d372f0e9aeab45d4346798ddc4533b17e36798bd5

Observation f05b22b4-7696-4abe-826e-4e0913b26c81 · outbound

This paper cites FlashInfer: Efficient and customizable attention engine for LLM inference serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer FlashInfer: Efficient and customizable attention engine for LLM inference serving

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.842799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.842799Z digest=sha256:f99c3b5606a8e027bb002d925b5480d1d1f11e4960b599f5936bdc2c088dde9a

Observation e64eb608-11df-4107-9838-aa70d32b32af · outbound

This paper cites Willcock, Suvinay Sub- ramanian, Felix Chern, Alek Andreev, Shreya Pathak, Felix X.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Willcock, Suvinay Sub- ramanian, Felix Chern, Alek Andreev, Shreya Pathak, Felix X

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.946243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.946243Z digest=sha256:0279df81f7617ef393a7e17215f6f9e606a51064b83c8a15e3ebec8545e70dac

Observation 0d251480-4d71-4901-8605-8624723641b2 · outbound

This paper cites Orca: A distributed 17 serving system for transformer-based generative mod- els.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Orca: A distributed 17 serving system for transformer-based generative mod- els

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.040565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.040565Z digest=sha256:aeb4dc4992454ba72335016959014f85c621b4e4396ae1c619725695ac732e0b

Observation 555c11d3-bd72-4d90-bf87-6e816a092f19 · outbound

This paper cites Stateful large language model serving with Pensieve.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Stateful large language model serving with Pensieve

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.124125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.124125Z digest=sha256:613c0cf34642b6f82edc81aa285240f2dd834ce988b2151cfa5063c6ecffd979

Observation 1192aa68-5c5d-4b93-8287-a82c84c26301 · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.211420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.211420Z digest=sha256:827f7c01a8ef753d3e62cb9abba184015072656731bcc8843bd73f290a9c4864

Observation 402fccba-1f55-4d3d-b334-3a0acd67166a · outbound

This paper cites Na- tive sparse attention: Hardware-aligned and natively trainable sparse attention.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Na- tive sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.286688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.286688Z digest=sha256:6dd404d5017ff0eb0da785e738c9ec0c0a450023c1a8358e766bbe0c5c562fbe

Observation 4c82d628-7573-425e-8245-af6f02bfb635 · outbound

This paper cites Rethinking database high availability with RDMA networks.Proc.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Rethinking database high availability with RDMA networks.Proc

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.375120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.375120Z digest=sha256:04a9a3dfc19a323ab1b1e655ec9117ae0949740bccce8b23b6845bb2598a41b0

Observation 377907d7-b3b5-4727-91b4-89842814a8f0 · outbound

This paper cites GPU checkpoint/restore made fast and lightweight.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer GPU checkpoint/restore made fast and lightweight

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.465535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.465535Z digest=sha256:6ab51d57cfd746f4b0f61b466166e451188f71ae2bddd9dc43d40945bf9d11d6

Observation e4208fbf-b923-4b7d-8b31-2468997733fb · outbound

This paper cites Jenga: Effective memory manage- ment for serving LLM with heterogeneity.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Jenga: Effective memory manage- ment for serving LLM with heterogeneity

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.526988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.526988Z digest=sha256:bd26c6da1e51c91f0ca68953dadc169e428e596e0f82dbca90ef8fbf59ba3cc4

Observation d81aa742-8004-4f47-8637-cb295962f17c · outbound

This paper cites BlitzScale: Fast and live large model autoscaling with O(1) host caching.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer BlitzScale: Fast and live large model autoscaling with O(1) host caching

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.589085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.589085Z digest=sha256:4a2af92724338024eb7abc3cbdc8a5b5986f979ed0608dda0dcab62565ec9a20

Observation 492e4f93-fd0e-4e3b-9c76-b72f1f4028fc · outbound

This paper cites HACK: homomorphic acceleration via compression of the key- value cache for disaggregated LLM inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer HACK: homomorphic acceleration via compression of the key- value cache for disaggregated LLM inference

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.661302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.661302Z digest=sha256:111793b751641c6d6086107ec99be0c5b22797a9183a1f55b51fa7798ce282e0

Observation fa1106db-6561-4d4d-9c62-25479bce809b · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Barrett, Zhangyang Wang, and Beidi Chen

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.735381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.735381Z digest=sha256:83fb11a20a92bca3c157415136564369b9bd701ace5ae211289a218a54fbe129

Observation 54039205-e481-42a2-873b-a180b6ef5095 · outbound

This paper cites Gonzalez, Clark W.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gonzalez, Clark W

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.788776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.788776Z digest=sha256:b60a02868187f16b4dad6fce393afedd26f4556b8165ec337182e95a0fa25bc0

Observation d86c767e-91d2-438f-83f9-fb8f8c5dcbe5 · outbound

This paper cites Dist- Serve: Disaggregating prefill and decoding for goodput- optimized large language model serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Dist- Serve: Disaggregating prefill and decoding for goodput- optimized large language model serving

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.869913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.869913Z digest=sha256:73f743fc5786c1949e2fbc137ade8d7432cc843a49540f17b6585a3bd9b325c5

Observation d5dbca3c-7b69-4390-b99b-ed9965171a49 · outbound

This paper cites NanoFlow: Towards optimal large language model serving throughput.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer NanoFlow: Towards optimal large language model serving throughput

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.946822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.946822Z digest=sha256:1ade1f3b22e1375aeff9f0122f84de4aa94b5a23631dc08c618f084841e4c8db

Observation b6bb2120-4bfe-4d09-ab2a-7f7d53b1ff4e · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 964

Resolution
parse uncertain
no resolver link, observed 2026-07-31T16:21:17.363791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.363791Z digest=sha256:999429d4501e46d7a2ad8c5a437c9b7e402f32341c302fb3005126c3e038443d

Pith citing papers

No inbound Pith citation observations are available.