Pith. sign in

Paper Citation Record · LEDGER

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference

As of 21 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2502.03589.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03589 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:32:01.436765Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact4
  • verified fuzzy11
  • unresolved38
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2ba7df7-36d3-4a78-9046-e1f3f217cef9 · outbound

This paper cites 01-ai Model Yi.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference 01-ai Model Yi

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:03.098044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:00.894208Z digest=sha256:b2fb8a8e7ae8be6e14b2660c1c4a4c2c9b0e17c6e849ee2c8685aacc33ad50ce

Observation 5854c212-5933-451a-8341-4562f7fa0dec · outbound

This paper cites The code of HKVQ.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference The code of HKVQ

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:03.084134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:00.899952Z digest=sha256:32b05c5681dc1413441f4744be7e22b8d4e21c836497199ffb0c5d2dfa1014e1

Observation bd4341a2-d886-43dd-9535-6e9a8572f1c4 · outbound

This paper cites Falcon-180B.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Falcon-180B

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:03.069761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:00.905385Z digest=sha256:3f9cfd269cd4c9db48718855c30ca18b24411d32cef24ff55d64a449193d2c98

Observation e5cb92bd-4660-418d-a5ab-40a8aeae0dcc · outbound

This paper cites GPT-4 explaining Self-Attention Mechanism.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference GPT-4 explaining Self-Attention Mechanism

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:03.055142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:00.910371Z digest=sha256:d0667dfd117d45b0888ef631d22e69b38f0fd06a6c097de5d54797f0e296be9f

Observation c21cbef9-aa4b-4df8-aa37-214429fbd23d · outbound

This paper cites Meta Llama-3.1.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Meta Llama-3.1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:03.040318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:00.939057Z digest=sha256:f4bea299eb681809bfe26411ddf1a5ac46a12781b28da5099a0f3e6b2a301fb4

Observation 9ff9a508-597f-4b3c-a97e-4bc61994a4bf · outbound

This paper cites Microsoft Phi-3.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Microsoft Phi-3

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:03.026102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:00.963268Z digest=sha256:b68c8e8f8f738f41a8efca4d98541d0710c7b422bad835767da4813567762412

Observation 163b85a9-07f8-4024-9df5-71603d79ce3b · outbound

This paper cites Mistral-v0.3.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Mistral-v0.3

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:03.011753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:00.984003Z digest=sha256:a7ae46438cb300d37c484195577ca2a2f787f8246ca61f3ec498d655517668ac

Observation 02973831-aee7-4f0b-bda8-4f53338cb54c · outbound

This paper cites Open Compute Project.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Open Compute Project

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:02.997340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.012004Z digest=sha256:5414e1fb5ceaa440c3909662b89b6d2a78832d745ef1aa4aeb39b62477c2ac8f

Observation 6dcc1684-1bd2-40c3-b2b3-f7cd9862d832 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.983111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.031821Z digest=sha256:6e42f6f0844e4da9534b3bbfba4f298040d019fb51ab2b7a999df19697a98f47

Observation 23e2105c-de90-403f-a26f-480b948be9ee · outbound

This paper cites Amazon Web Services.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Amazon Web Services

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:02.968849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.051883Z digest=sha256:98a082086b5c143c7cba19be43ef366092d492626375a6cda1c903051318e0f4

Observation 0e9fb43d-1b58-494f-86c7-1d6742b9abee · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.954404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.070817Z digest=sha256:3d22ac08e3aa57d72a53a9a00415631f63fd028d9634925ab9fcc2df9300e5a7

Observation 4acdba96-c182-4a45-8e84-30aa83e51501 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.939427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.090781Z digest=sha256:37b308b442927b45c31141fe699ac5e2c767625acfbc63e2ca0140d3f4b87bed

Observation 11043b35-3024-42d6-8f60-8a03c50dccc5 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Evaluating Large Language Models Trained on Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.111833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.111833Z digest=sha256:9b0f0655e99aa2aa33d26e48386f44b9138f46207ae97f0150ae1eee7c39dbdd

Observation 34f3ef10-7653-40f8-8a64-605c83bcb907 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.924597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.129220Z digest=sha256:b6cbfedec0e2ce64b2edb0f8f79b85d922333e8c9e1ede65b8f21fdbf6815e45

Observation 1ff33bc0-3d6e-4e1d-8462-35c0521e1580 · outbound

This paper cites A Discourse-Aware Attention Model for Abstractive Summarization of Long Documents.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference A Discourse-Aware Attention Model for Abstractive Summarization of Long Documents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.137495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.137495Z digest=sha256:74c68ec5385957e28a46898da355648e28e47590a4b072e41db514ca0c705a59

Observation 9c1b44c0-fe71-44a6-a672-d123ae50fcf2 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.909728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.145652Z digest=sha256:d399b9c0ae7e1d26f105586aa1d16ac134a0e8040df428d204cc75c218ecf249

Observation 3916f688-24a1-48c6-b73c-ea8ff5124090 · outbound

This paper cites Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.161788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.161788Z digest=sha256:ae714619e40844cff5280b7834d57343531f971af3940caeec5953deb4b42d38

Observation 40c9fab2-64d4-4312-9ec5-935173b45f97 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.167087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.167087Z digest=sha256:ae0e7bbdccf2897c5f93d7f54404cba1e76afc1a1534b84fd4f5057d6c1d2ee3

Observation fc836456-25c5-443f-8225-5952f92ded69 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.895532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.172720Z digest=sha256:c186498eaf85c96cd9c79b95ee5efec5128e813fffc76db96f06f69fbc581cb9

Observation 88ea88ac-368e-4b14-9b73-62d7d2e45d53 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.182854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.182854Z digest=sha256:e9564a941b279de393eacec268819785080dd8f52336b61ca9767bd8e2cccdca

Observation 205d1710-fe8e-42da-8476-8cb6c576a0c0 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.880794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.187928Z digest=sha256:52afbe2f5284a72bf066419b5b3c5ea42dcb1b3a823fbb0d799611afe101e66b

Observation 274878eb-54ab-4a3f-b47c-6cb7c37c1241 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.192677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.192677Z digest=sha256:9266bf1134a5e0aeb6f6618d9cedbd35274a898591c2662f6c858f4e71a63286

Observation f448ce07-40a7-4590-96ec-aeee0f670564 · outbound

This paper cites Mahoney, Sophia Shao, Kurt Keutzer, and Amir Gholami.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Mahoney, Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:02.865340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.197825Z digest=sha256:bde553ad612ce4110cfcbe4cbfe7cbbdc20e624c2210dd7fd862e895c6d2b6b4

Observation 08a01091-abe2-435f-a039-308069137136 · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.202215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.202215Z digest=sha256:49bb3937c0bdfbb8b5a3bf8f2d783dff96875d0fbe609955218e78fda856673a

Observation 7c56cfd0-de8d-4781-a1f1-ec29bb3f2ef0 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.206880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.206880Z digest=sha256:cf20fab996546a78941028d3d67fa11824e51b150463497da7b558876298c390

Observation 97c23686-db7e-4b51-8a8f-ec6db3e583c3 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.851131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.211610Z digest=sha256:9ad6131f8995e7a27ee44d78e641e1c9979c5a4dc88b2fe3156547d57f8ee2ff

Observation 278e1116-92ab-49ce-a207-f6e0189758c4 · outbound

This paper cites Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:32:02.836716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.215898Z digest=sha256:392535b306c45d1d30f04e4311b725ee7d808d8d8c479ef7d244a35a01f2e167

Observation e93c6383-835c-4d41-a09f-8d5f598c669e · outbound

This paper cites TurboAttention: Efficient Attention Approximation For High Throughputs LLMs.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference TurboAttention: Efficient Attention Approximation For High Throughputs LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.220321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.220321Z digest=sha256:ee75ea11397eecdceabfa29b79310e6a45e22c555a004e37ead595403b3630e9

Observation a312f7a9-8693-4c31-a9e8-20d62d99656f · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.224806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.224806Z digest=sha256:5b412fbb28bc999c8678f00cea09cd278da61bff8ff5a1576a02662f426f04d1

Observation afd68f26-a247-43aa-8f31-a41cd4897bed · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.822612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.229542Z digest=sha256:35b2b062fe856780fc782aa5fb8ee85ec2b7e8a81a9e74062dae1c016eec86a3

Observation a1c35189-21a7-4b4e-bfcc-825de178e8ab · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.808544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.233899Z digest=sha256:024295bdf01bd73086832ac40b93085a470b46191451b80d26ea445ac42a2fe3

Observation 34d1d726-cca1-4989-87c7-f656023d394b · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.246076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.246076Z digest=sha256:9f638166c86f99198268aa5ac4ddaec339f3e6c4a1cac64cc996f5cad6fb16c7

Observation b7313fc8-ecb5-4591-9dca-a8117548d204 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.785061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.299578Z digest=sha256:47dcdd85ac4c70fdbac347fb42fc57fcd3c89a932ca2918744b41ea03ad256e2

Observation a0410b61-cd43-4383-a6a3-a9a0328a870f · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.770500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.330702Z digest=sha256:8e3e3c6bdc3aa9b66a04549bd37817485b4dc490dfbaa32b8610515dc88d350a

Observation 6562a816-020d-4bda-af14-24cbae0127fd · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.347729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.347729Z digest=sha256:b223ffae66c947a1ecd2e33aea850a8744db09c4a07ebf70e144497c622c76b9

Observation af63ab59-ca79-4cf5-a240-478204840bb5 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 36

Resolution
malformed identifier
arxiv_id_nonexistent, observed 2026-08-09T04:32:02.199037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.355654Z digest=sha256:546e515806b231e85ca37f1739b990122bd10956da63ee55f4af7f400a64f3d6

Observation 9491ba51-c74a-4e9a-957d-15551c576bca · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.746448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.359758Z digest=sha256:ea63f20e88bdff261bb93c169435181df71c42eec414221683db94ff23db9ab1

Observation 677af7c4-5aeb-43c3-825c-e06c4b0a669a · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.732455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.363656Z digest=sha256:e12e97e37ca642b7663e32d654ddbb791bd4b1786452eca6cd868599a1b6ab8a

Observation 17dabb12-3d8a-4376-a7a3-08931c0a49bc · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.718185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.367667Z digest=sha256:8f4aad5178ed14078fb0ddc964b8c01f3130e05e97691c220054c46408039212

Observation faac764e-f192-4b76-a627-ff89abdeb381 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 40

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T04:32:02.031876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.371430Z digest=sha256:9195cd391c8b3337eb5257b263ca7d76cf750b87d3b3dff08b14296f3e631122

Observation 5f39808e-8b62-4c56-8038-1967e987706c · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.375623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.375623Z digest=sha256:be03f022787712e38903e400ea281611644154f7dc1d1adeef6438963dc03bff

Observation 40dd6b79-f0f8-4152-837b-ecda2eb40778 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.703991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.380570Z digest=sha256:a26263d1617e77307627c9198b1e6dd0aa903ee6491ff89f38570ee716dee2de

Observation 2047d041-1f96-424a-bcff-546d06e3ae79 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.689339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.385124Z digest=sha256:e6603f88adaf195f41460afc339ae99ae87f029c06179d6c04017eea916b85e4

Observation d9d133c9-8e81-4a92-9f7e-cdd4ed11e529 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.674377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.389791Z digest=sha256:d7ff5622d085eb1fbf9e5ff8a51c5e29ada6b667080aa8e7dfd229624faf126e

Observation 254e552f-d338-4896-9083-775b66e68b27 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 45

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T04:32:01.859716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.394372Z digest=sha256:c0117beb03dbe36e0b517c43124c74008a33892b19bb64100730d18ff47bff1a

Observation 307e8010-7f9a-4177-b493-6f9ae3976a48 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.659248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.398968Z digest=sha256:a748fbc858741a184686738d313cbd3a92c3b20f8673a00899e7a6c538a4d7f6

Observation cca5ef1f-3fee-444d-8c42-d78f9b5e852b · outbound

This paper cites FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-09T04:32:01.542534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.403532Z digest=sha256:0b0cd3acd396bbe8aab63635f247f1ee028c5ca66c7c9c0214391cc079c1fa2e

Observation cd6b8efe-e3fa-4474-a515-8ec7e9cc7aae · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.408211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.408211Z digest=sha256:29dc0f8eb5f02dc80b5180703e6e76bbd426936d84abcd6a7fe7882dcff27d58

Observation eea89336-f7b2-4247-a02e-c49c6117cfe8 · outbound

This paper cites Efficient Sparse Attention needs Adaptive Token Release.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Efficient Sparse Attention needs Adaptive Token Release

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-09T04:32:01.504847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.413110Z digest=sha256:bbce41af68ba751f80c74b98c407e06d39eddc0daa189532cc3af0af91272ad6

Observation 055b31a0-fd46-4bd4-bc6b-e6501e20a04e · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.644603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.418143Z digest=sha256:d3459b8b69683532852731e454810526355b78d4bbfe1d55542fe019d7f925d3

Observation b4a436e4-ddd6-4423-af60-93370e252956 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.630591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.427705Z digest=sha256:773209dd309a23fc5be145705842d3b593ab4c078b34e235dd61d4d826394004

Observation cbeba9c0-585c-4ec7-a3eb-e10292816068 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.616630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.432313Z digest=sha256:3694e3101a14e24856c6f954a2d43a127e70026a8a72c5b9357828abc48f5587

Observation 16274fcd-7937-4808-8ca9-47af324fe340 · outbound

This paper cites Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.422886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.422886Z digest=sha256:f6b7f4b195c3e71c4a1f3ee109875a2bb8d83c3a09fc1e4e27502284472006fe

Observation ad085163-5b7d-4e62-996f-caa301007be1 · outbound

This paper cites an unresolved cited work.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Unresolved cited work

Reference 210

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:32:02.602407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.436765Z digest=sha256:fa8cfa6238c6cee83add4f014d83f8a7c0e7bf20189890eb914de8564012a9d8

Observation 120be18e-c3b2-4a1c-b23e-b089ca252665 · outbound

This paper cites In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ’23).

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ’23)

Reference 2023

Resolution
malformed identifier
arxiv_id_nonexistent, observed 2026-08-09T04:32:02.438778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T04:32:01.272605Z digest=sha256:22ad2e8bcf1dcd985ee4f5e77f9d60391c25834cb8a66b67fb286265b353a44b

Observation e22e2b6e-089a-437f-8be9-ea3f7eafbbe0 · outbound

This paper cites In Proceedings of the 2024 Conference on Empirical Meth- ods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.).

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference In Proceedings of the 2024 Conference on Empirical Meth- ods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.)

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-09T04:32:01.178053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.178053Z digest=sha256:ff5e64ec53336dca40ae0f68a6d82f0dd1cd26201a3e63d6107cf2fc5fca954c

Pith citing papers

No inbound Pith citation observations are available.