Pith. sign in

Paper Citation Record · LEDGER

Palu: Compressing KV-Cache with Low-Rank Projection

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2407.21118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.21118 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:40:37.603820Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.121083Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f73c71fa-3690-4967-8845-19cad3674fec · inbound

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models cites this paper.

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:49:33.802941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T13:49:33.747672Z digest=sha256:ffeb11256dac028e7c76c3e5a87af1312db514a826036879bd4dc0969301cbc4

Observation e3967fdd-1b7c-46d2-ac7e-5fb0b7207041 · inbound

Attamba: Attending To Multi-Token States cites this paper.

Attamba: Attending To Multi-Token States Palu: Compressing KV-Cache with Low-Rank Projection

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:56:33.034655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:56:33.034655Z digest=sha256:e7f5b79fd21138c64839f76aa2ca4f22201f5e16a39924a66725b84adab0bd3b

Observation 323c0ab9-4f1b-4b74-ad4e-9bb33b0d543a · inbound

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals cites this paper.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.776013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.776013Z digest=sha256:69d39acb5fe85fe1c7255d41290e3e60add86027a2b5070ef4bcdcfe42aa4908

Observation 3ff7df90-6c6e-49d1-80ec-2aa6f8355212 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Palu: Compressing KV-Cache with Low-Rank Projection

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.985213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.985213Z digest=sha256:76baf55ad858bae008eb1ba4aaaced7518a258f269d686a87e32b1546dcb2590

Observation 9e72fe70-0f02-408e-ad53-19ad07a7a089 · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.189891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.189891Z digest=sha256:fdae7edb3dfd0c927e6b8117c905e846fa4a33e33bf06b1d810b71b0a1837d36

Observation def58ec3-b6a0-4d23-a0ba-3f4148078342 · inbound

TransMLA: Multi-Head Latent Attention Is All You Need cites this paper.

TransMLA: Multi-Head Latent Attention Is All You Need Palu: Compressing KV-Cache with Low-Rank Projection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T11:48:01.171526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:48:01.171526Z digest=sha256:4193d0edffa5ade1ebf5280658c6354c2302b17ce78792db7ba72ef39460e04d

Observation f857c088-56f1-40a0-b779-0614274d8b18 · inbound

A3 : an Analytical Low-Rank Approximation Framework for Attention cites this paper.

A3 : an Analytical Low-Rank Approximation Framework for Attention Palu: Compressing KV-Cache with Low-Rank Projection

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:04:57.131560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T15:02:33.864861Z digest=sha256:9f28c471cc23ec460b0766929ec4f9824eb8b08fc600ea4da631b41884505c78

Observation de377a09-f03b-457f-b1bd-0eb0dffa5fe3 · inbound

LatentLLM: Attention-Aware Joint Tensor Compression cites this paper.

LatentLLM: Attention-Aware Joint Tensor Compression Palu: Compressing KV-Cache with Low-Rank Projection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:25.552150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:25.552150Z digest=sha256:1da0e8c33158d3d3443b2493a4af3f2aecc4e08fa58c7e5087e047285694cdf2

Observation 7e5ea5fb-a49e-40fb-b73a-2d4c9cb73d6a · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Palu: Compressing KV-Cache with Low-Rank Projection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.031025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.031025Z digest=sha256:c438770cbae87c129652906fe50c17491adc4be1cb26e183584c7c32f1ab7c3a

Observation 484ef146-5291-4782-9dd0-fdc49ba24a4b · inbound

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering cites this paper.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.825323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.825323Z digest=sha256:abe92dcc41ca2793bd1f13f6b3f570961b620b311284146d4db78f85c02ae1b6

Observation 6b808613-58d1-4cf3-b466-da47f926ab65 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study Palu: Compressing KV-Cache with Low-Rank Projection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:26.941477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:26.941477Z digest=sha256:6d72d247719c8fbf2442e4d84bd49a1ee10455110d66e19282b15f2bcacd3ef7

Observation 77689f63-3157-4247-bc97-f1a0c49a8c57 · inbound

Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache cites this paper.

Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache Palu: Compressing KV-Cache with Low-Rank Projection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T01:10:37.292088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:10:37.292088Z digest=sha256:df23c976b5d651e36f6e50b76e045023dcb1a58fbc58410b65537f0f2bd98366

Observation 11c2edea-ccd7-4683-aea8-76ea40ffa21c · inbound

RCStat: A Statistical Framework for using Relative Contextualization in Transformers cites this paper.

RCStat: A Statistical Framework for using Relative Contextualization in Transformers Palu: Compressing KV-Cache with Low-Rank Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:40:37.603820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:40:37.603820Z digest=sha256:96b00680178fb4b01d5d462bff9bc319ffd93a6df35b45d283693ce8db09636b

Observation 72c9b27c-2c3b-4a5d-87d3-bc691b0d4124 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.670799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.670799Z digest=sha256:ea73002f2235119565fa155ab0286bcefc880df6d06fed308dc343a5bacc6a5d

Observation f929d593-9219-4713-9a8b-eef18834e01c · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.348432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.348432Z digest=sha256:0e408f52e5eee6bc60f13736514b9c77b983ec166d4a250ee7f4ed7a598b6291

Observation 98b382e6-6302-4125-9fce-9adea19682cf · inbound

OjaKV: Context-Aware Online Low-Rank KV Cache Compression cites this paper.

OjaKV: Context-Aware Online Low-Rank KV Cache Compression Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:26:24.651969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T13:26:02.980973Z digest=sha256:a66ff805b518ceb2f998b7d99da5ae36d9bf8f791823a48c5f788fde24e2e554

Observation 57baefaa-2193-4053-b8d4-b8587bdf6c13 · inbound

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction cites this paper.

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:09:36.971195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T01:09:07.983785Z digest=sha256:9a3d4d0aee4a7255f13dcc5922162b51a6dd331e1c52c4a8ed41eec9bb5e226f

Observation 7ea17c51-bd2f-45bc-9051-11e8fed10732 · inbound

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models cites this paper.

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:18.020672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T21:05:09.254086Z digest=sha256:71d30b4eb2bc01440ee967cba8eff57d82dce8c5380d2b9ee37ef3f25c1af7d5

Observation e152ac02-d2c4-459e-8575-17a5922a051c · inbound

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization cites this paper.

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization Palu: Compressing KV-Cache with Low-Rank Projection

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:50.983412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T20:18:04.392331Z digest=sha256:017a4559a5d16bd96b5b1dee3edab4823a0ec821dd0c2ab29ced20ed096fd4db

Observation f5475433-8212-4511-bf65-d1c982f6ba30 · inbound

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization cites this paper.

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization Palu: Compressing KV-Cache with Low-Rank Projection

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:18:16.527054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T12:16:07.797702Z digest=sha256:1004ebbbadebd014c473592730f99af7ab0eeb2ac185703dac0cf2f49ca89f42

Observation 9ab8b0ee-4d6c-4b4b-b1ec-b4cf79efc620 · inbound

Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference cites this paper.

Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference Palu: Compressing KV-Cache with Low-Rank Projection

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:43:23.770713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T07:43:18.828740Z digest=sha256:4266c2c7e6c3851124a15b82e1256aeb2ab1737d28b3befcfc9e742780479dd5

Observation 554a9a90-546c-4240-a260-d11ac706e328 · inbound

GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs cites this paper.

GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs Palu: Compressing KV-Cache with Low-Rank Projection

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.774202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T22:54:55.101568Z digest=sha256:b17711866fa3244fc472caee64061bd0a9aab4efbf8db556cb0420ee3f3159ac

Observation 06d2376f-b0a2-4a6c-a291-d84b8421dc2c · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation Palu: Compressing KV-Cache with Low-Rank Projection

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.508413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:35e842333b47e9958e436fa1c2ce77b7c0c8022259eecd1fa222a881e2f01463

Observation 90c87189-d919-4472-a996-0f980e90c98a · inbound

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models cites this paper.

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.707197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T19:09:29.347284Z digest=sha256:ca9ed59ebbd5b479f7cf55919d90e25fce02edb2286c59f325db7ca494423d4c

Observation af4941be-b6f9-4ab6-80a6-1de7e319cc81 · inbound

EinSort: Sorting is All We Need for Tensorizing LLM cites this paper.

EinSort: Sorting is All We Need for Tensorizing LLM Palu: Compressing KV-Cache with Low-Rank Projection

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.699258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T18:31:01.804061Z digest=sha256:d0541970ee6808a978395521f8e668d76efbb7217b17a537816c7e0356d573ed

Observation 69977081-c482-4a04-864e-82f935153758 · inbound

ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models cites this paper.

ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.348156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T13:09:37.979051Z digest=sha256:b0f620440ccd429ed7ef493c1fdb48d302b9feb3d50631b8a244b849f9518664

Observation ec9e31b7-358a-4438-9a9a-195ca299eee0 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Palu: Compressing KV-Cache with Low-Rank Projection

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T01:36:44.122209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:c880f6242dd03a7bcbcdc93b0481ddbf2137b36bb17525288cbf05f42cebef80

Observation 3e2573aa-faa8-4189-b696-6d4092c86a7b · inbound

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs cites this paper.

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs Palu: Compressing KV-Cache with Low-Rank Projection

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T06:31:50.083575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:31:50.083575Z digest=sha256:617b2c6200fdee5f232dc7fdbf49ed9942d96c346df5768682d381f454b1073a

Observation 7c90bfd3-a5db-4c70-8f38-0575b9e87fb7 · inbound

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge cites this paper.

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge Palu: Compressing KV-Cache with Low-Rank Projection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T06:50:14.941014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:50:14.941014Z digest=sha256:57f4d7e3c642f1000369ce82b43a17922e19795593bda16a4b1930f6de92a5b5

Observation d4431f5f-d32c-4e5f-9a5b-6ed319e2395e · inbound

DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation cites this paper.

DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation Palu: Compressing KV-Cache with Low-Rank Projection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T18:02:37.410810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T18:02:37.410810Z digest=sha256:5512c4ab03c1f705531ad25c946388761d4dd24fa7b974f9c7ae8062c7ddf300

Observation 9efb6cf7-0d03-46c7-bd1f-c19c47410427 · inbound

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding cites this paper.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.614843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.614843Z digest=sha256:697538c7f003b647df491dbbe6e7be0388319bb908a870c1ebe9adf37c28b7c5

Observation 0cefdd22-aad1-4e4c-80c1-cb14334d7f0d · inbound

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization cites this paper.

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T00:43:48.033830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:43:48.033830Z digest=sha256:eab106cb1c738ff2607cfbab24f782a40d66ee065a8b0324bfa37adc6d502fda

Observation cfef5d0d-dce5-4308-b32a-06239cc816f5 · inbound

Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes cites this paper.

Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes Palu: Compressing KV-Cache with Low-Rank Projection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T07:49:39.405147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:49:39.405147Z digest=sha256:9ef173af6e2e6150d01fb3b981b8aab706d0cdd7b7986cf0a8742c9ac3780dd6

Observation 064a8211-11ca-45ae-b1ce-7bd7c3b7bc12 · inbound

AnchorKV: Anchor-Residual KV Cache Compression cites this paper.

AnchorKV: Anchor-Residual KV Cache Compression Palu: Compressing KV-Cache with Low-Rank Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:52.922607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:02:52.922607Z digest=sha256:1ddb41684561d235e898cd52f8f9b026f41ce7bf0d3f85122b6ed206532ac486

Observation d3be1814-e7c1-4d18-aa49-b2d1acf584b7 · inbound

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval cites this paper.

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:35.285100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:35.285100Z digest=sha256:7281cd0265fa287b315b12d89d7e56f58303d41de9da3269647d004b21654c07

Observation b4baf7a2-63fd-442c-ba6c-a225c61f16dc · inbound

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval cites this paper.

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:54:42.694725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:54:42.694725Z digest=sha256:b1da49f83691a1d8f2a6947e7d1e39aa965813437040c12f841526791fd81691

Observation ebec126e-da94-419a-a482-322eae95ddbb · inbound

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference cites this paper.

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference Palu: Compressing KV-Cache with Low-Rank Projection

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:46:03.872365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:46:03.872365Z digest=sha256:218fcae916c32e0f6ece3ba6f71c7068309ac6e34f853d51bd3e708caedd0d0f