Pith. sign in

Paper Citation Record · LEDGER

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM

As of 5 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2505.05772.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05772 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T16:58:11.105511Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T13:40:32.154860Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact16
  • verified fuzzy25
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0a6df20-421e-413f-97cc-750445cc8415 · outbound

This paper cites GPT-4 Technical Report.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:01:48.639275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:5e38edc418380893498672361a2c6aadaa33aa953bfea9ca6c504e9ea1062580

Observation 89aa74ae-03a9-4b48-a18f-00d1d7171069 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Toolqa: A dataset for llm question answering with external tools

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.377125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:e39f06d56b2e0d92ec2fdf6b05863cfedfaaccec8232d68cd84d869497a2a0b0

Observation 4526d313-beb4-403d-8b94-36cdd01588ca · outbound

This paper cites Pythia: Ai-assisted code completion system.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Pythia: Ai-assisted code completion system

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.333294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:174229b6e6e8a0b97f1444530d570cec6cb81249baaef44a1b909b37eeba093c

Observation ae8f3319-cc17-4964-b833-289d45f85d79 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Code Llama: Open Foundation Models for Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:01:48.631980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:c93d5e6998ab8abf76b2ec5d3932f55ffc65f2a6fea04429b366773d882e2ec6

Observation ede85c5b-18c4-48b7-b4d7-1bb053793cc7 · outbound

This paper cites Reflex- ion: Language agents with verbal reinforcement learning.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Reflex- ion: Language agents with verbal reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.363753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:6bab604ea2abbed4c21a7f9f42b8db5f9f10431e2c15879da09a573a68e6ddcb

Observation 941d6a88-bfbf-4603-a507-4b48cca6fb0a · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Tree of thoughts: Deliberate problem solving with large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.346304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:9b44a5a6aab6ad7dfe50bc3cdf478254a1e84554253ae60874e26d226ff6f55b

Observation 21ed6cd8-7c11-4c00-b404-172f6148547f · outbound

This paper cites Pre-trained language models for in- teractive decision-making.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Pre-trained language models for in- teractive decision-making

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.337605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:737fd16a781f17ead7b9a783d28378f6b4fef9a80364292042de34f27e68e2d9

Observation 36049c4f-f4fe-4e17-b7c8-9ef2c856038b · outbound

This paper cites Efficiently scaling transformer inference.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Efficiently scaling transformer inference

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.341873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:f272c9f9417ab54187538e6e4e43ddea581afb6a6c29d168d34caf67012da075

Observation 06af12f3-6f77-4e15-a146-599423b98d93 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Efficient memory management for large language model serving with pagedattention

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.359684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:f011e72e14bc118e016186ff054a290bdb0c18f25b177c3de76e1d2ab6902203

Observation 1dfbf1f8-f989-4f41-8315-7d33ca7c7cba · outbound

This paper cites Computedram: In-memory compute using off-the-shelf drams.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Computedram: In-memory compute using off-the-shelf drams

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.355941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:1c42ecf35406e432e64a5652dc92119d4d61499a9e2690d13da993a2ddfb2fff

Observation 50ea3acd-516a-4b93-ac3c-d69b5d97c29b · outbound

This paper cites Newton: A dram-maker’s accelerator-in-memory (aim) architecture for machine learning.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Newton: A dram-maker’s accelerator-in-memory (aim) architecture for machine learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.372596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:df446e73af0279bea1af9d64317435c2f9619ccb8f5a60cfe26a4bf4200b4ed6

Observation 4c465b38-8349-4d7b-8606-9697ae0f32ba · outbound

This paper cites Pathfinding future pim archi- tectures by demystifying a commercial pim technology.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Pathfinding future pim archi- tectures by demystifying a commercial pim technology

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.367776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:d7868aaa04cea46295f2f05fea559b99d45e1aa5fc905d0242fc882783a0a947

Observation 155f3fca-325f-482e-a0a2-1b8b0fd96582 · outbound

This paper cites Accelerating neural network inference with processing-in-dram: from the edge to the cloud.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Accelerating neural network inference with processing-in-dram: from the edge to the cloud

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.350627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:f23915800ceafdb503d5b72542c3a540af10a632e587687ad42ec3e726a76be8

Observation 8936493c-a694-4ffc-897b-cea7b35adc62 · outbound

This paper cites Floatpim: In-memory acceleration of deep neural network training with high precision.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Floatpim: In-memory acceleration of deep neural network training with high precision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.420304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:98ea0b042b0d84a85017894bc1b0f8a56dc6d93e841e1c20902184fa2beac81d

Observation 1f184fd9-141a-40e2-a6e8-51bd32ed463d · outbound

This paper cites Attacc! unleashing the power of pim for batched transformer- based generative model inference.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Attacc! unleashing the power of pim for batched transformer- based generative model inference

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.424514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:3b26c83ec1e3df1ad439f8d9040c176f55f2667eaf3ff640a5926d81c236d289

Observation 8aa03bab-105c-47e3-99ba-d160b0985255 · outbound

This paper cites Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.436845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:ecb19d5846229887cd2200de3c8d20eb33b3e815f199eca09d077217139c4cb7

Observation f93a87a1-3539-4d8d-9c0a-4b6ad83f4f0f · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:01:48.552416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:ccccb36c94925108b55b503ae5e31e593511e138d80f5ebb031ed3074463d462

Observation 9f1d6902-e0fc-41e5-821c-492cf3bde11c · outbound

This paper cites Transpim: A memory- based acceleration via software-hardware co-design for transformer.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Transpim: A memory- based acceleration via software-hardware co-design for transformer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.398828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:42b38db734300b8cbb28d9d8d3c0382f9d6407302f4977a91c8172c5467b6e98

Observation 9d007ad9-5b83-4d92-9a62-fe0cc180a204 · outbound

This paper cites Lol-pim: Long-context llm decoding with scalable dram-pim system.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Lol-pim: Long-context llm decoding with scalable dram-pim system

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.545822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:9ee39f38e27fb4d4d0aa2f7b2b872e0365dab9de75c246518a8158383579a126

Observation 7231b18a-40a2-437d-856e-a0bcfd714dff · outbound

This paper cites PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.594637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:38f2b1786493bafedf0ff066f904c14e52ec45f35cdfb36ef7514b45e643c237

Observation f4275ed3-266c-4257-8746-7032d350c841 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.624426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:9aa1c21d94ab986a8b199dfff0183bd2555a427ba9379766fffcedbad135a4bc

Observation 0db3f7e9-fc34-4b3b-b4c1-d4730001bcc4 · outbound

This paper cites How long can open-source llms truly promise on context length?.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM How long can open-source llms truly promise on context length?

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.428396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:27827af1c0c2ffa127e53216d4eb303c745d7aeb50c75d4f130aa9df04139403

Observation b2af87e2-ed41-4321-8847-b2a07add210b · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM LongBench: A bilingual, multitask benchmark for long context understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.403037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:b89b959e27df36d407f31ac3824c75e2e51ba34086496d0b9c301a9534537ca3

Observation 65e06e91-dfd2-4e45-b4b2-eebe66bd1b12 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:01:48.586773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:8edf48a079a64037ecc06cd4a4bba443a638eee21c6409590cd2a46fa5fc2f55

Observation 9c1b9472-4040-40b8-a965-e6f115cdefa1 · outbound

This paper cites A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.612925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:fb81b1f2b35544061676a85b2609846901496101b2b19cc31c7203bb857c1ba0

Observation 98baa4ae-bc26-4d02-919a-07e843da4097 · outbound

This paper cites The narrativeqa reading comprehension challenge.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM The narrativeqa reading comprehension challenge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.411447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:672bf9937470252492ac3f3dff01323e48f250f0353cef401d1858495a32ce81

Observation a17e694a-e77b-45bf-b825-12561e36e8a8 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:01:48.558998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:9da8a4935339ec54587db5f56a63e8fe33ef5fdad3b51a034215c5a09a0dc27d

Observation 428ab48b-387a-4092-9ceb-7c96d016df93 · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.579348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:7300d3045656da8aa5207198cca8b3caca0a875b1763955e79a044830e7b6ccd

Observation ae8f505c-cb96-4055-b767-1e0a55506ca2 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.407532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:dd6587915d6bbc9cbfd4f167d403cc405eca0ac6ad88fdccbce95685a5023e54

Observation 7caf4020-cc34-49fc-a41a-94b14bf95d8c · outbound

This paper cites Longcoder: A long- range pre-trained language model for code completion.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Longcoder: A long- range pre-trained language model for code completion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.382078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:1d782bf2b3fbf1b6218eef80790f5c64e62520192744c0e52853efb5ffd27e76

Observation 46ce69cf-37df-4b9a-b60b-22475a9d62bc · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Compressive Transformers for Long-Range Sequence Modelling

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:01:48.565577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:7b465bdc00456e1eb0256d8a6811645a7f10bb927779398e1d4d20b75f820501

Observation 859f5777-a94f-4b71-b05a-95ee2b571523 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.432620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:580e0d1501da6816eaf4a2f7c332a01c5bb8b9c256ed7128de6f4c896640eed1

Observation 733b74c5-244a-427c-924e-d71070f9fb19 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.528269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:8cd4c779e3f6fa3ccbb629826e2af224c11361ca4bcb3b7c8c62cdd329154a64

Observation 03cb7c30-75f8-41c3-bd49-12f3696aa48b · outbound

This paper cites Accelerating bandwidth-bound deep learning inference with main-memory accelerators.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Accelerating bandwidth-bound deep learning inference with main-memory accelerators

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.391337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:78a7c2f1b5a8b5770d5f8e65b3ad2927aa55824741cd57c6d0dfa53a6ce10bce

Observation 16d14693-4eb8-49f9-9238-c7d78dd72f95 · outbound

This paper cites PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.603458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:8a47d68fd2b65b32be85a884a3a1c310aedcaf1e2933936ac3737f9ce79975cf

Observation a0793e62-3d74-4805-8c09-a6a6691cc991 · outbound

This paper cites Make llm inference affordable to everyone: Augmenting gpu memory with ndp-dimm.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Make llm inference affordable to everyone: Augmenting gpu memory with ndp-dimm

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.386567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:2095237a5a04411ef06628b1829218a01f06931a3d11af0dada570d2c9d40fae

Observation 5db33249-cd19-4fec-901e-8d790f2eb251 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.415931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:11300bf14e4cbf32ae3a16da89a43d0a2186e87c58eb0e8a082caaf6eb023563

Observation b08d30b4-d613-41cc-bb3d-92167b9d2c20 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:01:49.441191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:1502f16ca191f057ddd2a43eaf8348b2008cd103219e68ce1b56c5662ad8417c

Observation 53d99a0a-a4df-4861-b32c-47134b63434a · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:01:48.534979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:326e6e40d7c82d3afd702fbc1c34ddef29d5d0e9566620a4257924c9c433a88a

Observation 5c6b3ef3-574f-4e08-ad45-4712437638d1 · outbound

This paper cites ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.520894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:fb7803840251235c7ee1be9a4e501bc12583ed0e6ea6277d733686d1b8630fdf

Observation a98d89ab-9a03-43c4-a630-94002f6745c8 · outbound

This paper cites Squeezed attention: Accelerat- ing long context length llm inference.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Squeezed attention: Accelerat- ing long context length llm inference

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.618802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:36483ead5b3a343dcb88ef263b714f9de7937e31071aa4105efa8d4200670b45

Pith citing papers

Observation e1804785-eb12-4a0c-ba86-85da2072338b · inbound

FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference cites this paper.

FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T13:40:32.154860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:40:32.154860Z digest=sha256:4515b444d013946d357b9270f37ce4ceadcb5b4b5648980842480987110716a3