Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T16:58:11.105511Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2505.05772.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T16:58:11.105511Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T13:40:32.154860Z
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d0a6df20-421e-413f-97cc-750445cc8415 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89aa74ae-03a9-4b48-a18f-00d1d7171069 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Toolqa: A dataset for llm question answering with external tools
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4526d313-beb4-403d-8b94-36cdd01588ca · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Pythia: Ai-assisted code completion system
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae8f3319-cc17-4964-b833-289d45f85d79 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Code Llama: Open Foundation Models for Code
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ede85c5b-18c4-48b7-b4d7-1bb053793cc7 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Reflex- ion: Language agents with verbal reinforcement learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 941d6a88-bfbf-4603-a507-4b48cca6fb0a · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Tree of thoughts: Deliberate problem solving with large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21ed6cd8-7c11-4c00-b404-172f6148547f · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Pre-trained language models for in- teractive decision-making
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 36049c4f-f4fe-4e17-b7c8-9ef2c856038b · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Efficiently scaling transformer inference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 06af12f3-6f77-4e15-a146-599423b98d93 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Efficient memory management for large language model serving with pagedattention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1dfbf1f8-f989-4f41-8315-7d33ca7c7cba · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Computedram: In-memory compute using off-the-shelf drams
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50ea3acd-516a-4b93-ac3c-d69b5d97c29b · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Newton: A dram-maker’s accelerator-in-memory (aim) architecture for machine learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c465b38-8349-4d7b-8606-9697ae0f32ba · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Pathfinding future pim archi- tectures by demystifying a commercial pim technology
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 155f3fca-325f-482e-a0a2-1b8b0fd96582 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Accelerating neural network inference with processing-in-dram: from the edge to the cloud
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8936493c-a694-4ffc-897b-cea7b35adc62 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Floatpim: In-memory acceleration of deep neural network training with high precision
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1f184fd9-141a-40e2-a6e8-51bd32ed463d · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Attacc! unleashing the power of pim for batched transformer- based generative model inference
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8aa03bab-105c-47e3-99ba-d160b0985255 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f93a87a1-3539-4d8d-9c0a-4b6ad83f4f0f · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f1d6902-e0fc-41e5-821c-492cf3bde11c · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Transpim: A memory- based acceleration via software-hardware co-design for transformer
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d007ad9-5b83-4d92-9a62-fe0cc180a204 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Lol-pim: Long-context llm decoding with scalable dram-pim system
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7231b18a-40a2-437d-856e-a0bcfd714dff · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4275ed3-266c-4257-8746-7032d350c841 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0db3f7e9-fc34-4b3b-b4c1-d4730001bcc4 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM How long can open-source llms truly promise on context length?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2af87e2-ed41-4321-8847-b2a07add210b · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM LongBench: A bilingual, multitask benchmark for long context understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65e06e91-dfd2-4e45-b4b2-eebe66bd1b12 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9c1b9472-4040-40b8-a965-e6f115cdefa1 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98baa4ae-bc26-4d02-919a-07e843da4097 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM The narrativeqa reading comprehension challenge
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a17e694a-e77b-45bf-b825-12561e36e8a8 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 428ab48b-387a-4092-9ceb-7c96d016df93 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae8f505c-cb96-4055-b767-1e0a55506ca2 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7caf4020-cc34-49fc-a41a-94b14bf95d8c · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Longcoder: A long- range pre-trained language model for code completion
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46ce69cf-37df-4b9a-b60b-22475a9d62bc · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Compressive Transformers for Long-Range Sequence Modelling
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 859f5777-a94f-4b71-b05a-95ee2b571523 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 733b74c5-244a-427c-924e-d71070f9fb19 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 03cb7c30-75f8-41c3-bd49-12f3696aa48b · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Accelerating bandwidth-bound deep learning inference with main-memory accelerators
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 16d14693-4eb8-49f9-9238-c7d78dd72f95 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0793e62-3d74-4805-8c09-a6a6691cc991 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Make llm inference affordable to everyone: Augmenting gpu memory with ndp-dimm
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5db33249-cd19-4fec-901e-8d790f2eb251 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b08d30b4-d613-41cc-bb3d-92167b9d2c20 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53d99a0a-a4df-4861-b32c-47134b63434a · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c6b3ef3-574f-4e08-ad45-4712437638d1 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a98d89ab-9a03-43c4-a630-94002f6745c8 · outbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM Squeezed attention: Accelerat- ing long context length llm inference
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e1804785-eb12-4a0c-ba86-85da2072338b · inbound
FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.