Pith. sign in

Paper Citation Record · LEDGER

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

As of 7 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 2 inbound Pith citation observations for arXiv:2506.15704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15704 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:41.059939Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:03:08.617257Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:04:20.888331Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7e9b751b-e021-4306-9155-4dc11094e21f · outbound

This paper cites GPT-4 Technical Report.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:33.909659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:33.909659Z digest=sha256:014adf0965872e81aae215e53e3c568089b49072d98b94785394aee815c06a44

Observation 6905c1fa-97b4-4f8c-8379-e543693552b3 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:34.959797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:34.959797Z digest=sha256:205361725cb9c490b0d03032291c75ae8869a51b54d7c478745901b8ac2d42f4

Observation 4818a542-f672-4751-9d18-0e618417980e · outbound

This paper cites Loki then and now: the trickster against civilization.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Loki then and now: the trickster against civilization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:42.432776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:36.864775Z digest=sha256:3dddaedb9cb6dabd0c58dd466ca4f6c103243a0c6c49819ba34fe53ae2b6913c

Observation fbc23119-1404-43fb-bcbe-a736885ebc2c · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.033773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.033773Z digest=sha256:4d958255b644bde48a129a1a2481b32928d77534234faba56d902f84beb69022

Observation 1b64cb96-4555-4e9f-96da-9b6ea86f7653 · outbound

This paper cites A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.206520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.206520Z digest=sha256:b8ea7da81f0bb48de30d2e632b7953d35191b0ffc623442b318520076c1a1bb6

Observation 7c429e67-9fbe-4570-8558-ab2d3cd7d6ab · outbound

This paper cites Human-like episodic memory for infinite context llms.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Human-like episodic memory for infinite context llms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.366095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.366095Z digest=sha256:893d90c541159d3dd7ae9fb3c6c19b0a19d3da29194ff56e4db6d639396693dc

Observation a9c97e9d-4fe5-4fba-ba4d-0a63006dc951 · outbound

This paper cites FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.472671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.472671Z digest=sha256:18f3fb5d854eef32c1c99f941da4cadbded076f1dac614258fcc8e241b9fd4f4

Observation 550258a3-3477-4a84-96ab-1b4fb5775f14 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.591170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.591170Z digest=sha256:27d6d65f8e606eee5879e51e55e66f618869dbc90d190569179754f629d4f5d1

Observation 8dd36822-23d0-4a69-a797-88e0c5ff4a97 · outbound

This paper cites KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.745707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.745707Z digest=sha256:314b2799e66bdf27b6228bfab86abb0e9003764d0ce244820a6fc2a499e07565

Observation 2f11bb7e-adc0-419b-9367-506515bf257a · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.911487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.911487Z digest=sha256:1488258e09bb8ee8339359ddc783526174da2454d65699fcb71f85e1fbfd3162

Observation 9e46c474-80d9-4b1b-8ce7-4d9e4401fd52 · outbound

This paper cites NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.015826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.015826Z digest=sha256:6d9c2f1b70638dbb57bf993362333048bbf9c4a3635139819e510c2c4f59fad7

Observation a34eb7be-896c-4b89-b6ba-31cd8f00ef2a · outbound

This paper cites Compute Or Load KV Cache? Why Not Both?.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Compute Or Load KV Cache? Why Not Both?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.180039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.180039Z digest=sha256:96d88ebaf2438d53991a582ed632aaf7089ce0840cc78fa7c3c87eb97eb5401e

Observation c1bfc8f9-aeab-4816-b3b3-d2d996cf9159 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Efficient memory management for large language model serving with pagedattention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.354119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.354119Z digest=sha256:9f5344afea57f2718daf8bf26ed0ff581ab78f4b68ab52e2e00b74391af3e585

Observation 661a023e-a340-4058-832c-8704d5177c46 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.493837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.493837Z digest=sha256:d06c3644c381dd17072bc0c0d834b48458cf5acb1a050dd8161d055ff9c24e02

Observation eada5741-34d4-4ef6-a9c4-bf9fa63b8c1a · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Snapkv: Llm knows what you are looking for before generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:42.226835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:38.606433Z digest=sha256:c022b8aa3badc61269e46ef05a9e3650962f59dd49aa8f0578bdba64e78e608b

Observation 16d82528-4536-421b-acd9-37cd4177d06c · outbound

This paper cites DeepSeek-V3 Technical Report.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.773218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.773218Z digest=sha256:827fd2d1785b65c77a3f8f986cf798e9f2c633e58b8bf5992f28ca00abe8d383

Observation 977bba42-95d7-49de-a137-96176e3d0470 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.885050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.885050Z digest=sha256:e8ee7b0c63d4719a4e5b79ba418048546b582b1ef8573eee44c0cde4baa86661

Observation e878039f-0681-4e7b-9be6-f5d943c4b78e · outbound

This paper cites ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.010317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.010317Z digest=sha256:3d4885624a594000f5eaa82bb69bf1d9f1d8f9254707aaec725ba3e0f288c4f0

Observation bc97e210-8720-4018-962e-cd4d98de5a87 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.114620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.114620Z digest=sha256:3e9558df2bb0b2c2b9a4619a99c787e59d058ef4eb1700f246ecea399387f99f

Observation 1e514596-25c8-486d-96ff-73b94b3258e6 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.247507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.247507Z digest=sha256:1f0bd73a60992dbc6ad647fd8bc42378fb209235752950baecbe01e7f306b37b

Observation 952fef15-d388-47c7-aa55-f96d7c1b1060 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Code Llama: Open Foundation Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.414534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.414534Z digest=sha256:62814f7b9aeaf9532d9442f917bc2e001290dacb14e5fcb8bf8871145bd196df

Observation cc38a9bf-8c56-45a5-abae-122c51efc83e · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.514366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.514366Z digest=sha256:9b7133b6878d4d593320b37e91f17c6e639a75fa45f96beebc830c039d7531bd

Observation 9545cbf4-34d8-46f4-9560-9cacafa53065 · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.632271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.632271Z digest=sha256:933aaa934e79a0ae579fcba24645489427cafd3606aa530c2fa8eb335445873d

Observation 04f60f49-951a-40a4-a26d-2e8c88f2a8bf · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.786344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.786344Z digest=sha256:97f2d5862ed75bceebd41916fd89b744c29dfb25e9eb4ececfbfd2194a023e38

Observation bed9b6f2-dc40-4894-92bb-2aa4d92a1ff7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.932459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.932459Z digest=sha256:06ea5aebd934812f7bc8f206f4b018e84f69ed2cab10b74049d4f0e57261b38a

Observation 6a23d7f7-75e2-45b2-9981-415acf380d00 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.064295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.064295Z digest=sha256:4bd66d2e824bddc10a02b68dcc05600623ff662b7b3aef74528451a7fe2316c2

Observation 9b3bc751-1820-4855-b643-3b526a4061b9 · outbound

This paper cites Infllm: Unveiling the intrinsic capacity of llms for under- standing extremely long sequences with training-free memory.arXiv e-prints, pages arXiv–2402, 2024.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Infllm: Unveiling the intrinsic capacity of llms for under- standing extremely long sequences with training-free memory.arXiv e-prints, pages arXiv–2402, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:41.973919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:40.195808Z digest=sha256:b3a69dff4117d58ae4ccdce8ef308479bd43d0842a56f8fec3d2ca10b90a79e2

Observation aa115202-2d34-45d5-9d2a-faa4234d69e4 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.368474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.368474Z digest=sha256:e60121d807be5540a2828f0c4ace154d8a4f92eb5e6949ef105a3a7bc9dff26e

Observation d557360f-bcb7-4c09-9807-c01bc5effd5e · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Efficient Streaming Language Models with Attention Sinks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.524261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.524261Z digest=sha256:6191ff1c19d7a572398b21e3c8ad8f57f3595f771cf50cf9d2d19e8356cda475

Observation 23115589-32d9-4e42-b217-8997271304a3 · outbound

This paper cites Qwen2.5-1M Technical Report.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Qwen2.5-1M Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.663165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.663165Z digest=sha256:7f3932480f6db36ca3187fbc8775e3ff1f273ceff41f3b4607628e305e6222fb

Observation 5437d5b9-1181-4359-adaf-152295e111b3 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Orca: A distributed serving system for {Transformer-Based} generative models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.778601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.778601Z digest=sha256:d895f2c964f784dca123c5b5b80f00a2552e82e5ea73aa3b88f98b5db699b51d

Observation 0a5774c3-cbdc-4a02-aab5-6c6eaea9326c · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.900312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.900312Z digest=sha256:ced4e6335a07bc56030c39d3542338253ccf0db3519d3d4f9f6ba244f2700ba3

Observation 0057c0b2-26c2-4e71-9083-60f8c3ffe1db · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:41.662361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:41.059939Z digest=sha256:342abfc791994b3c12ff6e23d25bdf0db9d7001280c56b31fca54a4b6acaa802

Pith citing papers

Observation adc68189-911a-4ea2-bb71-b3f964fbeaf9 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.988048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:549bbb48cc1573ef29fa772c84a4e07e94fa11263f678b3a053e5843fd8df57c

Observation b692aa6d-3478-4390-a79c-b57fa3cd51c3 · inbound

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding cites this paper.

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:04:20.890041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T07:03:08.617257Z digest=sha256:3e8662af50c10d7d0fc3e484dcf2d9348d7524de2eb9a31209a660fdaceb0efa