Pith. sign in

Paper Citation Record · LEDGER

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection

As of 23 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2605.16839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.16839 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T21:19:31.263068Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact13
  • verified fuzzy15
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c49e1f1-842b-477f-9731-720187dbbeb8 · outbound

This paper cites OpenAI GPT-5 System Card.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection OpenAI GPT-5 System Card

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.902683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:d7e68d080e0b3adb513183df4b5d31dd53c80e9ff30bd45157aff1f31b363ba5

Observation 7dac6bb5-8975-4233-b01a-39f5d059e39e · outbound

This paper cites System card: Claude Opus 4.6.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection System card: Claude Opus 4.6

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.531284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:21f567b457cd780051467c8b08ded35602b526bf23eae4ec99ccbdbff01c5d56

Observation 427d16dc-dccc-4896-92ff-f12179a59d74 · outbound

This paper cites Gemini 3 Pro model card.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Gemini 3 Pro model card

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.528536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:c7a799b5da8e32cc9c2f5bebdb20cce4ee6c77c3264976574735e66647672bd1

Observation ae350d43-5571-4832-b100-698a4dfdb9fc · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Kimi K2.5: Visual Agentic Intelligence

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.966394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:a70e94ef3c0133fc4664783a79d0ce53feb94fe73f4a3f6a45dc48dc10cc111c

Observation 1140920b-090f-495a-98a0-18ee2b204c0b · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Deepseek-v4: Towards highly efficient million-token context intelligence

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.494555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:0972c212ec13d72885c0c908e47bdeae10ca8e07b20faefc351f04277c0ab0a4

Observation fb7488f5-483e-447d-a260-5b81a734ce55 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.960862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:1c2d570547f2910c14be6007551ccc854e3e778cc94b88e01f1c49f67d02cec9

Observation 4b255596-1b75-44a7-bf22-4b928072c54c · outbound

This paper cites Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve}.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve}

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.500331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:121e08405fbc98dfa165cf6538a4ef68a9eab4e1839c9519f9db9414f0457c9d

Observation d1c4bb5a-928c-4644-b225-9e432bc9ea47 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Efficient memory management for large language model serving with pagedattention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.490959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:3b7570f6690d62d0c4caf613ab785cf393e398205b58e3c52bcfb513f47fee72

Observation c44d499d-c017-4259-acd8-ab795910f747 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.521774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:cbbf72779c6a318bfc48043aa970fd00a55b3f8865325337912021c9081a2625

Observation 71394caa-a3ef-42bb-944a-e9e3ddcc6ea0 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344–16359.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344–16359

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.512809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:f2fc069e2494821e7f5be1fa002b4328019a9f098135464a9785d61ad2d7f399

Observation 88b246b5-f362-410a-804b-57dbe56635bc · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.973278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:db20f9a236926f44858817e05de62695a873a2366f2fba66c3eabc87e2ea9eff

Observation e8a9a4c8-9ffb-4b56-ae93-25d269a0e84f · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.Advances in Neural Information Processing Systems, 37:68658–68685.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Flashattention-3: Fast and accurate attention with asynchrony and low-precision.Advances in Neural Information Processing Systems, 37:68658–68685

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.518963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:78971db435a70a5d0e82e53820af5d59f201b878d67d1f9f692d30742c180ae5

Observation f77e0783-d5e9-457b-b771-9a5d613c3075 · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.525509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:1d2a3589b8e20677fe041de2b9b37bf681ed6eeae289ef496c665764ef0fafdd

Observation 691d2615-4b55-4092-9f23-e629ebced8cf · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:22:47.999879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:736efff061f9ac88ac731416d0aab5610c6684632f8b7e8e1b64681c3d42ae9b

Observation 4efd9c23-e8cd-46fb-98b2-a79596de833e · outbound

This paper cites Xattention: Block sparse attention with antidiagonal scoring.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Xattention: Block sparse attention with antidiagonal scoring

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.488224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:3176fde79e5b79b294f4a149841c209debaa5ebfb3b2476761eb91a829999320

Observation 05a6764b-b7de-422a-8305-57fc48fa1474 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:22:47.994012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:48ab1e3bdc1cda7e922acba4f82b86cae11102fb7f5976e9c8c5b9acf9196f22

Observation 11b447fc-e0cc-4902-9a33-bf158d7c27fa · outbound

This paper cites Flashprefill: Instantaneous pattern discovery and thresholding for ultra-fast long-context prefilling.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Flashprefill: Instantaneous pattern discovery and thresholding for ultra-fast long-context prefilling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:22:47.947282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:9449b383d0b8a303a550f4dc2180627650bb6b030f0717634217739dbe718191

Observation 962e510f-9fd8-4b6b-bcb9-c0ea03388e50 · outbound

This paper cites Quoka: Query-oriented kv selection for efficient llm prefill.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Quoka: Query-oriented kv selection for efficient llm prefill

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:22:47.980489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:6b017c1c947ca39ba215a5495c96b7796ab45876725a87bf438b0f37c7397b26

Observation 2cbc0011-0c9a-4350-af15-73391938e936 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.497264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:b545ea6c9b0a7b847aef19f0d23a6fc6bf5b77cb4b6ece870522e023f3c2f741

Observation bb7c0f7a-e9e5-43df-bf1f-8b221d277c35 · outbound

This paper cites The Llama 3 Herd of Models.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection The Llama 3 Herd of Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.914756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:85a35e344c124d5bc333e0d128e70615715c8e3c6d6ecaee077cf87e96592f9f

Observation 635520a1-6475-4869-b933-81c8b241ed8a · outbound

This paper cites Qwen3 Technical Report.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Qwen3 Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.908978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:fae9a28fc37484549e53ca0ca716b545d0ed44934439246110852716c0fe38c7

Observation c25436bd-a858-4593-9e92-77d60582b0c6 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.940654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:1f06cf8193836cf9664f687faeb6367e970af8416f0dc578a4033c6f25f8a580

Observation 954920ce-7e86-49df-9e2f-0c543f769b18 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.986335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:829d8edcd725651480b4cc08b1e50812d01623f3329e1b041ad9a9dda6ee33b1

Observation 89eaba84-a09c-4c15-9f4d-083ab0c63a24 · outbound

This paper cites Flashinfer: Efficient and customizable attention engine for llm inference serving.Proceedings of Machine Learning and Systems, 7.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Flashinfer: Efficient and customizable attention engine for llm inference serving.Proceedings of Machine Learning and Systems, 7

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.508994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:10f6374676a2dd3a675fe2473db38b74e3b05cbaa20353ca861f85e00460c2b0

Observation ac3ca924-8a70-4938-bbff-cf5ff5feb2c8 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.954338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:a334e235e99d0d6398ca6723efcae2587b4eec1daf1114390c43a60147a86ef9

Observation f9643c7c-9d23-4a0c-91b0-f5fc7581215e · outbound

This paper cites Native sparse attention: Hardware-aligned and natively trainable sparse attention.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Native sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.506159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:d24daf57b5d666cdeee60c375653b393bd526e7e1afdfcbee0f01e22d4187a19

Observation b9489a97-4056-4f77-a71b-cf452c776c4d · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Snapkv: Llm knows what you are looking for before generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.515735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:f7e7babcc9cc7d46c3d2b8cb294650742dd72defede233e4cd5584d02b6a0aa2

Observation 855511d4-d017-4326-8c52-d89fe0761a8e · outbound

This paper cites Quest: query-aware sparsity for efficient long-context llm inference.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Quest: query-aware sparsity for efficient long-context llm inference

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:22:49.503164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:61e07ec88c9dd13a0e0fd66c0282432d8e93b090ca37ea08cde9fb5113d22fb1

Pith citing papers

No inbound Pith citation observations are available.