Pith. sign in

Paper Citation Record · LEDGER

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2502.09647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09647 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:46:11.464371Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T13:17:57.627009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:21:24.216191Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e8dbd32-4477-423b-9cf2-da50db8eeae2 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.343368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.343368Z digest=sha256:51288d5cfee032aed8e27e6deb3e806c089034fe9040dd8ccd6db7258ade4d02

Observation a3ae2c5c-deaf-485d-b45a-53cead37c47e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.362705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.362705Z digest=sha256:9f6a5a901c255cc01f6ccb0bded46a6dfe51f6dc0f24c0fd1d9ea94e74a46e17

Observation 210a51f4-a9b1-4d1c-bb8a-fa88c98b3879 · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.366865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.366865Z digest=sha256:e8c5c7b7c8ce0828bb74ed022410f776b8428335309d82e7a96185a827957728

Observation 7e67fef2-d988-42db-b8e1-78c4b474c4bd · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.370692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.370692Z digest=sha256:e2924c527291b205981040fe2e4aea1c656746fd5005d394d27e29fc9ad3b6b1

Observation 4d1a736f-6e03-4e41-a732-64618d141386 · outbound

This paper cites The Llama 3 Herd of Models.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.374583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.374583Z digest=sha256:b61d2a7a57370e1f8aae2358c79d2457b270e735d64ece855e307d8f0d670282

Observation 7cdedb49-90d8-40ec-bf61-b7844b04b5ea · outbound

This paper cites When Attention Sink Emerges in Language Models: An Empirical View.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification When Attention Sink Emerges in Language Models: An Empirical View

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.379308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.379308Z digest=sha256:5a634580b1d5c168ed198b6005eead6e2aacdec8ee7db3cf3d8d8d9f042b92ea

Observation 1b23ff28-9c93-44f8-b1d1-d7b36a38705d · outbound

This paper cites Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.383126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.383126Z digest=sha256:d2196c1f93bed06b0a792a064da43cde1bbc72e641cee07b329f8f9d2729c545

Observation b0d34772-fb1e-462e-a094-5c7c29496126 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.391132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.391132Z digest=sha256:17aa32e0bd7e637af05bc100e0a761814166883bae0d7da848d518f3384441bb

Observation e5a635cf-9922-4b56-b3e5-eca0ad66b825 · outbound

This paper cites Mistral 7B.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.394908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.394908Z digest=sha256:14fa72f3dbb507c9cc7d4ee2c099ed817b3b84444256e019f44e737440104708

Observation 85ee0a17-706f-4b77-bee4-a04e7c72add4 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification SnapKV: LLM Knows What You are Looking for Before Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.398714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.398714Z digest=sha256:fc6eca0675368ca69af96cc8acf614cbbf2b777ab2891a270a4cdd1bf4fcecc2

Observation 2d344b57-1caa-4697-9de6-70177ce76f36 · outbound

This paper cites DeepSeek-V3 Technical Report.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.402637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.402637Z digest=sha256:17e45ed34a2c09a770d7498cad7e7004c0b1d589389c9717b7d248f98f4ee0aa

Observation 7085e3eb-6652-4322-a50b-b0d1814b0118 · outbound

This paper cites In-context Learning and Induction Heads.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification In-context Learning and Induction Heads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.407058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.407058Z digest=sha256:97554fae6292be42dff613c0364209df3bd1748fec77b482c69e71b9c6a5e9c4

Observation 10e161f8-1e4c-413d-8d14-22e51e458da2 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.415240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.415240Z digest=sha256:58e64330de49ae730471c827147897c6cf91e8e50d044719ddb1e1e4d521c399

Observation e9cacb04-604f-45c6-97e3-80b3e285eaeb · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.424138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.424138Z digest=sha256:201fc4b6cf863559028807cbc3bab1da04b895fd027f18bcebc682c3b3aec85b

Observation 2883a270-e414-4807-909d-47b065b6dd6b · outbound

This paper cites Retrieval Head Mechanistically Explains Long-Context Factuality.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Retrieval Head Mechanistically Explains Long-Context Factuality

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.433213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.433213Z digest=sha256:02e16fa79c26e19ee1f9a4ffdef81e2a4ac0034eb5774d788e767e92444d209c

Observation 139f1823-1281-4720-991a-971bac5260ef · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Efficient Streaming Language Models with Attention Sinks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.437333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.437333Z digest=sha256:19df065d75ec4b0aeec96c12602687fa0e9649e0f4afd491cb43e690154f3797

Observation 008df740-affe-4a11-a002-85932160f87c · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.441263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.441263Z digest=sha256:c4ae5ed8a783f2214f53a86ed3a9a94acb75b8a61047d77ab5b30119c600876b

Observation 02ec2ad8-0dd1-4c47-9c1a-0b9dbffcc7da · outbound

This paper cites Attention Heads of Large Language Models: A Survey.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Attention Heads of Large Language Models: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.446469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.446469Z digest=sha256:e0db9f87541b12ec12b7857f06d0a2a5c672af972dcf5d94afa984fbfa131d6a

Observation 3b08fee2-2923-4f2e-849d-c6c8b4f77bc1 · outbound

This paper cites (2024); Tang et al.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification (2024); Tang et al

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:46:11.773078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:46:11.451445Z digest=sha256:ae0a60706773d09ddad319450fe5a516e5ba6e225809338a4b5b3aa57b2e7552

Observation f95a754a-db31-4341-b0d2-d7808e626bf5 · outbound

This paper cites qa-1” and “qa-2.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification qa-1” and “qa-2

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:46:11.760733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:46:11.455913Z digest=sha256:42c365adc1d6f8a773ceab010f5c4693c202d6ab79596136052dc5498711793e

Observation b52d1065-85d1-41a5-b1a0-59cbbdd4747a · outbound

This paper cites an unresolved cited work.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:46:11.749185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:46:11.460284Z digest=sha256:49f314b58721a2b069b3045a0fd189edb9888f679e3570b6e01e70b620a29a52

Observation 654e4c44-559e-44a3-9442-9f1820fcbcad · outbound

This paper cites observation.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification observation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:46:11.737206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:46:11.464371Z digest=sha256:5453838dac1c53e012956a4dba76dfe66f7c8fac4d6b42bd75173b9c65279a71

Observation 870ee56a-545b-4041-89b2-ff0acca5654d · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.419712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.419712Z digest=sha256:75680923282b2933903d5433ddaced8d41c0a9a38223c8087d67527fd0ea50ef

Observation dc37a957-0548-46d9-ab98-bf9802b373e6 · outbound

This paper cites Efficient Large Language Models: A Survey.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Efficient Large Language Models: A Survey

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.428921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.428921Z digest=sha256:903478bd696e2e7c7b54cf72902a698e444400dfbe65434232189f1d8086e4c1

Observation decf2b0b-31c7-418f-b8d3-0a81247d7f17 · outbound

This paper cites Qwen Technical Report.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Qwen Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.358715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.358715Z digest=sha256:dd91be8bcc8127999d4dad5c5efda7da68b52743c2ff3376c6ac29e606a66fa0

Observation 00232cfb-b8d0-4ffd-92b7-ac3239210597 · outbound

This paper cites Transformers are Multi-State RNNs.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Transformers are Multi-State RNNs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.411181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.411181Z digest=sha256:312dd276eb9486db6ca772c95ac832c83c2068f3da12c0ff82d08737be9936ad

Observation 9d1b1eb3-2eb3-4236-a4f4-eef2411fa7f7 · outbound

This paper cites ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.348797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.348797Z digest=sha256:13b2ea82102dc8c5a7d007bbf65cecefa0d0c5cada3720e3148990232daca3c9

Observation 8ea75401-2c1b-4358-8536-c1d5d5a5fd8d · outbound

This paper cites Program Synthesis with Large Language Models.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Program Synthesis with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.353983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.353983Z digest=sha256:437ce82f65b88f2f668f451628ff4fe0ffb82959c9c331a3550ed986bc2da121

Observation fc685a8a-5258-401e-b729-4d8d8067d15f · outbound

This paper cites On the token distance modeling ability of higher RoPE attention dimension.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification On the token distance modeling ability of higher RoPE attention dimension

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.387098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.387098Z digest=sha256:8b9ea040b0751a278e7deb611e6eefa19ae7be81e26d2b770705ae81b874560e

Pith citing papers

Observation 2cb2a063-772e-4b2d-8e67-99c4953ce8be · inbound

Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding cites this paper.

Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:21:24.218559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T13:17:57.627009Z digest=sha256:82bff01df4be048107eb94fba20d7b2d434d32bfa8885d829fbe515d0697040a