Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:02.274620Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.02572.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:02.274620Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:45:48.458849Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T00:45:48.747511Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d00d6d16-d450-434f-9c50-f55eeaeb2bb2 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0264211-adc8-4dcf-9e69-f46c3df0a536 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28ab32b5-96e1-4cf7-8a59-a4db1e863739 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef753db-df26-4f3a-ad33-345fced6e436 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5258d014-ba2d-4a4b-b4e7-f30b0cc1acb1 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9867a5ec-f2d4-463f-bc37-ca11d9540298 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a455538b-71a8-48d8-88b4-fa44995695e9 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference MagicPIG: LSH Sampling for Efficient LLM Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e84205aa-0633-4f08-9907-660047ea07d1 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b527f51-05cf-4587-8bc5-3bf20c84f762 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5eca77-3f83-46e3-8a11-8324063aa0d8 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference HashAttention: Semantic Sparsity for Faster Inference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af753a9a-dbfa-4988-8115-d17d67229f61 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8629c365-fbe6-4024-9823-fe469e099adb · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Memory-efficient Transformers via Top-$k$ Attention
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19c23bce-5875-44f6-92fb-259f5d75ee64 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98153a66-8728-42ce-93cd-e39469ce9dac · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6999e4-445f-4533-874c-ec85146d0447 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1edc92a5-580a-4b71-85b4-741ef6a3ddd3 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02e562ac-2889-411f-b68f-268dde24cd2e · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 012433c7-92dc-48a1-aabe-e548db1efe5e · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 37306390-c055-4a70-999b-5d757f92dd6c · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 098a8c1b-004e-4148-b148-d46686046608 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference DeepSeek-V3 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b7360b8-110d-47b9-a140-000b9b0cab85 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9117e012-3f1c-47d6-9b79-58783242eea5 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4d96180c-7e34-4588-b55a-2bb8ac645924 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39d47bc8-9f11-4b4f-a4d5-29fc19b45eb6 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c9d52bb4-15be-4955-bc9d-3d23740d9950 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e03dd71-908c-4e4f-a12d-e05dcd4523fe · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db9cac96-1645-4357-b4db-31300544deb0 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Loki: Low-rank Keys for Efficient Sparse Attention
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2178ee9d-6c7e-4b91-bb98-be1e9a5b5e6b · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1537f08e-129f-4807-ac9a-fcca9ebf512d · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8c98f77-f5ad-4674-bfef-2c7c10eb8831 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e625ea-5c24-45e6-899b-70ef97c44005 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d50ea01c-306d-4f00-ad0e-a8fda0e12902 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 266518e8-ac32-4b3a-aa97-26293bc11f2a · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e3fd5a6e-c4bc-4898-9132-2d7c822408d2 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 536b9b5a-dd44-432e-93ed-2e855ea981e6 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa457d13-4aa4-42e0-b74b-cda988a847f4 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Efficient Streaming Language Models with Attention Sinks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8438d211-1c14-4880-88be-6b85ebcf3437 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823f57db-0166-4562-9bd7-c85d58cb7e38 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99128bbc-79ff-452d-8e15-744b868cf3bf · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06830067-bf8d-4a76-ac42-edd18430f3f6 · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation baecab88-b21b-4b40-ab88-44eae0a1654e · outbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1c847a-d352-4183-97c4-ee08663a825b · inbound
Training-Free Hashing-Based Attention via Binary Principal Components HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.