Pith. sign in

Paper Citation Record · LEDGER

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference

As of 22 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.02572.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02572 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:02.274620Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:45:48.458849Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:45:48.747511Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d00d6d16-d450-434f-9c50-f55eeaeb2bb2 · outbound

This paper cites online" 'onlinestring :=.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.349731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.349731Z digest=sha256:219b85c3a5d0e3b8a584d896199f2e6cf8046f388cf528ae0fdca059250a920b

Observation d0264211-adc8-4dcf-9e69-f46c3df0a536 · outbound

This paper cites write newline.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.459565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.459565Z digest=sha256:96788caea176d9b59347c1964ed010790d465cd70ee5fa8576d4ffabde388130

Observation 28ab32b5-96e1-4cf7-8a59-a4db1e863739 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.675401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.675401Z digest=sha256:71bcb8cebbdde600aa26e948eb7b72f1ccea1c872791ca6cb6f7fb33d050c445

Observation 4ef753db-df26-4f3a-ad33-345fced6e436 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.843590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.843590Z digest=sha256:e2383168da2833052866308dd3322abb643dd83f20d2dfbb44ab3f059c3eaa8d

Observation 5258d014-ba2d-4a4b-b4e7-f30b0cc1acb1 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.958618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.958618Z digest=sha256:42ea65db93d6f997e7317d4e8aa8f8f7abdabe7d712af3d2688f35708475b4c4

Observation 9867a5ec-f2d4-463f-bc37-ca11d9540298 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.070396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.070396Z digest=sha256:b7579c81f89833592579244764eb5c48daaf2ba1e19bd7d32a59dfbb0dc9598b

Observation a455538b-71a8-48d8-88b4-fa44995695e9 · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.308531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.308531Z digest=sha256:afc56763da510a0cefb47908eabee46c0a7b835a3b7a3e4f3913d3e7d81d96a6

Observation e84205aa-0633-4f08-9907-660047ea07d1 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.446766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.446766Z digest=sha256:027e07c048c73d87e8a8ef4ee32a37f0c6cad3b09f3825a4282dcdcd66a762cd

Observation 0b527f51-05cf-4587-8bc5-3bf20c84f762 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.692270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.692270Z digest=sha256:61c1fcec5e0c0658f04ca5f5d0db96340e281aa8e79c9a2c61ab8abbbcd0a870

Observation ab5eca77-3f83-46e3-8a11-8324063aa0d8 · outbound

This paper cites HashAttention: Semantic Sparsity for Faster Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference HashAttention: Semantic Sparsity for Faster Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.886175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.886175Z digest=sha256:b2c512a3495f9605b050cbad6da9deb9839e84ba8d34e9f8899ac5dee324dcc3

Observation af753a9a-dbfa-4988-8115-d17d67229f61 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:06.014550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.049157Z digest=sha256:1803ba25b19afa62c04da6f6bf86204331e713f0af4b9dad836be55743ad6d85

Observation 8629c365-fbe6-4024-9823-fe469e099adb · outbound

This paper cites Memory-efficient Transformers via Top-$k$ Attention.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Memory-efficient Transformers via Top-$k$ Attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.218497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.218497Z digest=sha256:3e2ca929cd372098c92c235e00cce1c8de8dbfc098a6379c6ace9cdf4c6cf6cc

Observation 19c23bce-5875-44f6-92fb-259f5d75ee64 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.316579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.316579Z digest=sha256:6f03701f5851f998954fa19e937ea34ca4ab51caf53d7ecfc751032a5051fe1b

Observation 98153a66-8728-42ce-93cd-e39469ce9dac · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.386831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.386831Z digest=sha256:1d432e88118c9a51a787e313ca8c0c5b1431533576e5d9dc88e74263b027f97d

Observation 8f6999e4-445f-4533-874c-ec85146d0447 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.794332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.455430Z digest=sha256:5bfde0ef4d8404d849a8663172869147d962963e55b7eb73378e47c250e93739

Observation 1edc92a5-580a-4b71-85b4-741ef6a3ddd3 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.546700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.604558Z digest=sha256:078f572dc067f329c43cf9bf800e1e6b2feb21a4561bba43a2807602a6234ffc

Observation 02e562ac-2889-411f-b68f-268dde24cd2e · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.706591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.706591Z digest=sha256:064522441397432902d7dda7dd138cdef366564c3a3cd49871acfc4a9f856e02

Observation 012433c7-92dc-48a1-aabe-e548db1efe5e · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.347082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.830867Z digest=sha256:1799818ec8df07ce52792ec6b4bcee46397d3702b1f674f04ca5635accef6e28

Observation 37306390-c055-4a70-999b-5d757f92dd6c · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.949662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.949662Z digest=sha256:e3a2f2fe6343b3d7526f249d66c9afb67571df3fe4adbec926d1a8ba7fd9378b

Observation 098a8c1b-004e-4148-b148-d46686046608 · outbound

This paper cites DeepSeek-V3 Technical Report.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.038860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.038860Z digest=sha256:1cc07cdae45c264c69d44e510954fb8012c5130765e4e89db3d3d76ca8549368

Observation 2b7360b8-110d-47b9-a140-000b9b0cab85 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.133693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.133693Z digest=sha256:5491b0eb0e7d41ef08760c5a7ee6db768838d6e0a437daab2216e719a89d775b

Observation 9117e012-3f1c-47d6-9b79-58783242eea5 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.128766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:00.262542Z digest=sha256:113b1aabf024fd8e03be9e458e001d6932d82723b8be3877b00c4b5aa5d77fb4

Observation 4d96180c-7e34-4588-b55a-2bb8ac645924 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.945624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:00.355214Z digest=sha256:c3588f317f397c076c582e188d88c8678041ea9f3caea79ceb1c83a85b708333

Observation 39d47bc8-9f11-4b4f-a4d5-29fc19b45eb6 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.729592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:00.450652Z digest=sha256:e97d3d1921816aff451256bd1716d238f249fd55a44c7aeeb579c1c659e2f2a1

Observation c9d52bb4-15be-4955-bc9d-3d23740d9950 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.637951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.637951Z digest=sha256:677d38dcd42007415514d44c38271c337578bd99cb3920a57cf850bd27f2332f

Observation 0e03dd71-908c-4e4f-a12d-e05dcd4523fe · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.727994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.727994Z digest=sha256:d4bc02b2b30ba2d07447dc0fd591d1e588e7308b8062b4f7bd14b09b23055734

Observation db9cac96-1645-4357-b4db-31300544deb0 · outbound

This paper cites Loki: Low-rank Keys for Efficient Sparse Attention.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Loki: Low-rank Keys for Efficient Sparse Attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.797982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.797982Z digest=sha256:2364a83861b811fc8e0a60e0fa34c1b37d08bc1e317712aa305224cfd3fca7ee

Observation 2178ee9d-6c7e-4b91-bb98-be1e9a5b5e6b · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.895653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.895653Z digest=sha256:258502a4e2e2239681870fb3fd7f0f8f7cd6907a5fb7066a9d12a998f743396b

Observation 1537f08e-129f-4807-ac9a-fcca9ebf512d · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.503621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.047307Z digest=sha256:8951420730b7d694e9d86cd83af3a64736a87e923facd97bd8b80140a23e6b95

Observation f8c98f77-f5ad-4674-bfef-2c7c10eb8831 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.145607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.145607Z digest=sha256:3f130fb251f5906ca54d8ac4d0f9528db60521536c13a59b3545086550181549

Observation 42e625ea-5c24-45e6-899b-70ef97c44005 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.299814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.299814Z digest=sha256:aa977c0f3f0c88797bb507861f84924ad72d0e9664cabc30c5f790ff6b9a3237

Observation d50ea01c-306d-4f00-ad0e-a8fda0e12902 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.278953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.430394Z digest=sha256:26aad494e9fc26d67c19e735fedb6849cea2f5d044061ea7382b739f11e69b05

Observation 266518e8-ac32-4b3a-aa97-26293bc11f2a · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.091956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.499338Z digest=sha256:eb721b00934b996bbfc3d6b6f21bd2ac0ba73e0b72362293c685c4225b915e79

Observation e3fd5a6e-c4bc-4898-9132-2d7c822408d2 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.842301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.598450Z digest=sha256:5e11e4c669f31e57b4e3d0ec506ed265002d8c0d4cee549a1817803434ea844e

Observation 536b9b5a-dd44-432e-93ed-2e855ea981e6 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.625450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.694965Z digest=sha256:71d672af7b8a703b55de999a72ec7cbf0894be4ec374c2b9e6d757131196300d

Observation aa457d13-4aa4-42e0-b74b-cda988a847f4 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.762361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.762361Z digest=sha256:5eb0d3eb34366c8aaacb53760e269c776ea678544715c41c7b11473a0b69c6da

Observation 8438d211-1c14-4880-88be-6b85ebcf3437 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.860825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.860825Z digest=sha256:730181e72b53459cfb78e420cd9aaf08cf944c95655e4d97687296e88f2092ef

Observation 823f57db-0166-4562-9bd7-c85d58cb7e38 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.351191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.933392Z digest=sha256:2df2e9390a92c7393bb9f732959124367182c758921d0356e6ca7919552d4af3

Observation 99128bbc-79ff-452d-8e15-744b868cf3bf · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.056187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.028935Z digest=sha256:beacde6084650dbd0291ed0d2154a793f9d7dbba1803fd14452afcb11f7b899b

Observation 06830067-bf8d-4a76-ac42-edd18430f3f6 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:02.786954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.149538Z digest=sha256:cc064eec8ede7cb4811883fa621adbda0c73559e2e8c2a950bd05b91c09a51a4

Observation baecab88-b21b-4b40-ab88-44eae0a1654e · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.274620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.274620Z digest=sha256:b9b58ddbc6c9eeda57697eead010811e87ff73ff5b2d6191c1aa3ef69ffd0b31

Pith citing papers

Observation 8a1c847a-d352-4183-97c4-ee08663a825b · inbound

Training-Free Hashing-Based Attention via Binary Principal Components cites this paper.

Training-Free Hashing-Based Attention via Binary Principal Components HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:45:48.754100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T00:45:48.458849Z digest=sha256:2e2e7abab9609311b88de8279b9368c5bbaa70637dbdcd8c0fa81c2e41a95737