Pith. sign in

Paper Citation Record · LEDGER

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.02572.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02572 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:02.274620Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:45:48.458849Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:45:48.747511Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d00d6d16-d450-434f-9c50-f55eeaeb2bb2 · outbound

This paper cites online" 'onlinestring :=.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.349731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.349731Z digest=sha256:c1ede7e3cb68683c1b76e7c2f89dba1f724f25c03156fe972e786c5553660548

Observation d0264211-adc8-4dcf-9e69-f46c3df0a536 · outbound

This paper cites write newline.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.459565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.459565Z digest=sha256:e0ce111f39fae248f96ce48ea3de950766294bdf026531e3de5502973930991d

Observation 28ab32b5-96e1-4cf7-8a59-a4db1e863739 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.675401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.675401Z digest=sha256:322650888384b374e327dbdc9fd376af3f771f1655cc1808895eab3c30d4245b

Observation 4ef753db-df26-4f3a-ad33-345fced6e436 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.843590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.843590Z digest=sha256:d268821620224dbe6d43e36915362769d132f41127fef39e1c44477f3ab2cc92

Observation 5258d014-ba2d-4a4b-b4e7-f30b0cc1acb1 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.958618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.958618Z digest=sha256:53222fb90d2b2c5f1c7e5cc05e9ce0084941213e106e58b99732af7c5274885b

Observation 9867a5ec-f2d4-463f-bc37-ca11d9540298 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.070396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.070396Z digest=sha256:a1c399a1853746dea9e55abedf823ea6ee7b1eaef21422b4e478c56717f42c36

Observation a455538b-71a8-48d8-88b4-fa44995695e9 · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.308531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.308531Z digest=sha256:e0292d5c57d7c28ffdd1d4fcab22582d603f14273997bc3c439f3d74e63e83e5

Observation e84205aa-0633-4f08-9907-660047ea07d1 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.446766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.446766Z digest=sha256:80eb43dfd9e03768625c93b409692e23304737b0f9d79e786fec233bdf7ad263

Observation 0b527f51-05cf-4587-8bc5-3bf20c84f762 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.692270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.692270Z digest=sha256:7870fda4e640c5c248023484ed01a3ccd5627c8c99e1ff237668ab5e046d0d84

Observation ab5eca77-3f83-46e3-8a11-8324063aa0d8 · outbound

This paper cites HashAttention: Semantic Sparsity for Faster Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference HashAttention: Semantic Sparsity for Faster Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.886175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.886175Z digest=sha256:d775a4376fea736ad8cc896fe172aea5d5edcd3115a07e8728f322fd0bd200c0

Observation af753a9a-dbfa-4988-8115-d17d67229f61 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:06.014550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.049157Z digest=sha256:d044cd27958faa466413405c13466ee9e6bf79bb0fb59c123972b8380384dc5e

Observation 8629c365-fbe6-4024-9823-fe469e099adb · outbound

This paper cites Memory-efficient Transformers via Top-$k$ Attention.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Memory-efficient Transformers via Top-$k$ Attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.218497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.218497Z digest=sha256:0fb275e5014649c0c9b5ee721fd813dd13e4fa42eb761c52e1e620738db10897

Observation 19c23bce-5875-44f6-92fb-259f5d75ee64 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.316579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.316579Z digest=sha256:e5f75b6fb1d8cb7ef6994fda3dcae127c4a8f5cd03ae04d4d7f5cbcfe94cc1d0

Observation 98153a66-8728-42ce-93cd-e39469ce9dac · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.386831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.386831Z digest=sha256:8dda2c487be041179c40683eb9079cb3979fd51fa4db75a9145f29062fe31362

Observation 8f6999e4-445f-4533-874c-ec85146d0447 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.794332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.455430Z digest=sha256:2b2254c69ce52acbb6bc6acdd442a5ee4f0b72d38ff299b6cbe51c4f62467b27

Observation 1edc92a5-580a-4b71-85b4-741ef6a3ddd3 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.546700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.604558Z digest=sha256:c60295741d8f38288072bb29964ec9532e1084215a9e099178a327ef93042020

Observation 02e562ac-2889-411f-b68f-268dde24cd2e · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.706591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.706591Z digest=sha256:24d2e516c19753489eff62b344817877669adea50cdc79ef8011751e5e165134

Observation 012433c7-92dc-48a1-aabe-e548db1efe5e · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.347082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.830867Z digest=sha256:faf0e2809a50a9c62c64c58f608499b711fbfd3844e729c39af8f6136e2e31a7

Observation 37306390-c055-4a70-999b-5d757f92dd6c · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.949662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.949662Z digest=sha256:284502fb66757a0e631aa821c12a828e9f5a60f3ee7ae6a86cc28fc84ad3e890

Observation 098a8c1b-004e-4148-b148-d46686046608 · outbound

This paper cites DeepSeek-V3 Technical Report.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.038860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.038860Z digest=sha256:8a06e5dbb95d000948cdb061912683b1bd5604f1290b5a6e5fbf9cbdf81ca18d

Observation 2b7360b8-110d-47b9-a140-000b9b0cab85 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.133693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.133693Z digest=sha256:73940a45adbcfc12856e402233c20f9c6e4d0e91532f5e9d20fb029ebe257ab0

Observation 9117e012-3f1c-47d6-9b79-58783242eea5 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.128766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:00.262542Z digest=sha256:abb4ee843d2ce66e5bd46e498c25d1be8b469d8b3d605bb8dbe3465783155500

Observation 4d96180c-7e34-4588-b55a-2bb8ac645924 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.945624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:00.355214Z digest=sha256:79521b4079ff468cbde6b64eefd5e9879018e185ad279d3c4875d376fee6a26d

Observation 39d47bc8-9f11-4b4f-a4d5-29fc19b45eb6 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.729592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:00.450652Z digest=sha256:106abf79c3f1f9af1426d0e2f94b309cd35eab1f05e6518a6b8c70e95430ace6

Observation c9d52bb4-15be-4955-bc9d-3d23740d9950 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.637951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.637951Z digest=sha256:703727017efbe5a645c7322151b032c9547ae63aeb3e40baea12c0596d68c8c8

Observation 0e03dd71-908c-4e4f-a12d-e05dcd4523fe · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.727994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.727994Z digest=sha256:93fafcf5b660c344cd7c20e2c4e63419bb95b7308bee48cbc659cb4736e895c8

Observation db9cac96-1645-4357-b4db-31300544deb0 · outbound

This paper cites Loki: Low-rank Keys for Efficient Sparse Attention.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Loki: Low-rank Keys for Efficient Sparse Attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.797982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.797982Z digest=sha256:ecc6c037a478a1395ea33574a6c0190a72959050aa08f5c2a6d62c7a8da011dc

Observation 2178ee9d-6c7e-4b91-bb98-be1e9a5b5e6b · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.895653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.895653Z digest=sha256:15ca253b6c7f63f8452d8cc4248618589e8bd3a0295fd1a4caef8ae9ce1b219c

Observation 1537f08e-129f-4807-ac9a-fcca9ebf512d · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.503621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.047307Z digest=sha256:9ab74343a8abf95bf83bbb2eb553ccd2a73232214c7df08b474d75bf7666a56e

Observation f8c98f77-f5ad-4674-bfef-2c7c10eb8831 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.145607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.145607Z digest=sha256:28ab32a89fc300cea3c392875880cbb3fd305b1aed53459ac672d54dfaadb2b5

Observation 42e625ea-5c24-45e6-899b-70ef97c44005 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.299814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.299814Z digest=sha256:f91b0d142bd2fee4f2035ab2b469b4e25803ee391d67a1756750f8da86b58e98

Observation d50ea01c-306d-4f00-ad0e-a8fda0e12902 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.278953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.430394Z digest=sha256:809988f3e18f0da9d4be964a716820bc6624ad8994928fb34c69b34e93185248

Observation 266518e8-ac32-4b3a-aa97-26293bc11f2a · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.091956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.499338Z digest=sha256:5f837e92bc03b3ac7f342abbe63290e8b80c3e81cd743f26b81fa6358d7d7b33

Observation e3fd5a6e-c4bc-4898-9132-2d7c822408d2 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.842301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.598450Z digest=sha256:fe54913adfdbca9b4848d63580ef23415abf3d5d6fe6da3608fdf689cca7e81c

Observation 536b9b5a-dd44-432e-93ed-2e855ea981e6 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.625450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.694965Z digest=sha256:d81fcdb77f1def489c6e5c241bfb676fc24ee973289c4f57ddf6991dc202cb8f

Observation aa457d13-4aa4-42e0-b74b-cda988a847f4 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.762361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.762361Z digest=sha256:dc30a3bddf1947902692e07828199c49fa990de45d4964025e840835843db895

Observation 8438d211-1c14-4880-88be-6b85ebcf3437 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.860825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.860825Z digest=sha256:c0e9ea808c421748966c13f9eb55b7fd4d780c67051d8089ba6b0c2011403c75

Observation 823f57db-0166-4562-9bd7-c85d58cb7e38 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.351191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.933392Z digest=sha256:bbf385f2683fd36c4343d46b762413891ea0b63ffcf22f7b977f0fbdb5740fb7

Observation 99128bbc-79ff-452d-8e15-744b868cf3bf · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.056187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.028935Z digest=sha256:abe9691deca618392b30ef3e3d6ccaf1395ee1220830faf446061d322a682f21

Observation 06830067-bf8d-4a76-ac42-edd18430f3f6 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:02.786954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.149538Z digest=sha256:f2bac83ab54961ca7674eb35e601a58d02a375acb036b47cf5fcdb07ae718281

Observation baecab88-b21b-4b40-ab88-44eae0a1654e · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.274620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.274620Z digest=sha256:d54663f7a064b42bacde25adcb35412ef2c721d33e6fa4c7eb37568e08b8db8f

Pith citing papers

Observation 8a1c847a-d352-4183-97c4-ee08663a825b · inbound

Training-Free Hashing-Based Attention via Binary Principal Components cites this paper.

Training-Free Hashing-Based Attention via Binary Principal Components HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:45:48.754100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:45:48.458849Z digest=sha256:0eb3edbbd54dd8d2c266391530c0b3a575a7bff5464f9e07384e0cf948286472