Pith. sign in

Paper Citation Record · LEDGER

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs

As of 17 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2509.11155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.11155 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T17:05:55.055812Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 93258213-1c11-4b1b-8b7a-c5439dfa1701 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.920316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.920316Z digest=sha256:1875014d6fb5aa3b5176158a8406d68cba333a257874815037420ed523d191ec

Observation 3068b344-1b2e-4afc-bcd8-2d5638db01b8 · outbound

This paper cites Longformer: The Long-Document Transformer.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Longformer: The Long-Document Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.924616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.924616Z digest=sha256:d67f56d3bcca9f0eb429781603a8701be44f322484cb4d274adc3e62e9ea7e49

Observation d459326f-8f50-416d-b648-27a8bc4187c4 · outbound

This paper cites Floyd, Vaughan Pratt, Ronald L.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Floyd, Vaughan Pratt, Ronald L

Reference 3

Resolution
verified exact
doi, observed 2026-08-04T17:08:30.591388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T17:05:54.928598Z digest=sha256:f0065cf2bfa929c75d3ec4800c5bb9f04437ce6e0b16f1b9e2cc93e9f2ee3eef

Observation c9f95e5c-ed68-46d4-8503-bc743d3c6970 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.932346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.932346Z digest=sha256:e99578eb3c1c5791a3d718ea71c4db4cbef6912e86ec1fcff94cf3b87ddc859f

Observation 83258d0e-bb77-4b48-8932-691bcba3fa81 · outbound

This paper cites NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.935394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.935394Z digest=sha256:2faf9d7f6520e2b369fe7e273dac78977af3d586b2835a3b8f5a9dd8cc3868a5

Observation 32712d33-f2e2-41ca-8a62-3d2626d1bf0e · outbound

This paper cites Rethinking Attention with Performers.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Rethinking Attention with Performers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.939066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.939066Z digest=sha256:2d9789381d9fbfba0ec33a17a3b42fa33800e591a368d848186930f093cae994

Observation 928076d6-bbac-4a8a-83ce-ec384ae1ca19 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.942617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.942617Z digest=sha256:f9de8e4db12948f193a4b87c67b314f96c92587c4333a3c56aab37ccac9d5428

Observation b72dc323-ea48-49b6-b2ce-a2dfc170fd58 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.945381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.945381Z digest=sha256:ff01531232d2a07679a5adeeb157be9f8b1bb7810aa2364d74d95a50e7d1bb33

Observation a4b44001-d106-4a33-9d79-225cde32a359 · outbound

This paper cites Adaptive Pruning of Pretrained Transformer via Differential Inclusions.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Adaptive Pruning of Pretrained Transformer via Differential Inclusions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.948805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.948805Z digest=sha256:65674d9c54e49c991f3ab0cf63537ce2573fc7fb95232e66329f41e6bf0a3824

Observation 681293ee-ff80-4652-932b-910106da141b · outbound

This paper cites Truth knows no language: Evaluating truthfulness beyond english, 2025.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Truth knows no language: Evaluating truthfulness beyond english, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.952490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.952490Z digest=sha256:158da72025a5acefefd7d66d4c3ea9649b6c29182d75d78cb1af707d63ff9627

Observation 6ae6a731-bbe5-4f09-b6b4-5ccebb72c198 · outbound

This paper cites The language model evaluation harness, 07 2024.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs The language model evaluation harness, 07 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.955204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.955204Z digest=sha256:4d1fd22f3528c3dcc150382f7afd755960a08a952178838bf5397d5cd1a3cc24

Observation a0d25841-e30e-4357-9079-b2337568756e · outbound

This paper cites Golub and Charles F.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Golub and Charles F

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.957823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.957823Z digest=sha256:db6e4f464d0c3a5f4a5dd66b73f634427a63a15c60faf62b356a9b8d404735aa

Observation 1cd2a967-6d22-4941-b422-aa169b0fbca9 · outbound

This paper cites Deep Learning.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Deep Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.960445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.960445Z digest=sha256:8c9db70c682a1251214e9c26291f46a62f1c68d398614f2d06e07b771fae309e

Observation 8d8130db-1d67-4ed4-9c44-159d67edf4e2 · outbound

This paper cites The Llama 3 Herd of Models.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.963867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.963867Z digest=sha256:f21c0fc1eff784e819b84da945be3f4a6edeb7af47e756ae49e4ac2957a5dae7

Observation f3a389cd-f447-45f4-9016-05d4dc3b863b · outbound

This paper cites an unresolved cited work.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.967207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.967207Z digest=sha256:2ae7b48f5c6c9fc6f35fff63dd66480ce5dcff9a753016d4b9132ceff37de11c

Observation 1169c6d5-c313-4985-82f0-b1e6f55915a4 · outbound

This paper cites Kv caching explained: Optimizing transformer inference efficiency, Jan 2025.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Kv caching explained: Optimizing transformer inference efficiency, Jan 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.970414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.970414Z digest=sha256:6f1386b688f878a3ab58d7b9e8e54aaeff701143f55843e05eb3a9a99e389291

Observation ed0ab0ff-6081-457f-8c89-c47ef52b9221 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.973004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.973004Z digest=sha256:b3ce30fb1c637fa8a24dcb8ee75d15151424b91313d9e73658e99913c3dd2b7d

Observation 73ee96cc-d960-4ef6-bece-fc5b621ed121 · outbound

This paper cites Principal component analysis: A review and recent developments.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Principal component analysis: A review and recent developments

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.975560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.975560Z digest=sha256:d00eea3fb17179d8ff249b91ddf0211c1d3c0d6e798be173386e761315252698

Observation c6bcbe9d-9e97-4110-91f6-33cf769d5d16 · outbound

This paper cites On The Computational Complexity of Self-Attention.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs On The Computational Complexity of Self-Attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.978185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.978185Z digest=sha256:9cd486740919984c3903d6711f7cdf8baa3d2b99ab1d188ccefac7f89268f1d8

Observation e58010b8-3eb5-4e19-a215-70c18dbe6969 · outbound

This paper cites A Comparative Study of Pruning Methods in Transformer-based Time Series Forecasting.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs A Comparative Study of Pruning Methods in Transformer-based Time Series Forecasting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.981663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.981663Z digest=sha256:689aaf8774bd2392f7e5ff0ef0e6ca02aea2747cbfd1f183f9253fead75847bc

Observation 9fcd80a4-4490-4be2-8fba-487cb7b61496 · outbound

This paper cites Reformer: The Efficient Transformer.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Reformer: The Efficient Transformer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.984317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.984317Z digest=sha256:c3619cd39bdeee0fffd425d9b32e5a8cdfffbeba781b58553d1acd465f814a63

Observation 05c3b78b-2afb-4170-bb3b-b8fc2e5f8b1d · outbound

This paper cites Klema and A.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Klema and A

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.986974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.986974Z digest=sha256:b16c918813f6e3b002cda56ef3e3a7c870fcb63e13656cf68801335f0f149878

Observation a77ceef9-f494-40d3-b3af-6307d68e9773 · outbound

This paper cites Tutorial: Complexity analysis of Singular Value Decomposition and its variants.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Tutorial: Complexity analysis of Singular Value Decomposition and its variants

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.989498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.989498Z digest=sha256:81b64419623821a9f8af3bdd835e09c28e3126b2746bdeb76bc33f3b926c9360

Observation 0855175d-1051-4a16-9406-04a69017e4bd · outbound

This paper cites Kivi: a tuning-free asymmetric 2bit quantization for kv cache.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Kivi: a tuning-free asymmetric 2bit quantization for kv cache

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.992615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.992615Z digest=sha256:0734acb20a0002d6414d16b454c77e7157590cef35574cc4b8f39d04c800ea3a

Observation f7c6b5c8-ffce-4e42-96d8-61bb809735af · outbound

This paper cites Massive multitask language understanding (mmlu) on helm.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Massive multitask language understanding (mmlu) on helm

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.995910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.995910Z digest=sha256:8ae860483edb8be04d75c1464e5b943d91f8b2571d7e621cddd3d9b63b49c1dd

Observation e96c7a5e-4f22-4b03-9269-f2de0b7960b3 · outbound

This paper cites Principal components analysis (pca).

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Principal components analysis (pca)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.998406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.998406Z digest=sha256:4f359cf7eaa361cf0fc1b697a5ff474c0fce49aa2a5eecd2f3d3bc3ee0934710

Observation 8d05378f-68e6-457b-a989-a0d7c3366fa7 · outbound

This paper cites Pointer sentinel mixture models, 2016.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Pointer sentinel mixture models, 2016

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.002034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.002034Z digest=sha256:9448c027fa26ec2bffae6040bebb246c0244c96f88f73e3d7fd3fe7df455a600

Observation fb0c36db-40a2-4220-8d31-0f9d7104b207 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs OLMoE: Open Mixture-of-Experts Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.004692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.004692Z digest=sha256:92f9b7dd45367e6474f150758539311d704127400f49b3e88ead3f6df7fcdcf4

Observation ecce2904-108a-4489-bac0-607d95f91640 · outbound

This paper cites Scalable-Softmax Is Superior for Attention.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Scalable-Softmax Is Superior for Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.007461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.007461Z digest=sha256:86e93e74eac08bc5f731e0c7d2d90bca444d622e44b8d2c903f1b45cb3e1ae8f

Observation f0bd9784-91ec-47c6-9db7-33bf7abcd744 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.010618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.010618Z digest=sha256:b6080c1aa9b406eb99fbfa06236e145393696211974d681a5e4695d747739bef

Observation eaedb5fc-a1d8-4bac-81a1-a6767fe9db9f · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.013832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.013832Z digest=sha256:c9d447687b2ca0e037abf2337a0b109aa0e8d87786ccae0c8a5e783396cf2d2f

Observation 3ee6bd84-b9bd-4f01-9314-7740af35a0a0 · outbound

This paper cites Roumeliotis, and Manoj Karkee.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Roumeliotis, and Manoj Karkee

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.016714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.016714Z digest=sha256:4bb7fbe025bb228e5c08d140a427a813426d4b7004098bb59a9b55ce94f12432

Observation 5388853c-a40a-42c0-b007-0138bf6d4b68 · outbound

This paper cites Eigen Attention: Attention in Low-Rank Space for KV Cache Compression.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Eigen Attention: Attention in Low-Rank Space for KV Cache Compression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.019176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.019176Z digest=sha256:9fbbfec13f5bfe60dc5addac6f92bd72d9da8ccc5ec5d015f89137131e2da5e3

Observation 49f7425a-66a5-403d-bdf2-06b232b5d774 · outbound

This paper cites Loki: Low-rank Keys for Efficient Sparse Attention.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Loki: Low-rank Keys for Efficient Sparse Attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.021776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.021776Z digest=sha256:1b5596252e7f87516550485cfd8b27a02ac70119815b89270f927ff9db8a48d7

Observation b87d149a-da18-4137-9bc9-9bb63f236dd8 · outbound

This paper cites Efficient Transformers: A Survey.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Efficient Transformers: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.024663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.024663Z digest=sha256:85a311b63857d50f72726008cfb828e9bbaf1e80280033e4be6f01f0f1a2c268

Observation 2526b1cd-9fd1-4bf5-8eaf-a70cb1a1c812 · outbound

This paper cites Attention Is All You Need.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Attention Is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.027257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.027257Z digest=sha256:816c548c5b1a627c4ce996b3cb54e8d99457e08d9477dfc175eba26d6338e2db

Observation 7399275e-57e3-46bf-a805-62c2ac65a583 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Linformer: Self-Attention with Linear Complexity

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.030959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.030959Z digest=sha256:f92ff482cbc023f104733138b42ca06797b96ba444626019d8ba5734637fd5a4

Observation 41752cb9-bf29-4457-9506-43f60f78cadd · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.034000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.034000Z digest=sha256:1570950c1f0a2dca4e8e125cddac472389d24ccab02ffaf66e8a755128fc124e

Observation b5c767e0-c209-4fa3-978b-bb9ab05e7a23 · outbound

This paper cites Chi, Quoc V.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Chi, Quoc V

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.036698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.036698Z digest=sha256:746dad8524577e99ca8c5aa7b1e59f695392e36eebe41a7e20ce07acc4209276

Observation 5898bd25-447c-4626-9270-3cfb753ad702 · outbound

This paper cites LazyMAR: Accelerating Masked Autoregressive Models via Feature Caching.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs LazyMAR: Accelerating Masked Autoregressive Models via Feature Caching

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.039620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.039620Z digest=sha256:aff4dec4be5cd59fcd6f144c97e6d43cbbec48baa77cd74008a41a476e91c994

Observation da446570-6873-4af9-b485-8273b2a09827 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.042420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.042420Z digest=sha256:9c9aec13e2f8bfa733f095d77392c78bfdd9811aa3ad2ffd98e3cdb095029aee

Observation 9edb87ef-4250-4a67-a241-c844189226b9 · outbound

This paper cites H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.044993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.044993Z digest=sha256:78482a7234a90428f66baddcbbd4858d7de01ec2eb3a4c79360406e022b2941d

Observation f076a612-3a6d-4de1-9f4f-f58f04997241 · outbound

This paper cites BlockPruner: Fine-grained Pruning for Large Language Models.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs BlockPruner: Fine-grained Pruning for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.047920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.047920Z digest=sha256:cb80ffb845b4d4ac58bfadf781c8cfe4a2fd30e9f966b6d6fb6545e1ade4381d

Observation 7cb230c1-48d9-4d10-ae25-ca3af94e7c7a · outbound

This paper cites Aligning books and movies: Towards story-like visual explanations by watching movies and reading books.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Aligning books and movies: Towards story-like visual explanations by watching movies and reading books

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.050852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.050852Z digest=sha256:5e66859ebdac16d3201b1aac6e2f912ff8f99164297025ba186e2cca3a763102

Observation 29910e74-dc77-4a5e-b3fb-574c031c8c31 · outbound

This paper cites Wikipedia-hindi dataset.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Wikipedia-hindi dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.053441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.053441Z digest=sha256:3cf871dd6a691f9c7467a61149e269cbd48be5446a7fdd9a421b2eda3da31616

Observation ff8aeaf9-d3ce-4d3a-905a-a4e8f777b3fe · outbound

This paper cites write newline.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.055812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.055812Z digest=sha256:4ac748acdd958e126e9925a91bbaca966c89a639f6de25742eef23b291aaf365

Pith citing papers

No inbound Pith citation observations are available.