Pith. sign in

Paper Citation Record · LEDGER

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

As of 5 August 2026, this Paper Citation Record lists 100 of 145 outbound references and 56 inbound Pith citation observations for arXiv:2306.14048.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.14048 v3

Coverage vector

measured 100 of 145 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T18:00:50.053377Z

measured 156 of 156 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:53:30.705096Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 145 outbound references displayed

  • verified exact59
  • verified fuzzy33
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

28
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation fced8819-25a8-4397-bd21-65975ce0bbf6 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models LaMDA: Language Models for Dialog Applications

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.150984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:5ecf945dcb0a600011c5dbe3885f921318cef5fa0171e7b188a0aad85dd2ade8

Observation 3723f9fa-8459-42e7-ba17-cd504a4ab633 · outbound

This paper cites Wordcraft: story writing with large language models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Wordcraft: story writing with large language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.585642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:c6d5f7b45911e0a3955158eeafdfc44cf4c9dcf8c90ed0598aa0f36495cff8ce

Observation 4f5a1443-9b46-4896-ab1a-eb93dc799768 · outbound

This paper cites Emergent Abilities of Large Language Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Emergent Abilities of Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.130973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:0531f29eee416298cc7ee23653496872da0fb49cd1cbcafe825ba3e93802698a

Observation d3ac7804-4d5d-4b0a-88cc-9f74bfbf764d · outbound

This paper cites Benchmarking Large Language Models for News Summarization.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Benchmarking Large Language Models for News Summarization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.137445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:105d5291cec2d35cf73ae0840babc908edd87751dfb7257fdcef0bffdc9bc68c

Observation 56958681-84c9-4f80-a94e-4e5c691163df · outbound

This paper cites Efficiently Scaling Transformer Inference.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Efficiently Scaling Transformer Inference

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.145712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:7ae4e92acf6c04ca2990e94222c79eb88f5e0c7710de4ac65190dddb623f1d15

Observation e43b65a3-9de0-43bb-8a64-3d1a37572feb · outbound

This paper cites An anomaly in space-time char- acteristics of certain programs running in a paging machine.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models An anomaly in space-time char- acteristics of certain programs running in a paging machine

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.556288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:150935efc3e013ff98545c824d796d74c5091d0661d0a1505aa8dd4c682c8e8c

Observation 3f23259c-2217-4900-a171-ea4efa633e41 · outbound

This paper cites Reformer: The Efficient Transformer.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Reformer: The Efficient Transformer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.403497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:44720e088442a17204623c29c0537d49a8ece0fbbb1b0441ba5e414d466944b9

Observation d04301a6-53f6-4b00-b006-2f96f081ba71 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.562913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:a279231c802eedd7733a55b4ce1e02d309e6d31a3b4392372b40d8247d35102a

Observation c0d05d41-8d81-40ff-ba0f-92bab4c0557d · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Generating Long Sequences with Sparse Transformers

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.407792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:8f9c7deed1f135fe3f7273bcdc5a6e5e5d632fc4229b8089b89d8a3c118b61d6

Observation e3ff51cd-f292-404f-a82e-ff5464fed06d · outbound

This paper cites Rethinking Attention with Performers.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Rethinking Attention with Performers

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.411525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:a89bf0e5b5b16c4cf46116e3c368b7c2d3d064f8f19bd6e5997bb0fce9f86231

Observation e29fab6c-c4d5-4e9f-9162-7fb3e13daf97 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.574127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:197ee7d4e5632f2740b273b5073c43d3d339ced88d7fb80ed137a081414afe51

Observation a0b8175c-e1ef-4d7f-bc10-19c7097ebbc8 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Fast Transformer Decoding: One Write-Head is All You Need

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.415420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:8d5424aefe8445811688a05d4cf7d8a1778a672f5d95bbc3eefa55586e9fd25b

Observation b6bc521f-0fb4-4bbc-9491-7624598cf83a · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models PaLM: Scaling Language Modeling with Pathways

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.419451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:cfba8a78dc9b907f42308bb45011944176df5285bc3a7f997977e3df21c25398

Observation d777afda-dd71-4d88-8f07-8d34078906b2 · outbound

This paper cites Learning to Compress Prompts with Gist Tokens.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Learning to Compress Prompts with Gist Tokens

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.423410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:5cdc859ed31d5f906303dbc653c00a5b6c25b7e3e51bfca3837c3b95dfb07be1

Observation bfe2969c-938c-4912-ad93-fad37f09316b · outbound

This paper cites A framework for few-shot language model evaluation, September 2021.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models A framework for few-shot language model evaluation, September 2021

Reference 15

Resolution
verified exact
doi, observed 2026-05-17T18:00:50.125201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:f77a7bd59859ac35d501686b2badfe9c21c8b5040b69712aff778b3a8ec3bf7c

Observation e7f06c6e-425d-4089-b4da-5e359e40c414 · outbound

This paper cites Holistic Evaluation of Language Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Holistic Evaluation of Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.427058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:39d3e8e30125797e99037ce5eb3cadeb72eab8415f885aa12f107272da16b178

Observation 39900118-976a-4971-8a9c-d9c4cbe878d9 · outbound

This paper cites DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.431176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:c923a583e0a771ce3d02f7c55d6f5875e0b12aef777935c6799cdae24318aa9c

Observation 808935d7-d2ef-4e8b-bcc6-8b3a25a05233 · outbound

This paper cites Hugging face accelerate.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Hugging face accelerate

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.594043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:bb22f94740c87e47b62ecefd961a84a88ec18988cbf0758cd4b088afa3ae2d95

Observation ad34b387-962f-4c5c-8a3c-889b798ae04d · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.435347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:a1cd26d93aed48cd2191df09d9b211f536c78f8241273fa5f7905f4e7bcf89c1

Observation c4f47523-3b41-446e-9406-cf3e323a1f4a · outbound

This paper cites SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.439584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:530f8b51aea7065f110129483da4bea7c84a5baec8e9f852279baa14e7f84da2

Observation f6d64c80-83b1-49e1-982d-3bf2c2709cd0 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models A Simple and Effective Pruning Approach for Large Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.443610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:56dcc0ba88b40268006ee735ad69d964167036275b41e8fbe5815bf075d6854c

Observation 11843369-3084-452a-be33-b5b9ae8df156 · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.448044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:10d5742319e456404763a22e9c96c6a63fc318f43e22711a18083b796ef7f2a7

Observation afcd2480-bf7c-4261-ae6a-cc7b2f109de8 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.452114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:eefadd7a4f5d80c719786fda2e182f3e19c9a37539914697c78c0e92bcf1f5b5

Observation a2561729-235e-490c-aa35-b56b9a8c7dc2 · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.456232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:4f9334e0be533992067ee791878f28776d7c2576c2f839ad9f45e15d98b4049f

Observation 29d83d7a-d3a3-4d48-b1df-85a29bb200fd · outbound

This paper cites ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.460651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:769c46e77da8eb0638ad815bb0d685c51f42c53953b3829492c371a358705f40

Observation e2894113-913a-4245-bfea-768e0d9661f8 · outbound

This paper cites an unresolved cited work.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-17T18:00:50.615585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:9d1320256e8ac1eb514dbbb8a16fc4c1aa173eccf7422de3bdd2a26e7bfb4466

Observation c79384c4-a832-403c-9ff4-e9d9df506b64 · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.465229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:204fb9ac1d26c2bdf7e016aa83743cc24d7003db2ec165031c816cd143b71e15

Observation 1d96aeed-1f0c-4b9f-89db-bfff5098ca6f · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.469602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:806673476b160aa643262b7b5cd13ddbf5c5c18c847cefb9b92505d945e5b0b5

Observation e171e07a-bf43-456d-94fc-05f4222d9d15 · outbound

This paper cites CoLT5: Faster Long-Range Transformers with Conditional Computation.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models CoLT5: Faster Long-Range Transformers with Conditional Computation

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.473738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:d0a350bb8a4c7721628fa9e1a4123f43029ed7e95e0d9284e29948c7bf0e938e

Observation 4da32423-fa58-431f-b363-29295125d054 · outbound

This paper cites Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.478057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:00ee844e8a250102c01c2c6ebc5f1b6ffa7ce9735787d78311eb2cf72cecf569

Observation 4b4a1bac-4f88-491d-9e00-a224064295ab · outbound

This paper cites Efficient Transformers: A Survey.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Efficient Transformers: A Survey

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.482223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:0891d5fa4436b4a7bccfb708377f720a0b78ffc8ad3f6d7c76042168b3c2c837

Observation 4fef8ba4-b828-43a2-bc4e-ce3ab5674094 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Spatten: Efficient sparse attention architecture with cascade token and head pruning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.632000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:f5a772cd8a99a49c7f81bb03beeca38e39a8a44d188032e61056c087b21c0a01

Observation 124d4f2d-c709-4ef8-9a2b-68823a89fb29 · outbound

This paper cites The lru-k page replacement algorithm for database disk buffering.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models The lru-k page replacement algorithm for database disk buffering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.634555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:f5899595f5096ab2c8e73ea77986ded5c91e851d01c904ba5b1d8f28601ef590

Observation a06881bd-7c08-4ab2-87c1-18cf95ca6c51 · outbound

This paper cites Lrfu: A spectrum of policies that subsumes the least recently used and least frequently used policies.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Lrfu: A spectrum of policies that subsumes the least recently used and least frequently used policies

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.637086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:bb93c0c43185d4fa09629a15ef42054024591a4c7944612334d3371c90c19dff

Observation e16670d8-40d1-4f20-8134-227a63d0a8e5 · outbound

This paper cites On the Expressive Power of Self-Attention Matrices.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models On the Expressive Power of Self-Attention Matrices

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.487379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:d2f5323afd6c48267f48a6c5aa9369868e7205f9df0bfbdc0ee6535bab82594a

Observation 3cec6851-35f7-4026-8f75-32e1079bb50d · outbound

This paper cites Inductive biases and variable creation in self-attention mechanisms.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Inductive biases and variable creation in self-attention mechanisms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.642232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:300697e072d69c32ba3be6ca833b68c53c8590252066c6ae8bfeb9088d713424

Observation adf41e20-c616-497d-973a-5514eef02777 · outbound

This paper cites an unresolved cited work.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-17T18:00:50.644639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:1fc5f05cadc7a2d6ae729312eefe574e2b48c30ef9fac12713b2505d032e1d87

Observation e4a8c8c6-f9b3-4021-aeb8-9c1db567286a · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models OPT: Open Pre-trained Transformer Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.491623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:acff7d4a0d2dd882e551088afb55b416ca4a95f05f77d6ec7f7629af64ee51f9

Observation 6060d875-e1e8-443b-a5fe-02b7e5d77e48 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.496470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:c4d0e44e1949d6f7ee8146a04571c66e21e4d601aa574372d34c2ef8194847a7

Observation b26b22c3-052a-48b3-b637-7204475c6d8b · outbound

This paper cites GPT- NeoX-20B: An open-source autoregressive language model.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models GPT- NeoX-20B: An open-source autoregressive language model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.652714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:2e27058b8860517730219bf1cc65876b7ce0157996c3ccb6586f821001e79371

Observation 757ed7cf-6910-41e4-b61c-aed7646caff4 · outbound

This paper cites Choice of plausible alternatives: An evaluation of commonsense causal reasoning.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Choice of plausible alternatives: An evaluation of commonsense causal reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.655275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:9c8b421b0dc3e8c1a89ee03d59f6d7bc4f573783f2d9b037754f9a38fe7ac396

Observation 04133704-8e88-4ac0-9c7c-4d32fde26b81 · outbound

This paper cites MathQA: Towards interpretable math word problem solving with operation- based formalisms.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models MathQA: Towards interpretable math word problem solving with operation- based formalisms

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.657985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:ff1498c8c3983c92c1d7663a848d8549b736020f4932ad897603b66cb13be5b0

Observation bf96a834-9bf8-40e8-9d2d-fc2496f3bf18 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.661003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:ff2230a836b84e2275819e053b80651945d0b395b5cb44d8c7ffd61103014aea

Observation cee97b64-2cea-4a2a-93ff-fc69196cea8a · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Piqa: Reasoning about physical commonsense in natural language

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.663867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:557d81e1fec0fd5f7d98a4209b583c0df2bdc088ef0f47987cb297a3f23ea79a

Observation 0dc83c89-93fa-4e34-895d-4a0f1158bf7a · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.501846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:611649ed3ce684b96a65300ab7328a3c5dabea388481fc84fd8900b633cfb1f4

Observation 8ae21149-5e5c-4a9c-9c35-e3540ed94964 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Winogrande: An adversarial winograd schema challenge at scale

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.669424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:e90aa72064e3ee1a45927da84c2b6054532c811c4d2d056e33192d7fa3b776d9

Observation 4f46ea0d-79d6-4bd0-878f-fc561fe7c2e3 · outbound

This paper cites Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.506635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:6f6db36bdeb4b09b9b158543e2deea5d68f6922de5b52a81b7a0b80adc78b69b

Observation 14aa63f3-2faa-467e-902c-0a4a6c849b10 · outbound

This paper cites Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.511095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:3d2e721164b737076d7041b78a91e95fa10377feabc1c42afc7946304f473e5a

Observation c7a107ec-3c6d-415d-b91e-5df8870c6d2d · outbound

This paper cites Hashimoto.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Hashimoto

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.677381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:7f7e81f81c4d73694beb751a6703b7ce4d1cad395834f5ba6e297a25239a8548

Observation 4e444dd0-ec7e-4aef-8704-9dd9576650f9 · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models P Xing, Hao Zhang, Joseph E

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.679904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:00b96d3212665667e2c96ddaf6ed0ecd9e6cdc9f0f4692d0ea29bc31ec7db143

Observation 6b837594-07d8-4e69-a558-167ac31b889b · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Efficient Streaming Language Models with Attention Sinks

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.515409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:a2bd46562fdb3d7698c9d6b5a64720a48d846c9cfa1fc4840c0fedfbccdd9495

Observation 1798dffb-aa43-4817-bb5e-e033a0ad4754 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.520156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:f835aabffcffe1aa47340023e19f4b6cf0ba624d67e63b51e1fee6964a5cde67

Observation d338e598-f15b-4608-9ea8-b1a674e6ac14 · outbound

This paper cites Compressive transformers for long-range sequence modelling.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Compressive transformers for long-range sequence modelling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.687807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:ba81d3f1829d5e966e443ea5bc0abaca13895e43b0d4b7d907f0e9995cc11bae

Observation 96dc0444-d38f-42ba-ad82-431d6e8032ca · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.525397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:fea0fa03625211a7e5986a8606bd4b9cd883809e481dd8182a0c5eefe42e14aa

Observation 2bda1484-707f-48c3-a2a8-72b7f260e5a6 · outbound

This paper cites Quantization and training of neural networks for efficient integer-arithmetic-only inference.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Quantization and training of neural networks for efficient integer-arithmetic-only inference

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.693427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:740d17ce310560cbfd923b20aaf47c1a7c149f217cffbdbe5d89d93880b65727

Observation c604098a-7ab9-4bf8-9b79-670b776bd44a · outbound

This paper cites Data-free quantization through weight equalization and bias correction.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Data-free quantization through weight equalization and bias correction

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.696167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:db55ecbffa6dcdb97e538c8cd9e6e4c1591b4c6be937f615296915da4ee8aaad

Observation 1b7d9ff6-7108-46ea-a452-4c2a6d13f6bb · outbound

This paper cites Improving neural network quantization without retraining using outlier channel splitting.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Improving neural network quantization without retraining using outlier channel splitting

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.698829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:fdc3a469853d070e51acba88844e4bb24888f20241eb2dc892ffdf7665af2403

Observation 87979139-e97b-4295-a87a-55ae02c0691f · outbound

This paper cites Pruning Convolutional Neural Networks for Resource Efficient Inference.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Pruning Convolutional Neural Networks for Resource Efficient Inference

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.530175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:6b912ed6c7c0543a8c07b0397753befefc196407807d23caee59efe4d1d7e90f

Observation 25be5d4e-4748-42ae-a472-95f6c5e6d6b7 · outbound

This paper cites Rethinking the Value of Network Pruning.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Rethinking the Value of Network Pruning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.533939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:5051a38865bbc682d502f0570423c0e252de114e872be9e9518903347006a11c

Observation a05d7b55-b29d-4282-8852-ffbf71d4d74e · outbound

This paper cites Filter pruning via geometric median for deep convolutional neural networks acceleration.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Filter pruning via geometric median for deep convolutional neural networks acceleration

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.570428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:1a5541e95c858c2ecdf08302b1c4038b110c951b8c023b53ae9c2c6414f8ae1e

Observation 98d11d86-9fde-4f7d-bfae-0aeb5bedddd1 · outbound

This paper cites Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.577163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:d70c034ea90a992257f22599c78634b4a8b3996bce5b02195dda12449fcedd94

Observation a009c789-57d3-490f-bf16-f8e6b17d94bd · outbound

This paper cites Distilling the Knowledge in a Neural Network.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Distilling the Knowledge in a Neural Network

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.537885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:5e9f3d6936351e103bfc5a0b6df680502b921e89a788d7e11dd4a078035f05d6

Observation 8aa88b96-59ea-429f-bb79-fd8912b2d21e · outbound

This paper cites On the efficacy of knowledge distillation.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models On the efficacy of knowledge distillation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.582813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:75d02a78189c5ece12f2545014b8b954b9afa41f6ad900faaef6a0b570dfa7b9

Observation a9712fad-1422-4898-8d0e-7367cbdc2d1b · outbound

This paper cites Distilling Task-Specific Knowledge from BERT into Simple Neural Networks.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Distilling Task-Specific Knowledge from BERT into Simple Neural Networks

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.542178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:44a23177fc7ccc88fbeb27d15b8eada1659e51440281c7a141bc7d863f69bb03

Observation c2c8b5e1-66a5-4ce5-9a41-efc7531c18d1 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Training data-efficient image transformers & distillation through attention

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.591420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:30cad13264bd3f72c53872f2d2e970125424756ba909a57f3167f83d6a043d23

Observation e8390ad4-7635-44b2-8896-d1cc38e39fa3 · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.596617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:7e038632592372895b80069f4d57130e101d1a69ab86d78def4758abb856202d

Observation 4c65e2f0-e722-4d16-84e1-9b0103f0ab98 · outbound

This paper cites Xlnet: Generalized autoregressive pretraining for language understanding.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Xlnet: Generalized autoregressive pretraining for language understanding

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.599309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:76a4776e6288a437ad74934205360c04700683a8185ba1622b08b678e159995d

Observation 4ab4ba32-40c3-4997-9fb8-2dd2edc23213 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.545885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:ce687ed499d0eccbb0ffb256df96379a5bb9ae899fede5295fb9374e36a9eb51

Observation 703db5f6-1c1f-4b45-b083-da9089f4ab36 · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.549413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:b56854c78e0ceb4fdcf3aa1c6327756d76b37343167b72067d2c36437a4b848a

Observation efd5aa8f-f9b2-41f0-a21d-cf3c29ae43bc · outbound

This paper cites Radbert-cl: Factually-aware contrastive learning for radiology report classification.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Radbert-cl: Factually-aware contrastive learning for radiology report classification

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.607408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:4e57ca373a538d4c38c9bbd2adfeac6bc49bc4afc26ba2c9f6e1a6cadbc34d72

Observation e76e8d3a-f596-4325-90de-809355d75a96 · outbound

This paper cites End-to-End Open-Domain Question Answering with BERTserini.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models End-to-End Open-Domain Question Answering with BERTserini

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.553208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:6ae3475be26da1ea79196a08771351ba481a140217979ac4552c45b2bbd8a22e

Observation 903a2c9d-2541-4e64-862f-7847b969b636 · outbound

This paper cites Cognitive Graph for Multi-Hop Reading Comprehension at Scale.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Cognitive Graph for Multi-Hop Reading Comprehension at Scale

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.157037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:d61ef9e16689f913cfdc2f41189b526547e9e34fea9efd8d5c8d7963c6d39a50

Observation 6fb9e946-9fdd-4fd8-84e5-501129283dac · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.164706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:48b3072feba4956fd5422ce5d95edca6706afa72d184af52672d71d243e8b562

Observation d25f4ad0-005d-4fdc-88b0-cc0173297700 · outbound

This paper cites Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.172211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:a1984eb0fb62b0cc39519491d7c09cfdd35dfb0896c43c152ad21436cca5c2f2

Observation dfeee84d-d826-487a-80c6-cc94337c6bac · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.178213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:03758144cb36ad915bc685fa419be2711fc9ca7d91d11db93effca0359cb0e96

Observation 7d9be3ae-ad1f-485c-92b8-b9df6d7a6523 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.626351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:4d684f6b6e1440ad43ffe2d6286260e1924f62c66f011019f31069a671313668

Observation 16a09e60-44b7-4f4d-b57c-4a5c073adfef · outbound

This paper cites Language models are unsupervised multitask learners.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Language models are unsupervised multitask learners

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.629254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:01fb7aee0474055d764e8dd6e6e5be92854db364f9212dc9c6e2a9c911ceab54

Observation 4aab7ce4-0c33-4628-910d-f61c63394e0a · outbound

This paper cites Language Models are Few-Shot Learners.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Language Models are Few-Shot Learners

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.185268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:2e3efa27c12d1da1485d2343de973518db043ef4618cbd2296586813679b6e5b

Observation d482c76c-b353-463d-9ad6-509f7517b362 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.191596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:18e72d35cbc01155c6dde9a640782dde7f91509a48eb5d41e10bc68ef6061d51

Observation 6ffdfec1-49ee-4537-b1be-83306ce40955 · outbound

This paper cites Why {adam} beats {sgd} for attention models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Why {adam} beats {sgd} for attention models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.650074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:93ff257e53c8a335d310f381425799b811464e8feca8379c25164c83e98b626c

Observation 760fe633-d1de-41b2-9955-934cca7e5fef · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.197627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:eb9dc977376ae911935480c97498b6a7ca335934fee630c4e6bacbf1e577106e

Observation 66847ad0-6d38-4560-ae74-b30dcff183ab · outbound

This paper cites Understanding the Difficulty of Training Transformers.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Understanding the Difficulty of Training Transformers

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.203271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:cd09fbfac7ef8ea1fa30f90bafc73341312594948797d981969655da4c2cd8d1

Observation cce7173d-6953-47a0-a782-e30cc4dd70f2 · outbound

This paper cites Sequence Length is a Domain: Length-based Overfitting in Transformer Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Sequence Length is a Domain: Length-based Overfitting in Transformer Models

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.209655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:1c23b98e7f1e6d40bea009c3bc979e4dc689111ec271d75244ca836887fce619

Observation 2dbe6057-7d62-4600-8a9d-bd05107a1a9f · outbound

This paper cites MixUp Training Leads to Reduced Overfitting and Improved Calibration for the Transformer Architecture.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models MixUp Training Leads to Reduced Overfitting and Improved Calibration for the Transformer Architecture

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.214877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:4b74a9280f528b983ce8c77b68f699b40d5eb0acf8921a632f857d4061b4214d

Observation 7104600b-ecd1-47a6-b49e-23e38e36686c · outbound

This paper cites Very Deep Transformers for Neural Machine Translation.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Very Deep Transformers for Neural Machine Translation

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.219314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:5d88b60d5a37d7ff0bb2b712e6aa74617a2fb8cac8390eb2a2afeeef6d1c21c7

Observation e471c91d-0620-49ac-8792-e22d2245e466 · outbound

This paper cites Optimizing Deeper Transformers on Small Datasets.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Optimizing Deeper Transformers on Small Datasets

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.223556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:f87f22b5d9c8fca60b244a747ace5527bc5046b5e3d402f1f139409e42f9c8f9

Observation 8d5c11b2-09c0-4f78-bf7f-96fdc1b07d41 · outbound

This paper cites Gradinit: Learning to initialize neural networks for stable and efficient training.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Gradinit: Learning to initialize neural networks for stable and efficient training

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.559772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:130482d75f58b5bf85a8467046c0f04bf715f105e74cda664d8bc497a6c5cb6f

Observation 0032da27-34b3-4abb-8f56-06ff0d7d7351 · outbound

This paper cites Adaptive Gradient Methods at the Edge of Stability.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Adaptive Gradient Methods at the Edge of Stability

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.228327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:dfd1e1589b0d7aa4956a0afc9092e3c60deca5931cbe881c099d6334eee3d142

Observation 1298970c-a84f-4672-b77e-b7ccfaab61fa · outbound

This paper cites DeepNet: Scaling Transformers to 1,000 Layers.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models DeepNet: Scaling Transformers to 1,000 Layers

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.233161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:26f95213aac5c539d1922defbecb9bb7f62c90868cab6274bfc5d1904aecafd6

Observation 9dfd760a-94af-45f0-a775-31971bab2e40 · outbound

This paper cites Unified Normalization for Accelerating and Stabilizing Transformers.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Unified Normalization for Accelerating and Stabilizing Transformers

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.238049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:130a8c3fa8ab6a492576efb26a45c7db3e67c4958812e60b3d26087881c1b49c

Observation 8f15ad49-108f-414b-a021-0a9611ce5088 · outbound

This paper cites Decoupled Weight Decay Regularization.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Decoupled Weight Decay Regularization

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.242148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:881dc40c8a8af184d3494c50825cea381b6422adac1e4150bb5f9a87d886459e

Observation 79a0ca9e-7588-42ed-94d0-45a741b57356 · outbound

This paper cites Texygen: A benchmarking platform for text generation models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Texygen: A benchmarking platform for text generation models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T18:00:50.604446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:02c4dd9c60d1200161e82ea0fba8143f0042cdd87321d27a7a3aad6ccc362b81

Observation c51fca5b-dc7f-40a9-9e39-e2dbd71c9b94 · outbound

This paper cites The case for 4-bit precision: k-bit Inference Scaling Laws.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models The case for 4-bit precision: k-bit Inference Scaling Laws

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.246733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:7ca297436bda542865f18bc5bbde8fb7dbe23612b37c39412d05af9ea5d1c45c

Observation 74fdee36-718d-4b7e-8127-48c05911f74e · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Lost in the Middle: How Language Models Use Long Contexts

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.250976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:974d7e83a640ced2e4eb6b700f545acbb4019e150d048495fdb82f1214ede874

Observation 9fa8c444-9829-4561-96ed-f8386aa535ac · outbound

This paper cites Improving Length-Generalization in Transformers via Task Hinting.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Improving Length-Generalization in Transformers via Task Hinting

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.255634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:fccdcd164f26470d0baba4e4f2f25494b31e7947afdf4f1c1be1cd57467f857d

Observation c781432c-cef1-452d-82e3-29929cc02250 · outbound

This paper cites KDEformer: Accelerating Transformers via Kernel Density Estimation.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models KDEformer: Accelerating Transformers via Kernel Density Estimation

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.259821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:3799cc110f855b7ace2c9f62709bad6a3bbe8122f20150029226ab20d0e2ac60

Observation 8fcd0d75-f433-47f3-9578-adac114f591d · outbound

This paper cites Fast Attention Requires Bounded Entries.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Fast Attention Requires Bounded Entries

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.263908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:4178b4bb79e5e46153cb2711e490e1ff81adf605660ebf0f7beee4063eddde2d

Observation ba6430af-b61a-43c7-a14b-c1bd2c02d4b1 · outbound

This paper cites Superiority of softmax: Unveiling the perfor- mance edge over linear attention.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Superiority of softmax: Unveiling the perfor- mance edge over linear attention

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.268087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:2b52e9c3186eea13b31620a2196e96b05d48496139c19ac78c6f419856db5ecc

Observation c20b00c4-9a48-4ab4-958e-49065139ff25 · outbound

This paper cites Algorithm and Hardness for Dynamic Attention Maintenance in Large Language Models.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Algorithm and Hardness for Dynamic Attention Maintenance in Large Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.272467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:a65ce0b4bf01196b8d1635c54c348f53c937fb27885f37a9abe0cc57608a071d

Observation bda7a126-8e8f-4c3f-ab0f-1e5290c0f36f · outbound

This paper cites Differentially Private Attention Computation.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Differentially Private Attention Computation

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.277018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:ebfd4fcb05b27483b4bc59dd77892b4fa375413d670579b1c0355937a70e0780

Pith citing papers

Observation feec9525-6ab2-4663-a649-4715699ec7d0 · inbound

Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs cites this paper.

Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 101

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T11:11:21.460613Z digest=sha256:a1a29e801e07d19e664be5b0c55739f0ec9abbfbada8cbd74a9e03c3eb9166bd

Observation 4a3d8252-327e-4618-adde-2316e5aa7c97 · inbound

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security cites this paper.

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 267

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:57:26.303195Z digest=sha256:2a51c4a9a1c34d1d4b675152bfa9875f61ecd20f048ba9552b0ec32ce464dd92

Observation 4e85d65c-a956-4369-a51a-b622d8611508 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:3e82d12f30e4f00ca88d57c75f80f51f6995b618f68f24f30b7caa750fdb4b43

Observation 7ef0a24a-5659-407d-ab9e-8d7e56027071 · inbound

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache cites this paper.

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:53:12.253243Z digest=sha256:9c5647806d3e624fdc4a096cc7c3ff1093995091952e584c085eb8d388600dab

Observation e0d93067-594e-4c67-b3e8-1669083c43a0 · inbound

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference cites this paper.

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:06:52.199866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T22:03:10.316005Z digest=sha256:9fdd0c99ed60b995340de6c8d86d50b89ebce20bea99046307ad76c57300bfca

Observation 451a3792-123c-43ed-b4ce-a1a3e8bebde6 · inbound

EvolKV: Evolutionary KV Cache Compression for LLM Inference cites this paper.

EvolKV: Evolutionary KV Cache Compression for LLM Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:30.705096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:53:30.705096Z digest=sha256:6355855e15415e0160146d15ac089aaea7c95c9e1962ee7521841c672393bd08

Observation 9edb87ef-4250-4a67-a241-c844189226b9 · inbound

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs cites this paper.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.044993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.044993Z digest=sha256:f2a74fdc307f5b85e80a7606a4c701a690e5d1dfcb6df6c0457e4359c0715710

Observation 3fedd7ea-92bc-4172-b79e-759a2ae14825 · inbound

StreamingVLM: Real-Time Understanding for Infinite Video Streams cites this paper.

StreamingVLM: Real-Time Understanding for Infinite Video Streams H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T11:51:33.345812Z digest=sha256:3cf57fdd043f4fea2677200b7c761761b8f1b60cffbe58f87b5f847cefc0430a

Observation 460379ee-66f2-47ce-81a0-e9e4a38deb76 · inbound

SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators cites this paper.

SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T02:00:39.309134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T01:59:30.583027Z digest=sha256:0672ee05f4d3d6946b8189d4523fbbe0c08366484ef16ab3d0b7650d28a9af38

Observation dc8779f5-f211-470a-889f-0b7d1b844a05 · inbound

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention cites this paper.

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:15:19.184970Z digest=sha256:472f7eaed0092b210e59d64123f3c642b6d07985d5f816e5cb6413ca537ab88f

Observation 612e23a0-667f-4f2a-bfe7-d5a5e68a4096 · inbound

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers cites this paper.

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:17:52.344313Z digest=sha256:1452ec2c78f63cb022be3f49f8d5ca8eb1ed09ab924c750da5cda70f760c7723

Observation d9778a5d-f037-45b5-b7ed-275fd99b1dc1 · inbound

Sub-Token Routing in LoRA for Adaptation and Query-Aware KV Compression cites this paper.

Sub-Token Routing in LoRA for Adaptation and Query-Aware KV Compression H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 20

Resolution
malformed identifier
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T22:49:19.948502Z digest=sha256:937193e38ca2dff7c5c8416970ac354adcb9100ce741756bce5b28036e5701a5

Observation 6dcc36c8-7868-417d-9fa1-a832b85e323b · inbound

StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k cites this paper.

StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T18:44:42.456111Z digest=sha256:c39f4c59ea96ca53abdf3c1be70f1b096b1d349bc01ecb52197ed0ef50741182

Observation 4d9dd0f4-fb30-44b7-9600-e0670087ea54 · inbound

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving cites this paper.

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:49:24.880528Z digest=sha256:6b380de3fe9640089186e0552a1872a10b60dad9f063ac47052fe376123cec6e

Observation f81c267e-9c5d-4df2-bddb-5efb08b6efc4 · inbound

Long Context Pre-Training with Lighthouse Attention cites this paper.

Long Context Pre-Training with Lighthouse Attention H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:10:38.610613Z digest=sha256:38cb2bba5f9d02d211ae70cd1b554510813aff986f7729460d5ad26990cd8720

Observation dfad0d42-ad98-4020-aff9-51d351f176ac · inbound

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache cites this paper.

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:15:35.871863Z digest=sha256:1600819f804b36c464243d54e9ed6b8348cf75ac125d434ff39ce7181d8555a3

Observation d1d313f6-5cca-4500-9a0b-e416e8a8d1de · inbound

MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference cites this paper.

MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 33

Resolution
malformed identifier
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:27:55.991919Z digest=sha256:5398ab9d0cfc2f32cd1f024ee2c3dbd8ae428758f5a289cd64b7427284e34afe

Observation 8e979098-7055-442d-a249-2c94874684ca · inbound

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes cites this paper.

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T01:53:36.091037Z digest=sha256:7d3f2711462a50899aa9629f9b0e0295f0abe9168296fe67624e30bd7b97163e

Observation 0f4dbb63-9d6e-4509-8090-2d4d91a71451 · inbound

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes cites this paper.

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T05:05:12.267070Z digest=sha256:2a703d55275babccb5601409448032216096bc7d627285423f1e6b2eba6cf339

Observation 70027ab9-2e09-4423-9509-cbeefd78c14b · inbound

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving cites this paper.

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:57:30.104609Z digest=sha256:8fcc81908e1384ac0af19fc088d5983737234e7ce8e8beef821e1a45a3b9bda0

Observation dea92b0b-2cbf-4eae-b4a6-dfbcbf9ba75e · inbound

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving cites this paper.

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.399312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:05:44.256565Z digest=sha256:fb511e1608d25c00cbdb506d0f5d290d774ecd28754b93659025617847d9b853

Observation 8da5e788-22d0-48ef-a47e-460901d27162 · inbound

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production cites this paper.

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.699870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:17:24.147248Z digest=sha256:b9b2fa836393a7ac16b6676a2203e3a2adb25a438adce754d9c6bcce084e7403

Observation e1d1c753-1418-44e1-8c4f-16c66ae0c223 · inbound

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction cites this paper.

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:13:15.965477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:13:14.295659Z digest=sha256:f69bdd46184fb134dc14cade76675b30338db3938562e539ba4ff705fe99ae30

Observation ec502985-6103-4f4b-9e68-e96be8bde6e2 · inbound

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference cites this paper.

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:33:43.310155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T20:30:33.208386Z digest=sha256:69d3ef9c1c59b4f1476ba5c84b31c320cac462854135b8ff56210f40a2f4c94f

Observation c4e11073-0340-4e64-b1c0-b709f818a6e9 · inbound

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression cites this paper.

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:11:06.211932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T05:10:53.133346Z digest=sha256:f011b7c30fce7879455642e61b979c14f5d9eddf699e8247147faf8194b92cde

Observation 50ad4cce-033e-47b6-893d-fc662222f88f · inbound

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression cites this paper.

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.893200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:28:08.750049Z digest=sha256:a7525a86fc4a218d4290162944661cf60982cf4c5047eb1f179b100cffc29f65

Observation 1c2dbe2b-80bc-4848-8abb-ef09bc172cc9 · inbound

Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference cites this paper.

Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:31:14.075975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T07:28:35.020478Z digest=sha256:1437b5f30dac4bfddf39f88a124306c5e4da54cd3fd7527aa038b7a4bada4ce9

Observation 1603240a-9912-424e-ab77-2db49318c433 · inbound

Tensor Cache: Eviction-conditioned Associative Memory for Transformers cites this paper.

Tensor Cache: Eviction-conditioned Associative Memory for Transformers H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:50:23.556008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-25T05:49:43.160574Z digest=sha256:aa36fb77e92c23701672ccd012f1deed9c1f75473c6045e8576d334611f1b001

Observation c04370de-e76b-4b16-8202-3badc1fea4f4 · inbound

MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference cites this paper.

MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:56:15.437700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T16:07:22.196982Z digest=sha256:82efb5da144d8093cecdb0e01f97649cd32e4caf7de3763a84a57b1b80e2f0a5

Observation 79088eaa-ddf6-4dc8-9e64-dce54cd198cb · inbound

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics cites this paper.

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:36:26.502044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:51:38.455604Z digest=sha256:57cbb1093e4b87d0d70df2f9cc00f3a5277d18d6da0e5542d6aaa3c9c8f497e1

Observation 528a05d2-c0aa-4009-bd68-960d11539a7f · inbound

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control cites this paper.

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 24

Resolution
malformed identifier
local_arxiv, observed 2026-07-02T22:57:26.313384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T18:34:20.645677Z digest=sha256:5787a70597bcfac1bada2456f17c88bfa1a6534e518f283f5466b5f44d356003

Observation e178fdff-db1e-43b1-b8e4-da8a5d514253 · inbound

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs cites this paper.

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:57:30.007634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T16:57:02.709967Z digest=sha256:3e3ee3b074933060753ba69cb81a9e198947f36deadd7f67ea069746aa86bbcd

Observation e68c5b3e-86f5-43f5-b8a5-4fc28c006116 · inbound

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation cites this paper.

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:16:16.031827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T15:37:34.129339Z digest=sha256:3c41dcde1f85e2bb84e463ff671a9daa02b7347ef2592f4bd8e221dfbc565f4c

Observation 158ebb36-a992-4fba-b6c1-450109e9e29d · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 119

Resolution
verified exact
local_arxiv, observed 2026-07-04T05:09:36.855256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:a8507e066202174e98c724d9413dcdb816a24a1237209594e79e42619e090599

Observation 449ab3a9-829e-4209-ba79-2af93b3ce3ef · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:12.058991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:12.058991Z digest=sha256:2a14915f1e1d78da15544d9e60b6b2d85f87b81daed224c25f5d87ed7bf2a900

Observation f3c1369b-9407-4761-857c-b34878ea43e1 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-04T11:09:46.366239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:82f23ac886b27147952045ea0dc513e87c41e5f11f9430f2f0eb5a6b05fc2c20

Observation c2f2d99f-2b15-4d10-955b-2c12a5c545c5 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:16.272956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:16.272956Z digest=sha256:d0f5e431e9385eaa3e97000f488589355b0834c747abdcdb96acbe459ecf7632

Observation c0a77d25-2fbb-4b60-8ebd-f1a608df33a8 · inbound

NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation cites this paper.

NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T19:03:51.729138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T05:00:15.336703Z digest=sha256:0beebaa41aee370ca94205a3088e4dc3e9f1ab9c8ceff50341dca62d687f27eb

Observation caca74ed-c51c-42e8-8073-eb52580cc987 · inbound

Self-GC: Self-Governing Context for Long-Horizon LLM Agents cites this paper.

Self-GC: Self-Governing Context for Long-Horizon LLM Agents H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:56:56.663672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-02T12:49:23.785904Z digest=sha256:55246178950978709c9a0c6bb90d15135a66602347a0f2bda1d522b6fecc5da7

Observation 518fc120-1384-4891-97de-dfdf2b68554d · inbound

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache cites this paper.

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T15:57:06.413990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-02T15:55:40.177742Z digest=sha256:96a192535a5c9b8be37f34b0be6ba3da495f68d7f89d9d54e4539a521a78fd1a

Observation c635fc2e-1e08-43e7-aa12-e6e9b082e89f · inbound

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets cites this paper.

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T13:48:20.293431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T13:46:18.862925Z digest=sha256:85fb5def4b8a34c25867321ae639704cdff5422dd6c208a3482b0b8b149c5217

Observation aadfd6ac-ad9f-4737-b53e-48ee66b6a92c · inbound

Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference cites this paper.

Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T12:14:57.342551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:14:57.342551Z digest=sha256:818faff16d1bdd53598dd5d87f899f538231551b3e58982af350d442de0b249e

Observation cbabba66-681d-49da-94ca-209ab0006f30 · inbound

Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference cites this paper.

Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T19:16:27.755940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T19:13:15.765652Z digest=sha256:e0647c74d6796336a0df96c357f561d41e85ce0a6748f480c4c829ecb3e52e07

Observation 746bd06e-4681-41d0-b4cd-74fc75a0c1c4 · inbound

Uncertainty-gated selection for block-sparse attention cites this paper.

Uncertainty-gated selection for block-sparse attention H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T22:15:14.580916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T22:15:14.580916Z digest=sha256:02389b8d32ef4337d1b330d9ff2cdce2d29b191592b813e759592aa893e4fa61

Observation 81edba1d-7309-4633-8653-88a48aef7398 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 152

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.151971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:a2293cb90d7a3bffcc1627cd037f7ad72c4e9f27cce6ae79c80e07e95f5c67e3

Observation a0f4ba99-644a-4282-85a3-61dd127e2510 · inbound

COBS: Cumulant Order Block Sparse Attention cites this paper.

COBS: Cumulant Order Block Sparse Attention H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T00:42:31.008284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:42:31.008284Z digest=sha256:79917bfb6a211d094ccc1a375776a93121a3f423ebcf4e9ebca9e6c16bcd3079

Observation 0ad4a6c5-85d4-4cb3-8215-f8d4301d8a90 · inbound

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention cites this paper.

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T14:50:03.831572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:50:03.831572Z digest=sha256:4cfce051bf682ec0a388e3a2a7dd6516cedcc910036734dedef113545afcc14d

Observation 6b3217b3-2592-4f26-b2a9-99c0ed3ca0af · inbound

Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams cites this paper.

Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T11:45:09.424370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:45:09.424370Z digest=sha256:d221a1e7ed6d111def55b53b84742b52e7b6b59354bb534af17d6b14aa511789

Observation f8da6e20-ac4c-47cb-a616-0b2d1539a9c7 · inbound

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration cites this paper.

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T09:50:33.867996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:50:33.867996Z digest=sha256:f5271d744c9938496df4ba7d87021a2fe0ebd56597d407d70326fc439e9fa32d

Observation afda2078-55a6-4cc4-8611-c8a56bff268e · inbound

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing cites this paper.

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T11:37:06.681833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:37:06.681833Z digest=sha256:bb1c49617bcce51aead29097a14c1b9f4cf7983f92b60d7159783f170d3084b5

Observation ef1acaea-990d-4edb-a626-0a6250f2280b · inbound

Error Certificates for KV-Cache Eviction via Randomized Design cites this paper.

Error Certificates for KV-Cache Eviction via Randomized Design H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T07:23:41.286874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:23:41.286874Z digest=sha256:d5a327c667513fbd338623dcf63e2cdc0b785f6f07f9891b10d9da6b8b41c8f8

Observation 55408385-768d-40a8-beac-d6fba92b60d5 · inbound

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV cites this paper.

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-30T15:53:12.297457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T15:53:12.297457Z digest=sha256:6d0f3d2ade1d8dc6d847eb32c61b2783d7995a781a1011e406f58a24dfb5acd5

Observation e8c939d8-4bd1-4484-ab15-52c782c9758d · inbound

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems cites this paper.

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T19:52:42.379821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:52:42.379821Z digest=sha256:4584ad4cca2873f1901aed3052820067ca9b05bd0b3873bb0b65088b0da51eee

Observation da916699-c32f-4568-b4f2-e95af5ce6ffa · inbound

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference cites this paper.

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T09:52:42.742462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:52:42.742462Z digest=sha256:fbcadd7fa1f5eb14efd53b0730931c6d7949087d3feecaa439edfbd5f1581bff

Observation f88a5294-26f9-47bc-a198-3f1b586d5ccc · inbound

Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding cites this paper.

Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T10:55:32.054239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:55:32.054239Z digest=sha256:b64b26661cc922e0d81b73f11b549bd47ee92ac04285aa7ef23f0b015ece2456

Observation 8167e744-39cf-4f6a-ae47-76757129dbbe · inbound

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization cites this paper.

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T00:43:53.157830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:43:53.157830Z digest=sha256:29c0027b1e388cbea3d9f51fd76f5aec8f39c533a98d390d8553af293665d71c