Pith. sign in

Paper Citation Record · LEDGER

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

As of 22 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 4 inbound Pith citation observations for arXiv:2502.08910.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08910 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:19:37.955444Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:11:17.706860Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T05:53:04.753054Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5da89157-b546-42a5-9486-966ef03a9ef9 · outbound

This paper cites write newline.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.781866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.781866Z digest=sha256:36b82aa8cc1094e979fd2f69e8ba2ce8af704118484991aa353ca28412239654

Observation 66c36de9-0f72-4e09-95b3-1422237f243f · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.788191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.788191Z digest=sha256:e383296573d7350a75dcca364a1003e43c4935b891c21be200e7f4628cbc2a67

Observation 1376c85a-8125-4214-8b04-6808982abe62 · outbound

This paper cites NTK - Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation., June 2023.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU NTK - Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation., June 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:19:39.112315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.794159Z digest=sha256:bb845a3a5264546f2dd780587740202dd8b1f5e5d758e6773836b0d24617b973

Observation 1bb8716c-3377-4043-beac-34b16c36cf78 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.800257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.800257Z digest=sha256:edb9ca1be8afb812324c547bff2cf76622b4580f1ebaa9a96abed4b432cb1fa6

Observation 26a7e9fd-9fcf-4035-8590-4d2117e44b35 · outbound

This paper cites Flash-decoding for long-context inference, 2023.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Flash-decoding for long-context inference, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:19:39.063296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.805815Z digest=sha256:f2ab7fafeff8a21b5e78c1a5acc8df188611fdbfc0f6238fc7fbbebbe61678ae

Observation 3545b50a-5b7b-4782-8228-be7689c494c1 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.811005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.811005Z digest=sha256:ccf827acae26df516557cc0e512de0ce0aadc9ed2b1220933c0a682fbe165da1

Observation 56f4ec7d-d063-45b8-b78d-a5cdf8372526 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.816306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.816306Z digest=sha256:8760edd8543b987b65278465a320263ec94ca07509783c0d9e73603c465bcbda

Observation f1e4bc0d-81d0-4354-bc7d-ea55a0ebc2c7 · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.822448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.822448Z digest=sha256:0a3381cc5beea08779b7ee3fdcffe6e13f3906b96e2cebd79e84af22a2411a28

Observation cf36afdf-078a-4509-a810-d3dd43910de1 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Gemma 2: Improving Open Language Models at a Practical Size

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.827717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.827717Z digest=sha256:94b56c5e19e40118145a4e1226aee291eda6a3f16352e97f240b5373fcefc7fe

Observation 42f8bdbd-f073-407f-b83e-93d7f6cc8a34 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.832723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.832723Z digest=sha256:8dc2f5ceadbcb2f69c9fe7153f0af8606c93b1db9434206c01f96eb912adb5a1

Observation 238d4a53-ec19-4a44-95f3-f91b2843baa0 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.838100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.838100Z digest=sha256:f5673e8160debc570fea5d3c66d0f920fa7d9fa4b777736d850ba25b4ef9bf68

Observation 1c5451d0-0d44-40d1-9102-6d299b70deb3 · outbound

This paper cites Mistral 7B.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Mistral 7B

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.843510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.843510Z digest=sha256:5cfa158d3489294d651b35fa2b54939c7e78e61bd7a4c5e0779a7725b5418f71

Observation 6ed5391c-3d37-4a5c-be11-bce0d9996892 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.848294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.848294Z digest=sha256:5a2747601e8860379485f1af521fea4f3ed55b37ca9ae4a4cc1865c9446ee297

Observation 5e5821af-de8f-4a27-ba7b-f7971d4cedff · outbound

This paper cites LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.854054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.854054Z digest=sha256:581144ebc699fb76022cd831d64d64b98f5423b90b6e61cdadc07c0f445e30b6

Observation 9d3eec75-e510-4ba7-8f31-ff15cca3b029 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.859198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.859198Z digest=sha256:4d6bafbbdf75e1c8c0f171789643bf912396f27795f17092f4379e732e9a0545

Observation 6d086886-eac4-406b-a35b-8fc72910edbb · outbound

This paper cites an unresolved cited work.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:19:38.997502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.864013Z digest=sha256:9949124e3f31d138d01ca608ebf0ed72e238d7724412b2ad7d28bf0a9fa03a82

Observation 383c3a34-8b05-4958-ac3b-0fea11b3874e · outbound

This paper cites an unresolved cited work.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:19:38.894823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.872254Z digest=sha256:7ad672eb679c9c23a06df379f191387c3e509428c27710f2922b7ed46dd8ecd8

Observation 1abf58e2-42c7-4a8f-906b-62e7cc0517d9 · outbound

This paper cites A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:19:38.625454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.877479Z digest=sha256:06c52a611f83a577171bea26a9afb7a5c03fb71f7af7190881f530684a7605c5

Observation 407523f4-dad4-40ea-be25-5fa859bf4d05 · outbound

This paper cites EXAONE 3.0 7.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU EXAONE 3.0 7

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.883272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.883272Z digest=sha256:35de84f501e4426c650faf0842fe14bfb9df6129a83e43d6e23b9d1f15727b8c

Observation 57decbb3-1e36-47ba-9f02-3d67d5c653fa · outbound

This paper cites EXAONE 3.5: Series of Large Language Models for Real -world Use Cases , December 2024 b.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU EXAONE 3.5: Series of Large Language Models for Real -world Use Cases , December 2024 b

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.888459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.888459Z digest=sha256:2561d7a418bbb705dee0c80954afd473b15a56d8c7edd8d87c61f5b1789b7b64

Observation c6b7e9af-e392-4be5-b046-466c5b80d87f · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU SnapKV: LLM Knows What You are Looking for Before Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.892901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.892901Z digest=sha256:c60af2f484b3f0628aec99df0175fdd913690fa9a455ed1a94f064d3a11f70b0

Observation 6c57521d-a509-44c7-ab93-8bc7d2dc9658 · outbound

This paper cites The Llama 3 Herd of Models.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.897728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.897728Z digest=sha256:48fd1d062340a22b6995b48a9c2564bd0b99c53f9a994f3365146802e3992e1a

Observation 92e6b6c2-8183-4301-8eab-e3804da67195 · outbound

This paper cites Transformers are Multi-State RNNs.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Transformers are Multi-State RNNs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.902581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.902581Z digest=sha256:5b29173db19aae83431465b4059a628827a394c609f94f52b42c6477513fff43

Observation b388a196-ed47-4fe7-9652-fd5cc6970cb2 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Code Llama: Open Foundation Models for Code

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.908106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.908106Z digest=sha256:3c3c1fe1d263527860d979b218e2e4b024871502e09308f9597ba8bcd572ca6b

Observation c08b7145-6752-4f1f-8046-99752f72a97f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.913396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.913396Z digest=sha256:ddeeb4eb7ef19e5e47a52cd52a202b5a6cc7f753273617cd68b3445dae7a1fa4

Observation 22c1a1c4-8617-42aa-bfba-9be1fd2a9751 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.918446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.918446Z digest=sha256:747a3e599d1a6c78d20b82a59ee42af4509d594d83fb63453b283868f72b8611

Observation d63b3f5e-75b9-4dbc-9aeb-73b44ba9cc58 · outbound

This paper cites an unresolved cited work.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:19:38.834399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.924374Z digest=sha256:2a54f92efc5f7b604c08dae34ed67f1cb428eef8305912a320376f553f5f46cb

Observation 87ec3dcb-ead0-453b-8e13-e762b07b64cf · outbound

This paper cites Attention Is All You Need.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Attention Is All You Need

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.928806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.928806Z digest=sha256:2ad711c3bb5ca436ff8ddfd33a6a0261c2b23ca23a41a46e88e18d5997c9e38b

Observation 48941070-f497-4cd2-986e-3dffe4170ca6 · outbound

This paper cites Training-Free Exponential Context Extension via Cascading KV Cache.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Training-Free Exponential Context Extension via Cascading KV Cache

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:19:38.069832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.933737Z digest=sha256:2110ec8b6b12f038d112b3c5be4c8a1be4f5edc0d48489f6f2d9d93c2d11a585

Observation 32ef12c0-7a4f-4fc5-b005-9a28ffecec08 · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.938758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.938758Z digest=sha256:32b1e92d2217dee4428c5e9422300292423637a055e68f0ed2e428a2b1f8697e

Observation d81550cf-d9b9-4065-82bd-9852e424cfb3 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Efficient Streaming Language Models with Attention Sinks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.944290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.944290Z digest=sha256:5ff86937d725a040aa69bd39c5794b71d0d65ad84c653cb0348f1549e6cfd353

Observation 7df593e2-e375-404e-be22-b30d877a80c0 · outbound

This paper cites $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.949529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.949529Z digest=sha256:bc04aeed7b380e88460b2a02e40e8863515f8b7689beef7b19045102c6215677

Observation 7d1d90c1-3251-4737-b299-aaed6e22a2c2 · outbound

This paper cites H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.955444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.955444Z digest=sha256:cad0a9aa13fa93da6469385259096d1b1e5445f3e28020d6468a5b2eafcd0c1a

Pith citing papers

Observation e97223e2-fc57-465d-8165-a03c84f75dbd · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.706860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.706860Z digest=sha256:d24480cf6106795421f0205d727d3868055f2f5ecc9b183db2d817c42d539df3

Observation ce79bcc8-386d-4afd-af53-3b4a5f0ec7d0 · inbound

FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression cites this paper.

FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:00:43.304573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:00:43.304573Z digest=sha256:cb97ab565e0a420fb9676c371192242ea04359a667a841ffb083d7123e519b1a

Observation be78e69b-d60a-44a1-b951-354d1b35ac41 · inbound

Controllably Efficient Language Models cites this paper.

Controllably Efficient Language Models InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:57.667935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:57.667935Z digest=sha256:b1d0b1994efd83990369ff0b21ad319966e49907e062e1592f53aed8922565b1

Observation 3836c4a1-58cd-44df-a149-fa7a038473eb · inbound

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents cites this paper.

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.754570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T05:50:14.266278Z digest=sha256:a0874d88004ec3c9c9eb36c6bbc25eae5d35d0a8fece9022a4fe86510edfee17