Pith. sign in

Paper Citation Record · LEDGER

Training Transformers for KV Cache Compressibility

As of 23 July 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2605.05971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05971 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T06:01:40.766843Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-22T06:31:00.163083+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact23
  • verified fuzzy38
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cbf67cf-30da-48db-9871-2b97be1dc410 · outbound

This paper cites Can Foundation Models Help Us Achieve Perfect Secrecy?.

Training Transformers for KV Cache Compressibility Can Foundation Models Help Us Achieve Perfect Secrecy?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.987477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:186f3ffe3f9e068e3b07dc2957b1c3c0c684000a1caac13ea99c328040921398

Observation aef4b139-e3c5-4177-9b93-3923e41c6c8d · outbound

This paper cites Longbench v2: Towards deeper understanding and reason- ing on realistic long-context multitasks.

Training Transformers for KV Cache Compressibility Longbench v2: Towards deeper understanding and reason- ing on realistic long-context multitasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.294421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:dc5cde81d3dee164f91caa0e5d25814d1be719f0c6c11f0d2c0743b64e14eb9b

Observation b4d5c7da-b73d-4ec3-a51a-c821f65cb593 · outbound

This paper cites Longformer: The Long-Document Transformer.

Training Transformers for KV Cache Compressibility Longformer: The Long-Document Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.990094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:d03e112fb558f197c93ca50d54a875a09ab57bda5ad19ad4ca0ceb278a67c3c5

Observation 3493bf0e-5511-4291-a28e-3e97eaa32b7a · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.

Training Transformers for KV Cache Compressibility PIQA: Reasoning about physical commonsense in natural language

Reference 4

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.793734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:43e27496a428747dbdf9a80d41be1eb75c0714e585b9410d4b97dd874310bfeb

Observation 7be11d93-61c2-4889-b04b-0795f160ba40 · outbound

This paper cites an unresolved cited work.

Training Transformers for KV Cache Compressibility Unresolved cited work

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.289944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:f1b25317a397cad795b33181978da89bd000a8f17bdb98c1db3d6edac841e756

Observation 9936c8c6-164c-45d0-93f7-f81526078cb2 · outbound

This paper cites PyramidKV: Dynamic kv cache compression based on pyramidal information funneling.

Training Transformers for KV Cache Compressibility PyramidKV: Dynamic kv cache compression based on pyramidal information funneling

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.287922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:5be57bfaab89217c7980764fcd67d2fc7c0e37f4b090d5028408abe045c8edea

Observation 66d1b34d-c5f2-42f7-8b5e-bc7920e8e19c · outbound

This paper cites Doc-to-lora: Learning to instantly internalize contexts.

Training Transformers for KV Cache Compressibility Doc-to-lora: Learning to instantly internalize contexts

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.959138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:98642ba3f9ef67066c227abffae5fe833e7f561f76664f0f533caab014872ca0

Observation a658111a-bcc5-49ea-9152-d1f6504481de · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Training Transformers for KV Cache Compressibility Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.981495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a2e0440b5cefa511d5b96e6057413c836124c6c01ab93a8785e908f9ea1a623e

Observation 1eba9c7d-de47-40c9-8bcf-90e758f08cbd · outbound

This paper cites Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass.

Training Transformers for KV Cache Compressibility Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.953195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:5748bd1fa9733c21ab1364e8082c3849352130dfaa1c9ee3eadcbf0e28146ae7

Observation 3c77b851-8b43-40e8-9fbd-2e4dda0c619b · outbound

This paper cites Adapting language models to compress contexts.

Training Transformers for KV Cache Compressibility Adapting language models to compress contexts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.285790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:05ce09c7a11878885976839719b5bb2de8530ea212c811455110ca919365e6e7

Observation d9c843de-99c1-4fab-9c12-28c8a18157f2 · outbound

This paper cites Rethinking Attention with Performers.

Training Transformers for KV Cache Compressibility Rethinking Attention with Performers

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.943990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:7f21052809ff809822680ae4266b65ca4519f94c136c464b196b2a07d1fc7388

Observation 4db62a8c-2dd1-455a-a48d-c04e2aa53630 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Training Transformers for KV Cache Compressibility Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.967996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:29dd79913373168114a9fccb92b925e308e180ba501c1616388e1a6c8195e85b

Observation f93f265d-ef37-4145-ac87-3a3901e4034c · outbound

This paper cites InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers).

Training Transformers for KV Cache Compressibility InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)

Reference 13

Resolution
metadata mismatch
doi, observed 2026-05-13T06:02:21.797274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2c0bdc0efb6949cbcd6277ab68d6d4bc890a1608a5296c1281eb633718e54b08

Observation f84f24d0-bfe3-49f4-9a63-c1def112e4f3 · outbound

This paper cites Approximation by superpositions of a sigmoidal function.Mathematics of control, signals and systems, 2(4):303–314.

Training Transformers for KV Cache Compressibility Approximation by superpositions of a sigmoidal function.Mathematics of control, signals and systems, 2(4):303–314

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.278883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:d89e521ab82940e3fdb1c6eb8c4aa275092f792da82bbdcec1e92bb8a5589b88

Observation dae1ed25-d9e0-412f-9bfd-c30c20856816 · outbound

This paper cites The centered convex body whose marginals have the heaviest tails.

Training Transformers for KV Cache Compressibility The centered convex body whose marginals have the heaviest tails

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.946969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:301d77612631f64c9ca5ee655b00ae9ede9e2a6d50f5b43df8aac2d23c229716

Observation 9e953d73-6031-42fc-9a30-726e1da7f055 · outbound

This paper cites Cartridges: Lightweight and general-purpose long context representations via self-study.

Training Transformers for KV Cache Compressibility Cartridges: Lightweight and general-purpose long context representations via self-study

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.283639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:b927833b51574a706ef42e2f090e986c60cec6328485617c7ce82033c6daa12a

Observation b5a6d708-345a-4c78-bccf-c14a19e9d834 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Training Transformers for KV Cache Compressibility Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.997183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:730b8ac0e2dbae07f7b84f1e904f827169cefed54e3f20cb7ab69f86fadd3170

Observation fa307a39-d96e-414a-895a-048eb46a3527 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Training Transformers for KV Cache Compressibility Efficiently Modeling Long Sequences with Structured State Spaces

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.984721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:c376ee4018df8ab2ec1538aa48b7251b17e588a7db7aeda56d99bcd2503b071c

Observation af3ae607-84bb-4fd7-bb85-7ca3fc615b42 · outbound

This paper cites Lighte- val: A lightweight framework for llm evaluation.

Training Transformers for KV Cache Compressibility Lighte- val: A lightweight framework for llm evaluation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.281046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ef505d8d592590dd5be421c0b4a7f9ba35533990536ba1963b0cfc39ed40139f

Observation 90d9039f-cf5f-484a-8221-1c221834db3b · outbound

This paper cites Delta-net: Real-time network verification using atoms.

Training Transformers for KV Cache Compressibility Delta-net: Real-time network verification using atoms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.296958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:beaca33026d96fc388be406dc1227a7a36031856f7814233c921f594851f096e

Observation 7eec02fd-e121-411b-a8aa-89dbac8dd5cb · outbound

This paper cites Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257.

Training Transformers for KV Cache Compressibility Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.274052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2ac4ed44da6f6f1bfebfa5b4d27846ad67588ac286cac6c56055c0260c1b5ff7

Observation 68a00469-4c6f-486a-9f23-033449813713 · outbound

This paper cites Dynamic Chunking for End-to-End Hierarchical Sequence Modeling.

Training Transformers for KV Cache Compressibility Dynamic Chunking for End-to-End Hierarchical Sequence Modeling

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:02:21.965287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2418cd961d9f0a768821b62a5c9ace3f7c28425c20c407db9e32843a14d6642a

Observation 9a326a7c-54ae-4f48-bf4c-838d0e8a0510 · outbound

This paper cites FinanceBench: A New Benchmark for Financial Question Answering.

Training Transformers for KV Cache Compressibility FinanceBench: A New Benchmark for Financial Question Answering

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:01:13.018607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:e5677c4ce0785706543a4a7df31a0dad63ff3454bb4ec31150475d6d5d855295

Observation 452ddc86-650e-4627-9d2e-738e064716b8 · outbound

This paper cites LLMLingua: Com- pressing prompts for accelerated inference of large language models.

Training Transformers for KV Cache Compressibility LLMLingua: Com- pressing prompts for accelerated inference of large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.271575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:688a14a34ada2902a3075823adb94d20af8ef4c0880a1402a51a4fde1d96abe7

Observation 147a14f9-4813-44e6-8fc8-36fa5f80ccbd · outbound

This paper cites LongLLMLingua: Accelerating and enhancing llms in long context scenarios via prompt compression.

Training Transformers for KV Cache Compressibility LongLLMLingua: Accelerating and enhancing llms in long context scenarios via prompt compression

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.276579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:6c0d9977597b1ef7a40c10ae259517b6f70234c7870fb5021d4dd1afdec71479

Observation 7b6c23f3-f5ff-4d25-a284-f34b0ee6cc09 · outbound

This paper cites Optimal experimental designs.The Annals of Mathemat- ical Statistics, 37(4):783–815.

Training Transformers for KV Cache Compressibility Optimal experimental designs.The Annals of Mathemat- ical Statistics, 37(4):783–815

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.269453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a08d6c18ce2c937020d75b765597a09922c36bf4a10bda33dce53a4f745cf309

Observation 7677d385-0243-47cc-85d4-c0dc0da6fdd1 · outbound

This paper cites Tchebycheff systems: With applications in analysis and statistics.(No Title).

Training Transformers for KV Cache Compressibility Tchebycheff systems: With applications in analysis and statistics.(No Title)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.264839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:3be307dc060f2ac7a02d017e7d16d39a749406af452929c5e6d3767e91464c51

Observation 14256ac9-4194-45c5-81a1-37e532fa0702 · outbound

This paper cites Chebyshevian spline functions.Siam Journal on Numerical Analysis, 3(3):514–543.

Training Transformers for KV Cache Compressibility Chebyshevian spline functions.Siam Journal on Numerical Analysis, 3(3):514–543

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.267004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:25e53dfac71869445f005886d4b0c1474ce1d135043851566ab302944b8f8384

Observation dc366d46-17fc-4c89-9a35-3a056296534a · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Training Transformers for KV Cache Compressibility Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.260987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:e91acda3fe7cdf2358f1b7a2898e142523306b7c94b51202bae496d0c686cbc6

Observation 265f9aee-8bae-4477-8744-baafd32ad3aa · outbound

This paper cites Kvzip: Query-agnostic kv cache compression with context reconstruction.

Training Transformers for KV Cache Compressibility Kvzip: Query-agnostic kv cache compression with context reconstruction

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.971227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:87f317bb2c2f202219e7366515e2e4faa7c5e95ca424589397eeb103ecbfbc9b

Observation 0c3ab234-0bc3-47c1-a830-35e95aba6680 · outbound

This paper cites Lexico: Extreme KV cache compression via sparse coding over universal dictionaries.

Training Transformers for KV Cache Compressibility Lexico: Extreme KV cache compression via sparse coding over universal dictionaries

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.254451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:74dbd61e789f2c4bb4cab847d07528f1b541709ea8527a6336fb0f5f7e416876

Observation d9c0c9a0-9b89-4abd-a86b-dc07dc7fb28c · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive NLP tasks.

Training Transformers for KV Cache Compressibility Retrieval-augmented generation for knowledge-intensive NLP tasks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.256522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a065e5e8f3e38af3c80288db68371263e18b859dc5a0f30057f446d1dbf2aa7e

Observation b36181d0-9a4f-491d-9dbb-4544c41064c2 · outbound

This paper cites Compressing context to enhance inference efficiency of large language models.

Training Transformers for KV Cache Compressibility Compressing context to enhance inference efficiency of large language models

Reference 33

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T06:02:24.258780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ae44df06720e6c1444588c4e27119e069064da99d19cbc6e5f2e4399f44a506f

Observation 0d344f97-1b95-470f-8156-b066ee485948 · outbound

This paper cites SnapKV: Llm knows what you are looking for before generation.

Training Transformers for KV Cache Compressibility SnapKV: Llm knows what you are looking for before generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.262982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:92d7a9a9541e5ca1bdffd457432dcda53182d35dc56977b5b2bcc18fccfedc73

Observation ced269bf-c1b2-4443-959a-8fa0189b7ab3 · outbound

This paper cites LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence.

Training Transformers for KV Cache Compressibility LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.993912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:3c1268b570dfdb44d2b5f677a33e8b81e47dd2922c8f184f5f9b172157c72228

Observation 89401e05-302f-4574-9d6f-221366601686 · outbound

This paper cites SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass.

Training Transformers for KV Cache Compressibility SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T03:03:41.308564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:3ac4fa4cf0b455740c24df0d8acae5f14825569f6997b09f34156c4f844231ae

Observation 628d164d-0462-476f-9b91-c95e952b6feb · outbound

This paper cites Pointer sentinel mixture models.

Training Transformers for KV Cache Compressibility Pointer sentinel mixture models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.292079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:6c796133682007843ae88d305a6f21f1d3501ecb9aac1f73f5100ff7d37623e3

Observation 8976502a-1931-45cc-a261-2796fd872c02 · outbound

This paper cites an unresolved cited work.

Training Transformers for KV Cache Compressibility Unresolved cited work

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.250037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a4956434dd4c3b33540b7e3c89eed25eb5482a1225a38d6aa46b6bbf075213f5

Observation bd3951d8-cdd7-41c6-9b01-43f233d1d814 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Training Transformers for KV Cache Compressibility Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 39

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.790714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:fc94477b2b92affa569e2aa80ef572aaa81b09ecbe35e6d074c2687ad000dea1

Observation e313e9fa-7d48-4e25-b6fe-6b3a50bef301 · outbound

This paper cites Learning to compress prompts with gist tokens.

Training Transformers for KV Cache Compressibility Learning to compress prompts with gist tokens

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.235100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:97d476d8e6082ef2d20d7315ac30df9faa5811a459695dd673249f896e31eb5b

Observation e5718b8d-1ccb-478e-8d9b-d9a69704d022 · outbound

This paper cites Using an llm to help with code understanding.

Training Transformers for KV Cache Compressibility Using an llm to help with code understanding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.220818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ecf3d6fe94e8e100b62a18f6ce8def1c986c581ada6661a222bd7d18a1e6ee22

Observation c076a8b2-29b6-4f6a-a674-b13ac4d9bd53 · outbound

This paper cites Context Engineering - Short-Term Memory Management with Sessions from OpenAI Agents SDK, September 2025.

Training Transformers for KV Cache Compressibility Context Engineering - Short-Term Memory Management with Sessions from OpenAI Agents SDK, September 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.223236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2a4e7ae72e2fd56a796001914932458f327307d99e87d37476510d34ed937a75

Observation 436b9803-9b50-47e7-8d5a-f49bbe585cbc · outbound

This paper cites Transformers are multi-state RNNs.

Training Transformers for KV Cache Compressibility Transformers are multi-state RNNs

Reference 43

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.803106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:802cccbc18c2bd0fc82196375f1e2239b5091a73bb382f837ba588f413e77c80

Observation 393bdcbc-4be4-4d49-99d4-95bf24436637 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849.

Training Transformers for KV Cache Compressibility The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.227937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:624bbf8c1f791f303c209bf72bdbbde23de21d5d2fdbd9090b40a9b604f3df65

Observation 4c76dc23-6216-4dfa-b833-1a6b6a2f153f · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

Training Transformers for KV Cache Compressibility Compressive Transformers for Long-Range Sequence Modelling

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:46:17.004740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:24c89830b21414648a5ee86c3b387b306d62b8fd9756f80b07995a17fc721279

Observation 40abb738-a348-46b4-aa94-045157a1adbb · outbound

This paper cites Effective context engi- neering for ai agents, September 2025.

Training Transformers for KV Cache Compressibility Effective context engi- neering for ai agents, September 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.218665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:7691306d92c557a9450e3253bff0dff429346754407238e873b2dee3b7f4a6e8

Observation 59cfac00-a8e9-4f9b-a166-4993e0673cf3 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , author=.

Training Transformers for KV Cache Compressibility Proceedings of the AAAI Conference on Artificial Intelligence , author=

Reference 47

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.786005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:24f30eb8e863dd01dcc562a3c8f249d7091e80785427612937bca261cf86d123

Observation 1d3aa75c-7184-4ffb-8779-6622470094bd · outbound

This paper cites Representational strengths and limitations of transformers.Advances in Neural Information Processing Systems, 36:36677–36707.

Training Transformers for KV Cache Compressibility Representational strengths and limitations of transformers.Advances in Neural Information Processing Systems, 36:36677–36707

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.213905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:e43f2f881ca34adbc1a6c9499907a4e3a27274ba0021d80d82190d9e99c7d080

Observation fe83771e-4d4a-471c-89e7-526400e0d2da · outbound

This paper cites Social IQa: Commonsense reasoning about social interactions.

Training Transformers for KV Cache Compressibility Social IQa: Commonsense reasoning about social interactions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.225530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:4183209bf58c7e55de791ca0645efac7763f935a3e52b081c50adb24ac9b02f2

Observation 3f357129-8531-4ade-8248-1d2c178f607b · outbound

This paper cites Social IQ a: Commonsense reasoning about social interactions.

Training Transformers for KV Cache Compressibility Social IQ a: Commonsense reasoning about social interactions

Reference 50

Resolution
metadata mismatch
doi, observed 2026-05-13T06:02:21.800181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:74611d8c2f26074f187ad8939c7caf2f7cc0ddba007305edf5c389e03523a9ef

Observation ce7ad80b-0f42-46bf-8ae6-71a8d92048bd · outbound

This paper cites QUEST: Query-aware sparsity for efficient long-context LLM inference.

Training Transformers for KV Cache Compressibility QUEST: Query-aware sparsity for efficient long-context LLM inference

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.252230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:3c68d90baa24186d2df778c6aad313a751e78ae8964cf26e97ddf82295646cff

Observation 3080a2c0-cb92-4ffd-81ed-a3fcae3bda5f · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Training Transformers for KV Cache Compressibility Qwen2.5: A party of foundation models, September 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.204011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ab61f0541f90b7edcd16cf5bc43fef13a6880f1eebd34410713ce2f6ca3c4298

Observation 8a2a4e3b-5597-493d-bd83-ac4725c3f36a · outbound

This paper cites Efficient streaming language models with attention sinks.

Training Transformers for KV Cache Compressibility Efficient streaming language models with attention sinks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.208978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:723e00b73d4b9f3c02c8a20169ce91e9acf8ac165730da7e91f5233d8cf47c8e

Observation 0bd4eb5c-44b2-47f3-b70f-859d0874db08 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Training Transformers for KV Cache Compressibility DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:49:16.947404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:6cfef63ae29aa5fcaf533c9def2d936d376dbab2aff41254c26ac73c67d27e36

Observation 6113bfcc-ac39-448c-a9d4-21b8a1d1bba4 · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

Training Transformers for KV Cache Compressibility Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:50:25.025022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:1e24cb0dbd534e394a79d5353ba24c2e1ec1d203e09adfc4f08fce18b9005a13

Observation dc536cf1-3d8d-4043-b146-f4033d7e12e7 · outbound

This paper cites arXiv preprint arXiv:2407.15160 , year =.

Training Transformers for KV Cache Compressibility arXiv preprint arXiv:2407.15160 , year =

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:22.000080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:e87608df76f52269b39f0a96d8ac313aadea31be2bf932de4df102d86fb71091

Observation 2bb92f07-362e-4e28-b0da-f1f7bf28eabe · outbound

This paper cites Deep sets.Advances in neural information processing systems, 30.

Training Transformers for KV Cache Compressibility Deep sets.Advances in neural information processing systems, 30

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.237288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:7619fd9346677b9088a1a30f776f8813b556888e0a9bc78891606787dcab0f6c

Observation 5cffa41d-2c43-4d58-931e-df3ed14a0102 · outbound

This paper cites Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297.

Training Transformers for KV Cache Compressibility Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.230077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:108bee6153ffdb82ac1c583f2fb962904774737d3f3a107cddf501e646ba17ad

Observation 6557247c-93a1-4764-9f98-24dfbadeb32a · outbound

This paper cites URL https:// doi.org/10.18653/v1/p19-1472.

Training Transformers for KV Cache Compressibility URL https:// doi.org/10.18653/v1/p19-1472

Reference 59

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.806038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2c4623cff5e236119648517a014da7e6376461fcf434bd91e616c32b399d1b87

Observation 78c43fa5-9a10-49a3-84fc-d3d4cf794a20 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

Training Transformers for KV Cache Compressibility H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.216232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:b0b3b3d128832bfca13e806a0a987230b625bb2ddfb838cf0ee0b09256e3d5d1

Observation 7444f49f-cb4e-4adb-83e3-5f870f1a0af6 · outbound

This paper cites Lifelong learning of large language model based agents: A roadmap.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Training Transformers for KV Cache Compressibility Lifelong learning of large language model based agents: A roadmap.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.232644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:47c3e3066659a4bc037ee6f65dabc76a95901234de17df39c0d04c14dbf7bcfe

Observation 7b4e5337-878f-4ec0-a002-6601849bdb1c · outbound

This paper cites Fast kv compaction via attention matching.

Training Transformers for KV Cache Compressibility Fast kv compaction via attention matching

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.244984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:3131c56d1c3bb65117811d9399ead5734471cdaadb7339f46daa2d29cc398ce2

Observation c6758ec7-0854-4dcb-8ee1-936b6165bf6a · outbound

This paper cites an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(70).

Training Transformers for KV Cache Compressibility an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(70)

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.247803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ba5cf7c1cc00f7d584da23c8b2396d2f646c30bbd44a55e21a92af40016f7d17

Observation 65356efd-af18-492c-b640-be94d32de4c7 · outbound

This paper cites an unresolved cited work.

Training Transformers for KV Cache Compressibility Unresolved cited work

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.206400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:f83a61d85be7c696e2f5c75f3642ebdad6c83e7888f39680b6113791ea7a5019

Observation 7ac604d7-f5e5-41b1-a5f0-4b3f2221255f · outbound

This paper cites an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(97).

Training Transformers for KV Cache Compressibility an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(97)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.211502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:bbe9a4f4464697e6c8a55a93f23fdc1ca50f361fc1928278c761c5e9437cd0b4

Observation a0c110ff-ee9b-45a9-8cad-83a1c8769821 · outbound

This paper cites By the same argument as in the proof of Lemma A.7, there exist functions ϕ:R d0 →R d1 andρ:R d1 →R dout such that for everya= (a 1.

Training Transformers for KV Cache Compressibility By the same argument as in the proof of Lemma A.7, there exist functions ϕ:R d0 →R d1 andρ:R d1 →R dout such that for everya= (a 1

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.950355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:494277df73c944c26fdd207fcfdab01af919cf583e0d1a470c1591c2e2e41352

Pith citing papers

No inbound Pith citation observations are available.