Pith. sign in

Paper Citation Record · LEDGER

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

As of 21 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2607.06827.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06827 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T20:30:11.127007Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact16
  • verified fuzzy34
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf29194d-1cfa-489f-9ed6-5e7190f31635 · outbound

This paper cites Instruction Tuning with GPT-4.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Instruction Tuning with GPT-4

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.449922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:f4bf3be0ffcea20f3a22952507458310e5aaa1b29664fa9bd237ef6418fe694a

Observation eea683b8-7027-4e02-bdaf-74fad92d9296 · outbound

This paper cites The Llama 3 Herd of Models.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs The Llama 3 Herd of Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.423879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:d46e84e796120a4293a72a2be8ef24890813c8ba43746563b0dbfa74a948bef8

Observation c43ec401-fa32-4f5c-a5b2-44a87e99bc8d · outbound

This paper cites Qwen3 Technical Report.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Qwen3 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.409134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:bdb89421786713eeea84a5715d1ed061b40c3fcabdfaf75a5d1bacf7728a6928

Observation 0d59d235-b32b-4ae5-8cbf-67d63b774509 · outbound

This paper cites Phi-4 Technical Report.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Phi-4 Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.420712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:2cc8ee01df6955f1d855e1677d81c9d9070fd9c13e7db2ad98883ef200531170

Observation bd8f82fb-f82f-4345-8fae-52185feaca23 · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs SALMONN: Towards generic hearing abilities for large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.713988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:a8cd5e0ad9e8ffa7730c200a11d9ef6cff6b47ccee195ba2f6db67e76557f65f

Observation 0eb31cc9-519b-4462-a17f-166f730a67b2 · outbound

This paper cites Qwen2.5 Technical Report.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Qwen2.5 Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.438595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:9f1a8e6231da8609db9f6c9bc2bcb169391202cbca6de647552bdd16fc7f80b9

Observation c9693dc2-8f6b-4e79-90e4-a082da1e3011 · outbound

This paper cites Alignformer: Modality matching can achieve better zero-shot instruction-following speech-llm,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Alignformer: Modality matching can achieve better zero-shot instruction-following speech-llm,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.711851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:c547fd51ed2265dfb4b882aa7f0d1483503d0d9e7c7f7635e3e67c94235ec826

Observation 3dea55de-3b2e-4df7-8ec7-13b8c4fc7a5e · outbound

This paper cites On The Landscape of Spoken Language Models: A Comprehensive Survey.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs On The Landscape of Spoken Language Models: A Comprehensive Survey

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.411923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:86f7c39842f88030d3f3658b82191cc8c4edcc45c9b45b29754e58e9f823c1e1

Observation eab642b6-1e4e-4f1b-9f97-bc41e9d4ecab · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.433275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:a79aa21b1a80e3f233af03cee842056c08a9ad38f855f682411b3cdf73f61072

Observation f01a2625-132d-4e4e-aa75-f02d060a22cf · outbound

This paper cites DeSTA: Enhancing speech language models through descriptive speech-text alignment,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs DeSTA: Enhancing speech language models through descriptive speech-text alignment,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.718414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:92c817ff69647539d5d44e2f465c4ae147ff8806be78d42bbadfb8b69714e89b

Observation 524c9804-a579-4ad3-95df-7c8496bf23dc · outbound

This paper cites Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.705608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:2b29c548cbb7c88e33bf02db3656a98c5ad28789022ee24804e787d92c66f523

Observation 39cc6a20-8cbc-4dab-9d92-a69221744adf · outbound

This paper cites Desta2.5- audio: Toward general-purpose large audio language model with self- generated cross-modal alignment,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Desta2.5- audio: Toward general-purpose large audio language model with self- generated cross-modal alignment,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.691218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:d273d176cab1d379a8d679f42cfa486a48a86c2d80b38217baea51b612ce9f4a

Observation 6db8e2bd-cf83-472c-9e4b-ace4cdc07a2b · outbound

This paper cites Prompting large language models with speech recognition abilities,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Prompting large language models with speech recognition abilities,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.701557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:fdff585562dfe59bd18af47a65e97923e4ef16c357fcfd157e06baa9a523e2d3

Observation 3fe8e6f2-f109-49ef-bf16-6fcddfd29a9d · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T20:37:34.398583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:71d0cbde9d911f2227706b11dd3e6be433bac31ef07e94eb9cccdb2c262bb9e6

Observation 3f7e0a52-5d58-43fa-926d-e60fa96128a2 · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.441426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:dd00313540e5e8a4f70bdcfdadc8763f7427eca8bc32e8c16d201e959b52f2b8

Observation 083a14fa-dd81-48bd-b888-c00565990f3a · outbound

This paper cites Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.699667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:cab1dec277c9fc5c079f48dfbaa80cdc7482493dc2fce38e6864aa8c0378fef6

Observation 2e081245-57f6-4186-a128-ad1f9b815477 · outbound

This paper cites Connecting speech encoder and large language model for asr,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Connecting speech encoder and large language model for asr,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.709409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:a975bbb86e5595191aa53547f958050b2b1c4f1b7958ad4b679cffadb1dc18ea

Observation 3354767c-0d74-4cf4-9ad5-3cdc80b03b4f · outbound

This paper cites Speech-xl: Towards long-form speech understanding in large speech language models,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Speech-xl: Towards long-form speech understanding in large speech language models,

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-10T20:37:34.402206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:23f88070ceeb1dc6d4e7f38796c2e172e867609dd111421cb825d4f9374e4b69

Observation 93a7ca83-4585-4a38-be2e-f3f596bf8444 · outbound

This paper cites Speech llms are contextual reasoning transcribers,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Speech llms are contextual reasoning transcribers,

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-10T20:37:34.427244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:15bd209760af1c1139103685699dca0411854027d4deaedc91b29d1ce5c1b838

Observation 7327d8b1-cd30-4a3f-a60d-6e21164a307c · outbound

This paper cites On decoder-only architecture for speech-to-text and large language model integration,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs On decoder-only architecture for speech-to-text and large language model integration,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.729136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:b34146b3a7ec4d987153ce9dc10ce97182fdc606befc11dde38d55719db921c5

Observation 6ac338bb-48c3-4918-b3fb-a060126c5936 · outbound

This paper cites Speechmap- per: Speech-to-text embedding projector for llms,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Speechmap- per: Speech-to-text embedding projector for llms,

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-10T20:37:34.417901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:4d797a799c3e8a663147d5e6613a8831e1ac11729b42907b5f17b657475435a9

Observation b2da3027-61ab-4387-bd63-21bfc808a640 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Snapkv: Llm knows what you are looking for before generation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.752754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:abefa112ded74d692c10075696d29fbeb7d95cb525ae927dc6661f3828b3b091

Observation 56f0e0a1-1c5a-4778-8880-38c2ef289fe3 · outbound

This paper cites Minicache: Kv cache compression in depth dimension for large language models,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Minicache: Kv cache compression in depth dimension for large language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.744945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:c2914a0c1d1ebec77ed8230783c678fb0fb432123b338c87ef0ca2afa35a259e

Observation 705beebf-fb9f-4fa9-9e07-f513173d625a · outbound

This paper cites Open asr leaderboard: Towards reproducible and transparent multilingual and long-form speech recognition evaluation,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Open asr leaderboard: Towards reproducible and transparent multilingual and long-form speech recognition evaluation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.754734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:f679bef340d7fa7b2cf2cf89856bb0d8a62376238e7d6e8bdddca4120abde689

Observation fb5ffe7d-8af3-4bed-96ea-283851fa926a · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Efficient memory management for large language model serving with pagedattention

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.742944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:bf27e6184438ac2e55cf414f72d5bd88304b577f9f29989ed843aa612c624d19

Observation 6151166a-3fd3-4b0b-bf07-b82ebaa178a9 · outbound

This paper cites Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.693481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:9552bed9d1f07614cf166ec711339a25aa312afe474e2f4dba99db883b3b5c81

Observation e7e404c7-3a2f-4810-9a5c-6fe5a7be0dbe · outbound

This paper cites Cjst: Ctc compressor based joint speech and text training for decoder-only asr,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Cjst: Ctc compressor based joint speech and text training for decoder-only asr,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.737191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:60ea5c60f512dbec47f41e26e2159208fce3e359e9b4a5c586f9880317bc1961

Observation 9ed1227e-f456-4693-b887-5afe4eafbb15 · outbound

This paper cites How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.414645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:09c1d5cb4f0f9b6c6a64b5e821998eebad4caf91100cfe7b870e3f4d7ce5aab4

Observation 6f6605fe-fcbd-4554-836d-5b64e9564ff1 · outbound

This paper cites Cif: Continuous integrate-and-fire for end-to- end speech recognition,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Cif: Continuous integrate-and-fire for end-to- end speech recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.695666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:6d3489b7041eeab0c480671c791d112e2b142db87212839e1d7271529e18dc35

Observation aba02e2f-c556-4b4f-8ffc-e855301556bc · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.741068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:cd11614bfe57b8d0f42c46aa9b3b0ce15c40d2c212a0ac1645fcff244989fed5

Observation 14c54e88-8246-4c65-b344-c9b593196b10 · outbound

This paper cites Distilling an End-to-End Voice Assistant Without Instruction Training Data.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.405190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:ea5909b2635271a36a01f655f9c2735de09e05a72e6a502b0ab82d46b8d766d7

Observation 3c3ff200-163d-464e-aa09-ffc4b311661c · outbound

This paper cites A survey on large language model acceleration based on KV cache management,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs A survey on large language model acceleration based on KV cache management,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.748991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:eb12ed8627915ddb091c56a5307f7a42d49d4688b79162ff75994da0ffa621ca

Observation c2e7dfad-44dd-419f-bb26-d97f8db5b63c · outbound

This paper cites H2o: Heavy- hitter oracle for efficient generative inference of large language models,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs H2o: Heavy- hitter oracle for efficient generative inference of large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.747054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:30169f32f6a6087ada2ab6897299800790174ac8c50fa12895ff03b2cf04cbe4

Observation 4113fdae-1bc2-41fb-aa09-0c8b85b1857e · outbound

This paper cites Efficient streaming language models with attention sinks,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Efficient streaming language models with attention sinks,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.724774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:d183d0dffd2160e56ebb4adc5abc6f5b884ca49163f9c4cdbd7e882bd21391d9

Observation ee2be897-38d7-4d83-bcc2-c0a7c1c5875b · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Kvquant: Towards 10 million context length llm inference with kv cache quantization,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.726897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:d44c9ad1765b4fa21587fbf76e42cf4090e7ed5b6e3871b1314068ca97e13dd5

Observation 3602818e-f029-493f-a7e6-d2610c0c9e6b · outbound

This paper cites Dynamic memory compression: Retrofitting LLMs for accelerated inference,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Dynamic memory compression: Retrofitting LLMs for accelerated inference,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.722786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:23b7a5986c6b0f8aa6574cf768d0e7d8cf1eb23dbc6a448d4ffc1ad8ff67f182

Observation 0dcc80aa-866d-4b6f-8816-53c67589c206 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.430154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:d523c4ddaf4e2d65d92ac99f6e34deebbe401d0426ce2719484841201601f004

Observation dc8867a3-fc92-451b-a903-ce4cb85dd9d9 · outbound

This paper cites Adapting language models to compress contexts,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Adapting language models to compress contexts,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.720595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:d115f07f425498315d340660a0620cc5251304ea0f5bf1d515e91995cbb85957

Observation b828db19-f4ab-4881-b9fc-778cb926758f · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.436040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:966646cbb77cd6ee040c450bdc392db3a7986758def1bef25cdb67757459ca99

Observation 6969e927-329e-4f93-9d8b-f5e11c135e37 · outbound

This paper cites GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.731112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:a9e27c724fe058c5e2d718c312ef1b55521ff97423a183fb0d4b19916ba2e9f8

Observation a262608f-fd45-4870-b55b-a47a2fde477a · outbound

This paper cites SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T20:37:34.447307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:bd1da2b263fb729a1cdfc57fbba68f8730ae036d9f05aaf821a95b7a344da33f

Observation e37a5331-43c5-4246-802e-547170de6731 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Librispeech: an asr corpus based on public domain audio books

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.703497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:da8cdd475d669f1ad0e2360b502cbf233aa3865a5d5696b5a81d97c28df74043

Observation c15ce9cb-2b5e-4118-a982-ba21284cf367 · outbound

This paper cites Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.707462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:39b971341f2f2b9152e62da30df68d1547b0d5e9f4ad8e2e460e9798a5dd3c61

Observation 399d939d-6696-48df-8dd6-5391eb85e01d · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Common voice: A massively-multilingual speech corpus,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.716163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:7eb3c4633428176f27a9be9ff998dfff6097773e8713e47a96414d0311994048

Observation c18b0848-e607-4b93-8dfb-38506e54df6a · outbound

This paper cites V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.735106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:2e109f8f6eef065a342c70e850c51de337e306872b820854cb456ca17234450d

Observation ec50311e-9717-4443-b3d6-a8e24d5cce5f · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.697696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:b3d4e097e4dbec40f281f13a02a85dc8f9fc0cf1d465ac50495d527dcb09d207

Observation a3d1013e-9dd2-4e52-a9c8-ace355b79206 · outbound

This paper cites ”earnings- 22: A practical benchmark for accents in the wild.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs ”earnings- 22: A practical benchmark for accents in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.733023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:e1bd3bc672c84329be6e011bd331195e7caa6d109adb2aac6883b84fa19c6f05

Observation 8496c9cb-6b13-40a4-b382-7d1572013902 · outbound

This paper cites The ami meeting corpus: A pre-announcement,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs The ami meeting corpus: A pre-announcement,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.739220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:96356fb8326c7f7e512081c203a11b0373e9cc219d8973b58c4e929ae3a84087

Observation fb2b9101-cb2b-4880-9cb4-6c6957330305 · outbound

This paper cites FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.444265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:d54b18ef6740f73ed17564e137a93934e0c18aa4fac6f64e022168487f9a2ed9

Observation b58aae29-800d-4348-878a-d7e8a6644d8e · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Robust speech recognition via large-scale weak super- vision,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.756790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:9c9cded60c798f91ed85cbc6839b9f9000f4cb8a535205a899524e4daec2963f

Observation 15296f6a-e947-4705-afe9-2de029452ce6 · outbound

This paper cites Conformer: Convolution- augmented Transformer for Speech Recognition,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Conformer: Convolution- augmented Transformer for Speech Recognition,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.750916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:c7b995582581903c570da264234542cd2bd9faab0a827903952e3f09675f1bee

Observation ed375f2a-eaa3-45b1-a3d6-2d3d93a3b60b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models,.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs LoRA: Low-Rank Adaptation of Large Language Models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:37:34.761787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:fb61d52881a581668ec4867a3ba997f0054715f14e4531508b70672bf0ac82e1

Pith citing papers

No inbound Pith citation observations are available.