Pith. sign in

Paper Citation Record · LEDGER

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs

As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2606.09366.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09366 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:26:32.594986Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact12
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4ba04366-f80d-42f3-b828-93d0309bccac · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.605027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:b9f96287b8232b023f4b1b694cb5be8f43e4a6be3ecb97a3f30016f0b3e02657

Observation 45bd6074-46ad-4a75-924d-cdddcf3a4728 · outbound

This paper cites Qwen2-Audio Technical Report.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Qwen2-Audio Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.607505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:09c5e9d363cb0af19bcc05765e706f4f149fefcb8ce8aec1a65411bd2323bcde

Observation 60e8cece-fda3-4ac2-9c82-259e6205bd84 · outbound

This paper cites Closing the gap between text and speech under- standing in llms.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Closing the gap between text and speech under- standing in llms

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.621346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:5267cc12bd32ccfc882e897cd0b80e15367036af34536413392f5fe1c0b66eb2

Observation 9e850727-45d9-46e5-b89c-dd7b39f449d8 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.611659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:3abedab8dad28eae262469b45177a80ce386206cd8b9b433d343c77e71271e36

Observation dc5cb171-33d0-47a5-a87e-03621ff2e003 · outbound

This paper cites Measuring audio’s impact on correctness: Audio-contribution-aware post-training of large audio language models.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Measuring audio’s impact on correctness: Audio-contribution-aware post-training of large audio language models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.618646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:b763897af5bbad5e91e28a6ce67a56b3c33d0c919b349d56d688f9a4bc90c975

Observation 621d1ba7-779e-417c-a2bf-bbbeba787d77 · outbound

This paper cites Kimi-Audio Technical Report.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Kimi-Audio Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.597792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:093e0348b8a25a71aefd76f8aad684caee4946565233f2f4e004e70785831d4b

Observation f31f6adb-b2a7-4358-a2e7-2e3c9ca1244b · outbound

This paper cites FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.615450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:a156de57385e46d1c0a9c71694a71b8acc306aec0d8643a7dca0544d3f05122a

Observation 575a00fe-35a2-48c6-9d45-dac408d68aaa · outbound

This paper cites PLOS ONE , publisher =.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs PLOS ONE , publisher =

Reference 8

Resolution
metadata mismatch
doi, observed 2026-06-27T16:31:02.727492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:527ca8daf583d8bdcbf58d8f44c1b099aff1fd03157486fe2268632929ee13d6

Observation af7227b8-3b48-4510-864f-4cd11dc8a227 · outbound

This paper cites Robustness assessment of large audio language models in multiple-choice evaluation.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Robustness assessment of large audio language models in multiple-choice evaluation

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:37:30.595468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:49e6f91841b9d52fc440dc62c13f1c1cb3889f50b70531ba5e787df8acee1965

Observation 8e179965-984e-494a-b90a-b91990573426 · outbound

This paper cites Available: https://arxiv.org/abs/2511.03310.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Available: https://arxiv.org/abs/2511.03310

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.597120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:fd174fe1128d44adef7b578a60e8689e5da2ea42a71b80a2da624aa7aff19047

Observation f5b65a5e-84a4-4f2a-9e20-ebb402c64e0e · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.599502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:4bda421de72edc2a0812a5f523eb23560c204a8a16b0357c7f4788e829eb52c2

Observation c55f5e85-e411-4c6e-b2e9-10faaf5f9493 · outbound

This paper cites LLaSM: Large Language and Speech Model.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs LLaSM: Large Language and Speech Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.580557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:1fd52044e8c8e947b910c25627cbfd6d228060ff6c33dbfad7be403690ae629c

Observation 01cba664-288f-4a84-a6ad-dabecc5f7b82 · outbound

This paper cites LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.589991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:8eda50d80888780d0be83bc629dd632fecd4e4214e48c6f5effdcbfd5917743d

Observation 0e312edf-c0eb-49ee-8d51-53de56662c24 · outbound

This paper cites SSR: Alignment-Aware Modality Connector for Speech Language Models.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs SSR: Alignment-Aware Modality Connector for Speech Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.589450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:d2701807a4edd3ce0ff6a408b6b09948f52ded37e9facb09190ad1a61cbc084c

Observation 1c3b0ca4-2a46-45af-a79c-ba6ace513038 · outbound

This paper cites Closing the Modality Reasoning Gap for Speech Large Language Models.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Closing the Modality Reasoning Gap for Speech Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.612695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:b5c5b411ceec1b027b00e56b20587bc21f474cd2dff4928cbbf21ce37b5c4a66

Observation b5ec2c3d-063b-4aca-a024-2f406df37cf0 · outbound

This paper cites MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.594742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:6f5505af3de8a48ab560e5f9426664a2c5252dbb1af64f416275a60e1ad4606a

Observation f6f9f46d-f6b4-434a-9928-6f2d81b15ade · outbound

This paper cites Qwen2.5 Technical Report.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Qwen2.5 Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.610025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:ba6823c0669cacd001237941d1964115db4ded3dc4cf77690fc1fa9b973134af

Pith citing papers

No inbound Pith citation observations are available.