Pith. sign in

Paper Citation Record · LEDGER

WavLLM: Towards Robust and Adaptive Speech Large Language Model

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2404.00656.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.00656 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:02:51.327669Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:37:34.396958Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b680d517-637c-45a4-abb2-09fcb9126e8c · inbound

A Preliminary Exploration with GPT-4o Voice Mode cites this paper.

A Preliminary Exploration with GPT-4o Voice Mode WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.327669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.327669Z digest=sha256:ba6cd07fbcf17a73228085c83bfe949500fc89eb64afc1116d19f21f558a25d1

Observation eb985253-43d0-4221-9e6f-15dd2dab079f · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.359286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:927000bd75f608eee85cc5ff2d043d7faace44f3b0155fb7388bfb13546dcf36

Observation a302002f-5c66-46e8-87fa-14ce7f01c409 · inbound

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits cites this paper.

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:49.525285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:49.525285Z digest=sha256:af28add52444ed7ea8c7bed1a7f508b0c5d0f274309c7581f805f4d0f638e32c

Observation 23f2d7f6-e5e9-40ff-8a66-6420ed41fd2b · inbound

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs cites this paper.

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:47.276785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:47.276785Z digest=sha256:f5a1289c3422ca94078c20b08e6a583ff51f52bfead8be0f49b87661dad3928d

Observation 39dc1cdd-a81c-44c2-89b0-8999099d6d37 · inbound

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models cites this paper.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:34.067293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:34.067293Z digest=sha256:a5d16b6e5a16068906c75f4ef1af0eb18cbe4321053b2579d98f695e8a338b46

Observation 9758e2b8-382e-4345-be9b-0970a5209c8d · inbound

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs cites this paper.

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:33.844445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:33.844445Z digest=sha256:8656c439b5e6c2d9fdf28c89e08f60d0b7fad75f34555b1830ae4584d583895b

Observation 0f9d2bf0-68e8-41e4-8fa5-9059d48c1758 · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.743399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.743399Z digest=sha256:c41e65f55df79a433600e5447029f2d9b40da90c467ab30cd6fe736ba6a305a6

Observation c481c263-ae91-4edd-b67e-aece03537254 · inbound

ALAS: An Automatic Latent Alignment Score for Audio Language Models cites this paper.

ALAS: An Automatic Latent Alignment Score for Audio Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.593442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.593442Z digest=sha256:12fae6bb70932ac472d11c1fa61877ad59b7313e7153804fe17156021e0dcf05

Observation 2d50605a-97c0-47c4-8eaf-39bcf667dc6b · inbound

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction cites this paper.

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:30.469502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:30.469502Z digest=sha256:ef57e4e717b8d8c1a6b03b189709b7e390e658e52c29e052da50346a67d257af

Observation 0901f8d5-522a-4fc3-9a6d-41b30ebce6f4 · inbound

Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition cites this paper.

Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:59.761074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:59.761074Z digest=sha256:cad6965cc5329948974dfc0e576948da3d295bf646fb6a36fc238a1180acec90

Observation e44347e4-7050-40ad-adc0-08546cffa511 · inbound

NAVER LABS Europe Submission to the Instruction-following Track cites this paper.

NAVER LABS Europe Submission to the Instruction-following Track WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:37:06.539254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:37:06.539254Z digest=sha256:2e4cd1a3162809fa73870eefa14f56a18e76a70994326a460ce5d0eebcd11c95

Observation 30911dfa-cd1f-4296-b876-2eec645a6286 · inbound

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM cites this paper.

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:49.527166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:49.527166Z digest=sha256:7884c1baa0ebac96e60d70b0c1d72e29bc3a28fc0bfbb39594525a26f0ac1a73

Observation 4fccb077-0b67-487e-996f-acac6bec9f86 · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.712860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.712860Z digest=sha256:cb6c34f70137f0b4b71115e7e262208b6d92237d04c4a985d5cf9a5e43132f13

Observation eddac622-e3c2-4186-a8bb-b36bb8e41ac1 · inbound

A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR cites this paper.

A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T13:03:14.524557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:03:14.524557Z digest=sha256:6952476424e648ef9ba6293ac0ce80083684470c047760dbc254dd85c8fb9f0b

Observation 3a65ff3c-010c-48c3-a432-af3a74704ccc · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.512556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:f33982f4ba5f38a51fc9e58b2d862c4191170bd9a77bc6d63d6bbb4b8ade3fda

Observation 9bf4cc41-8be5-499a-ab66-a0aa23b369a8 · inbound

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings cites this paper.

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:53:24.835833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:53:24.835833Z digest=sha256:717c4d85aca254faf9e4c3933363b26e6d982290787361ba2209c24a702040a3

Observation 972007dd-402b-4d7e-aa27-778be2800211 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:12.804002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:32636971dd991608a78a8d55a73bb4efe9b9e5df35e4e7bc24e3a1874b73ba1f

Observation 2dab21e0-b00b-4940-96c8-083110bdb8a5 · inbound

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval cites this paper.

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:39:58.380231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T02:49:35.819155Z digest=sha256:3c2b32f2d6814ef83da7edaf477049e3d93bb07497c0a13a88a2412aa98707aa

Observation 8e2f9176-2dbf-40c1-99b9-f4b58f4d0be4 · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.183604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T16:59:50.615875Z digest=sha256:3267050da24bd96b731b347b08516f4b3634d70163792c97f3096f793f34aedb

Observation c774a0a5-b0d9-4324-803a-1a86053f725e · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T04:34:50.946446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:34:50.946446Z digest=sha256:1aa7996c585ab78565903614470686597ece468ade683ece470e2f793a392f5a

Observation 63af21e0-5f45-421c-b25e-25dcd9573349 · inbound

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track cites this paper.

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:08:32.423186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T15:04:02.640015Z digest=sha256:7476db2e62a5921570c46b6316a484de15484c71209611c534e0738ac0f85928

Observation 6e4c7289-d80d-4a0b-906f-af201dded00d · inbound

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models cites this paper.

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T14:53:55.789976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-07T14:53:07.512543Z digest=sha256:3839aa4729ab150dca9e54a59ea4b35877c01562c22d7b1a0f95fe03899436f9

Observation f1d48f1e-066c-4610-8783-4b66cd7f5a09 · inbound

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task cites this paper.

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T04:47:56.963554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T04:47:56.963554Z digest=sha256:b0759a32f92a01cd0baab1979323aff9fbb9c72401a3f69a98adefbb1c428fdd

Observation afeac008-e22d-4945-a197-29c902ed6a3c · inbound

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task cites this paper.

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:30:02.391299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:30:02.391299Z digest=sha256:ce5e7c67ebbbfcad14d0bf5141f3defb43b34291be8d1ccc8fff891e4b30526b

Observation 3fe8e6f2-f109-49ef-bf16-6fcddfd29a9d · inbound

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs cites this paper.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T20:37:34.398583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:d18b364e39d8adddda8f8f138b7a5afcc9bad9789916a6976d8ccdcf2cf0bf1c

Observation 9d169afa-8bba-4e92-95e4-1b1b83d6a072 · inbound

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models cites this paper.

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T07:18:27.114855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:18:27.114855Z digest=sha256:52b405b3f382b85d1d5f5a6ca1b6c8a281eee9f10283127b89107862a2bb8a9d