Pith. sign in

Paper Citation Record · LEDGER

SpeechVerse: A Large-scale Generalizable Audio Language Model

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2405.08295.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.08295 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:32.222856Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 474d7e28-d8cc-4621-ae66-e608e21415dd · inbound

Qwen2-Audio Technical Report cites this paper.

Qwen2-Audio Technical Report SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:14:45.462747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:14:45.371564Z digest=sha256:0c2574fd3fc4a95c93062a492f9dfe7d8549de86cea7bb28bd94598af4f886c7

Observation 62792d3a-d3f9-48b5-badd-2459f82d3802 · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.291681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:041c02aeeb5cdd6e5db938e0830ce10197aca19454d816e1f04ba215d33543fb

Observation 6f095213-9732-49f6-9f03-231262498176 · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:54:03.293192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:bb092172f4592d95ed602cb72abfdf718939c76392ac3d594e080b68e6a4083d

Observation c9bfbf4a-251a-4423-a3d4-c7be0a6c46ad · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:07.987371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:35b5776bdba458edbaa85e1d11fcbda7adfa2120aa41185b372727803a8bdeb0

Observation 1e9316b4-b1fd-426f-8937-3caf555a31f6 · inbound

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving cites this paper.

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:32.222856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:32.222856Z digest=sha256:1346103909cd66577d8d96422e1bb519d8a7f980d8f51215ca6e0172c378c30c

Observation 443ccc8e-7d58-43a8-ba77-17d99a880be4 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:22:37.293820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:37a901f3f642676de1c35c65e9578498bdeb4c67011d51fdcd6aa2786efb1ef0

Observation 0a06db11-99ae-49f8-8285-e4a32689457d · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.082340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.082340Z digest=sha256:48eff37e344d1df7b71c0830394a1a4a5cc51e735596d1a1c5cd1315bb8e8737

Observation fdb3b362-ebbd-455e-8dd6-55c2a7f4ad4f · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:31.841304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:31.841304Z digest=sha256:1f774d40c07abc029ddc9538c7adb4b7aec430ae606ba9ed712d2e61fedf472d

Observation b91d2310-470c-4a35-a5e0-5ddf7a956343 · inbound

Your Spending Needs Attention: Modeling Financial Habits with Transformers cites this paper.

Your Spending Needs Attention: Modeling Financial Habits with Transformers SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T10:59:03.105773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:59:03.105773Z digest=sha256:68888c7c563e8137e95d2887886d1499717464005b2caab0c296ec8b51f74b42

Observation 89172613-8aed-4848-822a-cc168751ec30 · inbound

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition cites this paper.

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:01:21.059534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:01:21.059534Z digest=sha256:ddcba406f9a74353d74fafe6b26ecfb449b931a0b3042b819834bf76c73566c8

Observation f54078c4-3e0f-4f78-99d6-2b2fa9376b3c · inbound

TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation cites this paper.

TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:27:29.283691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:27:29.283691Z digest=sha256:09a4cf739439b525452c89fd3444f21e27a36a828653bab7604ce57b8746a7bf

Observation 44ed0d62-41c9-427e-9093-19ea7800fef0 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.516544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:967049959d35a17f12c4cc51c4bca61bfd32388834a7b077ad2b48f9287f33a1

Observation 24cce4b3-79af-445d-9dcd-fb45577f3dec · inbound

Direct Simultaneous Translation Activation for Large Audio-Language Models cites this paper.

Direct Simultaneous Translation Activation for Large Audio-Language Models SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T16:31:36.857609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:bb657bc8ace53224be803c91dbb6a92d2dd65f20665df1752ed1ff7e76adbea3

Observation 9a25b118-325e-4ad3-a8db-c625e6a19e72 · inbound

A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification cites this paper.

A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:48:45.994945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T00:46:08.921194Z digest=sha256:5628d63758279966d2213018138b89c6152808ce9b436d4f0ef97e012f3392b3

Observation 7950c080-4db3-4b22-9a33-64b1ead58181 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:37.503517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:37.503517Z digest=sha256:51420480bf818556457b1f8faba83de8ce022ece0448aba2864773ef6eaf5ad9

Observation f701f70d-084b-4191-93fb-9320c4c6b1fa · inbound

Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection cites this paper.

Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:40:51.416440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:37:56.416600Z digest=sha256:abaaa7a0a31f89b13311be7e05d0729514e357f29be637760a4d2d617d8c2ba6

Observation 9e4419e8-9bf8-4396-862e-f2a4021a6b9c · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:50:40.260650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:cdf39920a14d1e2673f53a1e61121bb3d4038a4d7176f9490ab05885dad8ba0d

Observation bb44bd26-3154-4e60-aea1-c6402245406d · inbound

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity cites this paper.

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:50.813861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:50.813861Z digest=sha256:4af7a393c658275a104fd781f7efaa3754ee6cc124fe016c66ba8c6ce1efd237

Observation 8875ffd5-8b02-434c-9207-648105f67899 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.282925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:1e15ea13fc8d1239ce11c29a0a39c78f3cbcef5bd75355ff1b8b3b5bec36271c

Observation bcc2be79-4c71-42fa-8a17-0086fe8db9dc · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:39:48.981456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:31cdf20f1943b5137d4bfdeeb90bda39a235ae3da496ecd9baf17a165afc8e34

Observation 60960240-19b8-47ba-b252-6a48a7d72167 · inbound

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding cites this paper.

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:54:45.172403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T06:54:41.345212Z digest=sha256:e9f1dc45800c02cf285a535f0db57a1e7f1eb07dca6cae5f0028e0776bd98d26

Observation ed66c343-5d3e-4173-8a20-25e1c6f7bd57 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.347302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:039908626b718398a919b39851111280b146f0c0e5a3e9341f4d60c944a8ea15

Observation 6754d537-dcf4-4b2d-b544-b7a8c614195d · inbound

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition cites this paper.

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:55:29.122979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T06:51:41.995108Z digest=sha256:e27adca9262900134a3dc4d6e033182af98973db03c0919fb32048be255f98b7