Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2305.11000.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:01:03.610454Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation beb4b7b6-4a69-4f99-a11a-32f8d391e5fe · inbound
A Survey on Multimodal Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 151
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ef7f953-8eaa-4ae6-af6d-0ac94bf9ebc7 · inbound
SALMONN: Towards Generic Hearing Abilities for Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e09f75fc-0890-4b05-b290-52f56213f9f7 · inbound
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 94aa5a0d-5450-42b7-8095-d05f2364977c · inbound
Qwen2-Audio Technical Report SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ef1432ba-bd84-420b-b16e-d9df22ac1ef1 · inbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 266e6da1-273d-4983-8c5f-bf1d8523eb6b · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f890245c-6531-48b1-a968-887b0e36d586 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 205
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9529745b-1880-4133-a879-50076ef4bab5 · inbound
Step-Audio 2 Technical Report SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 746e9178-8e30-4f6b-82c8-543a1934ebf2 · inbound
Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe02ded-932b-4f26-995d-3162e8dc2afa · inbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcec29af-ec7c-4da5-a927-e41fc3fa48d8 · inbound
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c53f719a-b5e5-4271-b8d6-e98932066d8e · inbound
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ece04e-a0b7-403f-9331-96d0bd98817f · inbound
Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994f3e9f-1d26-43fb-9e41-f1e9d6569115 · inbound
Group Relative Policy Optimization for Speech Recognition SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c5d885b-02fa-4c96-a6a1-5d549788bc66 · inbound
Enhancing Speech Large Language Models through Reinforced Behavior Alignment SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23427e32-03aa-4a5c-8c6e-7bfae1829283 · inbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be44166-4d2a-4af7-b3c4-8117fb1bf908 · inbound
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57fbd1f0-04d2-4e4f-8e7a-2cd30255b309 · inbound
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df67025e-725c-4547-a749-91886819c910 · inbound
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b089b6-57a7-46ab-8b68-20eebafca2cd · inbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76858a12-025e-480d-b9d5-3bfb75cceb6d · inbound
TSVer: A Benchmark for Fact Verification Against Time-Series Evidence SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 35d34721-1096-4ca2-9fac-63a653d4ea36 · inbound
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81dd535c-e9c2-4775-a761-8ac5e796dce7 · inbound
Two-Dimensional Quantization for Geometry-Aware Audio Coding SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1277395f-e0a4-4d68-a4c1-55f8ff23608e · inbound
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ebb5b19-0736-480d-887a-3505fb7b1439 · inbound
LLMs and Speech: Integration vs. Combination SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b944928d-db29-4572-823f-eab9987d6e4e · inbound
LLMs and Speech: Integration vs. Combination SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00bb5a82-a27e-4b68-bd2b-675f021571e6 · inbound
Neural networks for Text-to-Speech evaluation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c025e21-6469-489b-88a9-5d070b7a2c87 · inbound
ViLL-E: Video LLM Embeddings for Retrieval SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d395b289-6217-494c-9920-22cb2d766148 · inbound
Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d2c2d611-1a0c-464c-90cd-944333e05950 · inbound
From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0596067d-64ce-4802-8c34-b9805b42e38c · inbound
Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 43736d75-16dc-45c9-b0a9-72929fe69637 · inbound
Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f705df26-ba80-4391-bc28-1dfacd9459cd · inbound
Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 32eb141e-8641-4aa1-bcc2-3a42d0c5d095 · inbound
Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c25ac1d5-0dc0-4251-b1ea-41aa36e65816 · inbound
UniVocal: Unified Speech-Singing Code-Switching Synthesis SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cc14cf4-5788-4e9a-b665-2d17d061b13e · inbound
MOSS-Audio Technical Report SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9fed51af-1648-4b02-b58c-48c1fbe4630b · inbound
AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a2c7b6bf-0ab1-49e8-9801-8092192fcfc4 · inbound
Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e013a77b-8479-4a25-aee3-c1e3d9e81e31 · inbound
Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ad609d9-cf34-4c5d-8caf-03b6cd1cd960 · inbound
AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14fd51e6-a685-4433-ac04-f55e1d49f3de · inbound
Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 050462e5-03e9-4d0b-8e4c-9c0db67aac9e · inbound
HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d351f54-86e3-4725-925c-88f850bc0b3b · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.