Pith. sign in

Paper Citation Record · LEDGER

SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2305.11000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.11000 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:01:03.610454Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation beb4b7b6-4a69-4f99-a11a-32f8d391e5fe · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:41.981438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:94a1df62cfe95b1a47b6c7cdf947437de576a6f5c99fed5888975c0458098545

Observation 5ef7f953-8eaa-4ae6-af6d-0ac94bf9ebc7 · inbound

SALMONN: Towards Generic Hearing Abilities for Large Language Models cites this paper.

SALMONN: Towards Generic Hearing Abilities for Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:29:46.327448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T02:29:46.242983Z digest=sha256:f3e3980a857cc16450001d34d636ddba510e133cb7b9f0937589fe6a8bec5993

Observation e09f75fc-0890-4b05-b290-52f56213f9f7 · inbound

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models cites this paper.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.769454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:671e113ba8f1021331ce8fbcd308d91966e5b27be09d5bae09d48799459bb923

Observation 94aa5a0d-5450-42b7-8095-d05f2364977c · inbound

Qwen2-Audio Technical Report cites this paper.

Qwen2-Audio Technical Report SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:14:45.455080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:14:45.371564Z digest=sha256:7b91dacce433ae821c4d053b2852ee5e90c0768b4899eca5df61a36ee0b123b0

Observation ef1432ba-bd84-420b-b16e-d9df22ac1ef1 · inbound

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot cites this paper.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.549814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:2e3894b49fdcfafbd833b4df31f302698e91ea1683294c157c72cc6b0fbbab0e

Observation 266e6da1-273d-4983-8c5f-bf1d8523eb6b · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.796565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:e0f1fc54d1ce1dc1f188b7a4337a19f2ef06e92c5d41a096e8429a235493fdea

Observation f890245c-6531-48b1-a968-887b0e36d586 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 205

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.384876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:14461d52e669665630289fdb29b0d42e8d962aa54e4c48693a23be6cdcbc0d8d

Observation 9529745b-1880-4133-a879-50076ef4bab5 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.115148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:0edf88d0371f644d0443ead6ed57967eb7e05b92d7655127c97cedfde2414101

Observation 746e9178-8e30-4f6b-82c8-543a1934ebf2 · inbound

Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge cites this paper.

Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T20:01:03.610454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:01:03.610454Z digest=sha256:02dbbf9ccc850c61882ea2a5d9c606bcd05701fbfc96c5a7ad1983f91f583767

Observation abe02ded-932b-4f26-995d-3162e8dc2afa · inbound

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation cites this paper.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:02.039158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:02.039158Z digest=sha256:cf07e451e4c6b0d116e5c41c203b172f8374de7becd158e15ade9e90d7aaf598

Observation fcec29af-ec7c-4da5-a927-e41fc3fa48d8 · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:49.273789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:49.273789Z digest=sha256:8d805fd46007146f1bdb3dfd57b4ef3284e0681c6424cea0fc651155f92242c5

Observation c53f719a-b5e5-4271-b8d6-e98932066d8e · inbound

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation cites this paper.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.159489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.159489Z digest=sha256:b45bdf21119c03ce6f7e5c3ff2e3b8048645059b50daf88e27eb9c28eb41f0da

Observation 58ece04e-a0b7-403f-9331-96d0bd98817f · inbound

Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy cites this paper.

Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:08:08.991738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:08:08.991738Z digest=sha256:f8165960bf7c6950956993a74b72843ff1c12d80fd0842533c22b36f9fe89eb9

Observation 994f3e9f-1d26-43fb-9e41-f1e9d6569115 · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.309683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.309683Z digest=sha256:9c801a78f328fd55e9aa5e05a495015171856b8b7442d159d061f15b5583b883

Observation 9c5d885b-02fa-4c96-a6a1-5d549788bc66 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.529136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:0d2649a2372e448503b9f114631e3abbcb9a755a49094626cc7d35d26d6742cd

Observation 23427e32-03aa-4a5c-8c6e-7bfae1829283 · inbound

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings cites this paper.

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:53:24.021157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:53:24.021157Z digest=sha256:7e39db61531086d25194a9a88a8253f4d3e61d3c025d219f1b3ceef576e79007

Observation 0be44166-4d2a-4af7-b3c4-8117fb1bf908 · inbound

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training cites this paper.

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:53:41.496527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:53:41.496527Z digest=sha256:1247550e478ee0738a789b41eb8608abd916b18c31c1ee6ff169eeff923b71e3

Observation 57fbd1f0-04d2-4e4f-8e7a-2cd30255b309 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.341502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.341502Z digest=sha256:7d4be918076b0bdda00ee8b10513a408b2bfd2f5baf6ba094be4bc213a7c3054

Observation df67025e-725c-4547-a749-91886819c910 · inbound

Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models cites this paper.

Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:24.636694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:24.636694Z digest=sha256:bdf4363dc2fc1fa7c77a4be7d65c717e24a203cf4b7de4342bd5b156f0419b3a

Observation 47b089b6-57a7-46ab-8b68-20eebafca2cd · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.658307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.658307Z digest=sha256:c7bea41b71416d83b568d074e04466d5e9e89ac41549eaea840e283993eab326

Observation 76858a12-025e-480d-b9d5-3bfb75cceb6d · inbound

TSVer: A Benchmark for Fact Verification Against Time-Series Evidence cites this paper.

TSVer: A Benchmark for Fact Verification Against Time-Series Evidence SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:00:34.163267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T00:56:33.806425Z digest=sha256:fbb42c87025b8a3e2bfbf7466ae34add4985ce2aa174e3d3a964693d742aae8c

Observation 35d34721-1096-4ca2-9fac-63a653d4ea36 · inbound

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages cites this paper.

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:18:57.162741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T03:15:04.685150Z digest=sha256:6aa88b7351c5616f8aeada157ff21a942be9182f2b42583874cd02e35dc7583a

Observation 81dd535c-e9c2-4775-a761-8ac5e796dce7 · inbound

Two-Dimensional Quantization for Geometry-Aware Audio Coding cites this paper.

Two-Dimensional Quantization for Geometry-Aware Audio Coding SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:20:29.251157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T18:16:51.486807Z digest=sha256:a6432d86b45467c3b65ada6989fd2e8ad077ca9d310a53b7f2ac9cda7997bd4b

Observation 1277395f-e0a4-4d68-a4c1-55f8ff23608e · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.148297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:e7f17ede0690174459412aaad55d19122940c684a85226eb5ba2ee88fd8e5563

Observation 3ebb5b19-0736-480d-887a-3505fb7b1439 · inbound

LLMs and Speech: Integration vs. Combination cites this paper.

LLMs and Speech: Integration vs. Combination SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:45:28.328602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:41:22.138517Z digest=sha256:a039d1cb6d62a4d192360d4665ad43978d96712640db0f8597544049ad9ddc00

Observation b944928d-db29-4572-823f-eab9987d6e4e · inbound

LLMs and Speech: Integration vs. Combination cites this paper.

LLMs and Speech: Integration vs. Combination SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T20:46:17.286576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:46:17.286576Z digest=sha256:7adf188afa6d2af9d168182e02052a285463e591275ff612388cde462719b2e7

Observation 00bb5a82-a27e-4b68-bd2b-675f021571e6 · inbound

Neural networks for Text-to-Speech evaluation cites this paper.

Neural networks for Text-to-Speech evaluation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:49:54.660606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:46:26.884551Z digest=sha256:50ca306652056b8681e6d055ee40f602f1afd2d89899fb0bd4e087850e6819eb

Observation 7c025e21-6469-489b-88a9-5d070b7a2c87 · inbound

ViLL-E: Video LLM Embeddings for Retrieval cites this paper.

ViLL-E: Video LLM Embeddings for Retrieval SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.048536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:00:43.573409Z digest=sha256:d0feae9dbdedabb4eb49c0112f566b6d306cab2fad9980a51326f268ab862b83

Observation d395b289-6217-494c-9920-22cb2d766148 · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.175278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:f77567fc8ce6cd82257c096ca12d95ae0015cd78ca660b14ea2224fe5d06c702

Observation d2c2d611-1a0c-464c-90cd-944333e05950 · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.052994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:f8765a03e787225696dacd38bfcc8ef153a3f994e32e503404bc6359dc27c37b

Observation 0596067d-64ce-4802-8c34-b9805b42e38c · inbound

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training cites this paper.

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:18:07.280806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T07:13:29.041056Z digest=sha256:375851b41f4aba777c21660e1da8c3ee474d23eba2e3c13edd05a173b6fa5bac

Observation 43736d75-16dc-45c9-b0a9-72929fe69637 · inbound

Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation cites this paper.

Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:54:40.719565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T05:53:57.830880Z digest=sha256:0ab0b5117d2c0fda098dec52e34433e38ab77f2c0221b953da8c9bc866a2b95c

Observation f705df26-ba80-4391-bc28-1dfacd9459cd · inbound

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation cites this paper.

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T11:43:23.618393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T11:40:46.165547Z digest=sha256:2c27e74d3f1116b59878cb4c070088fd8fa24e414241ce6622977cbc8ff6abdf

Observation 32eb141e-8641-4aa1-bcc2-3a42d0c5d095 · inbound

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning cites this paper.

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-28T18:02:27.084300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T17:56:47.512549Z digest=sha256:e6b04fd91629651c7ffa1a3ec7d55aced9e53b1e0fea43f49984fd3e314abc4c

Observation c25ac1d5-0dc0-4251-b1ea-41aa36e65816 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.559539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:ae89573d674de19ddc67c3b5af55793237ddd39c7f38e4f4c26b6d8649eab920

Observation 3cc14cf4-5788-4e9a-b665-2d17d061b13e · inbound

MOSS-Audio Technical Report cites this paper.

MOSS-Audio Technical Report SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.093347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:f97a6dc265d1553918f3fbc1605cdd784bdb10ff8a425615b8424dc948de92a2

Observation 9fed51af-1648-4b02-b58c-48c1fbe4630b · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.406103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:54247f4fd35def4e125692548d127925e2d1d12c869ef4b52dc9371f8c4ca51e

Observation a2c7b6bf-0ab1-49e8-9801-8092192fcfc4 · inbound

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models cites this paper.

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:57:44.833050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T11:29:36.012349Z digest=sha256:b1d55da00ace743fa62b919ced8b6c9a92198a79e7c13dafed65c0428d61e3c7

Observation e013a77b-8479-4a25-aee3-c1e3d9e81e31 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.785270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:9a3131633b78d2c20731b563b43e0cabecc2185607965b3f27b5d5677758426c

Observation 8ad609d9-cf34-4c5d-8caf-03b6cd1cd960 · inbound

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? cites this paper.

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:29:39.437426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T13:22:12.541923Z digest=sha256:2e4d03d4364b90c68f353c78b78aa7be6d8bfc675a44a08defaa51c56e2f4fb9

Observation 14fd51e6-a685-4433-ac04-f55e1d49f3de · inbound

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? cites this paper.

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:30:08.198134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T20:03:10.858349Z digest=sha256:9ee1934fbe993970acc13c67e9e86ddfb0075394a7e5503b99e831f196f1fe9c

Observation 050462e5-03e9-4d0b-8e4c-9c0db67aac9e · inbound

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models cites this paper.

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T01:02:56.250216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T00:56:43.991936Z digest=sha256:d23c7f742ab526b3742c5ce8d428778bafd43036e3f220dc70cd3e386e33b34a

Observation 4d351f54-86e3-4725-925c-88f850bc0b3b · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.070955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.070955Z digest=sha256:fb25a2e8563bca9b83c05502b5a6988b59957c0d75efbcccaa116831b5e0cc81