Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:17:33.475816Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2505.13338.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:17:33.475816Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:17:33.349929Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-15T20:17:33.661042Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation df2a375e-9c0e-41e4-8847-09a411b397de · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Recent speech-LLMs, such as GPT-4 [1], Qwen-audio [2, 3], SALMONN [4], and MERaLiON-AudioLLM [5, 6], have demonstrated remarkable performance in handling speech-based tasks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5b3e0b13-e110-4c94-b1d6-097889e13798 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 44a15b1d-4201-4976-9b78-db81f5b07cdf · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation What is the content in the audio from the text transcript?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7c3ff951-5a41-4ef3-b9a1-8ddfa4e8a78d · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d1a5bcd2-cc09-493b-a09b-067d749df2e0 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c8563fb2-e5da-44cf-bd61-a3d7a911e0e5 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Our framework con- sists of pseudo paralinguistic label-based data condensation and LLM-based CPQA generation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 257f1e2b-b648-44d7-81eb-7239a32a713c · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a8c59dad-2984-4c64-b67b-03064de6c960 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation GPT-4 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4e56222-2487-41b0-b0f7-c939eed97885 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ec7c9d2-eead-4631-8149-034f16ba74cb · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Qwen2-Audio Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c02a3a55-2109-418b-a084-029faeb96724 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47227b3-c8cf-413e-916b-6b0f82b780df · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79ba7230-e884-473a-bbec-18385fee75db · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1a5cc0e-951c-4a5a-92d5-eb332a8eccf7 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation BLSP-Emo: Towards Empathetic Large Speech-Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcac46d7-5359-459f-a598-dae8ec7ca2c0 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5e1c955-edc6-462f-bfaf-0e48780e59b4 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cd1ed08-992a-426b-a7d7-721c2ccef8fc · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Advancing large lan- guage models to capture varied speaking styles and respond prop- erly in spoken conversations,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation adf22b68-6130-4741-bcf3-9340321f1889 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7090e5-9f39-421b-b88a-f8e163372e9a · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Paralinguistics-aware speech- empowered large language models for natural conversation,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6a3ad9af-4b75-44f7-8891-2fe02c825a7a · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cad0500-702b-4260-805d-dfb6d0be4ad9 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation AudioBench: A universal benchmark for audio large language models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 40ca13f3-6ee6-43b9-96a4-2eb4d6e4422c · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Dynamic-superb: To- wards a dynamic, collaborative, and comprehensive instruction- tuning benchmark for speech,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d7cda837-7c18-4c6c-822a-2c129c3e9a70 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation AIR-bench: Benchmark- ing large audio-language models via generative comprehension,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b88e64c0-7555-4a25-9946-fbe3044dbbf5 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Listen, think, and understand,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c5cccfd-59b3-42fd-aa6f-9f515ba723aa · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MMAU: A mas- sive multi-task audio understanding and reasoning benchmark,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8965e6db-6080-48ac-804c-ad5240d7ae40 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation IEMOCAP: interactive emotional dyadic motion capture database,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d012c2ef-7d75-4efe-8248-b05d1ac94826 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MELD: A multimodal multi-party dataset for emo- tion recognition in conversations,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a129374a-a85b-4739-941c-7dd619b5a975 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation What’s basic about basic emotions?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 85b7fc7a-0eb6-4d69-b73d-087401a787b0 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Theories of emotion,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2418c4d2-380d-4b51-8d7d-6902ebc7b32e · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation EmoBox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8bc8255e-2381-4b4b-b1f3-04acb1b2d49b · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Evidence for a three-factor the- ory of emotions,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5fc62db6-16d6-42cf-accd-f389cbbbabf2 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Goemotions: A dataset of fine-grained emo- tions,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 90ead142-5439-4f7a-bcd5-0cfd198265af · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation emotion2vec: Self-supervised pre-training for speech emotion representation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5d0fc0a3-4a9a-4c29-836a-9cf4ef2e8d60 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Dawn of the trans- former era in speech emotion recognition: Closing the valence gap,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 481c58b7-2554-485f-b02e-aeb833128bff · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Building naturalistic emotionally bal- anced speech corpus by retrievingemotional speech from existing podcast recordings,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3cbdafd9-b7a6-4a4b-ac0c-3a9ef94c2ea6 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation WavLM: Large-scale self- supervised pre-training for full stack speech processing,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26111fcc-14e8-4c9d-8386-3f5270689267 · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35d844f6-a36d-41b0-a4a0-63ae8a49271c · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation V oxCeleb2: Deep speaker recognition,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a6e8142-01c7-4b17-8889-90906e41acac · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation WhisperX: Time- accurate speech transcription of long-form audio,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e96a01a8-e897-4f96-8521-8b45b64ba23a · outbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation The Llama 3 herd of models,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5b3e0b13-e110-4c94-b1d6-097889e13798 · inbound
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.