Pith. sign in

Paper Citation Record · LEDGER

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2505.13338.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13338 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:17:33.475816Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:17:33.349929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:17:33.661042Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df2a375e-9c0e-41e4-8847-09a411b397de · outbound

This paper cites Recent speech-LLMs, such as GPT-4 [1], Qwen-audio [2, 3], SALMONN [4], and MERaLiON-AudioLLM [5, 6], have demonstrated remarkable performance in handling speech-based tasks.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Recent speech-LLMs, such as GPT-4 [1], Qwen-audio [2, 3], SALMONN [4], and MERaLiON-AudioLLM [5, 6], have demonstrated remarkable performance in handling speech-based tasks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.932117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.345621Z digest=sha256:309328579bfff4c7e8129ff77033893d70f380c4007cc337fb586d16940d2945

Observation 5b3e0b13-e110-4c94-b1d6-097889e13798 · outbound

This paper cites Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:17:33.666296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.349929Z digest=sha256:a86034542d4288a518b588e03e3176ee6f8c8e0ab671f9631c796efb1552331a

Observation 44a15b1d-4201-4976-9b78-db81f5b07cdf · outbound

This paper cites What is the content in the audio from the text transcript?.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation What is the content in the audio from the text transcript?

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.921636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.353733Z digest=sha256:84e7a96edc67e6725e5ce36a7e916b0f32081fe86832204cfa2a295d676d1e92

Observation 7c3ff951-5a41-4ef3-b9a1-8ddfa4e8a78d · outbound

This paper cites an unresolved cited work.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:17:33.911818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.357222Z digest=sha256:f2d6bde31aba9bc003e3693245137bd3ea9a36fc87f64872d4fdc9ff26f91d19

Observation d1a5bcd2-cc09-493b-a09b-067d749df2e0 · outbound

This paper cites an unresolved cited work.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:17:33.900926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.360522Z digest=sha256:834999352cba172a3abb0004c92162cd583eae7ef8baeeb15fb9643eb27b6a25

Observation c8563fb2-e5da-44cf-bd61-a3d7a911e0e5 · outbound

This paper cites Our framework con- sists of pseudo paralinguistic label-based data condensation and LLM-based CPQA generation.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Our framework con- sists of pseudo paralinguistic label-based data condensation and LLM-based CPQA generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.364103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.364103Z digest=sha256:7afee3ee871e565986bccd3563e93e5fd64d485d1b3d800fbaee6dda6a2274db

Observation 257f1e2b-b648-44d7-81eb-7239a32a713c · outbound

This paper cites an unresolved cited work.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:17:33.890556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.367822Z digest=sha256:401af2aa3d15cb52914cdb71be9b038a3cfeff855d0afdcff054df1c36b26590

Observation a8c59dad-2984-4c64-b67b-03064de6c960 · outbound

This paper cites GPT-4 Technical Report.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.371023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.371023Z digest=sha256:f9e50a2ec32dbbf7a5bba60c1c6a337918f9e0b9de2f184a852ed1564b22793d

Observation d4e56222-2487-41b0-b0f7-c939eed97885 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.374448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.374448Z digest=sha256:93435007d25cc05a68e855098c5adfcde783afa3684e41e3e03a1effed0544b3

Observation 8ec7c9d2-eead-4631-8149-034f16ba74cb · outbound

This paper cites Qwen2-Audio Technical Report.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Qwen2-Audio Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.377815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.377815Z digest=sha256:ed73d807332430d2e05a9a93348bdafba336f0efced05f9f98b19d09d657631e

Observation c02a3a55-2109-418b-a084-029faeb96724 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.380885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.380885Z digest=sha256:ff2f719ee02b0c787b189d146f6221e95819f94976799203965c5db6d5555018

Observation b47227b3-c8cf-413e-916b-6b0f82b780df · outbound

This paper cites MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.384225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.384225Z digest=sha256:b51a2ababb75acd60353e62caeceeb5b30342aba94746ca81f32502a1a17e0dd

Observation 79ba7230-e884-473a-bbec-18385fee75db · outbound

This paper cites MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.387816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.387816Z digest=sha256:021502437e0ae2b41ea0a7aa4f6d736438b59ef7a953a538c8f51fbbfbb227ea

Observation c1a5cc0e-951c-4a5a-92d5-eb332a8eccf7 · outbound

This paper cites BLSP-Emo: Towards Empathetic Large Speech-Language Models.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation BLSP-Emo: Towards Empathetic Large Speech-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.391551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.391551Z digest=sha256:6952ef5e3c040574177597abdc940d5537c796157a7e96f88b62bb55b72acac6

Observation bcac46d7-5359-459f-a598-dae8ec7ca2c0 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.394954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.394954Z digest=sha256:ec5df75021180b5254cfed8e7e9c7e3e95feb6db80be790ec34bdac0582df9e1

Observation f5e1c955-edc6-462f-bfaf-0e48780e59b4 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.399189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.399189Z digest=sha256:76efd5599415daeefa3eb7359272bb01607fa74e9ec5cf8f881eccb28919c761

Observation 7cd1ed08-992a-426b-a7d7-721c2ccef8fc · outbound

This paper cites Advancing large lan- guage models to capture varied speaking styles and respond prop- erly in spoken conversations,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Advancing large lan- guage models to capture varied speaking styles and respond prop- erly in spoken conversations,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.880417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.402172Z digest=sha256:d92db62cd805c752de70aeda16b00dd09be22f2914444cf05299bd4aa9b0c50d

Observation adf22b68-6130-4741-bcf3-9340321f1889 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.405082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.405082Z digest=sha256:271371db60e7e6c14f6c2f9ab8fb9767f3e51b383fc77a226f0d426a8fa74427

Observation bb7090e5-9f39-421b-b88a-f8e163372e9a · outbound

This paper cites Paralinguistics-aware speech- empowered large language models for natural conversation,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Paralinguistics-aware speech- empowered large language models for natural conversation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.868245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.409218Z digest=sha256:58e894df7c2a7f4534523c07a963f2c21739b8654a847ddfb024c6a4a6f94213

Observation 6a3ad9af-4b75-44f7-8891-2fe02c825a7a · outbound

This paper cites Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.412848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.412848Z digest=sha256:fa30d143ad0669af8b715c12d9bc8c7d608eaacecf52fcb0fb4c03cb6da26dcf

Observation 5cad0500-702b-4260-805d-dfb6d0be4ad9 · outbound

This paper cites AudioBench: A universal benchmark for audio large language models,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation AudioBench: A universal benchmark for audio large language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.858006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.415869Z digest=sha256:4d1cb9cc7ff5f38f88e06686a8192b72100bd248c747ffdef454a78ede470552

Observation 40ca13f3-6ee6-43b9-96a4-2eb4d6e4422c · outbound

This paper cites Dynamic-superb: To- wards a dynamic, collaborative, and comprehensive instruction- tuning benchmark for speech,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Dynamic-superb: To- wards a dynamic, collaborative, and comprehensive instruction- tuning benchmark for speech,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.847956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.418994Z digest=sha256:581b184f728b799df03e423308363de92d4b43bb7c6b02143a6578ad1a14a09a

Observation d7cda837-7c18-4c6c-822a-2c129c3e9a70 · outbound

This paper cites AIR-bench: Benchmark- ing large audio-language models via generative comprehension,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation AIR-bench: Benchmark- ing large audio-language models via generative comprehension,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.837167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.422609Z digest=sha256:b5fe5de6ff9ad36e31ad6488436128fc69f0e1daef84db24775046eb45e7952c

Observation b88e64c0-7555-4a25-9946-fbe3044dbbf5 · outbound

This paper cites Listen, think, and understand,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Listen, think, and understand,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.425641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.425641Z digest=sha256:52d6cc9b10466f7382339fcb2a096c820486692231c4eaa7ac160419478da3fa

Observation 3c5cccfd-59b3-42fd-aa6f-9f515ba723aa · outbound

This paper cites MMAU: A mas- sive multi-task audio understanding and reasoning benchmark,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MMAU: A mas- sive multi-task audio understanding and reasoning benchmark,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.818914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.428590Z digest=sha256:82845ecd9fd65df6b13bbbaeece00d1920461b1f5efd7a241860f8fe303fbbaa

Observation 8965e6db-6080-48ac-804c-ad5240d7ae40 · outbound

This paper cites IEMOCAP: interactive emotional dyadic motion capture database,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation IEMOCAP: interactive emotional dyadic motion capture database,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.807231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.431691Z digest=sha256:75f116571145bbbcaa392667b022f65ec0557fa623f410bf6f1d8fba264f070c

Observation d012c2ef-7d75-4efe-8248-b05d1ac94826 · outbound

This paper cites MELD: A multimodal multi-party dataset for emo- tion recognition in conversations,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MELD: A multimodal multi-party dataset for emo- tion recognition in conversations,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.797163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.434765Z digest=sha256:14e0d03031ce025caee25962f8f5d4599c0688e99031c1e9bd4a57a6d2ed07f2

Observation a129374a-a85b-4739-941c-7dd619b5a975 · outbound

This paper cites What’s basic about basic emotions?.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation What’s basic about basic emotions?

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.785652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.438307Z digest=sha256:e2d4c6273da9338c95ad1f82d143a8037aad58000ccd9432bd75c61c3ef04532

Observation 85b7fc7a-0eb6-4d69-b73d-087401a787b0 · outbound

This paper cites Theories of emotion,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Theories of emotion,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.775706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.441242Z digest=sha256:22c72823bd128bfcaf52bfb090cd89cb25253ef59e17858412b18f3438f14602

Observation 2418c4d2-380d-4b51-8d7d-6902ebc7b32e · outbound

This paper cites EmoBox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation EmoBox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.765848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.444355Z digest=sha256:b8af45c2457326c46acc2d685a19a545b68c9cde81cf56cdc49405e71f9602de

Observation 8bc8255e-2381-4b4b-b1f3-04acb1b2d49b · outbound

This paper cites Evidence for a three-factor the- ory of emotions,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Evidence for a three-factor the- ory of emotions,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.755021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.447515Z digest=sha256:8209074e83efb0121a6295bc64d49a0aae373fe73a5f5a4aec2597ed78e83c77

Observation 5fc62db6-16d6-42cf-accd-f389cbbbabf2 · outbound

This paper cites Goemotions: A dataset of fine-grained emo- tions,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Goemotions: A dataset of fine-grained emo- tions,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.745284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.450756Z digest=sha256:b06b8dc923310a35927b655aa9c30fdb2f20b0e696ac9dc498ed506757feedbc

Observation 90ead142-5439-4f7a-bcd5-0cfd198265af · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation emotion2vec: Self-supervised pre-training for speech emotion representation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.734700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.454003Z digest=sha256:ba2bfc96d5a6c9140219a1ed99447659b8b06a6f2bd3fd2a9ef08df02c6658f6

Observation 5d0fc0a3-4a9a-4c29-836a-9cf4ef2e8d60 · outbound

This paper cites Dawn of the trans- former era in speech emotion recognition: Closing the valence gap,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Dawn of the trans- former era in speech emotion recognition: Closing the valence gap,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.457390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.457390Z digest=sha256:b72158eff75207ac3c8ad54b453c1786a0d60cb11c792df99481f019e70d3ef8

Observation 481c58b7-2554-485f-b02e-aeb833128bff · outbound

This paper cites Building naturalistic emotionally bal- anced speech corpus by retrievingemotional speech from existing podcast recordings,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Building naturalistic emotionally bal- anced speech corpus by retrievingemotional speech from existing podcast recordings,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.718410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.460553Z digest=sha256:3adf09a9c6e0204ba980fecd2cd31766616be46d173ab1d057ab6ab14b9eff56

Observation 3cbdafd9-b7a6-4a4b-ac0c-3a9ef94c2ea6 · outbound

This paper cites WavLM: Large-scale self- supervised pre-training for full stack speech processing,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation WavLM: Large-scale self- supervised pre-training for full stack speech processing,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.463514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.463514Z digest=sha256:43786c44f6bb149b0799c9a151f70262c75734611597966a600c6283d7fffeed

Observation 26111fcc-14e8-4c9d-8386-3f5270689267 · outbound

This paper cites Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.701994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.466518Z digest=sha256:32c26ba1bde1e82774952422ecf4ee8c489b6d9f2087c9100d6f3a3d4a4e7e21

Observation 35d844f6-a36d-41b0-a4a0-63ae8a49271c · outbound

This paper cites V oxCeleb2: Deep speaker recognition,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation V oxCeleb2: Deep speaker recognition,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.469808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.469808Z digest=sha256:b04d441ff5fc1aed0dc6cd07300be9f2b7d34bc15114768f2893fab6b4b55c0d

Observation 3a6e8142-01c7-4b17-8889-90906e41acac · outbound

This paper cites WhisperX: Time- accurate speech transcription of long-form audio,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation WhisperX: Time- accurate speech transcription of long-form audio,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.686760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.472777Z digest=sha256:f9994dbc20b211bf9ff10399278ac492c9732b7cb6614ab843e3c26eb18a1349

Observation e96a01a8-e897-4f96-8521-8b45b64ba23a · outbound

This paper cites The Llama 3 herd of models,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation The Llama 3 herd of models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.676651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.475816Z digest=sha256:2021b56057c9c79aa5377b8ac3838a4a49eab0e7d6f1ddfabaca707e465c2d6a

Pith citing papers

Observation 5b3e0b13-e110-4c94-b1d6-097889e13798 · inbound

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation cites this paper.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:17:33.666296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:17:33.349929Z digest=sha256:a86034542d4288a518b588e03e3176ee6f8c8e0ab671f9631c796efb1552331a