Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:15.683108Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2505.23236.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:15.683108Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:10.171701Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:54:15.802818Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 299bf55c-e888-42e3-b1e7-39169c2fa4e4 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Despite decades of promising advance- ments, most research [1–4] has primarily focused on classify- ing speech into single, discrete emotion categories
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9f3eab2-c1fb-461b-9722-873266a0ca68 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition In contrast, previous research paid attention to either feature disentan- glement [6–17] or fine-grained emotion descriptor prediction [5, 23–26] alone
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3712781b-f7e8-434e-ad74-b333098e0403 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition In contrast, prior research employs external ASR models to generate transcripts before performing SER [12–17], which may lead to low SER performance affected by ASR er- rors
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d52dbc93-a0e9-4856-8893-20c920284adb · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47965573-3fe0-424a-96d6-35b3c71ee6fe · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Alternating Multi-task Fine-tuning A parameter-efficient fine-tuning method known as low-rank adaptation, or LoRA [27], is widely employed to fine-tune the LLM-based models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c68910f4-1435-4639-b716-444836c08e28 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition 14200220, 14200021, 14200324, Innovation Technology Fund grant No
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26923b06-ce5f-41c0-b26e-618f3bf866b5 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition 1, the architecture of our approach consists of a HuBERT encoder, two feature disentanglement blocks, two feature adapters and an LLM decoder
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4585293c-632d-4576-aea5-21e6d759d208 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition excited” with “happy
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b47ff2dc-0a09-458c-8e03-d54d1e487898 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Fine-grained SED features are dis- entangled from HuBERT SSL representations via alternating LLM fine-tuning to joint SER-SED prediction and ASR tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e3704ce-b016-4f73-bdcf-bc0477927b63 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Leveraging speech ptm, text llm, and emotional tts for speech emotion recognition,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 716a1476-f4f9-46f6-9363-dbd705f1728c · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Emotion neural transducer for fine-grained speech emotion recognition,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f662225-8bba-4bed-be4d-3c4337f3fdaa · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Generalization of self-supervised learning- based representations for cross-domain speech emotion recogni- tion,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94d8e0c1-0e4a-45f4-9df4-ae74b8e0a9d3 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Noise-robust speech emotion recognition us- ing shared self-supervised representations with integrated speech enhancement,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13c806fb-b9d8-4308-b06e-2c72e5de67df · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Emix: a data augmentation method for speech emotion recognition,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9aac4409-4fb3-4d54-90b8-195bc3af614c · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Secap: Speech emotion captioning with large lan- guage model,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 740cd773-f56f-4f3a-a9d8-b7540c1ba9fd · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Disentangling textual and acoustic features of neural speech representations,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3040545b-d819-40e2-a2e2-ee840debd6a8 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Variational information bottleneck for effective low- resource audio classification,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4a85e3a-78f4-4da5-908f-b9324d49f97e · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Multi-modal emotion recognition using multiple acoustic features and dual cross-modal transformer,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7702ef33-39a3-426b-a7ce-8948dfbd3e2b · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Wavllm: Towards robust and adaptive speech large language model,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bf1cdef-f5d6-49c1-8b1a-3573adcae58d · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Variational information bottleneck based regular- ization for speaker recognition
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59b65e9c-a181-48dd-b492-c7d643c712b5 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speech emotion recognition combining acous- tic features and linguistic information in a hybrid support vector machine-belief network architecture,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 624ceb2a-ce2b-4856-a431-8a132c52e55a · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speech emotion recognition with multi-level acoustic and semantic information extraction and interaction,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed8e0c42-30f3-4929-a271-fc278ac44fe4 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Multimodal emotion recognition with high-level speech and text features,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a8d77da-1007-4932-a324-c5433791a6ca · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Frontend attributes disentanglement for speech emotion recognition,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a882d128-dc8f-49e8-ad31-34f87f78b8d4 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speech emotion recognition with acoustic and lexi- cal features,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b906f673-8626-4d6a-8ee2-9e72177f15b3 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Salmonn: Towards generic hearing abilities for large language models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation afc23d35-1a44-4f0d-8a8f-97ef8a5fb68b · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3837274-6619-42bd-97ba-967e69161b12 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Qwen-audio: Advancing universal audio under- standing via unified large-scale audio-language models,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff9d09bc-26d4-49bc-88bd-29d2442f8e7d · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Enhancing multimodal emotion recogni- tion through asr error compensation and llm fine-tuning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3087c5ac-ee59-4db2-a8af-9f2cea5ff593 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Large language model based generative er- ror correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccce4f50-bd73-43ed-aa9b-42db47a99720 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Wavcaps: A chatgpt-assisted weakly-labelled au- dio captioning dataset for audio-language multimodal research,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 236dcc7a-36b2-4c24-a825-66e6e45d5d81 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Training audio captioning models without audio,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f2c1027-8a5f-4574-b1d5-0344ca928ad4 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Aligncap: Aligning speech emotion captioning to human preferences,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c85479cf-1aea-4d70-b757-9613e146ce3c · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Clap4emo: Chatgpt-assisted speech emotion retrieval with natural language supervision,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb537344-8ba7-4119-bf11-b9784529c642 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Lora: Low-rank adaptation of large language mod- els,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b977f26b-6c85-4318-86a5-00192b7efe7a · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition The information bottleneck method,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95e23167-2ebd-4a8e-a363-70e84235377a · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Deep variational information bottleneck,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59cd97ba-1cbb-434c-ba5f-e193df4c84ff · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Auto-encoding variational bayes,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b6ec395-0850-474e-a5ea-b1d4d6ecf308 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Bayesian learning for deep neural network adapta- tion,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fdb40e3d-b583-4f48-a330-d9b62c74d6d1 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Bayesian learning of lf-mmi trained time delay neu- ral networks for speech recognition,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7232bbc3-b526-43a5-a79d-00a6b2c17bdd · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9bdb7cdb-4b49-493b-b74c-8486e530d22f · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Distilhubert: Speech representation learning by layer-wise distillation of hidden-unit bert,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1fd39b26-c849-4e36-a363-4786a9f4459c · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Llama-omni: Seamless speech interaction with large language models,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28d1e536-a0d7-4aa7-8387-254eb5b1cb90 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition The llama 3 herd of models,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36a18816-7956-4c8a-bc2c-b5942ffe45ab · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speechcraft: A fine-grained expressive speech dataset with natural language description,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57555cf5-8236-49a1-9e33-98d395790851 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Gigaspeech: An evolving, multi-domain asr cor- pus with 10,000 hours of transcribed audio,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7dc4a628-7d37-4859-93f2-ed463caf19e9 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Emosec: Emotion recognition from scene context,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b489156c-9c89-4357-979e-e40cfd8ed274 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition IEMOCAP: interactive emotional dyadic motion capture database,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7b0602d-a9fc-4964-902c-23ed82a8486f · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Meld: A multimodal multi-party dataset for emo- tion recognition in conversations,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55dd8869-81ed-4273-ad07-36cb3b9553fc · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0bcd29b-7652-4683-b49b-bea1efd68a1f · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Disentanglement network: Disentangle the emo- tional features from acoustic features for speech emotion recogni- tion,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d661b436-49a3-4fa5-b7e8-336b0f139d53 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition How first-and second-language emotion words influence emotion perception in swedish–english bilinguals,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23b390d3-f5b6-48f9-9638-6921520dae5b · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Blsp-emo: Towards empathetic large speech- language models,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7775b9ca-23b7-4ccc-ba9a-e560c35769ab · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speech emotion recognition using decomposed speech via multi-task learning,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e7d8c3a-de36-4af3-985c-fc9407c83a38 · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Wavlm: Large-scale self-supervised pre-training for full stack speech processing,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c07f64d-df15-40c4-9fa8-2de47749a25f · outbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d52dbc93-a0e9-4856-8893-20c920284adb · inbound
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.