Pith. sign in

Paper Citation Record · LEDGER

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition

As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2505.23236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23236 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:15.683108Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:10.171701Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:54:15.802818Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy53
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 299bf55c-e888-42e3-b1e7-39169c2fa4e4 · outbound

This paper cites Despite decades of promising advance- ments, most research [1–4] has primarily focused on classify- ing speech into single, discrete emotion categories.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Despite decades of promising advance- ments, most research [1–4] has primarily focused on classify- ing speech into single, discrete emotion categories

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.195734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:09.919825Z digest=sha256:7a63be195449e918936608fcf2337917454ef6fceebb0518dab96ba9ab80016f

Observation a9f3eab2-c1fb-461b-9722-873266a0ca68 · outbound

This paper cites In contrast, previous research paid attention to either feature disentan- glement [6–17] or fine-grained emotion descriptor prediction [5, 23–26] alone.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition In contrast, previous research paid attention to either feature disentan- glement [6–17] or fine-grained emotion descriptor prediction [5, 23–26] alone

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.004534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.008381Z digest=sha256:aa6f16ed0fb144eabdabcd0e6727ca372721171ce04154223495d3f78a110b8f

Observation 3712781b-f7e8-434e-ad74-b333098e0403 · outbound

This paper cites In contrast, prior research employs external ASR models to generate transcripts before performing SER [12–17], which may lead to low SER performance affected by ASR er- rors.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition In contrast, prior research employs external ASR models to generate transcripts before performing SER [12–17], which may lead to low SER performance affected by ASR er- rors

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:26.823215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.078784Z digest=sha256:0814e453f71ce63fe37294b410b1ab79e73d115832afc830848706a61c5a180f

Observation d52dbc93-a0e9-4856-8893-20c920284adb · outbound

This paper cites Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:54:15.893612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.171701Z digest=sha256:36a3b9d8c3e1cff7e5ae944a8b08ceb774e3011f451b0f8fc6131c2a86956f26

Observation 47965573-3fe0-424a-96d6-35b3c71ee6fe · outbound

This paper cites Alternating Multi-task Fine-tuning A parameter-efficient fine-tuning method known as low-rank adaptation, or LoRA [27], is widely employed to fine-tune the LLM-based models.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Alternating Multi-task Fine-tuning A parameter-efficient fine-tuning method known as low-rank adaptation, or LoRA [27], is widely employed to fine-tune the LLM-based models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:26.664845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.279559Z digest=sha256:c7dcac7e650e0703bc51b4cc9ce662136a2399e8331b7e473492f4c889293218

Observation c68910f4-1435-4639-b716-444836c08e28 · outbound

This paper cites 14200220, 14200021, 14200324, Innovation Technology Fund grant No.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition 14200220, 14200021, 14200324, Innovation Technology Fund grant No

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.664004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.722341Z digest=sha256:45ec79a6c68821c3e6bdfcdb2d2e1e375d98d7d8b98e3754ed3765394c60738a

Observation 26923b06-ce5f-41c0-b26e-618f3bf866b5 · outbound

This paper cites 1, the architecture of our approach consists of a HuBERT encoder, two feature disentanglement blocks, two feature adapters and an LLM decoder.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition 1, the architecture of our approach consists of a HuBERT encoder, two feature disentanglement blocks, two feature adapters and an LLM decoder

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:26.344816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.437938Z digest=sha256:b60f9b975847b338fee1e11ba0ec22b1a6747f64794046c6bcb68985ad7bfa47

Observation 4585293c-632d-4576-aea5-21e6d759d208 · outbound

This paper cites excited” with “happy.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition excited” with “happy

Reference 8

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:54:26.126722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.525020Z digest=sha256:7facbc265045336db5fa71598f57f1da52718d1402aed363e2dcfa27b9db3c9d

Observation b47ff2dc-0a09-458c-8e03-d54d1e487898 · outbound

This paper cites Fine-grained SED features are dis- entangled from HuBERT SSL representations via alternating LLM fine-tuning to joint SER-SED prediction and ASR tasks.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Fine-grained SED features are dis- entangled from HuBERT SSL representations via alternating LLM fine-tuning to joint SER-SED prediction and ASR tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.890238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.608679Z digest=sha256:c4b2d8eec7f3809edf41025ff67e40b6d68a014e0f4befb28d60f8441bbba794

Observation 8e3704ce-b016-4f73-bdcf-bc0477927b63 · outbound

This paper cites Leveraging speech ptm, text llm, and emotional tts for speech emotion recognition,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Leveraging speech ptm, text llm, and emotional tts for speech emotion recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.377744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.694405Z digest=sha256:f655c5783b38bc8e7a0ff6054f1f4ee00ce29d4d1482973611869776bed7cd37

Observation 716a1476-f4f9-46f6-9363-dbd705f1728c · outbound

This paper cites Emotion neural transducer for fine-grained speech emotion recognition,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Emotion neural transducer for fine-grained speech emotion recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.426744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.821947Z digest=sha256:c6620ea749dc1a3b72580c2490292d8e97d9672e12ea0c13ebfef0dd09b5efa8

Observation 4f662225-8bba-4bed-be4d-3c4337f3fdaa · outbound

This paper cites Generalization of self-supervised learning- based representations for cross-domain speech emotion recogni- tion,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Generalization of self-supervised learning- based representations for cross-domain speech emotion recogni- tion,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.211711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.899986Z digest=sha256:ddaf3a4c43bb1288a4f42ef2aa0977a5729697ae8bdae09cb85174c20ca96654

Observation 94d8e0c1-0e4a-45f4-9df4-ae74b8e0a9d3 · outbound

This paper cites Noise-robust speech emotion recognition us- ing shared self-supervised representations with integrated speech enhancement,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Noise-robust speech emotion recognition us- ing shared self-supervised representations with integrated speech enhancement,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.981964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.004641Z digest=sha256:f769a749c0c7e28a24435d700a8e97d7ed94c53b1d807ae048d8d7e7ac9e2eba

Observation 13c806fb-b9d8-4308-b06e-2c72e5de67df · outbound

This paper cites Emix: a data augmentation method for speech emotion recognition,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Emix: a data augmentation method for speech emotion recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.750454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.096176Z digest=sha256:0e2414ad3500fd07acf1c383bfd78eb6faf8cf0cb20444dd280c4d7f12f68d84

Observation 9aac4409-4fb3-4d54-90b8-195bc3af614c · outbound

This paper cites Secap: Speech emotion captioning with large lan- guage model,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Secap: Speech emotion captioning with large lan- guage model,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.475284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.195114Z digest=sha256:8c2f5b5b958b5dbf72f1aaf303e52c23409ace0f28cbbad3aa8c27fde2ac5cd3

Observation 740cd773-f56f-4f3a-a9d8-b7540c1ba9fd · outbound

This paper cites Disentangling textual and acoustic features of neural speech representations,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Disentangling textual and acoustic features of neural speech representations,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.025148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.376928Z digest=sha256:7b465f4e5ed3251e5f6a034a82a1b550ff42bad154c5d06821ab34bc7d27b74e

Observation 3040545b-d819-40e2-a2e2-ee840debd6a8 · outbound

This paper cites Variational information bottleneck for effective low- resource audio classification,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Variational information bottleneck for effective low- resource audio classification,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.837850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.480535Z digest=sha256:dc6d571c682a985f2360520000f125ec55dec9e64ec80e92fb1d60c119fcad9b

Observation b4a85e3a-78f4-4da5-908f-b9324d49f97e · outbound

This paper cites Multi-modal emotion recognition using multiple acoustic features and dual cross-modal transformer,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Multi-modal emotion recognition using multiple acoustic features and dual cross-modal transformer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.595496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.555204Z digest=sha256:d9223306fa825248acdec5b1e403209ee065576279d8677019563d92144e566f

Observation 7702ef33-39a3-426b-a7ce-8948dfbd3e2b · outbound

This paper cites Wavllm: Towards robust and adaptive speech large language model,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Wavllm: Towards robust and adaptive speech large language model,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.033794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:12.971182Z digest=sha256:d3eeb66bbb385c4807bc3730445cc936e71bccdb82c9a1f7fa4dc837290d081b

Observation 6bf1cdef-f5d6-49c1-8b1a-3573adcae58d · outbound

This paper cites Variational information bottleneck based regular- ization for speaker recognition.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Variational information bottleneck based regular- ization for speaker recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.172317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.790971Z digest=sha256:c149755654ba0a822aca6a9ea58f3442867eb041fd5fdb51165bf04d4109f230

Observation 59b65e9c-a181-48dd-b492-c7d643c712b5 · outbound

This paper cites Speech emotion recognition combining acous- tic features and linguistic information in a hybrid support vector machine-belief network architecture,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speech emotion recognition combining acous- tic features and linguistic information in a hybrid support vector machine-belief network architecture,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.035813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.875837Z digest=sha256:352404f5a7bee653a46f2f14ee41ed2633c7dfe855c9302477c5d978410616e6

Observation 624ceb2a-ce2b-4856-a431-8a132c52e55a · outbound

This paper cites Speech emotion recognition with multi-level acoustic and semantic information extraction and interaction,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speech emotion recognition with multi-level acoustic and semantic information extraction and interaction,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.769508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:11.972808Z digest=sha256:cd39c7e53cb650d5bb34244f46f32bb627c2ad358df70a15f1cc6f459151f937

Observation ed8e0c42-30f3-4929-a271-fc278ac44fe4 · outbound

This paper cites Multimodal emotion recognition with high-level speech and text features,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Multimodal emotion recognition with high-level speech and text features,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.313076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:12.128689Z digest=sha256:fbe1e6722348e70b5a7d3cb4d55e01adff3f639306059d9c178bd2c293023739

Observation 5a8d77da-1007-4932-a324-c5433791a6ca · outbound

This paper cites Frontend attributes disentanglement for speech emotion recognition,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Frontend attributes disentanglement for speech emotion recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.101428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:12.234483Z digest=sha256:d25cc4752a67dce1cd07b554c8fb1c7be1d2941b14cf28afa83611b37b248d99

Observation a882d128-dc8f-49e8-ad31-34f87f78b8d4 · outbound

This paper cites Speech emotion recognition with acoustic and lexi- cal features,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speech emotion recognition with acoustic and lexi- cal features,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.934431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:12.531610Z digest=sha256:c44029d67f214a92e0e83c74ea0981e5cd51e7109d68d1bbbd18669bf3f9cf4d

Observation b906f673-8626-4d6a-8ee2-9e72177f15b3 · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Salmonn: Towards generic hearing abilities for large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.753545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:12.780806Z digest=sha256:49a79629ab002e55325957b9e60aa5fbf470a3c0451c68bd6b486504e926d91e

Observation afc23d35-1a44-4f0d-8a8f-97ef8a5fb68b · outbound

This paper cites an unresolved cited work.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:54:26.472177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.346016Z digest=sha256:da2e06eabeaaf79d966d85d3d7236a269e78db7427f81c6c7d3548effcbfdf65

Observation e3837274-6619-42bd-97ba-967e69161b12 · outbound

This paper cites Qwen-audio: Advancing universal audio under- standing via unified large-scale audio-language models,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Qwen-audio: Advancing universal audio under- standing via unified large-scale audio-language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.396711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:12.878612Z digest=sha256:512ef53386c7d2de41a670c2194faa57393c346d200403e2b06c6b3931dd7c3e

Observation ff9d09bc-26d4-49bc-88bd-29d2442f8e7d · outbound

This paper cites Enhancing multimodal emotion recogni- tion through asr error compensation and llm fine-tuning,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Enhancing multimodal emotion recogni- tion through asr error compensation and llm fine-tuning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:20.704496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.051955Z digest=sha256:e61d7c7b687458531b3e74d79c944b7478e0aa8fa44b6d242c98db6bef9ed663

Observation 3087c5ac-ee59-4db2-a8af-9f2cea5ff593 · outbound

This paper cites Large language model based generative er- ror correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Large language model based generative er- ror correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:19.935739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.128936Z digest=sha256:ad4c1aa26cc6622d0fcd0b86ef9d489483ae8038ddb1962732efc63cb2dd8e6a

Observation ccce4f50-bd73-43ed-aa9b-42db47a99720 · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled au- dio captioning dataset for audio-language multimodal research,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Wavcaps: A chatgpt-assisted weakly-labelled au- dio captioning dataset for audio-language multimodal research,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:19.774077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.224482Z digest=sha256:4c4f1e906cffd2cb501061b1fbd86821e3618dc26b5133e4103e33593b729891

Observation 236dcc7a-36b2-4c24-a825-66e6e45d5d81 · outbound

This paper cites Training audio captioning models without audio,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Training audio captioning models without audio,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:19.631531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.289487Z digest=sha256:fd86b454501954f01c14be133da82f737899e6baa0f75be1bc1aab327bdaeea0

Observation 8f2c1027-8a5f-4574-b1d5-0344ca928ad4 · outbound

This paper cites Aligncap: Aligning speech emotion captioning to human preferences,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Aligncap: Aligning speech emotion captioning to human preferences,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:19.458611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.369212Z digest=sha256:a7008e010022e44e0bfe99d2164066667e96d5c399bf68a04997567a8bbbcc70

Observation c85479cf-1aea-4d70-b757-9613e146ce3c · outbound

This paper cites Clap4emo: Chatgpt-assisted speech emotion retrieval with natural language supervision,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Clap4emo: Chatgpt-assisted speech emotion retrieval with natural language supervision,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:19.334182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.461108Z digest=sha256:bf083f8cd999d522c2d0f1223140255861f434806a42e25f69c6f7b55b11ed2f

Observation cb537344-8ba7-4119-bf11-b9784529c642 · outbound

This paper cites Lora: Low-rank adaptation of large language mod- els,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Lora: Low-rank adaptation of large language mod- els,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:19.240431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.593844Z digest=sha256:a4b98a2f4bbe1b1714bd611c26cb13c13ef2106dd9558ca067b0ffa0da98ecbd

Observation b977f26b-6c85-4318-86a5-00192b7efe7a · outbound

This paper cites The information bottleneck method,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition The information bottleneck method,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:19.139601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.669080Z digest=sha256:0747a7746fdec4c1e9fd3ec4fe3d2d374f4d52475aed1659608f72ff9e80d6b5

Observation 95e23167-2ebd-4a8e-a363-70e84235377a · outbound

This paper cites Deep variational information bottleneck,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Deep variational information bottleneck,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:19.038394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.753112Z digest=sha256:ff82b98ae21c047ffa705f88d76256cd4204f2cfbbba47fb690c1b51cb0f4be8

Observation 59cd97ba-1cbb-434c-ba5f-e193df4c84ff · outbound

This paper cites Auto-encoding variational bayes,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Auto-encoding variational bayes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:18.880438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.829934Z digest=sha256:2af689f7679b0d811b88396706da0a3c9f17e5aa119003c00c2fc254e2d2b7d0

Observation 0b6ec395-0850-474e-a5ea-b1d4d6ecf308 · outbound

This paper cites Bayesian learning for deep neural network adapta- tion,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Bayesian learning for deep neural network adapta- tion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:18.749934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:13.900142Z digest=sha256:4c180d552339da19e52bcd818a3edb2d9d95ec358bb8cc62028b0b2cb5d42f7b

Observation fdb40e3d-b583-4f48-a330-d9b62c74d6d1 · outbound

This paper cites Bayesian learning of lf-mmi trained time delay neu- ral networks for speech recognition,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Bayesian learning of lf-mmi trained time delay neu- ral networks for speech recognition,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:18.624336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.009523Z digest=sha256:9990cbe0302a0eccefe92ec95247fa3e2507fb54eac629a4f711a5b9d7717cd7

Observation 7232bbc3-b526-43a5-a79d-00a6b2c17bdd · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:18.483936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.119442Z digest=sha256:af64463ab886ff2872e811df630306871ad6540db3ea8068c29f010426c8dc98

Observation 9bdb7cdb-4b49-493b-b74c-8486e530d22f · outbound

This paper cites Distilhubert: Speech representation learning by layer-wise distillation of hidden-unit bert,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Distilhubert: Speech representation learning by layer-wise distillation of hidden-unit bert,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:18.295954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.253340Z digest=sha256:1838a3aeb13f79baef0f9a58b05da047c481b2a7236d0e87adc2cec30200436e

Observation 1fd39b26-c849-4e36-a363-4786a9f4459c · outbound

This paper cites Llama-omni: Seamless speech interaction with large language models,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Llama-omni: Seamless speech interaction with large language models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:18.076092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.362188Z digest=sha256:4fd9224bf82808aeecf3ac8e35452eeaa2ac78e986a72f3025c3a644af35896d

Observation 28d1e536-a0d7-4aa7-8387-254eb5b1cb90 · outbound

This paper cites The llama 3 herd of models,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition The llama 3 herd of models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:17.869159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.473295Z digest=sha256:a06cd15f48f35f5e48ad1036bf93891e70f1d77b1e909162efda9de1af5903f1

Observation 36a18816-7956-4c8a-bc2c-b5942ffe45ab · outbound

This paper cites Speechcraft: A fine-grained expressive speech dataset with natural language description,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speechcraft: A fine-grained expressive speech dataset with natural language description,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:17.708347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.569750Z digest=sha256:77bca1ac5e429eb378f10055f43ecc74e832d33c4859949d23e86209d954d1b9

Observation 57555cf5-8236-49a1-9e33-98d395790851 · outbound

This paper cites Gigaspeech: An evolving, multi-domain asr cor- pus with 10,000 hours of transcribed audio,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Gigaspeech: An evolving, multi-domain asr cor- pus with 10,000 hours of transcribed audio,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:17.535276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.677391Z digest=sha256:680095db66309b24875c775158d0bb132c35037494e6bd69e05e8d9a79a7e61f

Observation 7dc4a628-7d37-4859-93f2-ed463caf19e9 · outbound

This paper cites Emosec: Emotion recognition from scene context,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Emosec: Emotion recognition from scene context,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:17.365638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.750172Z digest=sha256:3c50cdfe98ea4867e89846115c6f9bdc09c6d237ef7a207de1475a09ff78bdb6

Observation b489156c-9c89-4357-979e-e40cfd8ed274 · outbound

This paper cites IEMOCAP: interactive emotional dyadic motion capture database,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition IEMOCAP: interactive emotional dyadic motion capture database,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:17.207672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.880607Z digest=sha256:dc915ab79600efeeedd14dd09f3a7fbc5aee69efcd6c48db0821c058e561fb93

Observation e7b0602d-a9fc-4964-902c-23ed82a8486f · outbound

This paper cites Meld: A multimodal multi-party dataset for emo- tion recognition in conversations,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Meld: A multimodal multi-party dataset for emo- tion recognition in conversations,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:17.009952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:14.959046Z digest=sha256:9dabdb6111a45abe006a0440c84f2b6e2567e7dd7fae0d9361fdbd7c5235fd6f

Observation 55dd8869-81ed-4273-ad07-36cb3b9553fc · outbound

This paper cites Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:16.836603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:15.042727Z digest=sha256:f436c327d8db24759490c51dd72efe5bc0bdb2067a8a0d71fc891a4f2f02573f

Observation d0bcd29b-7652-4683-b49b-bea1efd68a1f · outbound

This paper cites Disentanglement network: Disentangle the emo- tional features from acoustic features for speech emotion recogni- tion,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Disentanglement network: Disentangle the emo- tional features from acoustic features for speech emotion recogni- tion,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.246200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:15.161492Z digest=sha256:9e83bb47d26bfe2ca4aa1f47e232e190fc1e9869e57a505ce1000bb71875387a

Observation d661b436-49a3-4fa5-b7e8-336b0f139d53 · outbound

This paper cites How first-and second-language emotion words influence emotion perception in swedish–english bilinguals,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition How first-and second-language emotion words influence emotion perception in swedish–english bilinguals,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:16.649505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:15.268678Z digest=sha256:cbe5f8d0ccd8c339d37d8659f4c83e6ffaa45bd8ca8a2949f8696e7f10795b02

Observation 23b390d3-f5b6-48f9-9638-6921520dae5b · outbound

This paper cites Blsp-emo: Towards empathetic large speech- language models,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Blsp-emo: Towards empathetic large speech- language models,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:16.469881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:15.380851Z digest=sha256:5f6cad2436ff28d02b3a203fb2f5f0c084d1afc49497c5bc2834f2038ab476b1

Observation 7775b9ca-23b7-4ccc-ba9a-e560c35769ab · outbound

This paper cites Speech emotion recognition using decomposed speech via multi-task learning,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Speech emotion recognition using decomposed speech via multi-task learning,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.505430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:15.507673Z digest=sha256:cdab51eefe021a1621045697075b2a99382d06a8b19d63c2e25a70c332cf116b

Observation 4e7d8c3a-de36-4af3-985c-fc9407c83a38 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:16.263774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:15.581933Z digest=sha256:5ba430beb3c852c2511fe37d7afcd20602ac0f1a2bca065ab71c50827aa8ee6c

Observation 5c07f64d-df15-40c4-9fa8-2de47749a25f · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:16.098034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:15.683108Z digest=sha256:d61cf61ad5ad87ae08891a50338e140535a011d553377c791546fbc969bc435f

Pith citing papers

Observation d52dbc93-a0e9-4856-8893-20c920284adb · inbound

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition cites this paper.

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:54:15.893612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:10.171701Z digest=sha256:36a3b9d8c3e1cff7e5ae944a8b08ceb774e3011f451b0f8fc6131c2a86956f26