Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:11.848020Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.02088.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:11.848020Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:39:46.886300Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T11:40:12.056485Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cbc91fd8-2696-431c-83e7-60c230cb87b5 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Early SER relied on hand-crafted features but struggled with real- world generalization [2]
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0fd2927-ab92-4ea7-b6d2-df576f3a353d · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0e4715b-d7fe-41df-a432-fdb412559c61 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 The hidden states of the last layerLof the text encoder are denoted byZ L T (j)for positions j= 1,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 568c6db1-c9bf-445c-b9e1-92b773217c62 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33001ac9-da6f-4328-9c06-f8c518f3b824 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 We report results for unimodal speech models, bi- modal fusion with text, prosodic and spectral feature integra- tion
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a15c6334-fba5-4b40-977e-75b1be96e4a9 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Our evaluation of unimodal models demonstrated the strong performance of Whisper and XEUS, highlighting their robustness for SER in spontaneous speech
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a356d0a-2d75-4065-b968-ab684529c6e7 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 We also thank the Artificial Intelligence Lab at Re- cod.ai, the Institute of Computing, University of Campinas
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd61b554-8a33-449d-b640-1d6fbccfcf00 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Affective computing mit press,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95edbc09-aad6-4d1d-ae6e-d62d26767b08 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Iemocap: Interactive emotional dyadic motion capture database,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 575e1dbe-4310-41de-a8e4-6c02731c8573 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Every rating matters: Joint learning of subjective labels and individual annotators for speech emotion classification,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a988f1c-8399-4005-9333-b47543d5203f · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Speech emotion recognition using self-supervised features,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa89ce1-a259-4b5d-a722-c3b2640b180e · outbound
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8ca3961-213f-4b2b-91ec-7514c4909d14 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Speech emotion recognition with multi-task learning,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af2b8935-da7c-4af9-ab15-bbac4f956498 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Improving speech emotion recogni- tion using self-supervised learning with domain-specific audiovi- sual tasks,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc666eda-a686-4562-ac18-c5b4e6d3a2df · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Odyssey 2024-speech emotion recognition challenge: Dataset, baseline framework, and results,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4a34ba8-918f-4d92-a487-dfd354d892a4 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cafa7854-99b7-47db-b70b-92195a3418cd · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82513e37-8794-4a55-986b-5d009f0a68c3 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 814768dd-203d-4df1-bc62-9d272e49614e · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Robust speech recognition via large-scale weak supervision,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609fc84d-5d93-4fe4-a39b-a0eb1dd6b082 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Towards robust speech representation learning for thousands of languages,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb018c13-febc-43f5-9629-7294bac6d95a · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 A robustly optimized BERT pre-training approach with post-training,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa5f7fc5-f77c-4c38-8d0c-800414f835c5 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 SeqAug: Sequential Feature Resampling as a modality agnostic augmentation method
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c227393-77bb-4ef8-ad33-c932444ebe03 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing cross-language multimodal emotion recognition with dual attention transform- ers,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57f0aca7-3596-43b6-8b87-bf654988b716 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Ced: Con- sistent ensemble distillation for audio tagging,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfa2a6ae-afd7-4b88-a60d-6e728af99904 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adf8d24e-0b38-43d3-8eb2-5c3958ba2a54 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Graph attention networks,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73217c7a-5c82-49d8-8a29-7fffc3b96f74 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 GLU Variants Improve Transformer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e2d95a2-2ebd-46ec-8686-ddc1d47f9b94 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Espnet: End-to-end speech pro- cessing toolkit,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 571a2128-c921-4234-92c9-21618fd4f1e6 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Less is more: Accu- rate speech recognition & translation without web-scale data,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738e5e1b-7056-46d6-ab15-de24cbf70da5 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 1st place solution to odyssey emotion recognition chal- lenge task1: Tackling class imbalance problem,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2ef2ada-153e-4b7b-b0f4-8a58e4ec2f0f · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Fundamental frequency ex- traction in speech emotion recognition,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b59c1520-ded0-4b57-84eb-320e3a08162b · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Autoregressive neural f0 model for statistical parametric speech synthesis,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9bb7dae-4b5b-4664-94b4-2ea372393358 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Rmvpe: A robust model for vocal pitch estimation in polyphonic music,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5824415-23d1-4220-ac0e-c03afc0ea926 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing skin can- cer diagnosis using swin transformer with hybrid shifted window- based multi-head self-attention and swiglu-based mlp,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61f5b1ec-e191-4293-afb4-86cceefebbc0 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Searching for activation functions,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dec73175-a10b-4b58-8360-d52c414541d0 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Decoupled weight de- cay regularization,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072fc2c8-d40f-4cfd-8b25-d6b33dbf5e32 · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Panns: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b66fe9de-cb41-4a86-b269-35a9afcb68dc · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Focal loss for dense object detection,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75fdba96-0af3-4525-93be-145621894b8f · outbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Multilingual E5 Text Embeddings: A Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0fd2927-ab92-4ea7-b6d2-df576f3a353d · inbound
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.