Pith. sign in

Paper Citation Record · LEDGER

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.02088.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02088 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:11.848020Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:39:46.886300Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:40:12.056485Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbc91fd8-2696-431c-83e7-60c230cb87b5 · outbound

This paper cites Early SER relied on hand-crafted features but struggled with real- world generalization [2].

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Early SER relied on hand-crafted features but struggled with real- world generalization [2]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.900210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:46.828246Z digest=sha256:a79fdb4ae3bf2577e191bd9a3b891e11b75aee091dd9a3b9cb9306abca15cc7d

Observation b0fd2927-ab92-4ea7-b6d2-df576f3a353d · outbound

This paper cites Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:12.074820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:46.886300Z digest=sha256:0418c135d6970834bc145858f29fc5cc2f106454ac6679ae39e21a5d4a8d589b

Observation d0e4715b-d7fe-41df-a432-fdb412559c61 · outbound

This paper cites The hidden states of the last layerLof the text encoder are denoted byZ L T (j)for positions j= 1,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 The hidden states of the last layerLof the text encoder are denoted byZ L T (j)for positions j= 1,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.725410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:46.937316Z digest=sha256:3b0d6fd57855568f5c8a5a85a42034ff780b51724795a2d3644278b9fe759906

Observation 568c6db1-c9bf-445c-b9e1-92b773217c62 · outbound

This paper cites an unresolved cited work.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:40:58.649776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:46.962061Z digest=sha256:51fed33567ab22ce2a863ea8333c43afe75174e7eca5ad58a0793a4e509cd2e4

Observation 33001ac9-da6f-4328-9c06-f8c518f3b824 · outbound

This paper cites We report results for unimodal speech models, bi- modal fusion with text, prosodic and spectral feature integra- tion.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 We report results for unimodal speech models, bi- modal fusion with text, prosodic and spectral feature integra- tion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.509032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:47.019513Z digest=sha256:379e90ca8856cef0ea77b9e166e9542bf8b5fb875d7ec35510bec221af07fa61

Observation a15c6334-fba5-4b40-977e-75b1be96e4a9 · outbound

This paper cites Our evaluation of unimodal models demonstrated the strong performance of Whisper and XEUS, highlighting their robustness for SER in spontaneous speech.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Our evaluation of unimodal models demonstrated the strong performance of Whisper and XEUS, highlighting their robustness for SER in spontaneous speech

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.356842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:47.249227Z digest=sha256:2e79b78b95493e08444c0ee6dbfd48b5535422e734be7bcdfc1b492e486775c3

Observation 8a356d0a-2d75-4065-b968-ab684529c6e7 · outbound

This paper cites We also thank the Artificial Intelligence Lab at Re- cod.ai, the Institute of Computing, University of Campinas.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 We also thank the Artificial Intelligence Lab at Re- cod.ai, the Institute of Computing, University of Campinas

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.217464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:47.380433Z digest=sha256:04fadd4d491f1bf0959652c50c12a52876f97c468dbdda10e99e0e58d14c31fe

Observation bd61b554-8a33-449d-b640-1d6fbccfcf00 · outbound

This paper cites Affective computing mit press,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Affective computing mit press,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.093220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:47.456692Z digest=sha256:8228775a422997d936aac227f1cba95ac99f48affcf549209087fd5d5881f4e7

Observation 95edbc09-aad6-4d1d-ae6e-d62d26767b08 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Iemocap: Interactive emotional dyadic motion capture database,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:47.566911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:47.566911Z digest=sha256:f81a170115f38ba99246dedfd18ccca18a5bdf66944f575ae7a1f9aadc4f24f1

Observation 575e1dbe-4310-41de-a8e4-6c02731c8573 · outbound

This paper cites Every rating matters: Joint learning of subjective labels and individual annotators for speech emotion classification,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Every rating matters: Joint learning of subjective labels and individual annotators for speech emotion classification,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.052371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:48.879938Z digest=sha256:f4e8982a83e16f113065c3b2e147a9cbc2263681351c307ced33fa222e3613a6

Observation 5a988f1c-8399-4005-9333-b47543d5203f · outbound

This paper cites Speech emotion recognition using self-supervised features,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Speech emotion recognition using self-supervised features,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:51.667321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:51.667321Z digest=sha256:841cd1add6b1d57dca2729e8f5c5b308c593884d44ab20e8c2cf24d071f545fb

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.908096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:51.789972Z digest=sha256:9d2cb910a6e34d73abca3cee2135b9aec1279974d75f9592e1d5b740f23a5c75

Observation f8ca3961-213f-4b2b-91ec-7514c4909d14 · outbound

This paper cites Speech emotion recognition with multi-task learning,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Speech emotion recognition with multi-task learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.829289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:51.886757Z digest=sha256:d0b53d9c91b55a5760434e70a36bd51bbec09d013ff19961efdcb5f6dcc4d7e0

Observation af2b8935-da7c-4af9-ab15-bbac4f956498 · outbound

This paper cites Improving speech emotion recogni- tion using self-supervised learning with domain-specific audiovi- sual tasks,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Improving speech emotion recogni- tion using self-supervised learning with domain-specific audiovi- sual tasks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.755073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:51.999237Z digest=sha256:bf6dc7cc3175d6cb4d705697110a610dc660392a331f096beff9366396eccfed

Observation dc666eda-a686-4562-ac18-c5b4e6d3a2df · outbound

This paper cites Odyssey 2024-speech emotion recognition challenge: Dataset, baseline framework, and results,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Odyssey 2024-speech emotion recognition challenge: Dataset, baseline framework, and results,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.518830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:52.100807Z digest=sha256:049b7a1d4282ddf244d5b7fbcec5614c3e8f0301db4cbace2e7b929536b878f6

Observation e4a34ba8-918f-4d92-a487-dfd354d892a4 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.179588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.179588Z digest=sha256:ccb1111f923eabfd4d01bf633c167b74b24e9fe6e08fb83851f8aac8b873cda4

Observation cafa7854-99b7-47db-b70b-92195a3418cd · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.259720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.259720Z digest=sha256:420de1deccbcc35d4ae47fbf929c5114601900fac58d2f300a2259a3e80c66c7

Observation 82513e37-8794-4a55-986b-5d009f0a68c3 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.329379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.329379Z digest=sha256:da7833dcc3cc2c21595faeef8424a8d4f294776467666a0974c50454436a4bc2

Observation 814768dd-203d-4df1-bc62-9d272e49614e · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Robust speech recognition via large-scale weak supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.399609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.399609Z digest=sha256:2bda44ded3ee0cc5f790ade20520d37e527615053a40b96b45310eafc470f51e

Observation 609fc84d-5d93-4fe4-a39b-a0eb1dd6b082 · outbound

This paper cites Towards robust speech representation learning for thousands of languages,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Towards robust speech representation learning for thousands of languages,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.465169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.465169Z digest=sha256:8e3b46975638445402c1b5bc00a6da60b94086ae67c7d061ec5ffbba316a9a8f

Observation fb018c13-febc-43f5-9629-7294bac6d95a · outbound

This paper cites A robustly optimized BERT pre-training approach with post-training,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 A robustly optimized BERT pre-training approach with post-training,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:55.952387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:52.554243Z digest=sha256:c24db41f4bd27bde7993fe1766475d92439d7dbe0b856ddb32ff9a83eb9a259a

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:12.025604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:52.777099Z digest=sha256:190258e33c8f979cea17eda54e77e88b1b8098a7c7e6c98c94f0404d4b313191

Observation 6c227393-77bb-4ef8-ad33-c932444ebe03 · outbound

This paper cites Enhancing cross-language multimodal emotion recognition with dual attention transform- ers,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing cross-language multimodal emotion recognition with dual attention transform- ers,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:54.883253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:53.019255Z digest=sha256:9967d52ebb5f294566d46768f99b47162e65336afbd259b43141c49185b9b8a4

Observation 57f0aca7-3596-43b6-8b87-bf654988b716 · outbound

This paper cites Ced: Con- sistent ensemble distillation for audio tagging,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Ced: Con- sistent ensemble distillation for audio tagging,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:46.567879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:53.103227Z digest=sha256:e1508df0f4169de14a8f0d24ea71e92316923512ce46fff86273e293293bdf84

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:11.978878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:53.213116Z digest=sha256:94ea9f5277b24b9c59038e1e4167c58ac748dd7cff4a640a5c7d95df8abfa6ba

Observation adf8d24e-0b38-43d3-8eb2-5c3958ba2a54 · outbound

This paper cites Graph attention networks,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Graph attention networks,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:36.044619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:53.374984Z digest=sha256:94078c98c781144babab3ba81347c761abf55276b37115f0a16ebd37beecd41b

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:07.102687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:07.102687Z digest=sha256:59bf91a46356f69c07d9fae5f0eee51db1cd64af9c5d5c422442bfcfd81399ff

Observation 2e2d95a2-2ebd-46ec-8686-ddc1d47f9b94 · outbound

This paper cites Espnet: End-to-end speech pro- cessing toolkit,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Espnet: End-to-end speech pro- cessing toolkit,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:35.916452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:40:08.043740Z digest=sha256:2711c38c504aac033d92fc7d474e2a7874ed578f7ad63da1c318ad151255527f

Observation 571a2128-c921-4234-92c9-21618fd4f1e6 · outbound

This paper cites Less is more: Accu- rate speech recognition & translation without web-scale data,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Less is more: Accu- rate speech recognition & translation without web-scale data,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.312368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.312368Z digest=sha256:2e102b8ef231ff44c8a25ce3f5533d38d6c133341be9d64b2eceb4bff5fdf469

Observation 738e5e1b-7056-46d6-ab15-de24cbf70da5 · outbound

This paper cites 1st place solution to odyssey emotion recognition chal- lenge task1: Tackling class imbalance problem,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 1st place solution to odyssey emotion recognition chal- lenge task1: Tackling class imbalance problem,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:55.870025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:40:09.642774Z digest=sha256:af67dfc8e4068457d35e6ab3e60158ddf6509410f71bf8b347aa43bd70bd1469

Observation c2ef2ada-153e-4b7b-b0f4-8a58e4ec2f0f · outbound

This paper cites Fundamental frequency ex- traction in speech emotion recognition,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Fundamental frequency ex- traction in speech emotion recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:35.694011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:40:09.779236Z digest=sha256:582eb261908e4eb5ea2f0880ce66364bff1c58162e1da9014ca49f020d4a0012

Observation b59c1520-ded0-4b57-84eb-320e3a08162b · outbound

This paper cites Autoregressive neural f0 model for statistical parametric speech synthesis,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Autoregressive neural f0 model for statistical parametric speech synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:31.875345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:40:09.869009Z digest=sha256:2b0449e09ae61402c5fbd4d26aa03df96a0c1f9c3cf35c64d76bccc37b25f0ba

Observation f9bb7dae-4b5b-4664-94b4-2ea372393358 · outbound

This paper cites Rmvpe: A robust model for vocal pitch estimation in polyphonic music,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Rmvpe: A robust model for vocal pitch estimation in polyphonic music,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:23.794768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:40:09.900326Z digest=sha256:9680d3c44dbb11f81b3ba6fddae3d450dbfe41182da9b4f1a9f9c8415bf71ec4

Observation b5824415-23d1-4220-ac0e-c03afc0ea926 · outbound

This paper cites Enhancing skin can- cer diagnosis using swin transformer with hybrid shifted window- based multi-head self-attention and swiglu-based mlp,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing skin can- cer diagnosis using swin transformer with hybrid shifted window- based multi-head self-attention and swiglu-based mlp,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:23.219875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:40:09.918778Z digest=sha256:67832fea884d5803932020305b3c393e3cb80fcf6b3c170a9712aea16dd5ea12

Observation 61f5b1ec-e191-4293-afb4-86cceefebbc0 · outbound

This paper cites Searching for activation functions,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Searching for activation functions,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:12.252580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:40:09.980654Z digest=sha256:ba53c9fb5e45554b63495bdef2c8c47ce1834202e57f0b8f3c6d4cf3b0781b4c

Observation dec73175-a10b-4b58-8360-d52c414541d0 · outbound

This paper cites Decoupled weight de- cay regularization,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Decoupled weight de- cay regularization,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.029630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.029630Z digest=sha256:d1de05b2a743231ecfd30f777d0f529f4fa19556c05c361440f6780cd0c62c1a

Observation 072fc2c8-d40f-4cfd-8b25-d6b33dbf5e32 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.634578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.634578Z digest=sha256:db7b23a258e22090f77974cf1634af70ed920dd63b6f0b306dec9d735533841e

Observation b66fe9de-cb41-4a86-b269-35a9afcb68dc · outbound

This paper cites Focal loss for dense object detection,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Focal loss for dense object detection,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:12.140211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:40:11.768796Z digest=sha256:44cfe3c394e62ed7c437fd6833b678c1e3f2054b921b131e666b241887e14089

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.848020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.848020Z digest=sha256:158603bc2cf7784901b0b42af6de87b264f9678a179fbc444186eb4e9f494bd4

Pith citing papers

Observation b0fd2927-ab92-4ea7-b6d2-df576f3a353d · inbound

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 cites this paper.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:12.074820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:39:46.886300Z digest=sha256:0418c135d6970834bc145858f29fc5cc2f106454ac6679ae39e21a5d4a8d589b