Pith. sign in

Paper Citation Record · LEDGER

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2506.09556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09556 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:49:09.335889Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:49:06.392265Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T04:49:09.594801Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d16c6d11-83fe-4131-bab5-2ec70e60726d · outbound

This paper cites Categorical Emo- tion Recognition.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Categorical Emo- tion Recognition

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:15.822861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.275853Z digest=sha256:e3552a77f2fa58cd3c773b49d9b3f616ae1463662203b74cacb003791f763358

Observation 9badd1b9-a52a-469a-9d80-46af3ab08019 · outbound

This paper cites MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:49:09.666502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.392265Z digest=sha256:8624ec290adbed92e97379879d8da557167de4d0be37d7fe6cc54e5b764245e7

Observation 12d57c54-07d0-4e7d-ae7f-fe4e76ad3cb5 · outbound

This paper cites Overview The novelty of our approach stems from three key directions: 1) architecture, 2) data utilization and 3) training recipe.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Overview The novelty of our approach stems from three key directions: 1) architecture, 2) data utilization and 3) training recipe

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:15.724599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.486313Z digest=sha256:ea75a838042ada5bc8b43b64d34bbaa176b10c1d2f2bc9df81abc36fb9493844

Observation 0494fa24-9d83-4d75-81d0-814cf1cf5346 · outbound

This paper cites The Stage 3 meta-classifier uses batch size 128 and learning rate 1e-3 without M.MixUp.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions The Stage 3 meta-classifier uses batch size 128 and learning rate 1e-3 without M.MixUp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:15.598901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.572188Z digest=sha256:585ef60dece62345228d65d38912082e9b8324dc03e4f8a6b3a95cec9907f848

Observation dd6468ea-74ef-4c5e-8316-e55064c069c6 · outbound

This paper cites an unresolved cited work.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:49:15.480102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.645168Z digest=sha256:cfd6cdc7349db9dfc76d931d9e59f2061c0d0f24936f5f8eecd0bc3dda609b68

Observation 6a057bc7-41f9-428b-939f-0407ed8b9e8c · outbound

This paper cites an unresolved cited work.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:49:15.359673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.739106Z digest=sha256:7e30fdab8cd784d1aafef866a4d9576a86a60fe995c3d2cdd308390225b9499a

Observation 69592d8d-9fb1-4522-9740-adcab9a7dd85 · outbound

This paper cites an unresolved cited work.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:49:15.161502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.818159Z digest=sha256:c57c6accdeefa6e43abdae9ff2744d87f759ceb2040e71f81cc100f4cfacae64

Observation 0dad5a84-0db1-4a82-a23b-e883d17b77a9 · outbound

This paper cites Emotion recognition in human- computer interaction,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Emotion recognition in human- computer interaction,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:14.924694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.893534Z digest=sha256:2634eeba17426d94529c21952c56ab236df82c7f5648156a4648c6b0c53d04c4

Observation 36caba0a-ec5c-464b-a2d8-a4072a2a11fc · outbound

This paper cites Make patient consultation warmer: A clinical application for speech emotion recognition,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Make patient consultation warmer: A clinical application for speech emotion recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:14.793854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.956480Z digest=sha256:e82a4f4e925a1a52dbaa638ddfee011a7e164ea2fcc1503271574af3d8e2d3fe

Observation ddd14370-4c88-4304-98eb-0643ee3f4472 · outbound

This paper cites Integration of driver be- havior into emotion recognition systems: A preliminary study on steering wheel and vehicle acceleration,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Integration of driver be- havior into emotion recognition systems: A preliminary study on steering wheel and vehicle acceleration,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:14.628524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.031215Z digest=sha256:5721dc3230d1d6d7f39524c43560a7f29f5cdcbcf2e3ff2d8432b4b42a0d36b1

Observation e944c8ae-e12a-4a23-bc6b-7b3964a3cdb8 · outbound

This paper cites Call redistribution for a call center based on speech emotion recognition,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Call redistribution for a call center based on speech emotion recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:14.462141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.098384Z digest=sha256:a5a7f18a9d443422fe04bdbba901d42522530edff55535d4990f0bb0a31a53ec

Observation 26a9d0b5-8444-4c67-9620-331f874df050 · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions emotion2vec: Self-supervised pre-training for speech emotion representation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:14.265007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.181521Z digest=sha256:66c8eec4acbbaaed3c4546ece7c11c80344c5e1b51908b0c5d8eaf60b78db35d

Observation 1e4348d4-2ca5-4039-aea8-8ffe4ab2c932 · outbound

This paper cites Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:14.034592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.266815Z digest=sha256:e983f70c67035d4123e06c5d27ecbcd2208db89bcca829e7e4c4e958b6880d8e

Observation d1dd1b22-bb67-460b-9f0f-844433cf70d7 · outbound

This paper cites The interspeech 2025 challenge on speech emotion recognition in naturalistic conditions,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions The interspeech 2025 challenge on speech emotion recognition in naturalistic conditions,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:13.844521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.355321Z digest=sha256:2b053f451e7103757cc6cb97d6b2fc521dc9fbe6776b12c9d1d0f811c5edc72f

Observation 3e8f70a8-a254-4928-86a2-3f8df4e160f9 · outbound

This paper cites Building naturalistic emotionally bal- anced speech corpus by retrieving emotional speech from existing podcast recordings,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Building naturalistic emotionally bal- anced speech corpus by retrieving emotional speech from existing podcast recordings,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:13.649321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.443399Z digest=sha256:236a9d38956b7b8af050d6a3dd15a2c9c0f5b0616ca1e43aab523a53bd8fdee3

Observation 462c9e3d-d649-46fd-bb64-ddcb91d79659 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:13.366793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.548535Z digest=sha256:eae5a44942fd42c2c98d6aead8e4dc2eccf9330282775e74aaa364ba06e9a221

Observation 0a85acd4-f961-4969-a404-5ef36734388f · outbound

This paper cites Deep hier- archical fusion with application in sentiment analysis,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Deep hier- archical fusion with application in sentiment analysis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:13.125839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.679213Z digest=sha256:060ff63ec2f6ee1412ce46a45206122c79c6df3473b51565325b58ea2be127df

Observation c1d370ad-ad16-49ac-9a81-8fccc442f7cd · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Iemocap: Interactive emotional dyadic motion capture database,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:12.902567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.768729Z digest=sha256:f32a6d0b7fc664822f3285630a214b722bdf3cbdfbaf843a3a67b41175a922a0

Observation 8a41fd6d-e4bf-403f-970d-53cc5e70328f · outbound

This paper cites Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:12.558675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.908426Z digest=sha256:8c8258a848f900d1b77a4f21f29c43b2cab85c5c82164265465e536c51db0cb5

Observation 70fe6991-914a-48e8-9c93-61251e71b6df · outbound

This paper cites A survey of speech emotion recognition in natural environment,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions A survey of speech emotion recognition in natural environment,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:12.291713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:07.969095Z digest=sha256:7af4ac73bc8862403c990969793ef967735ea0c7f263d36d36d97a5cd2442919

Observation b60dca6a-993c-4f54-b86f-e1dc96de0065 · outbound

This paper cites WavLLM: Towards robust and adaptive speech large language model,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions WavLLM: Towards robust and adaptive speech large language model,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:12.041077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:08.021666Z digest=sha256:83c1eecf1b6bcdd59e071e8ad80a2ad79ed8f63954a04482baa429e58a358910

Observation 4d577b80-37f9-45ed-aba9-e89ed2092976 · outbound

This paper cites Soft-target training with ambiguous emotional ut- terances for dnn-based speech emotion classification,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Soft-target training with ambiguous emotional ut- terances for dnn-based speech emotion classification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:11.764424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:08.117496Z digest=sha256:7809cde37e26e84fe7101890d761a8eb849ec93b515526409222f770dcce8d8f

Observation 815d5e63-b387-41a7-b82a-e473abebfc6b · outbound

This paper cites Manifold mixup: Better represen- tations by interpolating hidden states,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Manifold mixup: Better represen- tations by interpolating hidden states,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:11.484741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:08.175330Z digest=sha256:133d510a60915474004fff8824d8288edd3431b6b9ee9f83da382ed86fb081f6

Observation b4770892-23f7-424b-aa38-1c9c8141cd32 · outbound

This paper cites Powmix: A versa- tile regularizer for multimodal sentiment analysis,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Powmix: A versa- tile regularizer for multimodal sentiment analysis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:11.266365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:08.234551Z digest=sha256:bc458dc77a580f5d85d36adc171c142c6366e7c6e360a6ef49bb58c954fcfd38

Observation 4d07d30a-53a2-4414-b0d1-c55a0835069b · outbound

This paper cites Speech emotion recognition with multi-task learning,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Speech emotion recognition with multi-task learning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:11.110759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:08.305637Z digest=sha256:a4013840438708831f5db43917ae6d43cca4185d2fb77d6ea8fee6f7a15a5f16

Observation a0b28d41-bf57-4949-aeb7-86e82cddf8fe · outbound

This paper cites 1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions 1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:49:09.548351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:08.398864Z digest=sha256:5a22e3abac6cbb3c6acbc1cc69a251891ac7aa09b2cbba7f525323c56b65d949

Observation 07c27be4-6c06-4774-b74d-2b0d0210aacd · outbound

This paper cites Meta-classifiers easily im- prove commercial sentiment detection tools,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Meta-classifiers easily im- prove commercial sentiment detection tools,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:10.972515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:08.528419Z digest=sha256:ac92d89a49c451c5364d14eefe18ba4a3a01fd5bd1357e7a46edb7633fb58e49

Observation d4f1aed8-ee61-4753-9e18-001f4ba3cb8e · outbound

This paper cites Attention is all you need,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Attention is all you need,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:10.786076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:08.636102Z digest=sha256:ee9ff9590bd56152866dab200265b4efcf267891f81eaef080877d1c16f28d2e

Observation 719d48d6-3b68-4008-9942-34204805d496 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:10.600445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:08.723671Z digest=sha256:60b82264f7beb72b9563668007d87ec23bb5e8e43f78c23e776c33a7f954522d

Observation d896ca3b-cecd-4987-a3a0-3f472e2848c9 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:08.809717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:08.809717Z digest=sha256:5a1584467960d209e068013259cbe38cf638d8aac29969c4aa8f54760c285bcc

Observation b7a0a019-5306-4013-80b2-77b90dbf5bfe · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Robust speech recognition via large-scale weak su- pervision,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:08.900960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:08.900960Z digest=sha256:82af98891ed77cfc93bd90c1da6ce8384a0e4d79060ab9dbf53864f22d508f5c

Observation 9d6105cb-5681-4626-9ecb-d3e791be2d64 · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Roberta: A robustly optimized bert pretraining approach,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:08.985412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:08.985412Z digest=sha256:b4c1548b2b5dcbfd50f9571c4c9bdb1833b5a111feecced3999207cdf08266c4

Observation f6fa6fb4-aab2-4edd-b0c2-793f97481112 · outbound

This paper cites Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:10.433802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:09.045868Z digest=sha256:a79e5039ef32855999dd7741abf69b883d580fcd66174e0fd5640cee24a82c5b

Observation cf9926c2-6b78-4f5a-8195-5296dfea9d97 · outbound

This paper cites Long short-term memory,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Long short-term memory,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:09.118064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:09.118064Z digest=sha256:6d1756659f49c6c13842f77cf628c3eefb94ff4fdd6598521bef7ce05f154938

Observation 344cae71-43c8-46cd-8ee7-93313d28eced · outbound

This paper cites Vicinal risk minimization,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Vicinal risk minimization,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:10.253030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:09.214257Z digest=sha256:799106e55d783188d6dcba1dd2369df4b160d682532de9b67f903227d9223c54

Observation 8add645e-4b2c-443b-b366-0b70293a58ce · outbound

This paper cites Decoupled weight decay regulariza- tion,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Decoupled weight decay regulariza- tion,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:10.084078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:09.272425Z digest=sha256:cc7a7f58a76eb9849a7a5f5f2a12b99a1d785a8460b0bbee0721b98b0b3aeca8

Observation b42fd3f9-8335-4f7a-8ef7-22e8bcfa5736 · outbound

This paper cites Py- torch: An imperative style, high-performance deep learning li- brary,.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions Py- torch: An imperative style, high-performance deep learning li- brary,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:09.907373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:09.335889Z digest=sha256:ebaa50646aad301587566166cb19eb5b0ca27b347648db8daaa98270848aaee2

Pith citing papers

Observation 9badd1b9-a52a-469a-9d80-46af3ab08019 · inbound

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions cites this paper.

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:49:09.666502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:49:06.392265Z digest=sha256:8624ec290adbed92e97379879d8da557167de4d0be37d7fe6cc54e5b764245e7