Pith. sign in

Paper Citation Record · LEDGER

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model

As of 23 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2505.20007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20007 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:07:25.935688Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:07:23.506269Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:07:26.629736Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85b01d42-30f7-49ea-b4fc-04c115e0578d · outbound

This paper cites an unresolved cited work.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:31.813394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:23.444560Z digest=sha256:addee5be445401d898e38e08568b20883e72933bf7dbc9016e9962f8f7aa8896

Observation 81b16ec7-6582-4c28-86f0-6b083741b15a · outbound

This paper cites Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:07:26.810012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:23.506269Z digest=sha256:6e599aa24a81db376f4a564a5425eaf9cd40368954d4d25c20ea6f96797b4ae5

Observation c62b6e52-2cd3-4500-8546-f3c16ae7c59c · outbound

This paper cites an unresolved cited work.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:31.609964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:23.625332Z digest=sha256:c2fee8d54cc77131139cddf6503d32360d58b6db6137eff8ec3c7974a4d98035

Observation a1410ac3-3a42-43f1-b9c4-34b3631ebd7d · outbound

This paper cites Dataset The provided challenge data consist of recordings from the MSP-Podcast dataset [23].

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Dataset The provided challenge data consist of recordings from the MSP-Podcast dataset [23]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:31.461658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:23.739049Z digest=sha256:dfb03b49b3ff9eb3a1c15ed5a83ddd4d02c2d9538ac370aa3a6205196c3f7d9c

Observation 9fff3cdf-f16f-422f-ae8d-60ccb7663244 · outbound

This paper cites Half of these models were trained us- ing only WCE loss, while the remaining half were trained with the additional batch balancing and SML loss.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Half of these models were trained us- ing only WCE loss, while the remaining half were trained with the additional batch balancing and SML loss

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:31.325643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:23.831691Z digest=sha256:77e18de25033dfa583420d18c0dd9a06d42b4802d8754cc05960026e300a6940

Observation f77ffcb1-d802-4a98-a13d-f96582f07e1d · outbound

This paper cites Notably, the pro- posed architecture benefits from the combination of multiple modalities, with its worst performance occurring when only a single modality is used.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Notably, the pro- posed architecture benefits from the combination of multiple modalities, with its worst performance occurring when only a single modality is used

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:31.163315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:23.887054Z digest=sha256:d7e368cec8c10d07e06e69d52649c7b087ea1017084391a8cfd2e7a9d17aee70

Observation 86cb8479-fc75-4ff5-8695-511c8774c39f · outbound

This paper cites It is also supported by FAPESP (BI0S #2020/09838-0 and Ho- rus #2023/12865-8).

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model It is also supported by FAPESP (BI0S #2020/09838-0 and Ho- rus #2023/12865-8)

Reference 7

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:07:26.534556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:23.921601Z digest=sha256:98c8e5232e74bace60cf8a8e88cdcf8b977cf90dc005fc6178af84a878c325e0

Observation f7f809d7-b7e8-49cb-a12b-e003f6c165ee · outbound

This paper cites Speech emotion recognition from voice messages recorded in the wild,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Speech emotion recognition from voice messages recorded in the wild,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:31.043328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:23.976538Z digest=sha256:679c7fa754a451aad70ef0fd77158beef906f5513f4546a6f56454d6e23df48e

Observation c2d137f5-ecde-482a-94f4-8723403ca5f7 · outbound

This paper cites Acoustic Emotion Recognition for Affective Computer Gaming,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Acoustic Emotion Recognition for Affective Computer Gaming,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:29.460374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.705467Z digest=sha256:a02595bb934b303403efd6736c8b57f13efc6024095d73027b45638043a3183b

Observation f97c11c4-f9da-4bbf-b0ff-d872d2b3b93e · outbound

This paper cites Speech emotion recognition using machine learning — A systematic review,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Speech emotion recognition using machine learning — A systematic review,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:30.911806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.191544Z digest=sha256:6b3c3796535a3548f103fbad7b3f407c11b7ea9a07f159f4fff48ee536b56875

Observation 7455f5c0-df4e-455c-a364-aacf3dd60a6d · outbound

This paper cites Speech Emotion Recognition in Neurological Disorders Using Convolutional Neu- ral Network,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Speech Emotion Recognition in Neurological Disorders Using Convolutional Neu- ral Network,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:30.705017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.272367Z digest=sha256:84cb7b28b14e85b27c8575043d0c84256b49dda0bdab167d213d9bd9a98d96f3

Observation aa7754ff-ee80-4b51-97b2-e64fd05fb0d5 · outbound

This paper cites Automatic Assessment of Depression From Speech via a Hierarchical Attention Transfer Network and Attention Autoencoders,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Automatic Assessment of Depression From Speech via a Hierarchical Attention Transfer Network and Attention Autoencoders,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:30.529284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.403263Z digest=sha256:8eb946e6e2cbf7696bdf0ed94f7222e2f0791f42e3ba93dcd1f382b72a070543

Observation f6b8a47f-37a0-4a9f-a109-23c45955205f · outbound

This paper cites Negative Emotion Recognition using Deep Learning for Thai Language,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Negative Emotion Recognition using Deep Learning for Thai Language,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:30.257569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.486837Z digest=sha256:9bf0436317dbd6826131e381040d4327e904ed49763c01d4f6f44817f5bc197b

Observation 36976519-c9c0-42c3-a844-50b01338b691 · outbound

This paper cites Negative emotions detection as an indicator of dialogs quality in call centers,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Negative emotions detection as an indicator of dialogs quality in call centers,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:30.023504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.546985Z digest=sha256:7c2a96af30438d0f0cede766e28198154aa1661dd1d962e29a9c70e3b9f14ff5

Observation 2c8ce4c1-fad4-4ff1-9edc-ed23994d1870 · outbound

This paper cites Using Paralinguistic Cues in Speech to Recognise Emotions in Older Car Drivers,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Using Paralinguistic Cues in Speech to Recognise Emotions in Older Car Drivers,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:29.861643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.587449Z digest=sha256:93cc2fe9477633f27d7ed17c6b83c61d4734cd87fda9b3c5c25b9e514022bb3b

Observation 257c00ae-18b7-4a35-acd2-3978261093c2 · outbound

This paper cites Affective Human-Robotic Interac- tion,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Affective Human-Robotic Interac- tion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:29.688898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.649319Z digest=sha256:de29e5e47de9db5fcd0275634e969f532725eb864c1d9c4b6604dc140f21e070

Observation 67bd16b8-eb3b-4c0b-b3cf-29a25965f6c9 · outbound

This paper cites Multimodal emotion recognition using cross-modal attention and 1d convolutional neural networks.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Multimodal emotion recognition using cross-modal attention and 1d convolutional neural networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:27.902950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:25.247125Z digest=sha256:5d7ef492418cebf6c74e986518b57e0f88071a1ada7ebb9e2c781e1d7146a0fe

Observation 438fe95a-3f1c-40d2-ba53-e49f31e543e6 · outbound

This paper cites Speech emotion recognition combining acoustic features and linguistic information in a hy- brid support vector machine-belief network architecture,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Speech emotion recognition combining acoustic features and linguistic information in a hy- brid support vector machine-belief network architecture,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:29.212818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.758493Z digest=sha256:2f52a383e6dcaca5095383fa86b71aac1a0658d010624ac987b0a1224c650ee1

Observation 30f96a3f-3b1a-4e22-bc5a-46487e6c690b · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:24.800948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:24.800948Z digest=sha256:793e7a84fbd059684e1ad1878cfc11398f5aa09e77a7ea2c9818d41720590c9a

Observation 11753f69-f739-4061-9930-c3d3f32e05c7 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:24.849791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:24.849791Z digest=sha256:ca171ecbf053cf3ed581681d42c820df2bca2ae19637b2df6d98ef11fa52e07c

Observation 548b8aa6-054a-4591-ab42-50cf786f7968 · outbound

This paper cites Multimodal Emotion Recognition,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Multimodal Emotion Recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:28.908352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.906079Z digest=sha256:8c94643a7f07f52435f8f0e1b7047c79c4a447d3c1897152d3766432531ca25a

Observation 92362960-d51c-4aeb-a1e2-e5be6b502f46 · outbound

This paper cites 1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model 1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:28.610132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:24.956717Z digest=sha256:989874736100652c62b902b6ddb572d25f4a3b581916d3f33d10db59f9c599db

Observation c417e31a-acc9-48b6-9e38-603fdb8651eb · outbound

This paper cites Naturalspeech 3: zero-shot speech synthesis with factorized codec and diffusion models,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Naturalspeech 3: zero-shot speech synthesis with factorized codec and diffusion models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:28.403961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:25.058412Z digest=sha256:a2df1b2da75945e82e56b06686dd9ae4eaaa99917b8615390e27e9ca8f6111e5

Observation 9efa3cfc-be32-4808-9d16-fd2946423efa · outbound

This paper cites Emotion recogni- tion through multiple modalities: face, body gesture, speech,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Emotion recogni- tion through multiple modalities: face, body gesture, speech,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:28.182628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:25.139441Z digest=sha256:096d18d74b8663fff4d1e76d359fb8c73bdcee908ae72b5fb7fe49fb79bd966e

Observation 719ee80e-8b33-41e2-b265-214c4917f5ad · outbound

This paper cites Amphion: An Open-Source Audio, Music and Speech Generation Toolkit.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Amphion: An Open-Source Audio, Music and Speech Generation Toolkit

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:25.853072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:25.853072Z digest=sha256:e1bd406c138c25653e600466606c4012c04f5540ab5a6c9ce2f7c29b7f632724

Observation 456ed2e1-8f2a-4837-82b3-a2bb1ed2a09b · outbound

This paper cites Using transformers for mul- timodal emotion recognition: Taxonomies and state of the art review,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Using transformers for mul- timodal emotion recognition: Taxonomies and state of the art review,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:27.772006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:25.308837Z digest=sha256:baa86c0fca9eac77e0cf00bc85945a0b31475eefef7be71e9f24d5420557dff0

Observation d96251dc-d414-423a-af73-20c24b5721a1 · outbound

This paper cites Deep neural networks for emotion recognition com- bining audio and transcripts,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Deep neural networks for emotion recognition com- bining audio and transcripts,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:27.647203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:25.379300Z digest=sha256:aaeb1b400df5986c17cb2dc85b58fdf4a2be898da58e897531bfc8ec47e23e81

Observation b90d3cd5-7c2c-433e-a73e-9ace91887ee0 · outbound

This paper cites Multimodal emotion recognition with transformer-based self supervised feature fusion,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Multimodal emotion recognition with transformer-based self supervised feature fusion,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:27.492322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:25.433282Z digest=sha256:f85f69f7eb49d4ff2ed1396b15abb80fd9e67fa799d396fb1fb07e0bc70e7689

Observation 9443443c-afb9-4378-8e89-84eef59d008b · outbound

This paper cites The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:25.500356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:25.500356Z digest=sha256:5c91a36aa90f6ade9c35baa6134db9783585f69c82cff364eeadf8f37d4dbb43

Observation 8eb908dc-9e08-4f78-9e9d-1596ea8d4e10 · outbound

This paper cites Stacked generalization,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Stacked generalization,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:27.207307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:25.591052Z digest=sha256:50b2e59e006c4529b7db6033d41bc3413881f56560d039277f4b739692767660

Observation 3280dc54-e0eb-480d-8e42-a32c1dc5b4d1 · outbound

This paper cites The interspeech 2025 challenge on speech emotion recognition in naturalistic conditions,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model The interspeech 2025 challenge on speech emotion recognition in naturalistic conditions,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:27.066600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:25.662426Z digest=sha256:83c8c8ffe29978abefa69613dea7ecca0cc324479bee68570a99c5a38d4445c4

Observation 9ec5e433-fa1c-449d-ad93-b4f637906853 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:25.726535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:25.726535Z digest=sha256:826a6b7a45bdb1aff03a82f3d162deb986d26277af354bb0cfe0bfecffe751de

Observation 01748e3d-fb67-41d2-b111-67b40c60117f · outbound

This paper cites Odyssey 2024 - speech emotion recognition challenge: Dataset, baseline framework, and results,.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Odyssey 2024 - speech emotion recognition challenge: Dataset, baseline framework, and results,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:26.966828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:25.935688Z digest=sha256:8b2bd50d6de6c859abb3eced518d74c1063d526be5842f64f0ee1a2b34c33e9f

Observation a07c5487-f316-4b0a-9500-72d1415fb878 · outbound

This paper cites EMOVOME: A Dataset for Emotion Recognition in Spontaneous Real-Life Speech.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model EMOVOME: A Dataset for Emotion Recognition in Spontaneous Real-Life Speech

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:24.074008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:24.074008Z digest=sha256:c4f7d84f7c4867b2ca564c93f28c091aae777c1aa9455e5c3d2203a53d6f05e1

Pith citing papers

Observation 81b16ec7-6582-4c28-86f0-6b083741b15a · inbound

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model cites this paper.

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:07:26.810012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:07:23.506269Z digest=sha256:6e599aa24a81db376f4a564a5425eaf9cd40368954d4d25c20ea6f96797b4ae5