Pith. sign in

Paper Citation Record · LEDGER

emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2312.15185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.15185 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:30:28.595710Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T18:28:48.480697Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8c1e784a-8609-4a47-b1bb-6dc4731fc21d · inbound

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS cites this paper.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:33:25.612916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:011f951d4ab2c23d3d49f41b4b0d47e99deb9336ba55ba290dc7420bcd66a365

Observation 4f3719f0-2809-4359-831b-ac0361df2d2e · inbound

ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models cites this paper.

ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T20:48:43.642661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:48:43.642661Z digest=sha256:b9e7c0b33db1d5b12cb4263ffbd3bbd69cb1466f4970679f45b165d5ae65c3e2

Observation 15e19581-d88f-41d5-8b30-29586fd03c7f · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.752127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.752127Z digest=sha256:da0dc12a20a3c98a691ec5449de64bb8c8ae1021c2705305047531b4b1bf0b1d

Observation e7f9a7a9-556d-458a-bfbb-362be4585dec · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.145680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.145680Z digest=sha256:09d0b9cb8d369ec0419b63f1b9b18e481b2faf031accbf76fa0b7ef713baf13c

Observation e3d606c3-af7f-405f-a4e5-839d6fb9decb · inbound

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis cites this paper.

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:30:47.727145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:30:47.727145Z digest=sha256:c6c47ce8363e81ea93ceb50b85aca8329c58a47c3891e71ccd2052f9cd51baf7

Observation 8a9fb71a-b4f9-4d32-9de3-63d071f9e117 · inbound

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion cites this paper.

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:28:04.187733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:28:04.187733Z digest=sha256:c15566bc1528fdcec06496164d1593c68ec531e1df5eaf8fea5e0e6ffeb64519

Observation cab6075e-e96e-4126-8053-974c7a93de7b · inbound

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios cites this paper.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.987198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.987198Z digest=sha256:065dd9083dd611f94df833617a419ce52c4801077bef9734263b4449ca860459

Observation 5a00d8ca-9701-4fe1-af19-c0862740addd · inbound

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation cites this paper.

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:23:50.605657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:23:50.605657Z digest=sha256:08f310a0edf0687f238b04bda114716d54e1e4202eab9c341c84dd4909700d47

Observation 49384579-9107-4709-b7a5-bcd2bcc0c23b · inbound

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts cites this paper.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.086034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.086034Z digest=sha256:92f47508ad76f824e3fb70bc0af024cbae3e5103c90891cd0ab45624b9097ec8

Observation dfe2489b-1aaf-48eb-819b-defc8d520a24 · inbound

Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data cites this paper.

Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:42.974249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:42.974249Z digest=sha256:72e8da6997dd1f46030f794d7debabaf307b5ab5e77542b616796896e6a3f99b

Observation 62909fd7-b90a-46aa-a123-24b808a9daba · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.866566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.866566Z digest=sha256:cbe273c8f97f2764955d8d112ea5c42c97045197148bcf8b267800688ec57202

Observation 75e6be26-a864-4409-b4b9-c61696ffcc7a · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.681995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.681995Z digest=sha256:cb2d1b8ee27c2299e72e1ac8e3f1f174a24964cef73929406efc41259abc7207

Observation 143e4c8e-f340-4307-b93c-822ef83fa9be · inbound

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model cites this paper.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.182473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.182473Z digest=sha256:3b6ea4d58cbe2fcde664602f4cc786b07ca2796b916974c2e3f11cd6e7f46365

Observation 974f1597-8794-4829-b9ec-2a0e2c86ae37 · inbound

Gender Bias in Instruction-Guided Speech Synthesis Models cites this paper.

Gender Bias in Instruction-Guided Speech Synthesis Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:11.621926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:11.621926Z digest=sha256:2e2e76da84e7efaffec88874ad1c4eb44c81ee87fe37a10dd3517f1a2bc53f18

Observation 4211e738-a583-401c-b5d9-771ac1f3df12 · inbound

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech cites this paper.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.850108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.850108Z digest=sha256:dedef72f407ace989bfec31c3d02ddfa8f7a2a3fdaac1ce9bfd161337edf2561

Observation 4f619372-433f-4e44-baf3-81c64b1af81b · inbound

EASY: Emotion-aware Speaker Anonymization via Factorized Distillation cites this paper.

EASY: Emotion-aware Speaker Anonymization via Factorized Distillation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:36.078119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:30:36.078119Z digest=sha256:ebb18e056afc2e4182f41f4fb9b73b1cae1d0e5465cd6b9c94d2cfc1639807a0

Observation fa03a2b8-92fd-4ae8-afab-9bee14f50ef5 · inbound

EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language cites this paper.

EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:02.058069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:30:02.058069Z digest=sha256:2b751fed2421986ff77dfa8c318268300eb2ba29ef34ac714e6f2b1c835c79a5

Observation 2a46fb90-8592-4e1b-9623-cba6f0a2de59 · inbound

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt cites this paper.

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:39.181667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:39.181667Z digest=sha256:61c0e120ffdb06435398c5309264412da3a2ccd58415aca2f7af02a1cca4119e

Observation 1283c4a1-509d-4616-a396-e1175b8febbd · inbound

Probing the Robustness Properties of Neural Speech Codecs cites this paper.

Probing the Robustness Properties of Neural Speech Codecs emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:31:39.523254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:31:39.523254Z digest=sha256:966e574215e7e0fe70415b8ab2a68adae92e2c24dbbd171602593cdda9ace6ec

Observation 3a23ebd0-dc44-4dec-b950-692437c41977 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:56.482687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:56.482687Z digest=sha256:69689317c524868761332b94951522e442eb306e1b1792ce867ce226f1dc5f8b

Observation 17b384e8-dc7f-4680-a281-53e03cb3b570 · inbound

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions cites this paper.

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:09.566555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:09.566555Z digest=sha256:d8a6ad3c31f542a0e99cc91db262bf0848622497107cff37bda49ca656c46c42

Observation 2c0e4eed-aa8b-4d57-84ec-3030fd4b73d1 · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.353059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.353059Z digest=sha256:f3f6cf353ff4f41caf9f211bd39589f86d821c56a400c68aa3fb2dbf8bb5bcf0

Observation eb7ebb49-ba2d-4e51-805d-2b1b66e5e02f · inbound

Optimizing Multilingual Text-To-Speech with Accents & Emotions cites this paper.

Optimizing Multilingual Text-To-Speech with Accents & Emotions emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:29.542314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:29.542314Z digest=sha256:987fe7d3141b6ae4eee66c86214eb17ae3ea4e6c6fa7d7b9454d9970d3fc6d4b

Observation 382fc309-fb3f-43a4-85a1-672827b718f9 · inbound

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding cites this paper.

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:57.081692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:57.081692Z digest=sha256:018716b43399c2d3491edcdbbc45ac4b221dfde0f534b8eff1dd0b5b7b66c3f0

Observation f780c7fd-ff9e-48e9-abaf-946cd67323b8 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.223838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.223838Z digest=sha256:1ee77e46e276bcca287687d18d878d7340ba8895e527dfcc7b4ffce7a7c46a93

Observation 301254d5-bccf-44a0-9c29-c9adc82369e6 · inbound

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors cites this paper.

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:50:25.414247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:50:25.414247Z digest=sha256:1ff96243bd50c7070913da3957f936672ccba24848268488459091ef6bf4f2c6

Observation fd2dd173-740f-4bdb-b45a-89eb39ad134e · inbound

Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment cites this paper.

Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:29:41.195435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:29:41.195435Z digest=sha256:dfe4ac40d5d6d4c6e8b0202d78eea9faa2dc396ccf2d5655ad2f2d0d209398be

Observation d904b568-df3c-4f04-9e7c-0eceb1425215 · inbound

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks cites this paper.

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:41:59.596971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T02:41:52.996457Z digest=sha256:bafbf06b72b9098641b5ed31992951472c94d33091f4671fc9dbba763b1aa96a

Observation f5dcc674-87c2-4a7e-a7b8-417df35ec02d · inbound

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody cites this paper.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.524538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.524538Z digest=sha256:a9f75f6fde3d318c38872a870d0b6068fadd5ccc8df4c41dd2a599dd20fd46cc

Observation 4f720ed5-cb13-4286-845c-a4e5e3b0d86e · inbound

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents cites this paper.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.444581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.444581Z digest=sha256:fc61e8dd26885133dee0c23c0a89769bb74481386806726a892d6421c1b29725

Observation 875a69b9-3d43-4033-95b0-c1a2e6d7ba3b · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.634188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T07:42:43.077644Z digest=sha256:0785dcf9c1a2c98cb63ae907d5b8510f5fde64c581a0ff66362acf61be15ac4f

Observation f2f20dca-d5f0-4a8d-8723-6e5aaa377292 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.123236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T20:56:37.533183Z digest=sha256:7e80dddf80a3a89269f0e2439fff8a50b8d5bf710bd4413642e73b76d9839563

Observation 26e9a64d-ee20-421a-8ce6-207d9ac59a60 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:49:44.597311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:49:44.597311Z digest=sha256:d3d40dbd01fcfa69be368e537bc99b264fae5dbe1b9b9dbabae50c3cfd780da1

Observation 1235f466-957a-4462-850b-41217ba13cd5 · inbound

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining cites this paper.

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T16:09:50.207908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:09:50.207908Z digest=sha256:bde75286f2e1f7d8ce624a685c5c115060988af2dd908423043b646feff10e4c

Observation 4331fecc-4aaf-471d-a446-b8f915a6a3d1 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.004189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:3dd6563b049b8178b2426838e7ef666491c5f34549c907b272ee63642cec2af8

Observation c34293b0-51c8-46a8-8778-04e736d509da · inbound

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses cites this paper.

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:01:26.002498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T13:13:49.855219Z digest=sha256:6ca1155514664bbc612fd7b02ca5766ba04ec629f311624116aab0b3845ffc2f

Observation f2cdfda3-401b-417d-b4a1-06c2b5264529 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.208024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:2138a47fa70a6030307a03a347174babdff0eebe74b8f485b60f0b917fc2bd63

Observation 3400357e-5d39-49d9-a2bb-3ad06fbcff50 · inbound

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection cites this paper.

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:48:04.313900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T05:47:45.259359Z digest=sha256:f992c7ad6e11a1cecbbf4abde15e833f97203f07355d2129a9fe36acd954bf81

Observation fb0a5d89-b9d0-46f5-9587-71f8ac923026 · inbound

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models cites this paper.

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:56:05.289172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T04:54:35.815868Z digest=sha256:3df69f553ed5f95e5a9b6c28f038c57b2f88221cb0c69111b5a2ce65915dc845

Observation a1f077c2-ab30-4e47-8e9c-e03bb13a4806 · inbound

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI cites this paper.

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:28:48.482501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T03:09:07.295810Z digest=sha256:37cfc59ec9c2fb269322565371e61fe5fa3adc573bf973c395d6edc7ae703603

Observation e9feee97-ee97-4e78-a2f2-09a872113081 · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.191735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:b333e69001360bfc941a530fec28bfb39504ecd76728d2987ffc57540e803b7f

Observation 0f87d94c-cb0a-4502-9e79-0d9d381219bf · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 189

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.155193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:c6c5562328fc2bb3ebd5993c91476eda9a65f383089fe20a7436f3fad1fef059

Observation e98932ef-6fc6-4774-b43e-c1653e5faa6c · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.459753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.459753Z digest=sha256:a54b70cc6e7f8e5aaa0c31458e7fa11c34da4f70683406f7511318efdde768c7

Observation 3e30aec3-6e10-456d-b113-a85825e87fb0 · inbound

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks cites this paper.

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:21.349519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:50:21.349519Z digest=sha256:094a41f7393603dff802b49a81a31631c63b1a8c65848e0a29324da07d6262a7

Observation 4c81d5b5-1e73-428e-941f-4a7c2c63a788 · inbound

Is Self-Pretraining really useful to improve diagnosis in medical Time Series? cites this paper.

Is Self-Pretraining really useful to improve diagnosis in medical Time Series? emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:38.489642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:38.489642Z digest=sha256:cf572ffa0611caeec4f63b08ae638ec275ae1d6d629b636322991ea51a8f5a54

Observation 91568008-84fe-4be4-84d6-d88c38f2ae7d · inbound

Is Self-Pretraining really useful to improve diagnosis in medical Time Series? cites this paper.

Is Self-Pretraining really useful to improve diagnosis in medical Time Series? emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:22.726620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:22.726620Z digest=sha256:c7a0f18b04ab420187b11e6c8bfe865e03cd194a762a592358cd3cc9014f85a3

Observation 73cdd2f1-b259-455a-af2d-98b2b7c70361 · inbound

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching cites this paper.

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:30:28.595710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:30:28.595710Z digest=sha256:e2800aebfaabb6701670c9730ae6da2dcf3cab1b7005381aa8ebae0a4a653e50