Pith. sign in

Paper Citation Record · LEDGER

emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2312.15185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.15185 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:34:43.524538Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T18:28:48.480697Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8c1e784a-8609-4a47-b1bb-6dc4731fc21d · inbound

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS cites this paper.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:33:25.612916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:3dcac72b61211e567ebd20ced413131a5e9cf89d8967cebbfad8d062ed8c7dd6

Observation d904b568-df3c-4f04-9e7c-0eceb1425215 · inbound

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks cites this paper.

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:41:59.596971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T02:41:52.996457Z digest=sha256:0e584a5fc7dc10e04f34798e930096aeb3a575b4491b9a7a89869aa9ebea8239

Observation f5dcc674-87c2-4a7e-a7b8-417df35ec02d · inbound

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody cites this paper.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.524538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.524538Z digest=sha256:d74d339fe3b08bc6c1e2208a1f6c454eedd757af34e72605cbf38d30ecd31405

Observation 4f720ed5-cb13-4286-845c-a4e5e3b0d86e · inbound

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents cites this paper.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.444581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.444581Z digest=sha256:fbeff6b2fb6d526153b231192dc461c70b368ac755255a21e58efd358fcbf525

Observation 875a69b9-3d43-4033-95b0-c1a2e6d7ba3b · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.634188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T07:42:43.077644Z digest=sha256:b2d08af9b2131367ded28e8c63667d42e574696d90a748abc7b5427614d4383c

Observation f2f20dca-d5f0-4a8d-8723-6e5aaa377292 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.123236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T20:56:37.533183Z digest=sha256:ce467b14388edff916f81483f97bed4ba1196ec5fadae04ea41394b5ce39c7d4

Observation 26e9a64d-ee20-421a-8ce6-207d9ac59a60 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:49:44.597311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:49:44.597311Z digest=sha256:8b3182993d60748ede835e28626bc80d5bc33ff079dce326064a79c89c988dbb

Observation 1235f466-957a-4462-850b-41217ba13cd5 · inbound

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining cites this paper.

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T16:09:50.207908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:09:50.207908Z digest=sha256:d657302e62666cf27101ce5df136c4bcf23d5b2c791cacf25f33700f94f8ab4e

Observation 4331fecc-4aaf-471d-a446-b8f915a6a3d1 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.004189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:ab7b2d5875c1d6f255c003e491c9fa9b5160f394f87a316e8012d156ce172714

Observation c34293b0-51c8-46a8-8778-04e736d509da · inbound

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses cites this paper.

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:01:26.002498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T13:13:49.855219Z digest=sha256:0f0e96ddf0a7cc43c43061f2302d051d6f53c591bb4875374349887328a5612c

Observation f2cdfda3-401b-417d-b4a1-06c2b5264529 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.208024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:50d2693b3661329d17ac4857052ee906f651632ae6f44f0a82c523e1a66484fa

Observation 3400357e-5d39-49d9-a2bb-3ad06fbcff50 · inbound

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection cites this paper.

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:48:04.313900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:47:45.259359Z digest=sha256:f8c74e69d142b65cf51ab7828f2dde0151ab40c9739437272fd3d7c458fe079e

Observation fb0a5d89-b9d0-46f5-9587-71f8ac923026 · inbound

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models cites this paper.

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:56:05.289172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T04:54:35.815868Z digest=sha256:847ca1fee8b965774ac9235272274e378b6737ca4400edd759490d6c64f28166

Observation a1f077c2-ab30-4e47-8e9c-e03bb13a4806 · inbound

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI cites this paper.

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:28:48.482501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T03:09:07.295810Z digest=sha256:008452c465c548411ffd2389010b4679fa26507e622db4ccb67b2862d9d5d751

Observation e9feee97-ee97-4e78-a2f2-09a872113081 · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.191735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:f2de4a4b1a82b4e34f73eba2da1f3bab45f8ebffc49567763280d92ce28f689c

Observation 0f87d94c-cb0a-4502-9e79-0d9d381219bf · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 189

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.155193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:32a9d1ae92d8b0d2e880b02e6b8b125b72f20c44e62474cac1ea7b3c1cd7d217

Observation e98932ef-6fc6-4774-b43e-c1653e5faa6c · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.459753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.459753Z digest=sha256:038f1320a1d282b4a1ee1310f4aab1431b0edbfb03582f987d3b7bb7478a1646