Pith. sign in

Paper Citation Record · LEDGER

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 2 inbound Pith citation observations for arXiv:2508.06890.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06890 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:34:43.671787Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T15:21:28.848595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:11:27.116919Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20926e1c-a02b-4a11-8558-dec56b9999a8 · outbound

This paper cites Emotional voice conversion: Theory, databases and esd,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Emotional voice conversion: Theory, databases and esd,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:47.562774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.120573Z digest=sha256:69d5b5bd4a2b23e4a8a9622667bcbafa5e033cc22eae0a16e3acf58b3294afa9

Observation 9875e92b-cf54-4384-84e3-fa3d1fe0ddf1 · outbound

This paper cites Training socially engaging robots: Modeling backchannel behaviors with batch reinforce- ment learning,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Training socially engaging robots: Modeling backchannel behaviors with batch reinforce- ment learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:47.302187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.210040Z digest=sha256:dab021eedc62daf2e37d4625bd4b5fe968184baa971534f463850115af247257

Observation 569bbde1-c154-4d48-a233-e9dd585397c9 · outbound

This paper cites Real-time speech emotion analysis for smart home assistants,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Real-time speech emotion analysis for smart home assistants,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:47.140382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.279725Z digest=sha256:04ba361e61cf07e70c79e4f18bd8fc5d1e10ac527e11b6bf5115333282af0bd0

Observation 799fd7fa-82b8-47cf-bfe5-052bae3554bf · outbound

This paper cites Pittermann, A.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Pittermann, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.934170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.356713Z digest=sha256:3332ab176cdfe28294292e5f979ba246881a2b5968da909b0a994947f616e3b3

Observation 8056ebe5-a7fe-434a-88d9-2bf869e08d2b · outbound

This paper cites Toward artificial emotional intelligence for cooperative social human–machine interaction,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Toward artificial emotional intelligence for cooperative social human–machine interaction,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.714801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.416266Z digest=sha256:ad69db5e69a8953c09a00863623fd1bc44d87318d1b0b438f06e99501bbb239b

Observation d084a43a-46cc-4540-9571-8507a4873817 · outbound

This paper cites Pavits: Exploring prosody-aware vits for end-to-end emotional voice conversion,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Pavits: Exploring prosody-aware vits for end-to-end emotional voice conversion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.519963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.491550Z digest=sha256:886e28fcc7363454b5b7c1a081502ee711faf03e7179e37b819bfca6401a1805

Observation 96917dc9-93b3-41d0-8a24-bdb7d82010c2 · outbound

This paper cites Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:41.548368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:41.548368Z digest=sha256:eed719172643a2d9f39eb63ac8ceb0a2c6924ebd92ea3ae24a3a756101cea280

Observation ad683cd8-97c3-4caa-a811-83c1199a57a7 · outbound

This paper cites Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:34:44.023541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.607221Z digest=sha256:dd25242d774f5bdea74e24de6bc2c35969bf33cfd674a2f1acaf55d18a6768d0

Observation aad03414-d5f0-44cd-9fcf-45afb8ed88bf · outbound

This paper cites Emotion inten- sity and its control for emotional voice conversion,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Emotion inten- sity and its control for emotional voice conversion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.358372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.694838Z digest=sha256:9819d1ef5c8fae128b17dfc87755956293c5ac642664435e432857bcd2ebd26b

Observation 6b2b43b0-48cb-49d9-9acc-10fe9346ae3e · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.193432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.760599Z digest=sha256:6dcb9462c04b399c73a900921e3927812aaae074d4b776f2628bed96615932b9

Observation 211ee3b7-2bc4-4751-9f7a-1e7d739a6f16 · outbound

This paper cites Emotional voice conversion with semi-supervised generative modeling,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Emotional voice conversion with semi-supervised generative modeling,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.024775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.783931Z digest=sha256:ff65361c1c7052dc615cb5e89022c3ce58129b4ce0435530a4b8daf425e18709

Observation 7bd4e329-4d2b-4c5a-909d-00d088e033f0 · outbound

This paper cites Speaker-independent emotional voice conversion via disentangled rep- resentations,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Speaker-independent emotional voice conversion via disentangled rep- resentations,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.844240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.859084Z digest=sha256:81e5f26393beb6c8ac12b58fcc29709c21bc347a51c74246a0e201d1245be285

Observation e9bc8915-48d1-43f4-8959-3a690a7847f4 · outbound

This paper cites Nonparallel emotional voice conversion for unseen speaker-emotion pairs using dual domain adversarial network & virtual domain pairing,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Nonparallel emotional voice conversion for unseen speaker-emotion pairs using dual domain adversarial network & virtual domain pairing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.664702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:41.941417Z digest=sha256:dbf1db62ef597d2e7be6bb3ff73cd60f3b20e6de6744e6303f015e27e4b46f22

Observation ba414906-ff95-4c66-8f02-5c0bf5a27a81 · outbound

This paper cites Enhancing zero- shot emotional voice conversion via speaker adaptation and duration prediction,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Enhancing zero- shot emotional voice conversion via speaker adaptation and duration prediction,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.493477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:42.000189Z digest=sha256:6942940912a1d62dfb8459bb6f0b4d6a9abfa9f946d5c1885db7cb04bfc22d2f

Observation 20bdcbc9-46bd-41a1-b92b-881aaaa60d92 · outbound

This paper cites Zero shot audio to audio emotion trans- fer with speaker disentanglement,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Zero shot audio to audio emotion trans- fer with speaker disentanglement,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.240270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:42.115007Z digest=sha256:ae4259f25cc79a1e854f9adc8939d108a8c5f7046e3b7cc954f995e4dd31915b

Observation 100315df-b02e-4db2-aff9-654018ffebb6 · outbound

This paper cites Multi-speaker emotional speech synthesis with fine-grained prosody modeling,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Multi-speaker emotional speech synthesis with fine-grained prosody modeling,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.027825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:42.182480Z digest=sha256:2d2f2456fd2cba53345ee352e538d789b1264cfe9c674485ec61cb24f1ff2627

Observation 9f1cda4b-3ec9-4fe2-86f8-a70cf40257b4 · outbound

This paper cites Attention is all you need,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Attention is all you need,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.258853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.258853Z digest=sha256:29091e3be92622316e844c282e0c859cf38fefd4e48f0c0a6f36f1a3d9d18b82

Observation a11c8ce5-6338-410c-bcf9-cee006ee5a8a · outbound

This paper cites Unsupervised domain adaptation by back- propagation,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Unsupervised domain adaptation by back- propagation,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.325027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.325027Z digest=sha256:f1c7bb3fecca7f2d77cd9c34e224859ca111ea7f536114ec2a5601a17bd27a47

Observation bf6675bf-e40a-4b2a-8e66-d8c90f8f04e5 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.391019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.391019Z digest=sha256:83d6bcb8a69f07c829189905f23f8001cab172e7cfc4a1cfab418142319ac3cf

Observation 801da7cb-fa6c-41af-8f11-f25e6968e1ea · outbound

This paper cites Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.833308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:42.453515Z digest=sha256:ea1297d0b147bb6a64d6aea0184323c142b3a96134fd842e7c1a92f2838a01d5

Observation be03c332-48be-41df-aaaf-f72c62f3b4fe · outbound

This paper cites Textless Speech Emotion Conversion using Discrete and Decomposed Representations.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Textless Speech Emotion Conversion using Discrete and Decomposed Representations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.520181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.520181Z digest=sha256:bf445e49becaf3151bd2fd7d72d4de8825c99598cfc15506ee6e093fa1bedfe5

Observation fb2b3f2e-cacf-44d5-b267-4036f96cc59d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.593208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.593208Z digest=sha256:e5cff942b518461ff65330ce630c05cc083e251d30412075f2390aaaebd76573

Observation 00cac1e2-1a7e-484f-9e59-5dc1273d5a1d · outbound

This paper cites Speech emotion diarization: Which emotion appears when?.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Speech emotion diarization: Which emotion appears when?

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.612588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:42.701389Z digest=sha256:c3096bfceabeba461791dd9a8413733defa0db0bdc225604231c6f8ec435897b

Observation 4643191f-1620-45d5-be46-ab5cccd2a862 · outbound

This paper cites Smoothing and differentiation of data by simplified least squares procedures.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Smoothing and differentiation of data by simplified least squares procedures

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.501286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:42.785584Z digest=sha256:42f6cbb159f9fcea83c09e8b1443c42bcabf46129eaa762ae1396035a0019671

Observation 9b6dc289-b4f0-4549-9aa6-dc52a1638dda · outbound

This paper cites Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:34:43.870586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:42.870526Z digest=sha256:ff7127fce79696b6628a3528f917f7987f650ca5b4959907999d6b489ad2f647

Observation a65bb56f-2d89-4a1d-acc0-5cc8eac11491 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.929455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.929455Z digest=sha256:7b47cbf9b15d243194cee80ba9af282e08eec08f987f970f4104d3ee91b4f634

Observation 0a3eed7f-de8c-4bad-a880-dc9c14addb56 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Librispeech: an asr corpus based on public domain audio books,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.988717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.988717Z digest=sha256:2d93a7143f505aec3ffc315119a28281a5fb29d422916a62e4de9a6e7d2d0bb5

Observation 8194931f-49f6-4562-b8c8-9f7815849fbf · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody VoxCeleb: a large-scale speaker identification dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.077499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.077499Z digest=sha256:8bb3b2db49d2f980ebbebd98d9e994215d7f7d35b966ca0dbf7bd9823752fd97

Observation e316a941-445a-4437-aca8-b5ff72e5b0ce · outbound

This paper cites World: a vocoder-based high-quality speech synthesis system for real-time applications,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody World: a vocoder-based high-quality speech synthesis system for real-time applications,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.132364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.132364Z digest=sha256:29a0ad8caa317245027df55e72aa21dfded439c68aaa6701e9a7aa3f5748e6ec

Observation 8545e167-d0ec-48e0-8495-0f55808c1266 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.378234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:43.226294Z digest=sha256:09fd0edf7d7ddfee5b33fd8156d53dad06e5f3a4fd68cd9dfc3bb8b7cfe5f84a

Observation 6e0497cd-f0cf-4150-acd0-997ad90397e5 · outbound

This paper cites Crema-d: Crowd-sourced emotional multimodal actors dataset,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Crema-d: Crowd-sourced emotional multimodal actors dataset,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.313933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.313933Z digest=sha256:899f0bb6febbda76b3416e87cbdc9fbbcab6c369d41ba936d10293fc1feaa73b

Observation f8e74ea5-68c6-46db-b60d-59dcc73f5f55 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Iemocap: Interactive emotional dyadic motion capture database,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.377951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.377951Z digest=sha256:5bf5b8cb5fe05be639a2d29e963a7f01380ad36099aa6f060de1cb7256543250

Observation 76398b06-ef06-4cd6-8449-401c4d65e55c · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Robust speech recognition via large-scale weak supervi- sion,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.467839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.467839Z digest=sha256:6983c0e26b4782a88973d39282421552991eb386e67008fc07f5ac9c1c70493c

Observation f5dcc674-87c2-4a7e-a7b8-417df35ec02d · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.524538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.524538Z digest=sha256:d74d339fe3b08bc6c1e2208a1f6c454eedd757af34e72605cbf38d30ecd31405

Observation 806a96a9-e1d7-4f86-9e2e-4f1b10d5b70f · outbound

This paper cites Pearson correlation coefficient,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Pearson correlation coefficient,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.593809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.593809Z digest=sha256:4aed11685267aa7ac2d2bb982b0ebdcb9027768efd7a0cb54dc28692b89c04bf

Observation b15cd2ad-0f72-4282-99ab-f5e65f92e585 · outbound

This paper cites Dynamic time warping,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Dynamic time warping,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.198017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T22:34:43.671787Z digest=sha256:7b8b2c2a229bc71a59c0cc4f5a6048369c262f0ecf8f09a65215ce1b15a8e2eb

Pith citing papers

Observation eddc379d-1f05-4fca-a08a-7a160911202c · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.119312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:7f5ed695e0bd7683421fde5e056d6ebbf223d66ef5c578f0b6d9b2c79ad1265f

Observation 0bae5d3c-20b9-4be2-9f76-b6bed6dfe19f · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.848595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.848595Z digest=sha256:7a65d90a88ec324768d0f543504abb75be25bab60f48800dc5a0e39ccb993a9f