Pith. sign in

Paper Citation Record · LEDGER

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 3 inbound Pith citation observations for arXiv:2506.04013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04013 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.365775Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.110093Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:07:18.246563Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73dae815-3c55-4526-84c8-912be14e0f2d · outbound

This paper cites Conventional VC models perform well in replicating speaker identity but struggle when the target speech is highly expressive.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Conventional VC models perform well in replicating speaker identity but struggle when the target speech is highly expressive

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.855135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.100974Z digest=sha256:8a25dc28346a63f491cfe8a4052e52e393be685904cf04f6cc54f3a9564011a4

Observation 60d15bb1-2e4b-439f-9dfc-c7c6aa33ce69 · outbound

This paper cites These models typically required text supervi- sion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion These models typically required text supervi- sion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.845772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.105713Z digest=sha256:a02cbdb2d212ab10173a797f0251c2b5074d0e11564e89a83dec0d9b47d5630f

Observation 2b05af34-6da4-4dab-be57-1f5d46dc096c · outbound

This paper cites an unresolved cited work.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:20.837000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.113786Z digest=sha256:61bf5479fc1f0a4436ec61e60b835e04e3d0f88326a2524b2af31ea54af0b597

Observation d66535c6-1087-46f6-aa49-3738da5f3ba0 · outbound

This paper cites All datasets are English and total duration is around 228 hours with more than 920 speakers.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion All datasets are English and total duration is around 228 hours with more than 920 speakers

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:55:20.825514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.117679Z digest=sha256:2b3039d2da1d9f1a4415b67415b0eb829786d2a62a77d9be0c3f5300478a74ac

Observation 89c442f2-6f5c-4f5a-8dd8-d074ee68083e · outbound

This paper cites an unresolved cited work.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:20.816313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.121966Z digest=sha256:e4dfd23726c0a6a4421dfcfde758db36b954b488e136f514bc5bb29397b405ec

Observation e262b5f6-fc04-414b-a4fc-43b864e1852d · outbound

This paper cites an unresolved cited work.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:20.805968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.125796Z digest=sha256:181416eb48e838fa7ee9a737fdaf22cceee2d120fb04d3880a5e7237dd55404b

Observation 92367e0f-9cda-4004-997d-3c19c61642e7 · outbound

This paper cites Styles2st: Zero-shot style transfer for direct speech-to- speech translation,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Styles2st: Zero-shot style transfer for direct speech-to- speech translation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.760631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.151933Z digest=sha256:3a51aa090febb9a03aebdfb39315a61c080cf2a4aaa1467ce0ec3a84d82fdd85

Observation c660e27c-596c-48f2-9e38-a14f95eef46e · outbound

This paper cites Chil: Computers in the human interaction loop,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Chil: Computers in the human interaction loop,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.796781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.129443Z digest=sha256:1aac7674e54c395ebb7cc458ac595512ae0ce7119ccd960319be34b46fba85a2

Observation ad0ef1ef-5d7f-465f-b9f1-e912a536f3a0 · outbound

This paper cites Towards an open-domain social dialog system,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Towards an open-domain social dialog system,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.788033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.132609Z digest=sha256:15517da3278f1288d678c4189562d0dd79c02bf7235e4195ca449778fe2ce4cf

Observation 25550670-15b0-48b7-9b2b-280828d1941b · outbound

This paper cites Face-dubbing++: Lip-synchronous, voice preserv- ing translation of videos,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Face-dubbing++: Lip-synchronous, voice preserv- ing translation of videos,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.779015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.136544Z digest=sha256:32aca3ca8cbde996fab4abc5e9a8ddd65dea08dd37c84075c74cd82e2d7a3718

Observation c9a77ed3-3ba0-4d54-8e9e-3620d532acd9 · outbound

This paper cites Findings of the IWSLT 2024 Evaluation Campaign.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Findings of the IWSLT 2024 Evaluation Campaign

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.139760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.139760Z digest=sha256:ab4c3599985a3c78e0b6de27d041a60c32bf74e12c2d1f754199e9305bb41754

Observation ba314579-786e-4c6b-a2b0-691e19b159d0 · outbound

This paper cites Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.110093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.110093Z digest=sha256:2913861f2dc309c8fcb9ddba7bf69dc9f556ffe754dd35a49fab320e266fd55f

Observation 0ab7ae64-8fa2-4efe-ba84-aec402de6f5e · outbound

This paper cites Simultaneous translation of open do- main lectures and speeches,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Simultaneous translation of open do- main lectures and speeches,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.770020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.144697Z digest=sha256:022e1edec56b60cc073c841a266bdc031202ab78ec3b2238a60757d9182c9f42

Observation c9f3b311-7eea-48d9-9bb0-b849850a81b0 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.148186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.148186Z digest=sha256:54e317fb4415047d7a9f2ad91704c80485d28ddd365fa3c7593e2054fe9860c5

Observation 5a98814b-9b96-409e-8a88-75df122dd784 · outbound

This paper cites Limited data emotional voice conversion leveraging text-to-speech: Two-stage sequence-to- sequence training,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Limited data emotional voice conversion leveraging text-to-speech: Two-stage sequence-to- sequence training,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.750351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.155032Z digest=sha256:21e4b674b408ecedb1c558a7076f97ff802175be95971788de6e9d6e632864de

Observation 28febc5b-ccf0-467a-95d5-5131f4e549e8 · outbound

This paper cites Nonparallel emotional speech conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Nonparallel emotional speech conversion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.740558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.158385Z digest=sha256:ce74508093a92ff7ee58a57e3b689d213355e9a17215b772c6e1f39484c2e611

Observation c9294002-d084-43b3-b41f-78731dc9f955 · outbound

This paper cites Nonpar- allel emotional speech conversion using vae-gan.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Nonpar- allel emotional speech conversion using vae-gan

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.730923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.161449Z digest=sha256:3c16af3a67f095c02b7eca588342410c8f0d96052feec3e19ec72afb85a821e7

Observation 90cacd4e-433b-4991-b40f-d1fb14fc9821 · outbound

This paper cites Textless speech emotion conversion using discrete and decom- posed representations,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Textless speech emotion conversion using discrete and decom- posed representations,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.721511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.165248Z digest=sha256:2e6897f681d062358d1751c7c85e296610313f382484045e15266ca58bdfdd36

Observation 3e6b91b7-04d0-456e-8398-c26d7a5e7f41 · outbound

This paper cites Hiervst: Hierar- chical adaptive zero-shot voice style transfer,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Hiervst: Hierar- chical adaptive zero-shot voice style transfer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.711516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.168157Z digest=sha256:2d826128ca04a1b3402acc512efa76d974ddc66e96cc68041dabb4ab2c6f61f6

Observation c5850056-bcc9-495f-af8a-9b5d27e67111 · outbound

This paper cites Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.701184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.172054Z digest=sha256:2285004c91a587b019db8f6d111993c5702f42f3c3b6ac5d0500e775e4cb2d4f

Observation aec52e0c-9e09-44f4-8d98-c38be0db8dcc · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.690149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.176032Z digest=sha256:78088449033e2ecb948676bf25a62a9882e6352dfc3bae4762ae8097a25542d6

Observation da85ec13-5aee-4693-b038-fcfb0199455f · outbound

This paper cites Freevc: Towards high-quality text-free one-shot voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Freevc: Towards high-quality text-free one-shot voice conversion,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.680603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.179658Z digest=sha256:3186af4547300b9c64b128895c992a5f38aa576eaf1b787dac01d4b6dcd421d5

Observation c3e35a3f-36c7-449b-b665-11c3fdf4c2da · outbound

This paper cites mhubert-147: A compact multilingual hubert model,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion mhubert-147: A compact multilingual hubert model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.670869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.183562Z digest=sha256:727e58fa15ac20dff7062dcdbe0b14bb3be82ad8e241f6894fbeff2e8a2710de

Observation 10c8d12d-7494-4972-a593-9ecf7d8b2fb6 · outbound

This paper cites V oice privacy- investigating voice conversion architecture with different bottle- neck features,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion V oice privacy- investigating voice conversion architecture with different bottle- neck features,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.661268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.186998Z digest=sha256:375dad192b3b20ed250e293420fb81002f8122931d112949bce5660fddb3ecc7

Observation 92ec07ed-625a-468a-9f3f-fbd34e7cb781 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.651408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.190765Z digest=sha256:36dca59b72d9d5b1fb8f7a113f716e96021aae62ca2a6b8bc6f7fd75a5f4677a

Observation 4d272469-5516-4d6f-8533-831cf5137352 · outbound

This paper cites Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.194507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.194507Z digest=sha256:dc76aa6dde98652cb491d3c20cd76bc5fabcbad58fe26cd65e39934f1fbb0ff6

Observation 5910ef84-2008-41a0-a9c6-2d41bbdbac57 · outbound

This paper cites Emotion intensity and its control for emotional voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Emotion intensity and its control for emotional voice conversion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.636704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.197913Z digest=sha256:560b58da0cb8714dd1564ac1962afb93cce9f4157009454491701b1515af53ae

Observation 20b2f702-bbf1-430c-b7e3-baaee31197d3 · outbound

This paper cites Accent conversion using pre-trained model and synthesized data from voice conver- sion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Accent conversion using pre-trained model and synthesized data from voice conver- sion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.627065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.294744Z digest=sha256:4393b6228b1b9e4b1ead22b581bd784bdeda9172293e77eb5a86b17681e42af9

Observation 2e0dc039-9d98-47b0-96f2-9f98dbf965b3 · outbound

This paper cites Improving pronunciation and accent conversion through knowledge distilla- tion and synthetic ground-truth from native tts,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Improving pronunciation and accent conversion through knowledge distilla- tion and synthetic ground-truth from native tts,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.616973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.298990Z digest=sha256:3b84cd0f5f7b8253c40e08703a5d5fecaf676d00ef985bfa5a7d6b975050d747

Observation e8cf2831-0a5b-4443-aa8a-5b00fb0bde59 · outbound

This paper cites Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.607359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.302169Z digest=sha256:a90d69dd285cce300bf23cc51d02c567c6d5538ef86da5dc7a4f3403112c544f

Observation 9d68f57f-7d99-40b1-9cb3-0fcb5d3b92f9 · outbound

This paper cites EmoCat: Language-agnostic Emotional Voice Conversion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion EmoCat: Language-agnostic Emotional Voice Conversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.306160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.306160Z digest=sha256:bfd8a78551ae996866bab41dabbc7b18b2259f9ef433d9abaf68fd554e20db4c

Observation 95c0fc03-f874-48e2-8517-fe0d1d91aed8 · outbound

This paper cites V oice conversion with just nearest neighbors,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion V oice conversion with just nearest neighbors,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.597489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.309572Z digest=sha256:c9157328a56b19509e8b6cf9807fdc27ad9f9eaffffc94ab9c9eecdda9a8b06b

Observation 03777a17-2dc4-4fa3-8d3d-4eb9d74b8b6a · outbound

This paper cites Disentangling prosody representations with unsupervised speech reconstruction,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Disentangling prosody representations with unsupervised speech reconstruction,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.587981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.312653Z digest=sha256:0b2ee7addb162ba4f7a2b017405281c4d1679cfb200f3d015cb9a6b7ff9490da

Observation b804b6f8-c9b6-4dea-a3cf-412b9b014202 · outbound

This paper cites Using joint train- ing speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Using joint train- ing speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.577501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.315921Z digest=sha256:02557836908bf3561d46a0fc671d2c5cc544666a721711d7f283530bed13fba4

Observation 70f5d6af-002d-4d63-b689-720df816fb73 · outbound

This paper cites X-e-speech: Joint training framework of non- autoregressive cross-lingual emotional text-to-speech and voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion X-e-speech: Joint training framework of non- autoregressive cross-lingual emotional text-to-speech and voice conversion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.566994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.319244Z digest=sha256:30e9b032f8c2bb7589fe4ce365b357765b898f24a22015ee548706b6171034e3

Observation 2613bdb2-d8e8-4fd1-acf3-8c897598eabc · outbound

This paper cites Zse-vits: A zero-shot expressive voice cloning method based on vits,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Zse-vits: A zero-shot expressive voice cloning method based on vits,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.556372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.322515Z digest=sha256:3fd733817e10b772d481ebaf3e43c53253317033850a86a29cd9591fae5b182b

Observation 465a0c8c-bb1e-4e75-8021-eb5edcc90a2c · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.325815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.325815Z digest=sha256:62367263fce508a6837032dcda3e3ce0d7e2de7cdd2c4e32c2d4f924ba83354d

Observation da1d2bd3-e7b9-4952-b276-cfb28b0a60f6 · outbound

This paper cites VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Design.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Design

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:55:20.427901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.329181Z digest=sha256:edebbef1d7548461079163b7f9022c86c8c62b86ada98f02c52fa4576628951d

Observation cdd4abe5-1b3c-4335-b5e3-418d8ea4ea1a · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.545193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.333182Z digest=sha256:1499b59b5f8daf9c6d6d95e12c43651b0d38343764baafe358e1c886c63f5025

Observation 2f69c47b-1579-421b-8de8-6ab66a58c7f5 · outbound

This paper cites Libritts: A corpus derived from librispeech for text- to-speech,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Libritts: A corpus derived from librispeech for text- to-speech,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.534957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.336222Z digest=sha256:3bc03af7fc92d40e5fc3c51fdbda6b18beb183ac480337405f711ef5f9e44851

Observation fe8a54c9-162f-4886-a27f-b26498b66e42 · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.524960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.339190Z digest=sha256:932e64a1f39a4618e2044e7b131c54771a370407384824fe2f4b5aef3b0c30c7

Observation b1d1b093-dbc2-468e-83ea-b92226d3d12a · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.342184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.342184Z digest=sha256:b1a41c454614a73c31ea7ff6d5a0656a5dd6033747b351bc9a77b18b4b648faf

Observation a5e2b151-c0e1-4e2c-beb8-922275a56e54 · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.345371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.345371Z digest=sha256:f28a7ac65d7a1ac8f3f6ef5ebc663013925514094a36bb5538f6d1a6ff4d1b54

Observation 76475748-f2ec-45c9-b57d-617c56837d54 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Robust speech recognition via large-scale weak su- pervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.349892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.349892Z digest=sha256:8b5eeab79405e9f088f23ec4648a8794136a68331cc2f69eb96126413fa8581b

Observation 2c0e4eed-aa8b-4d57-84ec-3030fd4b73d1 · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.353059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.353059Z digest=sha256:caca6b2023da8d082291e7deb79dbe80eaf271fca8456a613c49878f068559bf

Observation 4933fb26-7b1b-41ba-a702-96cae9263b1a · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.508696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.356312Z digest=sha256:9111d2b1a1b2b42d4c33665641b8af364e4fd16090fdac261b35f731435aa488

Observation 90739c25-a623-4060-a345-6fcf22a17a06 · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion A comparison of discrete and soft speech units for improved voice conversion,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.498394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.359499Z digest=sha256:a233ec736f699fbf2ff3ecab5a080b765b0ad09ea44792868f5d0a02d36ec291

Observation 48e54d4e-5efa-4b12-8455-e53b9a662115 · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Scaling speech technology to 1,000+ languages,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.362733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.362733Z digest=sha256:42abda4ff9866935f68de64db19633a1ded03160831ca8b094a36174198d4ff8

Observation f0698fca-027e-419d-a48e-4d851c7edf2b · outbound

This paper cites A database of german emotional speech.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion A database of german emotional speech

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.482825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.365775Z digest=sha256:bad24b196730ac05d122d8ec3c54e0e24a04719546dfd7c85767494c3540fcc5

Pith citing papers

Observation ba314579-786e-4c6b-a2b0-691e19b159d0 · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.110093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.110093Z digest=sha256:2913861f2dc309c8fcb9ddba7bf69dc9f556ffe754dd35a49fab320e266fd55f

Observation ca35d848-c3ba-4ab9-bd33-71cd9388e510 · inbound

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection cites this paper.

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:36.608042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T19:11:30.638672Z digest=sha256:ab21e4f66ba49bc2a9f2467a274de3d33af030649b3eb6606a7ef1c2740e54de

Observation ad6d5a5a-38a2-4068-a167-03f84c166bfe · inbound

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026 cites this paper.

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026 Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:07:18.248088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T21:39:11.265338Z digest=sha256:7bd374f6c488d55ce3cbc33109c4cace93cbf0a6c87bf70a73b4bdef15944379