Pith. sign in

Paper Citation Record · LEDGER

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

As of 23 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 3 inbound Pith citation observations for arXiv:2506.04013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04013 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.365775Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.110093Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:07:18.246563Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73dae815-3c55-4526-84c8-912be14e0f2d · outbound

This paper cites Conventional VC models perform well in replicating speaker identity but struggle when the target speech is highly expressive.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Conventional VC models perform well in replicating speaker identity but struggle when the target speech is highly expressive

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.855135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.100974Z digest=sha256:e017a83025b392e00c56d1454e18f37a361bc5c6fdfec6cc6e770f1549693ace

Observation 60d15bb1-2e4b-439f-9dfc-c7c6aa33ce69 · outbound

This paper cites These models typically required text supervi- sion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion These models typically required text supervi- sion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.845772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.105713Z digest=sha256:90c8bf608016ea4c50e4ce587e040c9ae64055df2be45be3a0a69dd3570786f4

Observation 2b05af34-6da4-4dab-be57-1f5d46dc096c · outbound

This paper cites an unresolved cited work.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:20.837000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.113786Z digest=sha256:f88bc40f8b6a8736b8130db329675b536e35b9be634f111eeee544b58bd22001

Observation d66535c6-1087-46f6-aa49-3738da5f3ba0 · outbound

This paper cites All datasets are English and total duration is around 228 hours with more than 920 speakers.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion All datasets are English and total duration is around 228 hours with more than 920 speakers

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:55:20.825514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.117679Z digest=sha256:0b0d09c2eed3be74594f4c31420fad33875788f9ba76de23039476dff1de05b9

Observation 89c442f2-6f5c-4f5a-8dd8-d074ee68083e · outbound

This paper cites an unresolved cited work.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:20.816313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.121966Z digest=sha256:70d1dfce9b97abf917b7b9cd869aa1d09722eba8530d9af0e1c47896bd5fc93c

Observation e262b5f6-fc04-414b-a4fc-43b864e1852d · outbound

This paper cites an unresolved cited work.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:20.805968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.125796Z digest=sha256:768cc0c6587474ad754fdb10ed5c1849b64ece9016a996e58a88932a12f7b28c

Observation 92367e0f-9cda-4004-997d-3c19c61642e7 · outbound

This paper cites Styles2st: Zero-shot style transfer for direct speech-to- speech translation,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Styles2st: Zero-shot style transfer for direct speech-to- speech translation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.760631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.151933Z digest=sha256:5d6a2c767c192192afaae2439da292bb745dc2ee824316cfa435cc6b1cec8757

Observation c660e27c-596c-48f2-9e38-a14f95eef46e · outbound

This paper cites Chil: Computers in the human interaction loop,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Chil: Computers in the human interaction loop,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.796781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.129443Z digest=sha256:1a6575ad5c1749d24f8e440fd34bbe8a65824b91c34d583b71dd9de5489d5429

Observation ad0ef1ef-5d7f-465f-b9f1-e912a536f3a0 · outbound

This paper cites Towards an open-domain social dialog system,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Towards an open-domain social dialog system,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.788033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.132609Z digest=sha256:35f1a68b9e603c4dba2fe217f853a6f9936298998767298d633ced9ede7deca1

Observation 25550670-15b0-48b7-9b2b-280828d1941b · outbound

This paper cites Face-dubbing++: Lip-synchronous, voice preserv- ing translation of videos,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Face-dubbing++: Lip-synchronous, voice preserv- ing translation of videos,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.779015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.136544Z digest=sha256:5ee2a2949d78d947ac511322b78021ea28a591531b6f3b7c16453d39c0f59415

Observation c9a77ed3-3ba0-4d54-8e9e-3620d532acd9 · outbound

This paper cites Findings of the IWSLT 2024 Evaluation Campaign.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Findings of the IWSLT 2024 Evaluation Campaign

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.139760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.139760Z digest=sha256:b3290d06e0c3deb23bad1c3639d5bc439a569a5dd8064db6f4718a4019e1e03a

Observation ba314579-786e-4c6b-a2b0-691e19b159d0 · outbound

This paper cites Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.110093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.110093Z digest=sha256:de03527b95bbf70124b9b1837d81ef01ec999cd6e39cff665a59800ef4c681f4

Observation 0ab7ae64-8fa2-4efe-ba84-aec402de6f5e · outbound

This paper cites Simultaneous translation of open do- main lectures and speeches,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Simultaneous translation of open do- main lectures and speeches,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.770020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.144697Z digest=sha256:432a2078457e96937a047433ba92d2882c5e2ab9bfb01d6b363bf565bfbcfa6e

Observation c9f3b311-7eea-48d9-9bb0-b849850a81b0 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.148186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.148186Z digest=sha256:99fe3e9437dedbb0a814c86a870807cf003e7c66853608f2113d4739a4efb10a

Observation 5a98814b-9b96-409e-8a88-75df122dd784 · outbound

This paper cites Limited data emotional voice conversion leveraging text-to-speech: Two-stage sequence-to- sequence training,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Limited data emotional voice conversion leveraging text-to-speech: Two-stage sequence-to- sequence training,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.750351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.155032Z digest=sha256:ba46143d9aee4bea5ff24474dcfa44b114b8444df6f818dd766685f431d01a51

Observation 28febc5b-ccf0-467a-95d5-5131f4e549e8 · outbound

This paper cites Nonparallel emotional speech conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Nonparallel emotional speech conversion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.740558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.158385Z digest=sha256:cd4f05e33a6d099fb8825e2abf2530e0d4701af05b9cc18d669eaa2701ef8144

Observation c9294002-d084-43b3-b41f-78731dc9f955 · outbound

This paper cites Nonpar- allel emotional speech conversion using vae-gan.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Nonpar- allel emotional speech conversion using vae-gan

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.730923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.161449Z digest=sha256:6126146e4753d8f4b2c12b2e78c246e9448632a342a222d41e45fc783954ec00

Observation 90cacd4e-433b-4991-b40f-d1fb14fc9821 · outbound

This paper cites Textless speech emotion conversion using discrete and decom- posed representations,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Textless speech emotion conversion using discrete and decom- posed representations,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.721511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.165248Z digest=sha256:7bd5838b41e3228da94305cd5340c5b383b78587fb85c1f343bdf0367af9e932

Observation 3e6b91b7-04d0-456e-8398-c26d7a5e7f41 · outbound

This paper cites Hiervst: Hierar- chical adaptive zero-shot voice style transfer,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Hiervst: Hierar- chical adaptive zero-shot voice style transfer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.711516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.168157Z digest=sha256:5ea29f8f3f00e49355bb3f90b15852d687a940fb477381f027ffbdca1b44d907

Observation c5850056-bcc9-495f-af8a-9b5d27e67111 · outbound

This paper cites Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.701184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.172054Z digest=sha256:b4a203b7233b6055198ee2c39d1695da1f32dc83e244dad2a995f10284c4ca39

Observation aec52e0c-9e09-44f4-8d98-c38be0db8dcc · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.690149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.176032Z digest=sha256:846dd7ce349e33542632da321cfb304639bcd7c000046d66efacfa6e7783d180

Observation da85ec13-5aee-4693-b038-fcfb0199455f · outbound

This paper cites Freevc: Towards high-quality text-free one-shot voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Freevc: Towards high-quality text-free one-shot voice conversion,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.680603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.179658Z digest=sha256:027589bb8a159fd5f3b6a0a8e132b7f1578422d309154f52915fe3d956b96878

Observation c3e35a3f-36c7-449b-b665-11c3fdf4c2da · outbound

This paper cites mhubert-147: A compact multilingual hubert model,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion mhubert-147: A compact multilingual hubert model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.670869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.183562Z digest=sha256:b711f12b25de7716ecf8d245fc9bf9cd221027f53585af87ec608089b7e48fa8

Observation 10c8d12d-7494-4972-a593-9ecf7d8b2fb6 · outbound

This paper cites V oice privacy- investigating voice conversion architecture with different bottle- neck features,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion V oice privacy- investigating voice conversion architecture with different bottle- neck features,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.661268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.186998Z digest=sha256:1c9d3282530a8ac9a72325c548aea14ca3692b49d88439f1ce2173dab7d8b88e

Observation 92ec07ed-625a-468a-9f3f-fbd34e7cb781 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.651408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.190765Z digest=sha256:1a67e3fff8d63067709545849b0e80a3399fa653a1cf02ebf2f660d1b62ea036

Observation 4d272469-5516-4d6f-8533-831cf5137352 · outbound

This paper cites Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.194507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.194507Z digest=sha256:16c990f3ecf0690e80601f12a2cf70d435336981231e4e64f58eeac83890da58

Observation 5910ef84-2008-41a0-a9c6-2d41bbdbac57 · outbound

This paper cites Emotion intensity and its control for emotional voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Emotion intensity and its control for emotional voice conversion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.636704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.197913Z digest=sha256:4d8fc6d3bd7bbedffc330fe9c347c42be71dce37eaaa779e5b9c1aabf927812f

Observation 20b2f702-bbf1-430c-b7e3-baaee31197d3 · outbound

This paper cites Accent conversion using pre-trained model and synthesized data from voice conver- sion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Accent conversion using pre-trained model and synthesized data from voice conver- sion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.627065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.294744Z digest=sha256:e47b2b1f3e3fd9a96dd05e7b0db1fed0d2be0c68ddc871137c08fb5a42ba40c5

Observation 2e0dc039-9d98-47b0-96f2-9f98dbf965b3 · outbound

This paper cites Improving pronunciation and accent conversion through knowledge distilla- tion and synthetic ground-truth from native tts,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Improving pronunciation and accent conversion through knowledge distilla- tion and synthetic ground-truth from native tts,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.616973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.298990Z digest=sha256:0ace596595a798ade3f0fb4490372316a6b74a5f0b033dd6c950a22f9ce8deaf

Observation e8cf2831-0a5b-4443-aa8a-5b00fb0bde59 · outbound

This paper cites Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.607359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.302169Z digest=sha256:376a8c2e319f6af93fcc58c2b3011439d88d8bc308c6d57e7e9f7e7f4d91e9fc

Observation 9d68f57f-7d99-40b1-9cb3-0fcb5d3b92f9 · outbound

This paper cites EmoCat: Language-agnostic Emotional Voice Conversion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion EmoCat: Language-agnostic Emotional Voice Conversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.306160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.306160Z digest=sha256:7636def50165f400b275c5905ee5c5638f97c4f69d874813e9d5f74158fb8077

Observation 95c0fc03-f874-48e2-8517-fe0d1d91aed8 · outbound

This paper cites V oice conversion with just nearest neighbors,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion V oice conversion with just nearest neighbors,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.597489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.309572Z digest=sha256:03f425009cfe4852011be67f7e4fe50bd0aa9c8c19ad1a2eded6c947d0c8c09d

Observation 03777a17-2dc4-4fa3-8d3d-4eb9d74b8b6a · outbound

This paper cites Disentangling prosody representations with unsupervised speech reconstruction,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Disentangling prosody representations with unsupervised speech reconstruction,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.587981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.312653Z digest=sha256:bd9a57145fcf0688e37750218c4c8557bf4b85847d555ede1ffb21292e87bddd

Observation b804b6f8-c9b6-4dea-a3cf-412b9b014202 · outbound

This paper cites Using joint train- ing speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Using joint train- ing speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.577501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.315921Z digest=sha256:90c1f6c0b8a0345792ff1745af35dc64cba486943ceaa91652f4f3989be47d84

Observation 70f5d6af-002d-4d63-b689-720df816fb73 · outbound

This paper cites X-e-speech: Joint training framework of non- autoregressive cross-lingual emotional text-to-speech and voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion X-e-speech: Joint training framework of non- autoregressive cross-lingual emotional text-to-speech and voice conversion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.566994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.319244Z digest=sha256:139d417fec57ee6d0d941716e6185dff69311f9a2d37aecd88d0c85f7f72e89f

Observation 2613bdb2-d8e8-4fd1-acf3-8c897598eabc · outbound

This paper cites Zse-vits: A zero-shot expressive voice cloning method based on vits,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Zse-vits: A zero-shot expressive voice cloning method based on vits,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.556372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.322515Z digest=sha256:149ea8e9b370db9d078b490b242582434c6403db56fa88a597d6f03290951ded

Observation 465a0c8c-bb1e-4e75-8021-eb5edcc90a2c · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.325815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.325815Z digest=sha256:d462367b8cc2183e5b889c22ea734692d4a62338f15800537e0df27cb101dd48

Observation da1d2bd3-e7b9-4952-b276-cfb28b0a60f6 · outbound

This paper cites VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Design.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Design

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:55:20.427901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.329181Z digest=sha256:3a79756f901e868bd3b9cd4e7da104aac34222cda6d248302a377c9123e48f4a

Observation cdd4abe5-1b3c-4335-b5e3-418d8ea4ea1a · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.545193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.333182Z digest=sha256:d22e29d87926ca0db777003fd7622bed076f1cff2f07207b2618fb70a68b47f8

Observation 2f69c47b-1579-421b-8de8-6ab66a58c7f5 · outbound

This paper cites Libritts: A corpus derived from librispeech for text- to-speech,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Libritts: A corpus derived from librispeech for text- to-speech,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.534957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.336222Z digest=sha256:dbbfdae324bcfc937d65e684d3265098fb9efe94af6b2de1545fd193fd5aa8a9

Observation fe8a54c9-162f-4886-a27f-b26498b66e42 · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.524960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.339190Z digest=sha256:7af8c7038009b708f210a707593de7cba0bf3037663f396093b9253b411921b0

Observation b1d1b093-dbc2-468e-83ea-b92226d3d12a · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.342184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.342184Z digest=sha256:5953eb85ebd306980746fed6aeee28ea6605481695b58434019d72a7ae3cffa7

Observation a5e2b151-c0e1-4e2c-beb8-922275a56e54 · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.345371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.345371Z digest=sha256:7486cd7a982bf867da20f00ed90f01649dffb3d1df424061379ced845086dde0

Observation 76475748-f2ec-45c9-b57d-617c56837d54 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Robust speech recognition via large-scale weak su- pervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.349892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.349892Z digest=sha256:00a7bbc49e1805f84d42c5f5e9e03d87bd870d56d0aa18c8824d171e728d3daa

Observation 2c0e4eed-aa8b-4d57-84ec-3030fd4b73d1 · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.353059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.353059Z digest=sha256:f3f6cf353ff4f41caf9f211bd39589f86d821c56a400c68aa3fb2dbf8bb5bcf0

Observation 4933fb26-7b1b-41ba-a702-96cae9263b1a · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.508696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.356312Z digest=sha256:4207180d98fd834cb5e661f688ad8864eac9daffa40a9bf89e28623af47686ab

Observation 90739c25-a623-4060-a345-6fcf22a17a06 · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion A comparison of discrete and soft speech units for improved voice conversion,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.498394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.359499Z digest=sha256:8fc7bcdcb7e6f296077d224a9f47d4daf7539ac4cd19b53871ea4ed4c4cb145d

Observation 48e54d4e-5efa-4b12-8455-e53b9a662115 · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Scaling speech technology to 1,000+ languages,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.362733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.362733Z digest=sha256:4897091cfb8cd65fb3f43af32df0d4886d5dc8d1cd613d7a130a6313bb7afdaf

Observation f0698fca-027e-419d-a48e-4d851c7edf2b · outbound

This paper cites A database of german emotional speech.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion A database of german emotional speech

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.482825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:55:20.365775Z digest=sha256:a608183ea9ec6e56b1dc965e999f00aa95af8f6c0a488312cf8d123916ecf752

Pith citing papers

Observation ba314579-786e-4c6b-a2b0-691e19b159d0 · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.110093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.110093Z digest=sha256:de03527b95bbf70124b9b1837d81ef01ec999cd6e39cff665a59800ef4c681f4

Observation ca35d848-c3ba-4ab9-bd33-71cd9388e510 · inbound

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection cites this paper.

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:36.608042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T19:11:30.638672Z digest=sha256:ca1feeb33fcac4e8acbc60aeca4cac070e12c088be7edc47454489dcc5269ded

Observation ad6d5a5a-38a2-4068-a167-03f84c166bfe · inbound

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026 cites this paper.

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026 Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:07:18.248088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T21:39:11.265338Z digest=sha256:d34a17cc6f12270867b719498eab38078eada59680409684a17219013df3d5de