Pith. sign in

Paper Citation Record · LEDGER

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training

As of 19 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 3 inbound Pith citation observations for arXiv:2501.04416.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04416 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:38:27.559518Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:57.241564Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T06:00:36.712633Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved5
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1df7a0b5-d9fc-47da-b9fe-f511f760defc · outbound

This paper cites Autovc: Zero-shot voice style transfer with only autoencoder loss,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Autovc: Zero-shot voice style transfer with only autoencoder loss,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.984045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:26.782695Z digest=sha256:d657db557fa7bacb8b6780b020f182e0a1d2a875a259d780b1c9e56b8a184e9a

Observation d5e025b4-12bb-4de3-99b2-75e414909a26 · outbound

This paper cites LM-VC: zero-shot voice conversion via speech generation based on language models,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training LM-VC: zero-shot voice conversion via speech generation based on language models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.891618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:26.845881Z digest=sha256:9920791504dac16d0cbf2060f131b06df46077c29c7db8d86b1c561c12023f22

Observation 45f85feb-2479-44f8-bcbc-164613bc1013 · outbound

This paper cites Multi-speaker and multi-domain emotional voice conversion using factorized hierarchical variational autoencoder,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Multi-speaker and multi-domain emotional voice conversion using factorized hierarchical variational autoencoder,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.881186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:26.880840Z digest=sha256:dcd1aae190a69440808c5a031efb13a9258d3ac1929fe3fdb440bb7c8d3b5209

Observation c92eb10f-b57e-4391-a1df-f6e41954f6e9 · outbound

This paper cites Un- paired image-to-image translation using cycle-consistent adversarial net- works,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Un- paired image-to-image translation using cycle-consistent adversarial net- works,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.870545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:26.885044Z digest=sha256:45ae415bc12ce2b1d5763b0ba66ba9d1113b56ed3ce4fbe8cb13f5308cf6005f

Observation ff3d329e-625e-408e-b179-c0ad19ca6c09 · outbound

This paper cites Stargan: Unified generative adversarial networks for multi-domain image-to-image translation,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Stargan: Unified generative adversarial networks for multi-domain image-to-image translation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.824001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:26.889107Z digest=sha256:643ccaf8c38043ed35d255931eb6e067a0d85a19598489eb2bd966a7caaf6d63

Observation 3c21a3e6-b082-4ddf-8e8f-2c549c34542b · outbound

This paper cites CV AE-GAN: fine-grained image generation through asymmetric train- ing,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training CV AE-GAN: fine-grained image generation through asymmetric train- ing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.653702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:26.892933Z digest=sha256:3c39ee093288962c64c85804b4c673bd2f5de84ea0203d0f042db9ad2de4ef90

Observation f0bf70aa-16e5-443c-bcd5-33d9b4846442 · outbound

This paper cites Dis- entanglement of emotional style and speaker identity for expressive voice conversion,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Dis- entanglement of emotional style and speaker identity for expressive voice conversion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.642700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:26.939420Z digest=sha256:e1f432323a80c2b907b5933a35b3cb155650c5da593c22f319ca88b798f5d03b

Observation 71c13269-88fb-46bc-8134-d4f37fdf57f4 · outbound

This paper cites DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:38:27.068728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:38:27.068728Z digest=sha256:2fd29b78422cc78d44ca92fce8374dbcc405dd89e31b8101dc36087a4ca24bd9

Observation ebb65c78-68f9-4535-93ba-3f922a831b4c · outbound

This paper cites Delivering speaking style in low-resource voice conver- sion with multi-factor constraints,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Delivering speaking style in low-resource voice conver- sion with multi-factor constraints,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.630637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.131149Z digest=sha256:181a3c917a42760931a02beeaa2f9012ae95f216798e75630730b5f104448790

Observation 8d839277-c962-4bff-8a62-c94d3bc21b89 · outbound

This paper cites Speechsplit2.0: Unsupervised speech disentanglement for voice con- version without tuning autoencoder bottlenecks,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Speechsplit2.0: Unsupervised speech disentanglement for voice con- version without tuning autoencoder bottlenecks,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.619230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.134918Z digest=sha256:64561ac4c20ecdf132508f94e84a5ab65a583d02d0c359cef16e1ec74b761017

Observation 025ed1cf-b6ae-4d85-86bf-6bf929519ce9 · outbound

This paper cites METTS: multilingual emotional text-to-speech by cross- speaker and cross-lingual emotion transfer,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training METTS: multilingual emotional text-to-speech by cross- speaker and cross-lingual emotion transfer,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.608502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.139289Z digest=sha256:32054593c178687b7bad33c58c49abdaf2cfac5f79935f4abda0967dc2d2cb6e

Observation c9fe511f-ed9b-40b2-a23c-d02eb8393644 · outbound

This paper cites X-vectors: Robust DNN embeddings for speaker recognition,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training X-vectors: Robust DNN embeddings for speaker recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.515621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.143642Z digest=sha256:64412a42d92e2dae5637dc21e9f24be0990cf70fcafdd3ad4ed743a1eceba3e0

Observation c0c7190d-bc8c-48d9-8c95-28aa2269c5af · outbound

This paper cites Towards improved zero-shot voice conversion with con- ditional DSV AE,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Towards improved zero-shot voice conversion with con- ditional DSV AE,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.474355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.147011Z digest=sha256:ab73961a8c86f7437f00d83b3fa32a37e60c5dea6b52b7cc9dff46c57d394d5e

Observation be2425b9-c148-4d31-8b20-688670926d6d · outbound

This paper cites SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:38:27.150840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:38:27.150840Z digest=sha256:7c8bf93e754a8855608b6597cbe521cf51edd5fdbf4ae95f5364fad62de4b115

Observation 5501901c-2b1e-43a2-aee2-62b84516628c · outbound

This paper cites Audiolm: A language modeling approach to audio generation,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Audiolm: A language modeling approach to audio generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.461424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.154452Z digest=sha256:3f2a8cc4faf295ced9497c4b749337cfb1557b6503f127ad9fc5f1d48a0edd21

Observation e4880a84-9395-4c86-a20c-eae6afe12ed3 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:38:27.158207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:38:27.158207Z digest=sha256:36cd29279750b32ee595126b6d00c01abe5cb5982e8cb34f4b9bd37f86cb04cf

Observation 6a4c0181-e05d-4e68-b28c-dcf13b4b7d8f · outbound

This paper cites High Fidelity Neural Audio Compression.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training High Fidelity Neural Audio Compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:38:27.162462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:38:27.162462Z digest=sha256:abb97836d344e195afb9a78e60a6584143a194f5967be1261844192972617a71

Observation 0080efa5-e570-4399-9453-3477c3169c63 · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.449645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.228670Z digest=sha256:33b6a7e8d30274c0c9c951c4a5784e209fc5514676a638559ee525bac198b82d

Observation 960db646-fca0-4de2-b32d-543c76c5f376 · outbound

This paper cites Make- an-audio: Text-to-audio generation with prompt-enhanced diffusion mod- els,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Make- an-audio: Text-to-audio generation with prompt-enhanced diffusion mod- els,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.363183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.329892Z digest=sha256:387e4983477efa95da841776b77ae0f61cd51b5ee5e2ee1359100635ceac1ace

Observation d9b8deb6-9169-493a-983d-d6978419faa3 · outbound

This paper cites Unsupervised speech decomposition via triple informa- tion bottleneck,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Unsupervised speech decomposition via triple informa- tion bottleneck,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.298553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.362328Z digest=sha256:52ac47499715472b94bdfe09ef4b8246917d3216afff1ee6f9c620162ab5837c

Observation 2ab3d7b2-b886-4df1-a761-1ff9efb53034 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training SoundStorm: Efficient Parallel Audio Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:38:27.366566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:38:27.366566Z digest=sha256:57c69bbe0c89271fdfec0a464e5667c0a5e4dd42de99b56c747977a22fc4a2fd

Observation 4579a6c1-2436-437f-bf65-76a6c560d63e · outbound

This paper cites Neural discrete representation learning,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Neural discrete representation learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.285838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.371245Z digest=sha256:f458d7336ea6d49a3e2cd6104ea99b6bb504aac5c25b97bffa96321b43f37bed

Observation 7be335ad-dc9c-43d7-8fff-da3dfe7f1a9c · outbound

This paper cites One-shot voice conversion by separating speaker and content representations with instance normaliza- tion,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training One-shot voice conversion by separating speaker and content representations with instance normaliza- tion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.271748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.374999Z digest=sha256:d62c17d901b0a7163ca6d29dc7a01458669709c8aace1228445d064ae956c5ac

Observation 20abf3fc-0e61-4064-95f7-5bf32eb75867 · outbound

This paper cites StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:38:27.690224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.382817Z digest=sha256:9b9e60d5dd7a4c37487d25c470409c3109408d0bb5445d170ac260e52df7b80f

Observation 1bf60e37-ae4e-4948-ad92-2d1210e06185 · outbound

This paper cites Unsupervised domain adap- tation by backpropagation,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Unsupervised domain adap- tation by backpropagation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.246333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.388188Z digest=sha256:1acc48913b2eddff3e561027377ebd2a7586287a1588ce411a03913d612d559b

Observation 5c30abd7-b04a-407c-868f-5fbcbe03c9de · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training MLS: A large-scale multilingual dataset for speech research,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.055376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.397332Z digest=sha256:ddcf9a937f8ec64db9c4dea23b9d073dca0202b23e46a7a9dbc777436c926213

Observation 2bdae9fb-c9db-414d-9c44-a4a62fef10d9 · outbound

This paper cites Token-level ensemble distillation for grapheme-to- phoneme conversion,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Token-level ensemble distillation for grapheme-to- phoneme conversion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.043300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.401501Z digest=sha256:748a97f32914c1d09a22c9423712a9fa9563a72c9d0f6faeeba8c77f47da6c42

Observation b973d6e4-6f0e-4eb4-aefc-9fb13f247a60 · outbound

This paper cites Style tokens: Unsupervised style modeling, control and transfer in end- to-end speech synthesis,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Style tokens: Unsupervised style modeling, control and transfer in end- to-end speech synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.032611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.468495Z digest=sha256:45f9b0aa6c9047551181da5aa4cb916fe6d67dd52073c9bf698bf76040413813

Observation 05ca7bce-a19c-423a-9cd6-00f7a091d7b4 · outbound

This paper cites Emotional voice conversion: Theory, databases and ESD,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Emotional voice conversion: Theory, databases and ESD,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.020819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.550656Z digest=sha256:96ce41498c42e16717bd054f2cf3a48698f9462a17c9d967a00804b9745ef5d1

Observation ed8c9e68-868f-4a98-a0bf-159634135bc6 · outbound

This paper cites Visualizing data using t-sne,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training Visualizing data using t-sne,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.007498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.555253Z digest=sha256:df16c8cfe8a79445f7a82c77ef3147e99779575503ded0d2f49f348e62f120fb

Observation 36de0868-4381-4f39-95f7-c42d84f6605f · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation,.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training emotion2vec: Self-supervised pre-training for speech emotion representation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:27.895176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.559518Z digest=sha256:1a51a18188d2ffbbdc8221031304ad2068953481d1db0eb4b1bb5fd1b527cedc

Observation f42d3be1-0067-486f-afd4-70609f15727a · outbound

This paper cites 37 of JMLR Workshop and Conference Proceedings , pp.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training 37 of JMLR Workshop and Conference Proceedings , pp

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:38:28.110573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.393542Z digest=sha256:9b4155e1fc1cf810e5864e286e18baa6df9d3d2efaac6b0ed343e780c21b8074

Observation 426fdf7d-81c9-4c57-8aca-23013801b639 · outbound

This paper cites 664–668, ISCA.

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training 664–668, ISCA

Reference 2019

Resolution
parse uncertain
raw_fallback, observed 2026-08-10T21:38:28.259490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:38:27.378725Z digest=sha256:3f0fcf402c373cdbc7a6ffa3ce9de89b89ef64e59f14c92ab48ee589af8b2b07

Pith citing papers

Observation 76b09370-129b-436f-93aa-eec1a6c3e74e · inbound

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model cites this paper.

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.241564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:57.241564Z digest=sha256:59b523864aa23b9dd6fe7ad9ff9a9362b0a27596c5334f7ab18cd0a35f3b2418

Observation c3b103a8-05ae-4853-9180-1373feef1466 · inbound

ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization cites this paper.

ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:26.433423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:26.433423Z digest=sha256:661d12db982b1a7621500bef33f4dc86226f664f51baebf68f317d7b4d4c3fc0

Observation 02e20413-4c61-451a-b3d6-628509b92948 · inbound

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection cites this paper.

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:36.715260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T19:11:30.638672Z digest=sha256:80b439d423c64b2b2fae2b4541b70ea4b5e96bcb5b69ff2c2ad707a40d490b25