Pith. sign in

Paper Citation Record · LEDGER

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.16691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16691 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:58:58.610208Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:31:12.289334Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0f6a7af-550c-4223-a65f-d92468978971 · outbound

This paper cites YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.253851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.253851Z digest=sha256:59c481d7f4bcf4af1f33725cf8bd34dc401ddbee35b729aa6bc73e4b76d3d748

Observation bf85f6d3-9313-4d5c-9dff-88dffc891119 · outbound

This paper cites Towards Robust Speech Representation Learning for Thousands of Languages.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Towards Robust Speech Representation Learning for Thousands of Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.363483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.363483Z digest=sha256:d6b00cefe8c72bf2c4ef567932709fc592c42a882c8e4f65f54acbdab4e7c9ea

Observation 13a8820e-17dd-4d1a-95bf-a9e3cef8c4c1 · outbound

This paper cites Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.506395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.506395Z digest=sha256:f685ca89355b0155eaf11a3c37b2cc079af7ccd368b27fe6d97555e3eaabcc6d

Observation 1adbc139-7c13-48a6-839f-2e209a6be972 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.651626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.651626Z digest=sha256:e22c1a9f38720a853c4b568bd513b98a112ee1415d203e47e39b48eae5afd801

Observation 5ad2fdbb-073f-4f29-a04d-568c9628d2a4 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.747744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.747744Z digest=sha256:2e527e040ff5edbb3d0dfaf3d9786186cf4482b09a3d0d9dda603a67a53f964e

Observation e23d19c3-e4d0-4c5b-9b03-a0f3341ab500 · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.060521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.060521Z digest=sha256:bbbb8a41981a9e88925625280e08441b603f5db147f5c53846735ff0ac9a5125

Observation 228b822e-6e8d-4a88-ad92-ce19c9c4def3 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.464868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.464868Z digest=sha256:fb87726b06626f1a4103d9c754f8338fed051201590d51435385dae141e24a4b

Observation 1c534e74-3d06-4aae-92a9-1ff10d953b22 · outbound

This paper cites SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:58:59.160738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:58:57.566166Z digest=sha256:658a3e25cfcc127e58c2ec4100438116e9abc68a209eddcbc76bdb4c49076d97

Observation 0b9c0f93-9ca5-4759-9b6e-400d1b5d827d · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Zero-shot Voice Conversion with Diffusion Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.726231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.726231Z digest=sha256:d48fe664aed87fb4409d0b63ade868affc1fd03a162f5eec6d40c2a6a96a0917

Observation b0bae56f-0cdf-4a18-b815-56dfb596753d · outbound

This paper cites In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:59:00.048756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:58:57.886020Z digest=sha256:5474492c11a8a4ea8846220bdfb77902b788f15d4b7d5e2e8f530d27fc85eab7

Observation 47aa4c1a-aabe-44ef-9e16-7d104fc9573b · outbound

This paper cites Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.976936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.976936Z digest=sha256:7091864c153755bf7cc0f55ccfe2e4a7411a7f397d5973a7940a02b0eb92f269

Observation e4d833d4-08ac-4fa7-8743-bdc62c644824 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.106744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.106744Z digest=sha256:85ed415997db6b7092a096575a4e78b6f52f51d72a94fc973c910329192ae39a

Observation ee506292-b81c-45b4-8312-f35493fea5df · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.237399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.237399Z digest=sha256:3792b2bc65f58420d8130907adb618eb9722b2adac721646305cf3485883005d

Observation 493a615c-a7b9-4057-bea5-f2569222ea15 · outbound

This paper cites StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.358652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.358652Z digest=sha256:01ccf2d57197dd8b738b39ed647aed761d6fe8bf490d62ab116bf7876b0ba9d3

Observation 74f6896e-0381-4b69-8a07-941b1548c3bc · outbound

This paper cites Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.610208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.610208Z digest=sha256:053c222824bccc27181b26db70511861f1725cd739a8fe96b7ef20e32854de57

Observation 495d6f42-2df5-42fd-a52b-50529f4c5772 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.472965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.472965Z digest=sha256:374eef9bb46f5e5165d0b0493e8d1f3153faf443b17b5165aa1fd3d2688283df

Observation 49016cc0-d922-442c-819a-db771d2bb707 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Common Voice: A Massively-Multilingual Speech Corpus

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:55.919987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:55.919987Z digest=sha256:c7dc6c948de6cc0b72760bdbad5c6e5665614aa787479a49b11f1c3b79a324ab

Observation 3efc365d-704a-4714-8d27-54e70b726133 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.191969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.191969Z digest=sha256:27c5d7b0aebe3cda3fd822ddd6f9867f5331d3ecb0947e6b166e25fd966837ee

Observation 8ffef98b-e581-4b4d-b210-3f777faee46b · outbound

This paper cites Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.149069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.149069Z digest=sha256:828baa476f5b5cc10135d59746716326be40b95a547247eebcdd2c6cf15c2c31

Observation ba9d2b6d-03b8-49d0-ad90-c6c82a1dff5c · outbound

This paper cites Voice Conversion With Just Nearest Neighbors.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Voice Conversion With Just Nearest Neighbors

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.015567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.015567Z digest=sha256:79c74190bf50f089c718029c11c649fcf86ffc5807c57b1fa8ff3ab299acb14e

Observation 56111811-5b5e-4c57-af0d-eb834e641a8d · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.891964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.891964Z digest=sha256:a655055b2b11aed5b3f69e7d4f205c1af67274ec2fc1af72d85b3ed9729e9b42

Observation 27f11978-8d2c-4b53-bb74-d756d3c3387b · outbound

This paper cites AdaptVC: High Quality Voice Conversion with Adaptive Learning.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion AdaptVC: High Quality Voice Conversion with Adaptive Learning

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:58:59.436990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:58:57.316479Z digest=sha256:3ba20bf0a9ddc8d8242e357de7ef45a76877a47878cfe394d5e7ee48ad4693cb

Pith citing papers

Observation ca3a5311-a37c-4fe1-a337-b0f3b844df9c · inbound

Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil cites this paper.

Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T00:31:12.289334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:31:12.289334Z digest=sha256:ed7acebd72fcd3e59d1f4e1d0e723c74723b44f7414bad3cbb7791c84143cb70