Pith. sign in

Paper Citation Record · LEDGER

GenVC: Self-Supervised Zero-Shot Voice Conversion

As of 11 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2502.04519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04519 v2

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:30:27.279088Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T12:21:48.698333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy55
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71c1e84d-e24a-42b8-a8bb-bf8fafaa6425 · outbound

This paper cites AutoVC: Zero-Shot V oice Style Transfer with Only Autoencoder Loss,.

GenVC: Self-Supervised Zero-Shot Voice Conversion AutoVC: Zero-Shot V oice Style Transfer with Only Autoencoder Loss,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.786485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.923708Z digest=sha256:9a3127693da0239c8e81f955186f89e5a32eb5f2424d77c9b7c50095dae7b4de

Observation 57116bc2-e5d8-408b-95e9-061b836bf82b · outbound

This paper cites GAZEV: GAN-Based Zero-Shot V oice Conversion Over Non-Parallel Speech Corpus,.

GenVC: Self-Supervised Zero-Shot Voice Conversion GAZEV: GAN-Based Zero-Shot V oice Conversion Over Non-Parallel Speech Corpus,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.770998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.928673Z digest=sha256:a2198ce49873a1474cddbdd4f3bbe6be4af50b9cc0d0581155637caa51b603b3

Observation df37264a-c034-4dc6-b631-090ebe334686 · outbound

This paper cites SIG-VC: A Speaker Information Guided Zero-Shot V oice Conversion System for Both Human Beings and Machines,.

GenVC: Self-Supervised Zero-Shot Voice Conversion SIG-VC: A Speaker Information Guided Zero-Shot V oice Conversion System for Both Human Beings and Machines,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.756520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.933327Z digest=sha256:1af6fa718a8e25e08098f73d7680bf00822d22ea8cced8711c823655459a61a1

Observation 3295ab77-01da-4aaa-aecf-fabe67f77fe5 · outbound

This paper cites An Overview of V oice Conversion and Its Challenges: From Statistical Modeling to Deep Learning,.

GenVC: Self-Supervised Zero-Shot Voice Conversion An Overview of V oice Conversion and Its Challenges: From Statistical Modeling to Deep Learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.741973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.938863Z digest=sha256:7c13ba403ba289e7530ad44f481f2f025bbf8be3bd5749616e3a920f490972ce

Observation 64126789-70a5-402e-adbe-610c91a006a8 · outbound

This paper cites Prosodic Features for Speaker Veri- fication,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Prosodic Features for Speaker Veri- fication,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.726198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.943697Z digest=sha256:5c8428aca9bdd52db626f7a24499a75ba7277d584920b5ee23ccc2db1e67959f

Observation 476bdcb9-b688-495a-b790-e32c5dbce96d · outbound

This paper cites FreeVC: Towards High-Quality Text- Free One-Shot V oice Conversion,.

GenVC: Self-Supervised Zero-Shot Voice Conversion FreeVC: Towards High-Quality Text- Free One-Shot V oice Conversion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.711160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.948572Z digest=sha256:2564868c17ef42e5659aad7d911dacdf0d94886f64d7d28b3515e2d2a5717213

Observation 538bf111-925b-44b5-8b3a-4e36cf41c6e4 · outbound

This paper cites The Database and Benchmark For the Source Speaker Tracing Challenge 2024,.

GenVC: Self-Supervised Zero-Shot Voice Conversion The Database and Benchmark For the Source Speaker Tracing Challenge 2024,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.697041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.953894Z digest=sha256:a8138ea1d55fb2c085375798315dd9d6036cde7cdfa9602e29bda1681b73a21c

Observation de760ea2-d4a8-4d40-a297-0217ed2dfa9c · outbound

This paper cites NeuralVC: Any-to-Any V oice Conver- sion Using Neural Networks Decoder For Real-Time V oice Conversion,.

GenVC: Self-Supervised Zero-Shot Voice Conversion NeuralVC: Any-to-Any V oice Conver- sion Using Neural Networks Decoder For Real-Time V oice Conversion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.681244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.958510Z digest=sha256:ebcaa05cfa4383f9f70bd74cfcd58ea71716a7a261401b6dad85484a6036343a

Observation 06902400-f3bf-48cb-9471-d92026eb4610 · outbound

This paper cites Identifying Source Speakers for V oice Conversion based Spoofing Attacks on Speaker Verification Systems,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Identifying Source Speakers for V oice Conversion based Spoofing Attacks on Speaker Verification Systems,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.665812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.962975Z digest=sha256:4e102d98951fe762477a65a36f4f1a877bd5662b04c6b3123f12678b8ff99afa

Observation 93174113-860c-4897-b38a-40e4f4548454 · outbound

This paper cites Privacy Versus Emotion Preservation Trade-Offs in Emotion-Preserving Speaker Anonymization,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Privacy Versus Emotion Preservation Trade-Offs in Emotion-Preserving Speaker Anonymization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.650766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.967484Z digest=sha256:55531ae29bec53a29c9906e426d873b02421f9645b1a3fcb8a3d84cae1183ec2

Observation ae623d40-0a59-4f2f-8eb9-5e9288abd926 · outbound

This paper cites Zero-Shot V oice Conversion with Adjusted Speaker Embeddings and Simple Acoustic Features,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Zero-Shot V oice Conversion with Adjusted Speaker Embeddings and Simple Acoustic Features,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.637204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.972138Z digest=sha256:1b9b71c1cb7e26d316bda340ab07890e868a6aeec735ed59b8e627ad38ba58fe

Observation 9de9da97-adc0-4f55-9dfa-719cfdff486d · outbound

This paper cites YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,.

GenVC: Self-Supervised Zero-Shot Voice Conversion YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.622650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.976480Z digest=sha256:34de75719cfc2e5cf76c3426083d4c88b7a3d2658be3c0c2e7c641cf9f1a71e5

Observation bbf74823-d13d-4fd4-ba91-14c7037c60c9 · outbound

This paper cites Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.607510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.980919Z digest=sha256:f6f513e28ec0f793888fe76e5515bef02a91df7898c99e68da454d81047b132c

Observation 283487b7-4969-4266-b073-6e2aedeb86a4 · outbound

This paper cites NANSY++: Unified V oice Synthesis with Neural Analysis and Synthesis,.

GenVC: Self-Supervised Zero-Shot Voice Conversion NANSY++: Unified V oice Synthesis with Neural Analysis and Synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.592358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.985319Z digest=sha256:318900af9861c4a51aa1ac103e86511e9cfa2f7824e9e263c516967623fdd27d

Observation 4e19c775-5283-49ee-b401-6287a4341f94 · outbound

This paper cites Better speech synthesis through scaling.

GenVC: Self-Supervised Zero-Shot Voice Conversion Better speech synthesis through scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:26.989835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:26.989835Z digest=sha256:8ed71d204ad62a4d781f7916d0dbf5446c56f2cc41d7c072bc7d5a9e125a68f4

Observation 4c70f03d-fd70-4823-82d4-4c41599edd4f · outbound

This paper cites LM-VC: Zero-Shot V oice Conversion via Speech Generation Based on Language Models,.

GenVC: Self-Supervised Zero-Shot Voice Conversion LM-VC: Zero-Shot V oice Conversion via Speech Generation Based on Language Models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.577280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.995136Z digest=sha256:897434f9ea38ed54fbe7529edfeb1fdcd167ca6d92c45a55c1ffeb67effee4f0

Observation 26d1928a-3ae2-404b-83e9-36498cc8c7cb · outbound

This paper cites AudioGen: Textually Guided Audio Generation,.

GenVC: Self-Supervised Zero-Shot Voice Conversion AudioGen: Textually Guided Audio Generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.561580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:26.999490Z digest=sha256:54313b0c24d6b0a9fd06c95dee0b0174f82cd9bad677abaafd00f86269384e41

Observation a7b5bfff-4835-4fca-9494-232bfbd952d1 · outbound

This paper cites Towards audio language modeling -- an overview.

GenVC: Self-Supervised Zero-Shot Voice Conversion Towards audio language modeling -- an overview

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.004124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.004124Z digest=sha256:b04efb9801f3481f409bdd0619772d46caaf479d2aa3570efa362aa18620b3be

Observation d04fdc0e-588b-44ab-8ec1-331074856051 · outbound

This paper cites AudioLM: A Language Modeling Approach to Audio Generation,.

GenVC: Self-Supervised Zero-Shot Voice Conversion AudioLM: A Language Modeling Approach to Audio Generation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.545081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.009093Z digest=sha256:efbd32cf1571369685ad90e846aac59bd53a6125d5d6f1f94990e5ca95f8df23

Observation 206007ce-f73a-4dd6-94e2-9d81f43803f1 · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec,.

GenVC: Self-Supervised Zero-Shot Voice Conversion SoundStream: An End-to-End Neural Audio Codec,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.529707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.013711Z digest=sha256:edf92267b680ae783be05b3cf255809cc8d73872494f8697246050937707d34b

Observation 883ded18-f916-4399-98d9-a62732267a98 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.513984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.018358Z digest=sha256:abec134e4342776f80a05ea6c1b929358617801c5514b65e0e619949395407b1

Observation f221b66a-5587-46ab-8e7c-6fc5c67c127e · outbound

This paper cites V oiceCraft: Zero-shot speech editing and text-to-speech in the wild,.

GenVC: Self-Supervised Zero-Shot Voice Conversion V oiceCraft: Zero-shot speech editing and text-to-speech in the wild,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.498449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.023068Z digest=sha256:9fd076e537673d1639240e661b5f6562e645f0effe32759f813a8e1040fc2cee

Observation c5f122a8-dc71-4a6f-8a70-27c097e0ed36 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

GenVC: Self-Supervised Zero-Shot Voice Conversion CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.027812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.027812Z digest=sha256:dd16702bee393ddf802264f4941f4aa16c35f3c0054675ba7c2cc7e63be864b9

Observation 3d79ea97-d09b-43cd-ab5b-75a2af4c6e69 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

GenVC: Self-Supervised Zero-Shot Voice Conversion CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.032838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.032838Z digest=sha256:fa83d4441863eadb3307dca46722d22b1475f87e8ca58aade21c752311d6682e

Observation 87c2e8c0-a97c-4489-87a2-d93220f7b470 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

GenVC: Self-Supervised Zero-Shot Voice Conversion Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.038490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.038490Z digest=sha256:380ad7c302c293f723494c0b79f3171772b968b92356e81f0f85845798396ee2

Observation d06cc4bc-1dc0-4259-bfc5-8c0e0d9f6c57 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.482999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.043717Z digest=sha256:27109a7751490974b942317d15a4322aa7157bb257948cfbe05bed696b550035

Observation 3c7bba64-613b-4124-b151-b901c08c3609 · outbound

This paper cites High Fidelity Neural Audio Compression,.

GenVC: Self-Supervised Zero-Shot Voice Conversion High Fidelity Neural Audio Compression,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.048979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.048979Z digest=sha256:9baaf8bd3e3232f0618b02f018cac9a4345499bdf73f137b29a3bd6a52dcad6f

Observation 4584d383-7a8d-4aa6-adec-bfdb1e2a0d79 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

GenVC: Self-Supervised Zero-Shot Voice Conversion Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.053709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.053709Z digest=sha256:917cd46bd2dc94bfc8b90b966cb12102d8b33be26e25c296807798a832deb453

Observation 2e6b6df4-1043-4cc5-8546-9e5b37c4b2af · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

GenVC: Self-Supervised Zero-Shot Voice Conversion VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.058372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.058372Z digest=sha256:d7853bd8230625f55160b36b2c25aea749f2fec0df7fbbc791cb332ab52e4201

Observation de3460c9-6464-4313-8607-a64dd04f09bd · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

GenVC: Self-Supervised Zero-Shot Voice Conversion CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.062791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.062791Z digest=sha256:c31798164d04147a05c2d94d919b593d42133f54ada46abe2b1d14b7f688923f

Observation dc341715-b904-4c60-a1c7-d36307712231 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer,.

GenVC: Self-Supervised Zero-Shot Voice Conversion MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.457575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.067170Z digest=sha256:60eb725f9e9743da8fcf162ba4297712dc0e5643fd2445a2a23be5b000800fd5

Observation f7010b0d-f354-4d29-956a-7f9edc90c6e6 · outbound

This paper cites Simple and Controllable Music Generation,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Simple and Controllable Music Generation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.442369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.071252Z digest=sha256:e6708cb260f413d35656209d333b1c8efb5a41f9aae9b9ad13dbfad212ddb89c

Observation 89e60a65-2387-4e5f-aa28-4397eca88b5d · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

GenVC: Self-Supervised Zero-Shot Voice Conversion Moshi: a speech-text foundation model for real-time dialogue

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.075348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.075348Z digest=sha256:bf72849256a897300b68c0f94916f091c515040fb8fb8f37ca2a24b68a0118a7

Observation 017cc5e1-ca13-49a4-b0eb-c714a05140fb · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,.

GenVC: Self-Supervised Zero-Shot Voice Conversion XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.427706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.079438Z digest=sha256:92ea1299857d848561afaf847fd44196dafbce262d8241530374ad4be4de6b71

Observation cf02f2e7-e96a-4201-ba1b-a3c5f0534f42 · outbound

This paper cites StreamV oice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot V oice Conversion,.

GenVC: Self-Supervised Zero-Shot Voice Conversion StreamV oice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot V oice Conversion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.413035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.084111Z digest=sha256:b9e2e2cea7b46e339d19fc00ec138501ccc5aa66536740489bfa78790ce00ca8

Observation e316393a-9101-4af5-8dc3-630a32aa095a · outbound

This paper cites Vevo: Control- lable Zero-Shot V oice Imitation with Self-Supervised Disentanglement,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Vevo: Control- lable Zero-Shot V oice Imitation with Self-Supervised Disentanglement,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.397784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.088859Z digest=sha256:8f0b8aac11a8131cbb87fecacf085a63e646ff79f1dab5ad1b789c4c667fb05f

Observation b59e1905-62e8-44d2-91c3-ab4a9795d566 · outbound

This paper cites The VoicePrivacy 2024 Challenge Evaluation Plan.

GenVC: Self-Supervised Zero-Shot Voice Conversion The VoicePrivacy 2024 Challenge Evaluation Plan

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.094409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.094409Z digest=sha256:2e0a2dd00f5d321405d231a599b6266956ceb33818bb767804043ae921421cfe

Observation e8faea19-f188-4800-ae26-702bcc22c750 · outbound

This paper cites Self-Supervised Speech Representations are More Pho- netic than Semantic,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Self-Supervised Speech Representations are More Pho- netic than Semantic,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.382328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.099365Z digest=sha256:22a73e61c734bd727e37a539be3e1a802180b5775be8fd530863167dcb0f066c

Observation 53d349fd-6a88-4d19-92ee-87da1b6168dc · outbound

This paper cites S2VC: A Frame- work for Any-to-Any V oice Conversion with Self-Supervised Pretrained Representations,.

GenVC: Self-Supervised Zero-Shot Voice Conversion S2VC: A Frame- work for Any-to-Any V oice Conversion with Self-Supervised Pretrained Representations,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.366410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.103917Z digest=sha256:f07191714a1fb9592f546bdfb581c0420e91b9b768df42f207c19830f73aa94a

Observation 6ab9a73c-379e-4d93-9b55-975ec331535f · outbound

This paper cites S3PRL-VC: Open-Source V oice Conversion Framework with Self-Supervised Speech Representations,.

GenVC: Self-Supervised Zero-Shot Voice Conversion S3PRL-VC: Open-Source V oice Conversion Framework with Self-Supervised Speech Representations,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.349822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.108491Z digest=sha256:5f198557b7cde47af9b0c3c297fac50cce779b1d9f1b73aa209be479806a7460

Observation a16894b2-ec46-47e3-81ce-2c0ba66f305e · outbound

This paper cites Neural Discrete Representation Learning,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Neural Discrete Representation Learning,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.331698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.113146Z digest=sha256:a1b57772c8d6dc99f94043f2bd882fdbde1e3242a37200365032733936e52a75

Observation 30a846b8-d196-4590-949b-9d622174ba4c · outbound

This paper cites Attention is All you Need,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Attention is All you Need,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.313805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.117803Z digest=sha256:a6a741250792faef9943149d5518f2296495c74b01c958c8217d11a35a8872e3

Observation f26daa61-61b2-4ba1-bac6-0a2fead6521e · outbound

This paper cites Flamingo: A Visual Language Model for Few-Shot Learning,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Flamingo: A Visual Language Model for Few-Shot Learning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.298619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.122721Z digest=sha256:0693cb03040316f778091956afca80aa462be563ec8ade80db8f422619fdb6b1

Observation 7288cac4-2c4e-4d43-afa4-b221ae96f723 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers,.

GenVC: Self-Supervised Zero-Shot Voice Conversion NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.282304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.127444Z digest=sha256:c11d55bba7bcb6c15a92fa75f675bdac8c55c3d5b43bdfba5a6e5171a0d8faf9

Observation 4fd34713-3c3d-48c1-a1fe-7319b6d7fe23 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,.

GenVC: Self-Supervised Zero-Shot Voice Conversion HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.266549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.131978Z digest=sha256:14e258904f988464431d78265be01e5b01d6f04e32e5c9c2647f465332cad2f7

Observation 77a7b491-dd80-4067-af6a-a7683c040642 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,.

GenVC: Self-Supervised Zero-Shot Voice Conversion LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.250786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.136838Z digest=sha256:a49216f90909d7ec174650e1a26f7af559c523baa20352de33ae1f27664fbc32

Observation 21fcf77c-46f6-4cff-8ffd-0aed48e9c707 · outbound

This paper cites Common V oice: A Massively-Multilingual Speech Corpus,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Common V oice: A Massively-Multilingual Speech Corpus,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.142012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.142012Z digest=sha256:5744549560b58bc280920c7157a9df55f932fb80d56194c96418fbd08c3629bd

Observation ddf739fd-5418-4521-b8e7-fd27cd1818d8 · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research,.

GenVC: Self-Supervised Zero-Shot Voice Conversion MLS: A Large-Scale Multilingual Dataset for Speech Research,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.146476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.146476Z digest=sha256:51521c5839ec4dc34f65c82eb7831f7b23e81b30af49516b10fc11ce309d450f

Observation d3d017ef-8948-4d91-9071-165606f983ca · outbound

This paper cites CMU ARCTIC Databases for Speech Synthesis,.

GenVC: Self-Supervised Zero-Shot Voice Conversion CMU ARCTIC Databases for Speech Synthesis,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.211890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.151196Z digest=sha256:9b10d35e0394c892eca748391badecf0037233696e2fb652b2a0a743cb2648d5

Observation 48910e9d-6469-4166-9e87-0251622825f1 · outbound

This paper cites The EMIME Mandarin Bilingual Database,.

GenVC: Self-Supervised Zero-Shot Voice Conversion The EMIME Mandarin Bilingual Database,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.195744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.155742Z digest=sha256:a125231cfae7df6c9603816affb1920492cacb66283155e0875ab3a57222a304

Observation f0fc4360-e347-43b2-b436-a216be764e81 · outbound

This paper cites Librispeech: An ASR Corpus Based on Public Domain Audio Books,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Librispeech: An ASR Corpus Based on Public Domain Audio Books,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.179990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.160700Z digest=sha256:76deeb3f8f109dc69172f63c36cc89c037544b0341980c8d28be3bcef2280383

Observation d3f548f3-a30c-4472-82a9-f95c79720c30 · outbound

This paper cites V oxCeleb: A Large-Scale Speaker Identification Dataset,.

GenVC: Self-Supervised Zero-Shot Voice Conversion V oxCeleb: A Large-Scale Speaker Identification Dataset,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.163119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.165129Z digest=sha256:fbf500aef5f74b610de8fcb201c1324970767e64feef5c401738fdf51099ce41

Observation 1eac42d6-8919-4482-86c3-ed284ed169d8 · outbound

This paper cites V oxCeleb2: Deep Speaker Recognition,.

GenVC: Self-Supervised Zero-Shot Voice Conversion V oxCeleb2: Deep Speaker Recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.148092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.169658Z digest=sha256:19aa82e9d7810ddc605c56676fa7503a42704aa971767259f2e0e42195517c58

Observation d2f1e9d4-05c6-4b38-b143-3270e332a156 · outbound

This paper cites ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers,.

GenVC: Self-Supervised Zero-Shot Voice Conversion ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.132448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.174264Z digest=sha256:74bf002b52e9396509de922986415cee753352871b53afce24c1ffde8ed6db76

Observation 3e0a27f5-c905-4ff7-bf7b-7f44d10c0a7b · outbound

This paper cites Language Models Are Unsupervised Multitask Learners,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Language Models Are Unsupervised Multitask Learners,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.116030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.178754Z digest=sha256:4ff361846817e51196bdf0024592d2ef8f57921ce8942c6809d74f9ad3b7003d

Observation 1f6270ec-ed66-4316-94af-48eef01aa0de · outbound

This paper cites Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity V ocoder,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity V ocoder,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.099154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.183079Z digest=sha256:d922548a3f4f92d391d2ceb035c79ee762352023bd90a51fee08608298e9e386

Observation f4122fc2-ac26-4c4d-ad20-377b459695d2 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre- Training for Full Stack Speech Processing,.

GenVC: Self-Supervised Zero-Shot Voice Conversion WavLM: Large-Scale Self-Supervised Pre- Training for Full Stack Speech Processing,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.187548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.187548Z digest=sha256:ffc7f2632a78dfb606b328d99308fe7370833a0a4c7c302422ad92ed3c63f1ab

Observation 65b9d42d-a89e-4127-83db-e8c489112a3e · outbound

This paper cites The T05 System for The V oiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech,.

GenVC: Self-Supervised Zero-Shot Voice Conversion The T05 System for The V oiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.071545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.192099Z digest=sha256:974f1d6b28b21d1c2a69ca4c7b0dcfd1ba653dd04ad18073ed3fc72c04b28c07

Observation 0908d85b-52f9-4b4e-a5a8-e5168c5ac5d8 · outbound

This paper cites Bias and Statistical Significance in Evaluating Speech Synthesis with Mean Opinion Scores,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Bias and Statistical Significance in Evaluating Speech Synthesis with Mean Opinion Scores,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.056193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.196870Z digest=sha256:7296f9b5e5bebf15a331c2fcd7a654f5b71045ddc00f42f77f5dedb5e753b3b8

Observation 73d23412-d759-4870-b8a1-8e9756cca0da · outbound

This paper cites Good Practices for Evaluation of Synthesized Speech,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Good Practices for Evaluation of Synthesized Speech,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.201566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.201566Z digest=sha256:7ea23db91a3ecce91dbc637f1408b8ec00bdcba489e03372f8223529440b926b

Observation f2f18524-0dae-410b-90d5-34ab5563589d · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,.

GenVC: Self-Supervised Zero-Shot Voice Conversion ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.039189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.206324Z digest=sha256:b6527901589542d40153118d3f4662a3d3012e31b7c6db5458febdc2943b8372

Observation 98089578-ca5b-461d-adeb-6d92b72c1c06 · outbound

This paper cites Scaling Laws for Neural Language Models.

GenVC: Self-Supervised Zero-Shot Voice Conversion Scaling Laws for Neural Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.210988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.210988Z digest=sha256:084226e48add99b9f9400cc58ac13b34b6aea92b29dcbd6e49e4e40fea7c0940

Observation 7f04a0d9-cb12-4ceb-9e0f-57512511c733 · outbound

This paper cites Deep Residual Learning for Image Recognition,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Deep Residual Learning for Image Recognition,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.024116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.215698Z digest=sha256:9f3d0fe07c93942b08d687142db9e0b3c1236447c1fedd28ce8f2bf2fd0c22d7

Observation 2c692cd9-628e-45a2-97e9-f3fb3be45556 · outbound

This paper cites The first two encoder convolutional layers upsample the input dimensionality to the DV AE hidden dimension of 1024.

GenVC: Self-Supervised Zero-Shot Voice Conversion The first two encoder convolutional layers upsample the input dimensionality to the DV AE hidden dimension of 1024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.008877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.220159Z digest=sha256:80ca2b994671008fafc87d5881f0e16bfa7a89508f2a76868a1b60efa3737242

Observation a61af82e-6815-44e1-b56d-b39355c0d6b8 · outbound

This paper cites The learned queries attends to all input frames, transforming them into fixed- length representations.

GenVC: Self-Supervised Zero-Shot Voice Conversion The learned queries attends to all input frames, transforming them into fixed- length representations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.993036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.225587Z digest=sha256:f4272286a1a4889e3b21558420fb0cc775a391afa93e0f631a60bee1ae6170fc

Observation b770d178-6fed-4271-937f-a964500f4b1f · outbound

This paper cites 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5 and 5 are allowed on the half-point scale.

GenVC: Self-Supervised Zero-Shot Voice Conversion 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5 and 5 are allowed on the half-point scale

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.979268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.229649Z digest=sha256:a8510fae0602331e9e55b4bdb97de622c8bef9a8bf0632ba5865c61547255b3c

Observation 68f8817e-affe-45a1-81d0-288b84771820 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.965523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.234565Z digest=sha256:0d323ba8f2e3293c2d4762d92e6360d0f3efedef1ab34d0e3ea2b831a6d744e8

Observation bcc41a1f-95d0-4356-8017-833399b18f64 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.951564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.238691Z digest=sha256:39a32d374df5deb839bb785d86b49a4965d9fbb1e55de45b3f731639a8ed65c6

Observation 2e924d13-7264-4f1f-a34e-8a8377acfc2b · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.937513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.243255Z digest=sha256:1609a02e9bb76dbfdbdbf29b472ed85f191b9da7fcfedeceed6c17ccf90945d8

Observation 9531351d-38d9-4636-9ab3-127abfa78418 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.922807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.248147Z digest=sha256:d5453d1b7ab535f32610ab53fd993ddbee31891809dcee69568e34d91c5c09e5

Observation d6f96f98-5b1f-4c1d-b957-4431079b99a0 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.907383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.252499Z digest=sha256:8f24c7b96f525b124665443d8578c9b4321c949bb9098a7a8655d2b7d431589d

Observation 0a08ca7b-02ba-4ae9-8f00-e10c947f5196 · outbound

This paper cites Use a scale from 1 (Bad) to 5 (Excellent), with increments of 0.5.

GenVC: Self-Supervised Zero-Shot Voice Conversion Use a scale from 1 (Bad) to 5 (Excellent), with increments of 0.5

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.891406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.256544Z digest=sha256:6e60796d99f8769b8a9cb0d566efcedb2979a6ea2f595514033bcda92229574f

Observation c38cca02-6785-4edc-9958-72c691439cb4 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.874432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.260729Z digest=sha256:93943158b45354bb22c1a1da5ec625b5094ca77b68080c2ca1621e997935b96e

Observation 649cbab6-ba49-4ae5-9f27-1b4e48a71dfd · outbound

This paper cites For example, the speakers may differ in gender or pitch (e.g., a high-pitched female voice vs.

GenVC: Self-Supervised Zero-Shot Voice Conversion For example, the speakers may differ in gender or pitch (e.g., a high-pitched female voice vs

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.858754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.265152Z digest=sha256:c7585302a6920f1d6b351c89f6866372d1adb436260f101610ee7875ab0129bf

Observation ec066b31-e376-4c15-a4a1-43c746ce1885 · outbound

This paper cites It’s clear the speakers are the same gender, but their voices are distinctly different, and their speaking styles differ somewhat.

GenVC: Self-Supervised Zero-Shot Voice Conversion It’s clear the speakers are the same gender, but their voices are distinctly different, and their speaking styles differ somewhat

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.841437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.270024Z digest=sha256:734b44258c9c6bf052100a4ce6e7e8d92dfcbfb12777d341631588b9659016aa

Observation ecda95be-b217-4955-b8f7-8c711b69230b · outbound

This paper cites The speaker voices are close, and the speaking styles largely match.

GenVC: Self-Supervised Zero-Shot Voice Conversion The speaker voices are close, and the speaking styles largely match

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.824860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.274365Z digest=sha256:47719318457ea793cbe1ceedf0528f1ae65948ef828b59254cd499c1e60ad2a7

Observation 98c69fa1-f049-493a-a3cd-c5e23bbccbb9 · outbound

This paper cites The timbre of the speakers is the same, their speaking styles match perfectly, and the environmental background is similar, including any acoustic noise.

GenVC: Self-Supervised Zero-Shot Voice Conversion The timbre of the speakers is the same, their speaking styles match perfectly, and the environmental background is similar, including any acoustic noise

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.809162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-08T22:30:27.279088Z digest=sha256:47d87af8bd5c86f337e000336fbb3c196f8b314a028b64f12eb9de08a5c431bd

Pith citing papers

Observation 7860c175-f3de-4ccb-9b1b-5cb692542da4 · inbound

Universal Speech Content Factorization cites this paper.

Universal Speech Content Factorization GenVC: Self-Supervised Zero-Shot Voice Conversion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T12:21:48.698333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:21:48.698333Z digest=sha256:451437419135f6f157f8cad7d8a4faa3098c4ee92658582617f3cd68f9b25039