Pith. sign in

Paper Citation Record · LEDGER

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

As of 9 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 7 inbound Pith citation observations for arXiv:2502.07243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07243 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.369265Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:21:59.575812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.691071Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1234a951-ef09-4289-a4bb-218485f1202d · outbound

This paper cites Neural discrete representation learning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural discrete representation learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.088321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.088321Z digest=sha256:bc74e8b1a753c6eafd8f715ae8c5852914a578f70b88a523713b6101c756cb69

Observation 9621dbf3-776a-4578-87e4-6d5cc7130f2d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.092101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.092101Z digest=sha256:a8db9b2864e8c3e304a920fe11c6ad378e2712f59457772c596b9a3cbd86a536

Observation cdcc9f2f-cdc6-46b8-8c46-646cb2c70740 · outbound

This paper cites An overview of voice conversion systems.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice conversion systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.095768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.095768Z digest=sha256:e19f6b30880661a46ae725fa8f8279d1f44da4de4c6c229d2ca341175a9aac49

Observation c384d50e-393a-4c38-9e76-05d4c1c59737 · outbound

This paper cites An overview of voice con- version and its challenges: From statistical modeling to deep learning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice con- version and its challenges: From statistical modeling to deep learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.099143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.099143Z digest=sha256:bfecaef24b179fb5852c4d115e79b69cd3ab154ff69ed092e8d02093cb087289

Observation 2ef66d9b-3efe-4cc5-9fdb-32a19f32fb3b · outbound

This paper cites Foreign accent conversion in computer assisted pronunciation training.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Foreign accent conversion in computer assisted pronunciation training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.102330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.102330Z digest=sha256:ae24414368f848eb4e8cbf56919dd0c46754244df21bf039f8e900a49d237ae7

Observation 9f2d5106-5bcc-4539-92f2-5c08a4174d82 · outbound

This paper cites L2-ARCTIC: A non-native english speech corpus.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement L2-ARCTIC: A non-native english speech corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.105470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.105470Z digest=sha256:659ab8785a9f9c761f21a92a624ed25f5738bf73809c0bc87bc64ba4b880e5a9

Observation 84010b50-fc81-409d-b51c-3e124781da88 · outbound

This paper cites Emotional voice conversion: Theory, databases and ESD.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emotional voice conversion: Theory, databases and ESD

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.108882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.108882Z digest=sha256:bcdf43d37d406a9bed0ddf3258056874f73f62d96e69c5b38b4f6718a8773b71

Observation 8f4141dd-7b92-4bc3-a3ed-6c8dade41c33 · outbound

This paper cites Neural Text-to-Speech Synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Text-to-Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.112021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.112021Z digest=sha256:39bbe6af22dae263ce6ca2393b0c6f94f47cdc3d42c89dcba3374fa23b27c93e

Observation bbf8051d-f030-4f39-8e3f-6990403e10e7 · outbound

This paper cites Converting foreign accent speech without a reference.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Converting foreign accent speech without a reference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.114754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.114754Z digest=sha256:4ce9c5ad99ec5b3bbd4f617f345cee585b6f954d61a22b78e36de9efdb4eaf1a

Observation b5ddc4a1-f58f-452b-9dab-64b5718a6ba1 · outbound

This paper cites Sahidullah, Aur ´elien Bellet, Marc Tom- masi, and Emmanuel Vincent.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Sahidullah, Aur ´elien Bellet, Marc Tom- masi, and Emmanuel Vincent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.117635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.117635Z digest=sha256:4602060dd4f9eb7a260b6124955307d197d9d638a4f68dc74ef788d268978404

Observation 05a4d2f1-477d-49e4-9574-6aa2484f14a0 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.120659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.120659Z digest=sha256:3221286dd3c35bdb05c6cfc5f7bcf9d0c5aa02b1b63d15f1eb2f0b25298da14c

Observation ab9f82f4-1ab7-4ad4-82d3-3aab0706d44f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.124096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.124096Z digest=sha256:e287484e547dda8ebad352bfc0ca1dd55707107599892febaf4cd0586bc6e834

Observation 11b705f3-918e-4255-a214-222525ab2f0e · outbound

This paper cites Maskgct: Zero-shot text-to- speech with masked generative codec transformer.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Maskgct: Zero-shot text-to- speech with masked generative codec transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.127638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.127638Z digest=sha256:a6bbc353a12a280c9dcf49de3635c1479b60aa3d1a6808e393469e9f9d50f44c

Observation 07ff5752-9eaa-4c8d-a501-508a918d48fe · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.130748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.130748Z digest=sha256:25975864563580a438a597fce7040d32f51983fdb3b728c954e758d7f23e3cef

Observation 6a2d694b-e80e-42a1-9d7c-b13db2eb9f4b · outbound

This paper cites Speech resynthesis from discrete disentan- gled self-supervised representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speech resynthesis from discrete disentan- gled self-supervised representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.133715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.133715Z digest=sha256:8e9622415424204bcccbcd397040a227d9b6bf0472ada412769a7c62705b1fcf

Observation 6e9e5b21-1953-45c9-9c85-9baf1e1642fb · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.136859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.136859Z digest=sha256:65a417afdf6fd71ecf67d9e488eb1e16b44e14f6100074d5d3216058a49c3560

Observation 53b580d5-50d5-4831-b8b8-aab0451b14bc · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.140473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.140473Z digest=sha256:9e814a35b62d14c3725adb4562b48e5f52ce731ce175511761ba06b61f06bf63

Observation 48cc366c-9511-4c1d-b45c-122c4c30fd9f · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:20.024364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.143740Z digest=sha256:dbefd24144e80c68a3445cce92c667bbf08b286d0cbd90ded7e2fc8a725baf55

Observation b16bd86c-730f-47cb-a719-d59fc2465f63 · outbound

This paper cites Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:20.015767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.146950Z digest=sha256:c050785bf9bd402f368a766c181969f989410ca8a609cc845445f56ce48c3b93

Observation b4e0abf5-deb8-4310-a6aa-de5bb23ebfa5 · outbound

This paper cites VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.150092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.150092Z digest=sha256:fb81d015a88b168c67fa21b828e10a58e8b3a396264b78a922dd5973260a6c4c

Observation 6e1cff98-0b03-451c-8ab3-a3b8d7261506 · outbound

This paper cites Convert and speak: Zero-shot accent conversion with minimum supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Convert and speak: Zero-shot accent conversion with minimum supervision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:20.006904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.153241Z digest=sha256:8802c21f27b91123f58f7642618ffea7a851c637b80da2939c5a9f83f6d31b2f

Observation d63fbd8f-c1e4-4798-99a1-4a417dabffc5 · outbound

This paper cites Au- tovc: Zero-shot voice style transfer with only autoencoder loss.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Au- tovc: Zero-shot voice style transfer with only autoencoder loss

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.998020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.156236Z digest=sha256:b5fd60401e248f34b416ac7cf27fd1c7fe0c2ad8941d00d0adbcdd706bc5a0de

Observation be8810d0-db54-481e-abb2-683a25f231af · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.158875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.158875Z digest=sha256:94b92c2c01dd13910e075767aeb5402eb106a3f25300766953e38f3439b2903e

Observation 712c1ffd-5225-45d3-bd18-6aba84d44a47 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.162381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.162381Z digest=sha256:df82bd587f03672fb17471562624e28fc27bc17fea02523c8d97ab4a32561c3c

Observation 54dc202b-f995-44b5-9425-fc2300fe4d46 · outbound

This paper cites Better speech synthesis through scaling.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Better speech synthesis through scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.165401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.165401Z digest=sha256:1423561de31df12e1209750d0a200e7bb98ddcd46d9b4940d3bf9ee0850a1437

Observation 8dd7fa07-0c42-4fbc-9577-7def92b7ad11 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.168802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.168802Z digest=sha256:002838c1af1969d4b753b7590c51933eeb1caf12b8591b7c9f107a506f38c6bc

Observation 9bf9cba5-4724-401d-84a2-5dc3a569a6c9 · outbound

This paper cites V oicebox: Text- guided multilingual universal speech generation at scale.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicebox: Text- guided multilingual universal speech generation at scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.172375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.172375Z digest=sha256:c78b9d02e903ef309ff4c31b68fdc4e04f821ff490b5225ffad6aa97b4383af3

Observation 8fc2728d-5e2a-48df-ab02-34693677b7c2 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.984343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.175536Z digest=sha256:41b93005ec9250188ceaf63775b1b9c5aacbc6385885dbd48f98a1fea634735a

Observation ff1696aa-260f-49fe-be8d-3e0e641aa098 · outbound

This paper cites V oice- preserving zero-shot multiple accent conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oice- preserving zero-shot multiple accent conversion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.975383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.178687Z digest=sha256:bb98599d504be2e3223ac1fec5dd51ff2d551dffe68f6d3e9eed5915aee9dc71

Observation 9d8b0e25-1f48-424f-a484-258480a55ff0 · outbound

This paper cites Schuller, and Haizhou Li.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Schuller, and Haizhou Li

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.966785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.181727Z digest=sha256:cdd3ea59f98fcd5c1af06ceb28c1df2743ad567e795ebe4e765cfabe187698df

Observation e4bf249b-c83e-41c0-bb00-de4227ca1539 · outbound

This paper cites PA VITS: exploring prosody-aware VITS for end-to-end emotional voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement PA VITS: exploring prosody-aware VITS for end-to-end emotional voice conversion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.958284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.184577Z digest=sha256:e39decb42ca83702c2191b61b57fe9d6f64043ac5ea93ed60481c17fbc189765

Observation cb915b3c-08b9-4090-8a15-d8a50bf05a15 · outbound

This paper cites Transfer the linguistic representa- tions from TTS to accent conversion with non-parallel data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Transfer the linguistic representa- tions from TTS to accent conversion with non-parallel data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.949309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.187544Z digest=sha256:84bb30a33bd5821d650a3946b450c220defffa2a08b9afed32d4408851654edd

Observation a52ec863-cd90-4e3b-97da-87c94d9798dc · outbound

This paper cites U-style: Cascading u-nets with multi-level speaker and style modeling for zero-shot voice cloning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement U-style: Cascading u-nets with multi-level speaker and style modeling for zero-shot voice cloning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.939715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.190541Z digest=sha256:18d74312b0353ad8f18983021182fa68c51cbe1587a31f0ecf3570cbf6c520d8

Observation 32c230e2-b170-4780-855b-e79bba4ccccb · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.193725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.193725Z digest=sha256:740d50492cab30091232ddfaabe04eafac9e4cd857a40c75743acfbb494e899b

Observation 6baec9ef-0a8b-435f-909f-835a1966cab0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.196997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.196997Z digest=sha256:9863627c22568f6093d8319b02a4fe981d703ed03b91dad06bdcbbdb8bcb13e6

Observation ffb292e5-e1b2-4a92-b961-b30131fcd04a · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.200113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.200113Z digest=sha256:ca1dd70af696ff09863891c7180990c2c3fc5e995b9411803ea5431dc671a5bc

Observation f59cd892-b39e-4fb8-b11f-f5882fcac97e · outbound

This paper cites Scalable diffusion models with transformers.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scalable diffusion models with transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.920850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.202942Z digest=sha256:560df17e56663865eee3f247c1f985d0fa50758e15908e878398180153b39610

Observation f9b274d1-daa9-4b5a-85f6-5411061b25d3 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.205829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.205829Z digest=sha256:07e69173bfabe1b7825373eb03d46529f9880199fa47a5b4f0c7425abe6a9b3b

Observation 02e7d8be-5b9d-4760-aba5-1cb3debb871c · outbound

This paper cites One-shot voice conversion by vector quantization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement One-shot voice conversion by vector quantization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.912239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.209096Z digest=sha256:e6f395c1451b1dcb6f5cb660b9cb98b765226d8b97ee2adc08ca44d7b94811a7

Observation fff07ec3-4590-4663-93c6-cc10cca13954 · outbound

This paper cites Unsupervised learning of disentangled speech content and style representation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unsupervised learning of disentangled speech content and style representation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.903838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.211965Z digest=sha256:02795677f196ed7a69a194cbfb4593fc3d50e8ed7e6ec1943d622e5f3d96f976

Observation 0b2c7c32-6e36-488c-843a-6e54cf7b2674 · outbound

This paper cites Text- less speech-to-speech translation on real data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Text- less speech-to-speech translation on real data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.895972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.215144Z digest=sha256:044d70ff66745e2797c6c94286e1e0ad3dab9e00ce6cedcfba19532cea464250

Observation f2dd42d1-ebc7-41b9-9bc5-b15b453909e6 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.888168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.218507Z digest=sha256:96aeae09dbef154cf235cba959add19512e384fc483f75968a0ad88a40fdea32

Observation cce790f6-3be1-4ef6-8e45-4148352c4a91 · outbound

This paper cites A comparative study of self-supervised speech representation based voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparative study of self-supervised speech representation based voice conversion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.880694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.221950Z digest=sha256:bb984fb3359498f23fb0963506372d32b13f9c37958c655384a0e736943731f3

Observation 7e0f126a-39a0-4447-90e0-55372264de06 · outbound

This paper cites Leveraging diverse semantic-based audio pretrained models for singing voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Leveraging diverse semantic-based audio pretrained models for singing voice conversion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.872428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.225095Z digest=sha256:d9d0d3d5064b65a12c4421fc4c221fdd0f2680a9d1b97ddbbc3aef5306dbaa8d

Observation c6debe6e-8d3c-4c3c-9386-b7b172cb9bbb · outbound

This paper cites Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.864166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.228370Z digest=sha256:055293193515ba83b6aeb2d6279be11f787413ea993efba991673131bfa625ac

Observation b21ccbbb-9ffe-4d95-b1a1-946d7527d093 · outbound

This paper cites Stargan-vc: non- parallel many-to-many voice conversion using star generative adversarial networks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Stargan-vc: non- parallel many-to-many voice conversion using star generative adversarial networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.855792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.231460Z digest=sha256:f8e13e671e2369ca6404d155a6bbc18cc4d58da576d9a9fb809ab1e763fdb21b

Observation 7ade7335-7583-48a0-bb39-d1564a094f0a · outbound

This paper cites Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.847440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.234763Z digest=sha256:e22e9df6abb7ca71fe60b92ea1337f833fff3e5fa2ddbdc8e2bef38feafc06fa

Observation e579a085-1702-4f69-8e7a-d3a41f7d7b54 · outbound

This paper cites Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.839236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.238179Z digest=sha256:bdbee38a7a12ccbd12c5bdae8841b4f3548387e0d998cf96e72432b25ee90df0

Observation 0117c55c-388b-4907-b30b-5a76b2feb2ed · outbound

This paper cites Tts-guided training for accent conversion without parallel data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tts-guided training for accent conversion without parallel data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.830985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.241259Z digest=sha256:ad177aa2896917aa5009e105a2e45c9807a9f1c098c07f56a5cadf9c5b89af53

Observation efd90fa0-40a8-41eb-84b7-b3095e3c6b4f · outbound

This paper cites End-to-end accent conversion without using native utterances.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement End-to-end accent conversion without using native utterances

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.822725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.244394Z digest=sha256:9bc9a1f3cba47f4bc427912cbb677e3f69cf0b151d65ab00869e77056b3892b2

Observation 2333a5bb-d67a-49cb-a334-c7bb3dee15cc · outbound

This paper cites Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.814630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.247887Z digest=sha256:c8baa4d7e88166ad5b43dc034ef7d051c93ca1524f5cecea213a2bd07705bc53

Observation e933a787-f410-4603-b343-601a48a25b3f · outbound

This paper cites LM-VC: zero-shot voice conversion via speech generation based on language models.IEEE Signal Process.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LM-VC: zero-shot voice conversion via speech generation based on language models.IEEE Signal Process

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.806051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.251307Z digest=sha256:5a2d8b7cf132d347e2ef93f4e8912431e4b375573c4b9b475da8fbd4beceb7f7

Observation 49e1554d-4460-4704-a07a-acb5915ffb54 · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.254583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.254583Z digest=sha256:95c220c0f063eb8ab40cfc00e030bd98a70a9e550c5e46e6f626d470760dbb22

Observation 616a3740-187f-4e40-9ae3-7ad81e2b7bca · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparison of discrete and soft speech units for improved voice conversion

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.797639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.258005Z digest=sha256:1dadcf909a899e95dd4a836f3cb64dfe308f2f7609590cc4d188121ece66b731

Observation 7f46fa70-3f63-4b79-80e6-1882ea7259c6 · outbound

This paper cites SEF-VC: speaker embedding free zero-shot voice conversion with cross attention.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement SEF-VC: speaker embedding free zero-shot voice conversion with cross attention

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.789762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.260858Z digest=sha256:6ed2fc7357edb15ac0e00f20d5ad0e765ed6c564c5fb0d46d025c59897cb1fc6

Observation 019df985-128f-4a50-9923-fc61b67370ae · outbound

This paper cites Neu- ral analysis and synthesis: Reconstructing speech from self-supervised representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neu- ral analysis and synthesis: Reconstructing speech from self-supervised representations

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.782063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.263434Z digest=sha256:3c49fa2cd483207c94a40d7e5615a545fb9de937e6c08d72ca80e735c7b19db3

Observation 8d86c200-b042-4825-9fd0-b3f4f953501e · outbound

This paper cites NANSY++: unified voice synthesis with neural analysis and synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement NANSY++: unified voice synthesis with neural analysis and synthesis

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.774318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.265980Z digest=sha256:7bde40b74cf43c11259b363131ee5ad4214c59f38bb342f110c7c3ae2d833ce6

Observation def82cc8-225b-4949-9917-0a981577767f · outbound

This paper cites Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottle- necks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottle- necks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.766187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.268522Z digest=sha256:ad4c9d1914442e0f245ae3c19f6241b3bdc205ade86cfb611b65c54f0207a2d8

Observation f0e34687-921c-4eed-87e3-61b6123eb262 · outbound

This paper cites Cox, Mark Hasegawa-Johnson, and Shiyu Chang.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cox, Mark Hasegawa-Johnson, and Shiyu Chang

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.757329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.271441Z digest=sha256:a07488bb37102af44574e81644850c88593447aa930bfe8b6fbf465ecda19ad1

Observation e0381617-1b1b-4451-9a2e-b8ba001f3301 · outbound

This paper cites Multi-speaker expressive speech synthesis via multiple factors decoupling.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Multi-speaker expressive speech synthesis via multiple factors decoupling

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.748797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.274184Z digest=sha256:5e2244c693c237bc7c841531debdf9ff8cefdf934744428669e09e295834d383

Observation efdb67c7-5795-43ff-878b-ab42b2040bc5 · outbound

This paper cites CLUB: A contrastive log-ratio upper bound of mutual information.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CLUB: A contrastive log-ratio upper bound of mutual information

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.739916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.276826Z digest=sha256:4ebe2fa78cb303e775649396f265d9ceb90c03622b0ea181fee7762f2f8010b8

Observation 44448173-2d8b-4098-ae9b-fd6b8fe82fd4 · outbound

This paper cites Repcodec: A speech representation codec for speech tokenization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Repcodec: A speech representation codec for speech tokenization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.731258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.279359Z digest=sha256:64809012c8d621bbc9b2488e1e6d69df6fcb3ce7d2b2479dc2eb646d7c9a6227

Observation ee6f654c-e8bf-43d3-92ba-67f5aa894cb3 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Soundstream: An end-to-end neural audio codec

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.281773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.281773Z digest=sha256:47eae18f35d71706d1c12c5ed03847fa41c20d3b5f79a64266e5a4dd7c3dd246

Observation 3a7660a3-b4de-4b7a-852e-2f8cbe5b0437 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Wavlm: Large-scale self- supervised pre-training for full stack speech processing

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.717869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.284565Z digest=sha256:a2025fce3fbe30393125472d84dca7bcef55ec79ea5bb22e1b765de95bfc815b

Observation 04a83850-4d53-4914-b52d-5c58321399a7 · outbound

This paper cites ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.709198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.287428Z digest=sha256:2e434ef97e59d38270112aadc6b5ee8ef7e3f3d0e7dfe10c84ba5f1d22fa85b7

Observation cb1b2cf2-77a4-4c63-8175-dbb8f8527d1a · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BERT: pre-training of deep bidirectional transformers for language understanding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.700442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.290292Z digest=sha256:6bfb18b00a36851a475607af2f13d3950e7086e860cf6d129c5f014d87fd6c8b

Observation 1caf02e0-b17b-4418-ab8a-847aa87da851 · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Bigvgan: A universal neural vocoder with large-scale training

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.687235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.296292Z digest=sha256:806559469d57bcffd2f022edaaecf2bb7b17c5ee13e1d2ed3bcedfcae3d5df28

Observation 97818603-d8b1-4009-b416-4ffbedb85cdf · outbound

This paper cites Tyers, and Gregor Weber.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tyers, and Gregor Weber

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.678910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.299049Z digest=sha256:eea99bc62b5c26a9b7d0b4c0a9e01b87373cdb9c09d7698cfca5f98468f9ea5b

Observation 7a190bd8-a972-40ec-b49f-710944c8370f · outbound

This paper cites The singing voice conversion challenge 2023.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement The singing voice conversion challenge 2023

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.671037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.301962Z digest=sha256:74f29206b7c125a98546ce95a83ef648bc32ccd49761330a9a0c20e11f24223a

Observation 41e9396b-97de-4d23-847c-cb904661bfc9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Robust speech recognition via large-scale weak supervision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.662748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.304779Z digest=sha256:9cf3e082adc4d48fc9c8acd2e3ebef967686baf5c456b340f82e27320fa8fedf

Observation f129dbbf-ba2e-4d4d-b419-611e0c960a33 · outbound

This paper cites Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.654569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.307589Z digest=sha256:ec1d6fa7e00fc120fbb57ecaa634028ee9f9b68db05a87eea8313e7b59986f0a

Observation d69f8ee8-2721-4124-a54a-686965e816ef · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.647207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.310413Z digest=sha256:0df20ada761bc7312fb1ef77ad72a9ac464102720535153b721bb8f7ddf2c50e

Observation 08f94af4-4fca-4428-8029-ea535e194cae · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.639684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.313298Z digest=sha256:aedd702f427b44b761d046192617fae2487e0ca48ed705212ddb7f1fa1302d4f

Observation d2b6dc8d-aa5c-4335-a1a2-5d2e740ad999 · outbound

This paper cites Libri-light: A benchmark for ASR with limited or no supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Libri-light: A benchmark for ASR with limited or no supervision

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.631954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.316297Z digest=sha256:2f5d6ade02e78fb68bb132f5891bb5655ec0151a912703203afa53352117b879

Observation 91b2c8cf-ee67-4ac5-a632-a16b75a59656 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Librispeech: An ASR corpus based on public domain audio books

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.622739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.319210Z digest=sha256:e888291eb1fcd8b3a9bbed503eb148236ec6512d85c0ced79311740f73a3b16d

Observation cc567952-405a-4834-803e-a1f8637b199a · outbound

This paper cites Amphion: An open-source audio, music and speech generation toolkit.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Amphion: An open-source audio, music and speech generation toolkit

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.613920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.322124Z digest=sha256:c63a33360cb30be163c7a07c78fbbc8c7cbbfc711ed77f9eb8e81f6e21b39311

Observation 35599e59-9f1a-431a-b55f-394a892cbc4b · outbound

This paper cites V oicecraft: Zero-shot speech editing and text-to-speech in the wild.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicecraft: Zero-shot speech editing and text-to-speech in the wild

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.605013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.324971Z digest=sha256:6c9e2fdf89169235d9ce7f2ecd12fa03523324575eb9fb17a87a552a41a4e5fc

Observation d4d9660f-dee0-4fd3-836c-94376ab66820 · outbound

This paper cites Overview of the Amphion Toolkit (v0.2).

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Overview of the Amphion Toolkit (v0.2)

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.327889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.327889Z digest=sha256:667312a639d47002b2e594a21e1f698a4da6ae6acd757e8e0650d09053a59183

Observation f007b9ae-b6b3-4469-a20f-6ff698ee87a3 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.596064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.331052Z digest=sha256:530045c952991a3efbf00387e4898b2a6f73f8622d88abcbc414d7b2af475ebe

Observation 8ea99c5d-0d73-46ab-a699-2c5b5fd2d259 · outbound

This paper cites Decoupled weight decay regularization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Decoupled weight decay regularization

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.586922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.333766Z digest=sha256:5b94a164485612f86fa39362d4e1f96a626c2f2e2fca81684cc6f7810bc29e39

Observation f4437809-355c-4f4e-9839-d4877cffa94a · outbound

This paper cites Kingma and Jimmy Ba.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Kingma and Jimmy Ba

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.336676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.336676Z digest=sha256:843ed96465a933140b0162b695cce105d1eca5ff00da0430548c20ac794ae094

Observation ef3c1bcf-7fce-40a9-a5af-1a4ec8b76b51 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Classifier-Free Diffusion Guidance

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.339612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.339612Z digest=sha256:03e627f982d6c1d2023b827d6db44812d7b36a6f51abdb5baeb79bee17cc1481

Observation 71139633-d1f4-49e7-9afe-ed8423788c6e · outbound

This paper cites Scaling speech technology to 1, 000+ languages.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scaling speech technology to 1, 000+ languages

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.573407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.342687Z digest=sha256:10a39c6b50925081cd031526c6bce86d4b2b9aa9500e7156a504ed2bf990dff3

Observation fab7ebe3-0ec2-4843-ae1e-2de7f9ddbfb5 · outbound

This paper cites Conditional variational autoencoder with adver- sarial learning for end-to-end text-to-speech.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Conditional variational autoencoder with adver- sarial learning for end-to-end text-to-speech

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.564600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.345544Z digest=sha256:9b954a4693065570a548e8a8594cf0212a2b91fa81ec64c4578ca6d4dd0db878

Observation f0be682c-099a-45b8-9852-3d4430a7214b · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.555736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.348402Z digest=sha256:2fa7e338d7a2a603afe60b0b07351dcbb6118fb500b3235fb687d921f28be7c7

Observation ed0f2eb6-4b91-4dec-9f18-5b4edbf627ae · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.351240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.351240Z digest=sha256:64a633ced8a9e3aa1adc42c4261795c9776b49ed38dc70a18c5031491f4cd7f4

Observation d2eb39b9-9991-480f-9f29-ec8fef4334f2 · outbound

This paper cites BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.541564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.354382Z digest=sha256:a3508032d2a8744c4eaeba96600dd9037ef1cc24742719c5ce14c23fa348fa55

Observation c7454f28-2d70-4bc5-869a-936e0d8b3b9b · outbound

This paper cites CSTR VCTK Corpus: En- glish multi-speaker corpus for cstr voice cloning toolkit (version 0.92).

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CSTR VCTK Corpus: En- glish multi-speaker corpus for cstr voice cloning toolkit (version 0.92)

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.532562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.357239Z digest=sha256:b583f606ffbeeaf752e5e0e39397188e402f0c3cee998bb31a14f84bea97a326

Observation 0b0598f3-d986-426f-9a83-05f99e54586e · outbound

This paper cites High Fidelity Neural Audio Compression.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement High Fidelity Neural Audio Compression

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.360299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.360299Z digest=sha256:4c0864e668f294dc9214fbca920c83354c09c8435e0c7be4f47f09e466df73c8

Observation 050584c9-9979-4017-a8a1-00ce234971ad · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement MLS: A large-scale multilingual dataset for speech research

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.523449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.363655Z digest=sha256:a15b4476278878a71278b8cd4854ba96b05a173250a6f26429b8148962faf210

Observation b8aac263-107c-4a80-9934-b90cdf3a6e72 · outbound

This paper cites Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.515151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.366424Z digest=sha256:92273d52a058059e5a166cfaff659210625879e006bce7a5af0f94f7a715b38b

Observation 06cef095-3980-4f5f-b5c1-4fef04f485f6 · outbound

This paper cites bit” vs. “bet.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement bit” vs. “bet

Reference 92

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T13:26:19.505856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.369265Z digest=sha256:f8345d3c8983ece58a9f951ad78a80145af326b1da6847b3bc839466abde6d38

Observation 7687b7be-def6-4817-b4ce-463f6fe28086 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 4186

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.293571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.293571Z digest=sha256:96849bef0ed69369c5d64591b97583e16dda39982f99980e948d513960c81dda

Pith citing papers

Observation abde2652-5844-45b5-bc1b-6b14c8aafb41 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.575812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.575812Z digest=sha256:946f118541a22a22726a8d092165d6264c2cd24d19e35b7263b93e8725c3ae6f

Observation 64546f18-8417-4c2a-b48b-0557f261eb3f · inbound

Entropy-based Coarse and Compressed Semantic Speech Representation Learning cites this paper.

Entropy-based Coarse and Compressed Semantic Speech Representation Learning Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:07.305867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:07.305867Z digest=sha256:3b9010613a8acfeb63eb653b727f22d251070ba1ce27d9730e200f04548cccf0

Observation 80a7b18a-b5da-4ab4-997e-a3a9b98d8558 · inbound

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck cites this paper.

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:50:51.780513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:51:38.030059Z digest=sha256:f543dacec8a0a9a83242494d95ee7bf668411dab2561cb540ccdb79e0640db6c

Observation 148ee6a2-99f8-4893-a50c-40d16a870387 · inbound

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora cites this paper.

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:01.383468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:57:07.894455Z digest=sha256:03938a930a2f0a73c66a571ad2c37dafdb924f165dc975489db9cb75ea481d04

Observation aef88493-6709-4379-9fe4-f2919aaf5837 · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:39:38.692633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:69c5cd0219f415a7dd14292d43005053158b0e6b06529c88ec7c005d4d803be5

Observation 6c019909-a819-4aef-b337-348e77d88fbb · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.092101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:fcd953f299bac6b6722e1e4f53a9917bde3984fbb81e43eaa679ffd1c36a817b

Observation cffc3372-92ee-43d3-99f7-c5a05a951d70 · inbound

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization cites this paper.

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T22:31:42.563915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:31:42.563915Z digest=sha256:94fc052ce69ebbaf007e422620cab9f2e70b55cb1a80912fde6fd23b5a5e0685