Pith. sign in

Paper Citation Record · LEDGER

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

As of 17 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 8 inbound Pith citation observations for arXiv:2502.07243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07243 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.369265Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.522433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.691071Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1234a951-ef09-4289-a4bb-218485f1202d · outbound

This paper cites Neural discrete representation learning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural discrete representation learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.088321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.088321Z digest=sha256:ef507dd59c16e58b8d05bc28a2aa54e8ef7d3fbf54e6d8ef3fe769a6c47f13c8

Observation 9621dbf3-776a-4578-87e4-6d5cc7130f2d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.092101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.092101Z digest=sha256:1a3a982e16a2090078e9e85c81116860a598868c71752e5285bc6ce9ba779b6d

Observation cdcc9f2f-cdc6-46b8-8c46-646cb2c70740 · outbound

This paper cites An overview of voice conversion systems.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice conversion systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.095768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.095768Z digest=sha256:d3e6c2e2d79dda846f8763ef250226c9bfc857f01daec778597e2dbcc1c2a1c5

Observation c384d50e-393a-4c38-9e76-05d4c1c59737 · outbound

This paper cites An overview of voice con- version and its challenges: From statistical modeling to deep learning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice con- version and its challenges: From statistical modeling to deep learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.099143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.099143Z digest=sha256:9cedcaa1416634ef0be5c70eb238c239e5e048a33e05550ae248dd97d2f5980c

Observation 2ef66d9b-3efe-4cc5-9fdb-32a19f32fb3b · outbound

This paper cites Foreign accent conversion in computer assisted pronunciation training.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Foreign accent conversion in computer assisted pronunciation training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.102330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.102330Z digest=sha256:97602aa5295fa1b5b8a9085741cc406ae015569ed52358406776a3e93f52a0f1

Observation 9f2d5106-5bcc-4539-92f2-5c08a4174d82 · outbound

This paper cites L2-ARCTIC: A non-native english speech corpus.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement L2-ARCTIC: A non-native english speech corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.105470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.105470Z digest=sha256:e78f8fcfeef45ae07251502f6f1ccfd7666535a28e586d31a649270d93efcfe5

Observation 84010b50-fc81-409d-b51c-3e124781da88 · outbound

This paper cites Emotional voice conversion: Theory, databases and ESD.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emotional voice conversion: Theory, databases and ESD

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.108882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.108882Z digest=sha256:8550951c1d85c8b5e56ecbe65c9a402954b35e8c077488719584c6fbbf36a510

Observation 8f4141dd-7b92-4bc3-a3ed-6c8dade41c33 · outbound

This paper cites Neural Text-to-Speech Synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Text-to-Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.112021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.112021Z digest=sha256:e0d9110c696abc89d0f6f16b6f02ec670a6d40343f028d906ae0737a133d3f1e

Observation bbf8051d-f030-4f39-8e3f-6990403e10e7 · outbound

This paper cites Converting foreign accent speech without a reference.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Converting foreign accent speech without a reference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.114754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.114754Z digest=sha256:f41abbb48407ab39e3073c8cf5ad0b584538b2b6e214cc7ea73fd92f85f39bab

Observation b5ddc4a1-f58f-452b-9dab-64b5718a6ba1 · outbound

This paper cites Sahidullah, Aur ´elien Bellet, Marc Tom- masi, and Emmanuel Vincent.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Sahidullah, Aur ´elien Bellet, Marc Tom- masi, and Emmanuel Vincent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.117635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.117635Z digest=sha256:5b480a6f74fd3232b0248b03459e62d2c455eef5750ea956735d6dc4ac7774a2

Observation 05a4d2f1-477d-49e4-9574-6aa2484f14a0 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.120659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.120659Z digest=sha256:d70422a46dd3797c127d5c9e88a92ce02d759238d1dfeb9a985673aa8112dbbe

Observation ab9f82f4-1ab7-4ad4-82d3-3aab0706d44f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.124096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.124096Z digest=sha256:30eb6999f8c93530ca1fc6c248e6b6557724b9b565e134c3aa194dd46987e12c

Observation 11b705f3-918e-4255-a214-222525ab2f0e · outbound

This paper cites Maskgct: Zero-shot text-to- speech with masked generative codec transformer.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Maskgct: Zero-shot text-to- speech with masked generative codec transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.127638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.127638Z digest=sha256:13d74bb59062cbb636d92205bc3fece40451ca3730eda6a143e86735bde218d5

Observation 07ff5752-9eaa-4c8d-a501-508a918d48fe · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.130748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.130748Z digest=sha256:bf056760e9e1ad0ae5c6d5d03b1c9d54e226ef41b8935caab6bfb39ec23d5df7

Observation 6a2d694b-e80e-42a1-9d7c-b13db2eb9f4b · outbound

This paper cites Speech resynthesis from discrete disentan- gled self-supervised representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speech resynthesis from discrete disentan- gled self-supervised representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.133715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.133715Z digest=sha256:b1d5ce0236d1e7cf4445c7e3a45c6f40ef4fead679fb5913a97a8cd355d2f9b6

Observation 6e9e5b21-1953-45c9-9c85-9baf1e1642fb · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.136859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.136859Z digest=sha256:6eb28cc328fa0f1b6a8bc85769b809222c6a09692e510bcfd94d74061578ae32

Observation 53b580d5-50d5-4831-b8b8-aab0451b14bc · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.140473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.140473Z digest=sha256:514d96c751bc276212c4f4307edeea13e8bf0c2dfdd2be37a61731a11bd9c66b

Observation 48cc366c-9511-4c1d-b45c-122c4c30fd9f · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:20.024364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.143740Z digest=sha256:8b616366120c9f23342b6b5a2bde4a184d82bc07b6d6f94518b739a7735962c3

Observation b16bd86c-730f-47cb-a719-d59fc2465f63 · outbound

This paper cites Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:20.015767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.146950Z digest=sha256:16a886aaad5169aa099fae3a79b06c8186b34bca1241f4a4b01c7ecbb81531ba

Observation b4e0abf5-deb8-4310-a6aa-de5bb23ebfa5 · outbound

This paper cites VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.150092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.150092Z digest=sha256:7636c5288929499ce135d26df82bd6b61699e811550f5521b1a2d8ac5118ef7b

Observation 6e1cff98-0b03-451c-8ab3-a3b8d7261506 · outbound

This paper cites Convert and speak: Zero-shot accent conversion with minimum supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Convert and speak: Zero-shot accent conversion with minimum supervision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:20.006904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.153241Z digest=sha256:7bfe39c284656572337e864c57d02498db69496e2dbe443c05dc2c1288c289a4

Observation d63fbd8f-c1e4-4798-99a1-4a417dabffc5 · outbound

This paper cites Au- tovc: Zero-shot voice style transfer with only autoencoder loss.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Au- tovc: Zero-shot voice style transfer with only autoencoder loss

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.998020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.156236Z digest=sha256:d2f9bf70d3179cfb0ee6dd4fb8e91f2948ddac17c1952679effec3693643f03c

Observation be8810d0-db54-481e-abb2-683a25f231af · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.158875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.158875Z digest=sha256:021a3e3eb738f5008dac3708aefccff51ef6429462f68e68515d7f11820ab567

Observation 712c1ffd-5225-45d3-bd18-6aba84d44a47 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.162381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.162381Z digest=sha256:9695c36fb3e1100d404c9191052805533341e679d677b1c9da7877de7c0f1cc4

Observation 54dc202b-f995-44b5-9425-fc2300fe4d46 · outbound

This paper cites Better speech synthesis through scaling.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Better speech synthesis through scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.165401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.165401Z digest=sha256:6bb3b92116161d31a932d00c06b4bb859f7288368cc16c4c1bdee1dadc0ce806

Observation 8dd7fa07-0c42-4fbc-9577-7def92b7ad11 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.168802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.168802Z digest=sha256:cdb6d508e9df339c8e6b6e79f1c04d52ffda273681a4df619da23c03b8af1d33

Observation 9bf9cba5-4724-401d-84a2-5dc3a569a6c9 · outbound

This paper cites V oicebox: Text- guided multilingual universal speech generation at scale.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicebox: Text- guided multilingual universal speech generation at scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.172375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.172375Z digest=sha256:386ff3d488b99699d5818a6cc96c6d1dd5c9a20f00f9129359b794e8bcc33e00

Observation 8fc2728d-5e2a-48df-ab02-34693677b7c2 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.984343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.175536Z digest=sha256:78a3f937ad99cd4cebd8731cffab05a4d72a07d5e5c2f8dded10beb5ed1355ae

Observation ff1696aa-260f-49fe-be8d-3e0e641aa098 · outbound

This paper cites V oice- preserving zero-shot multiple accent conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oice- preserving zero-shot multiple accent conversion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.975383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.178687Z digest=sha256:cc2268d921dbf2e6ef9da483e26ef1722e672b66901a43d5083eccd5a75f7581

Observation 9d8b0e25-1f48-424f-a484-258480a55ff0 · outbound

This paper cites Schuller, and Haizhou Li.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Schuller, and Haizhou Li

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.966785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.181727Z digest=sha256:d16cbcce86a5be294108ca73b974b2798f1413a93b200561bc5344754a117316

Observation e4bf249b-c83e-41c0-bb00-de4227ca1539 · outbound

This paper cites PA VITS: exploring prosody-aware VITS for end-to-end emotional voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement PA VITS: exploring prosody-aware VITS for end-to-end emotional voice conversion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.958284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.184577Z digest=sha256:16fddc277241d264771b7c29695e791e645a5f35b1210c9006945eac3f03bc15

Observation cb915b3c-08b9-4090-8a15-d8a50bf05a15 · outbound

This paper cites Transfer the linguistic representa- tions from TTS to accent conversion with non-parallel data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Transfer the linguistic representa- tions from TTS to accent conversion with non-parallel data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.949309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.187544Z digest=sha256:017ce64fbea11cb6ce140a2be56c5a4c69d06ffdaac96ce3c4bdc22a49933a4f

Observation a52ec863-cd90-4e3b-97da-87c94d9798dc · outbound

This paper cites U-style: Cascading u-nets with multi-level speaker and style modeling for zero-shot voice cloning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement U-style: Cascading u-nets with multi-level speaker and style modeling for zero-shot voice cloning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.939715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.190541Z digest=sha256:afbc810f50c11a6ed1062daf201b6e7307e19981c1a78b67fc37368681b6930c

Observation 32c230e2-b170-4780-855b-e79bba4ccccb · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.193725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.193725Z digest=sha256:421815d3bfcd51f0dac0f2046ac6d96163251a06fecffe775c9cf4ce438fb216

Observation 6baec9ef-0a8b-435f-909f-835a1966cab0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.196997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.196997Z digest=sha256:c80927cffb04a1f334df952f39f890cdb0b89958467fd23f0803f04b42e29c71

Observation ffb292e5-e1b2-4a92-b961-b30131fcd04a · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.200113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.200113Z digest=sha256:9116f166cb0c740367e2322591df4062c54c1f4e14c10b9e230efc48e2ac1aa6

Observation f59cd892-b39e-4fb8-b11f-f5882fcac97e · outbound

This paper cites Scalable diffusion models with transformers.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scalable diffusion models with transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.920850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.202942Z digest=sha256:cdcb801c58cb776c8001c06b097cff0532ae11c354f4f7a85b63db441f9546fd

Observation f9b274d1-daa9-4b5a-85f6-5411061b25d3 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.205829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.205829Z digest=sha256:a3e65f447c0b725119543c019f83dc0eb2b1a78ee1d4a48bf6b1b452cd390772

Observation 02e7d8be-5b9d-4760-aba5-1cb3debb871c · outbound

This paper cites One-shot voice conversion by vector quantization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement One-shot voice conversion by vector quantization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.912239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.209096Z digest=sha256:19c8c82a1cd6e5c1cc40af16556f4d29bbadb0deab525e312a5b058516be7354

Observation fff07ec3-4590-4663-93c6-cc10cca13954 · outbound

This paper cites Unsupervised learning of disentangled speech content and style representation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unsupervised learning of disentangled speech content and style representation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.903838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.211965Z digest=sha256:2a08ebcba7ff29a4e4b3c621d7566cef4c8499d382a572f2204caafad687b8fd

Observation 0b2c7c32-6e36-488c-843a-6e54cf7b2674 · outbound

This paper cites Text- less speech-to-speech translation on real data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Text- less speech-to-speech translation on real data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.895972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.215144Z digest=sha256:a919c6c0cb6c5a06b90f071aa8a99722a83227c416a8d6f17eec63f04792b9b5

Observation f2dd42d1-ebc7-41b9-9bc5-b15b453909e6 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.888168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.218507Z digest=sha256:488fa97bbed1d36ae6adfd3a36cec1d8246410b50b4c743b7eef695862c5c278

Observation cce790f6-3be1-4ef6-8e45-4148352c4a91 · outbound

This paper cites A comparative study of self-supervised speech representation based voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparative study of self-supervised speech representation based voice conversion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.880694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.221950Z digest=sha256:7f9fc12cb95c86b1173fd32d6434a256d38b18eae737af36f997fe996be0b3b3

Observation 7e0f126a-39a0-4447-90e0-55372264de06 · outbound

This paper cites Leveraging diverse semantic-based audio pretrained models for singing voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Leveraging diverse semantic-based audio pretrained models for singing voice conversion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.872428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.225095Z digest=sha256:a94b42d5fc8f45ab9eb420a422dbd85265e08d229f2cda1fbacdf4c499ce3b1c

Observation c6debe6e-8d3c-4c3c-9386-b7b172cb9bbb · outbound

This paper cites Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.864166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.228370Z digest=sha256:eb8366d21873768059d376b6bb1b276f94cdf696f846d3940ba7c1edddeb4dc2

Observation b21ccbbb-9ffe-4d95-b1a1-946d7527d093 · outbound

This paper cites Stargan-vc: non- parallel many-to-many voice conversion using star generative adversarial networks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Stargan-vc: non- parallel many-to-many voice conversion using star generative adversarial networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.855792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.231460Z digest=sha256:8ed3f96ef44236fbacf3d41a330a699d8bbf12c6286a90503680725e17dc0f29

Observation 7ade7335-7583-48a0-bb39-d1564a094f0a · outbound

This paper cites Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.847440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.234763Z digest=sha256:b2e077b03fc72aeee4fa9e316608f700783aa98458b37e68d435f1396a7bf79f

Observation e579a085-1702-4f69-8e7a-d3a41f7d7b54 · outbound

This paper cites Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.839236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.238179Z digest=sha256:95798c13fda2936803f906e3e158fcdef621e6d2cb4f88c59c60a925c4be652a

Observation 0117c55c-388b-4907-b30b-5a76b2feb2ed · outbound

This paper cites Tts-guided training for accent conversion without parallel data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tts-guided training for accent conversion without parallel data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.830985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.241259Z digest=sha256:a74f49e450506ccc8b29c3ce452bd097c786394c8dc4e7f9b98e060c10f11f97

Observation efd90fa0-40a8-41eb-84b7-b3095e3c6b4f · outbound

This paper cites End-to-end accent conversion without using native utterances.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement End-to-end accent conversion without using native utterances

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.822725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.244394Z digest=sha256:e346c7486b12ed7cd42d933c9aae859b2ee9c42900ee9ea184cb014c20ad5400

Observation 2333a5bb-d67a-49cb-a334-c7bb3dee15cc · outbound

This paper cites Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.814630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.247887Z digest=sha256:632880d655cbfa299b33b8c1352367aca1ba14d7155e4a56b7f77a2a59bb1e4a

Observation e933a787-f410-4603-b343-601a48a25b3f · outbound

This paper cites LM-VC: zero-shot voice conversion via speech generation based on language models.IEEE Signal Process.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LM-VC: zero-shot voice conversion via speech generation based on language models.IEEE Signal Process

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.806051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.251307Z digest=sha256:d9eaa19e56c900287716cb184a2088ab9646ac47dc942494514a1e4c28761c54

Observation 49e1554d-4460-4704-a07a-acb5915ffb54 · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.254583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.254583Z digest=sha256:a69e2df17071b3294dd75876ce7473158a67dfa3441280fd8ebf5961b11d67d2

Observation 616a3740-187f-4e40-9ae3-7ad81e2b7bca · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparison of discrete and soft speech units for improved voice conversion

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.797639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.258005Z digest=sha256:b33fd124b36ddcdc1c3cfe017ec40306b64d11a00ac09db63e823979a10e24a9

Observation 7f46fa70-3f63-4b79-80e6-1882ea7259c6 · outbound

This paper cites SEF-VC: speaker embedding free zero-shot voice conversion with cross attention.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement SEF-VC: speaker embedding free zero-shot voice conversion with cross attention

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.789762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.260858Z digest=sha256:59ea3813c13ae027d7cc7f3df96221269e9aea72d9b4460c10a377762ff0c971

Observation 019df985-128f-4a50-9923-fc61b67370ae · outbound

This paper cites Neu- ral analysis and synthesis: Reconstructing speech from self-supervised representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neu- ral analysis and synthesis: Reconstructing speech from self-supervised representations

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.782063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.263434Z digest=sha256:166f435ed02d93f380a9cd9a8dd90bd3f0aa2f0b4804d412ec7e9ada2bfc4889

Observation 8d86c200-b042-4825-9fd0-b3f4f953501e · outbound

This paper cites NANSY++: unified voice synthesis with neural analysis and synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement NANSY++: unified voice synthesis with neural analysis and synthesis

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.774318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.265980Z digest=sha256:33243926ca2bd627295d82246bc3d320a34c4a5b09980f7a60eb6201a8bb9f66

Observation def82cc8-225b-4949-9917-0a981577767f · outbound

This paper cites Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottle- necks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottle- necks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.766187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.268522Z digest=sha256:681ed30fff1e520d77ea0039d86f26d6867b4ed7877fbd47fb024d99bea2ceca

Observation f0e34687-921c-4eed-87e3-61b6123eb262 · outbound

This paper cites Cox, Mark Hasegawa-Johnson, and Shiyu Chang.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cox, Mark Hasegawa-Johnson, and Shiyu Chang

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.757329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.271441Z digest=sha256:a203bd59dd59a16c4d0e02a873a3bfdefc8b4d6d54d06864c1bfe1b305805a45

Observation e0381617-1b1b-4451-9a2e-b8ba001f3301 · outbound

This paper cites Multi-speaker expressive speech synthesis via multiple factors decoupling.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Multi-speaker expressive speech synthesis via multiple factors decoupling

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.748797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.274184Z digest=sha256:526dfb5adbf7f002632cf71f4bad07658b56bc1b49e131f89ae62c92f4df6da3

Observation efdb67c7-5795-43ff-878b-ab42b2040bc5 · outbound

This paper cites CLUB: A contrastive log-ratio upper bound of mutual information.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CLUB: A contrastive log-ratio upper bound of mutual information

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.739916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.276826Z digest=sha256:ceaae9ced967b60913e17b60c65bfc225db1a3aecfaaef6a4bb75f60d7e06371

Observation 44448173-2d8b-4098-ae9b-fd6b8fe82fd4 · outbound

This paper cites Repcodec: A speech representation codec for speech tokenization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Repcodec: A speech representation codec for speech tokenization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.731258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.279359Z digest=sha256:af0f5e4ddda8e5926285d1d49dfc45b2eedf01ad911aa8305e16870c5a1573bd

Observation ee6f654c-e8bf-43d3-92ba-67f5aa894cb3 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Soundstream: An end-to-end neural audio codec

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.281773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.281773Z digest=sha256:b5133a13948c84558ca1101ec91e2a578262715a0f04f0fb9c8ef9ab5b354fc3

Observation 3a7660a3-b4de-4b7a-852e-2f8cbe5b0437 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Wavlm: Large-scale self- supervised pre-training for full stack speech processing

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.717869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.284565Z digest=sha256:1d015235f7cf303e2e2ee55ad79b19312657000d10fdbf116576c11c8192565c

Observation 04a83850-4d53-4914-b52d-5c58321399a7 · outbound

This paper cites ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.709198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.287428Z digest=sha256:67357129e4e47533958102a9102902b6cf56519e65fe2dac8d5e002e8aab61b0

Observation cb1b2cf2-77a4-4c63-8175-dbb8f8527d1a · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BERT: pre-training of deep bidirectional transformers for language understanding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.700442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.290292Z digest=sha256:5b821d5fa13d5a9b1f771c3b87498d2539cb891e4b38991fa9d8ca1ec79bdf92

Observation 1caf02e0-b17b-4418-ab8a-847aa87da851 · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Bigvgan: A universal neural vocoder with large-scale training

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.687235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.296292Z digest=sha256:0105a2391ad06316acddc6ce4b26b1d9be36cd8dadee386f7b3c112714259fc1

Observation 97818603-d8b1-4009-b416-4ffbedb85cdf · outbound

This paper cites Tyers, and Gregor Weber.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tyers, and Gregor Weber

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.678910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.299049Z digest=sha256:86cf1f4d1a201271fa651ee7a7442f6d7fbb15fd054c4c60aca4bd775da21757

Observation 7a190bd8-a972-40ec-b49f-710944c8370f · outbound

This paper cites The singing voice conversion challenge 2023.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement The singing voice conversion challenge 2023

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.671037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.301962Z digest=sha256:bfff9010cd46bce5a851ba7c44979e001e8837f4c1a4e8af365c896ebf053338

Observation 41e9396b-97de-4d23-847c-cb904661bfc9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Robust speech recognition via large-scale weak supervision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.662748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.304779Z digest=sha256:6d207cd23b2896d33992a34d61d6d48f927321d5164a1e988f08b50600443c65

Observation f129dbbf-ba2e-4d4d-b419-611e0c960a33 · outbound

This paper cites Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.654569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.307589Z digest=sha256:94d813df3eba3a980ba0ce42346b9cacb2d405de69a5ec60c626113ef84859cb

Observation d69f8ee8-2721-4124-a54a-686965e816ef · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.647207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.310413Z digest=sha256:dd8240d360b2c8c330e60d82d07a7692f649c36c929c77713d3670650c01ef88

Observation 08f94af4-4fca-4428-8029-ea535e194cae · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.639684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.313298Z digest=sha256:e5bf403118765a1b0bb0a586aa33b464c639a2aaad14ae00ce2924aa2d2b715a

Observation d2b6dc8d-aa5c-4335-a1a2-5d2e740ad999 · outbound

This paper cites Libri-light: A benchmark for ASR with limited or no supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Libri-light: A benchmark for ASR with limited or no supervision

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.631954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.316297Z digest=sha256:4b582b771a30bc47e04fd9c5f3fb70ec668bff247e70b9bb19d8d25739a13089

Observation 91b2c8cf-ee67-4ac5-a632-a16b75a59656 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Librispeech: An ASR corpus based on public domain audio books

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.622739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.319210Z digest=sha256:3004a6717e84bcf83634284c2cffef3179f0db090a6dbcdc887a1cd5e8cd646c

Observation cc567952-405a-4834-803e-a1f8637b199a · outbound

This paper cites Amphion: An open-source audio, music and speech generation toolkit.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Amphion: An open-source audio, music and speech generation toolkit

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.613920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.322124Z digest=sha256:b5592fe09aad336cf90e2ee42fbe3ab8e007c5f91535aa984b0155770f5a02cc

Observation 35599e59-9f1a-431a-b55f-394a892cbc4b · outbound

This paper cites V oicecraft: Zero-shot speech editing and text-to-speech in the wild.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicecraft: Zero-shot speech editing and text-to-speech in the wild

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.605013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.324971Z digest=sha256:e71d9a22aa8f6ce64d7b72c45e69d12b3bb34931a8de42982b4445527b8a423b

Observation d4d9660f-dee0-4fd3-836c-94376ab66820 · outbound

This paper cites Overview of the Amphion Toolkit (v0.2).

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Overview of the Amphion Toolkit (v0.2)

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.327889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.327889Z digest=sha256:74fd499d894f2887791ade7c3b8fe89e2bc6cf1eff114ec91c46a1e2a016152b

Observation f007b9ae-b6b3-4469-a20f-6ff698ee87a3 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.596064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.331052Z digest=sha256:d265c81a3a9e9725c3101288f537f595b67881791f50ded63ec825d5805bb8af

Observation 8ea99c5d-0d73-46ab-a699-2c5b5fd2d259 · outbound

This paper cites Decoupled weight decay regularization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Decoupled weight decay regularization

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.586922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.333766Z digest=sha256:1a9d1c8dede87c80f452dd3b9b2e2a532251c4587f47701d56e3b4a5d8885e58

Observation f4437809-355c-4f4e-9839-d4877cffa94a · outbound

This paper cites Kingma and Jimmy Ba.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Kingma and Jimmy Ba

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.336676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.336676Z digest=sha256:5c8cce81071bf6153a8ac486680b28187b717fc243d1ea8be83a2a01625f2079

Observation ef3c1bcf-7fce-40a9-a5af-1a4ec8b76b51 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Classifier-Free Diffusion Guidance

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.339612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.339612Z digest=sha256:b8a866f56882bddfaf975789c8ba358a3d47b6036a1355a3709fb3d742a70fc6

Observation 71139633-d1f4-49e7-9afe-ed8423788c6e · outbound

This paper cites Scaling speech technology to 1, 000+ languages.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scaling speech technology to 1, 000+ languages

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.573407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.342687Z digest=sha256:089440548b3ba15c8e5cdcd5bb8ea5df7e60f64022977ff3180b2dfabcf77bf0

Observation fab7ebe3-0ec2-4843-ae1e-2de7f9ddbfb5 · outbound

This paper cites Conditional variational autoencoder with adver- sarial learning for end-to-end text-to-speech.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Conditional variational autoencoder with adver- sarial learning for end-to-end text-to-speech

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.564600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.345544Z digest=sha256:809243bb2015b0f4efc8e44be8b0f168d5bf2d1fe981222eddddfed760d717bb

Observation f0be682c-099a-45b8-9852-3d4430a7214b · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.555736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.348402Z digest=sha256:435daec046c298febb89ec64148cb684a7de2444663fdbe2429429dcfde69304

Observation ed0f2eb6-4b91-4dec-9f18-5b4edbf627ae · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.351240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.351240Z digest=sha256:91f86d628db474af8e8d5b061412be8f94325f3e152e291480bd273e4a5294e1

Observation d2eb39b9-9991-480f-9f29-ec8fef4334f2 · outbound

This paper cites BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.541564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.354382Z digest=sha256:3a2fe1afc30f4a0807ac2249aed86fe5fc5f53e255c6149a67eb9d0945cad0dd

Observation c7454f28-2d70-4bc5-869a-936e0d8b3b9b · outbound

This paper cites CSTR VCTK Corpus: En- glish multi-speaker corpus for cstr voice cloning toolkit (version 0.92).

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CSTR VCTK Corpus: En- glish multi-speaker corpus for cstr voice cloning toolkit (version 0.92)

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.532562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.357239Z digest=sha256:7df730ead76fb12861b9704855ff09b92e8a5dcb2105566c8d23b44a050bc95c

Observation 0b0598f3-d986-426f-9a83-05f99e54586e · outbound

This paper cites High Fidelity Neural Audio Compression.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement High Fidelity Neural Audio Compression

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.360299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.360299Z digest=sha256:7fa8cdf702d2cee0d40e8ab9de474de30f22634da228fbd89ba6f4a4014e3f1f

Observation 050584c9-9979-4017-a8a1-00ce234971ad · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement MLS: A large-scale multilingual dataset for speech research

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.523449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.363655Z digest=sha256:1f211d3a11073f2d27646ede1e8227ac920c0b571abfa8b29d930a413816d86c

Observation b8aac263-107c-4a80-9934-b90cdf3a6e72 · outbound

This paper cites Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.515151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.366424Z digest=sha256:2c7ec632c3cda4dc8fa9fc96d3c6c83008c9301124bd54932f98e18371d53bd0

Observation 06cef095-3980-4f5f-b5c1-4fef04f485f6 · outbound

This paper cites bit” vs. “bet.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement bit” vs. “bet

Reference 92

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T13:26:19.505856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T13:26:19.369265Z digest=sha256:39c819ccd7511966706cbb8511a29462c33f3319439aea9d35d038fa98483803

Observation 7687b7be-def6-4817-b4ce-463f6fe28086 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 4186

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.293571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.293571Z digest=sha256:85676b2e74ce4dc793bf40bdf1e6514507be1dab1ffa0627a9e6032965368181

Pith citing papers

Observation abde2652-5844-45b5-bc1b-6b14c8aafb41 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.575812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.575812Z digest=sha256:06494f8b5da9fff2688be221461baf9f74ab55b371036db9d89b245f384f17c4

Observation 64546f18-8417-4c2a-b48b-0557f261eb3f · inbound

Entropy-based Coarse and Compressed Semantic Speech Representation Learning cites this paper.

Entropy-based Coarse and Compressed Semantic Speech Representation Learning Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:07.305867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:07.305867Z digest=sha256:63a2991d88f45eda97f3d6dd5ac679ad65bf005204bb1ffa44ad76c57f2b20eb

Observation 80a7b18a-b5da-4ab4-997e-a3a9b98d8558 · inbound

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck cites this paper.

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:50:51.780513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:51:38.030059Z digest=sha256:a880b4c1bdbac13995dc1fe9a43eca5983af85f9511fdcf92cb9ac5b6aa3d14d

Observation 148ee6a2-99f8-4893-a50c-40d16a870387 · inbound

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora cites this paper.

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:01.383468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T14:57:07.894455Z digest=sha256:bed329b950d87ef0b4a3a15fc651df946026e876bbd6a2fbcce372900966157b

Observation aef88493-6709-4379-9fe4-f2919aaf5837 · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:39:38.692633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:0ef1df016322215006d48bb2ba2db7a6e63f864c7ce412757e5ea2c08b8557d6

Observation 6c019909-a819-4aef-b337-348e77d88fbb · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.092101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:fb5ddbd614fba8caff86f076aa83c9b5d76a0ce79d0a547a794ca66f86833c42

Observation cffc3372-92ee-43d3-99f7-c5a05a951d70 · inbound

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization cites this paper.

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T22:31:42.563915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:31:42.563915Z digest=sha256:e432e48adfd682c1b59495ed38eee39e73aac8a054a92063afb9aa2158d0adb5

Observation 86d3fbae-e589-4fea-8288-7939cd8d6701 · inbound

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances cites this paper.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.522433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.522433Z digest=sha256:ea62e8002871a3709145497b07cf22710387944d74174b4b48a790573e801de3