Pith. sign in

Paper Citation Record · LEDGER

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

As of 9 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2507.14534.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14534 v4

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:57:23.386903Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T21:17:48.886421Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:16:12.319512Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cedb4dbb-0fa5-4623-bdf8-6a18757a6d5d · outbound

This paper cites StreamVoice: Streamable context-aware language modeling for real-time zero-shot voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion StreamVoice: Streamable context-aware language modeling for real-time zero-shot voice conversion,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.317774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.058155Z digest=sha256:c0bfae497164b0d1bbd872c958d500ee8ef057c6d14c840e8377276ae8e68838

Observation 77e58ac2-15a7-4164-87f8-b751dd5abf80 · outbound

This paper cites Vqmivc: Vector quantization and mutual information-based unsuper- vised speech representation disentanglement for one-shot voice conver- sion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Vqmivc: Vector quantization and mutual information-based unsuper- vised speech representation disentanglement for one-shot voice conver- sion,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.301228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.124148Z digest=sha256:0adfd9d804bc47d3f95644c7e6199595432d5386afb43b9e3394e86c260d5724

Observation 6e9ee150-7279-4d55-a59c-157e9a33a042 · outbound

This paper cites Autovc: Zero-shot voice style transfer with only autoencoder loss,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Autovc: Zero-shot voice style transfer with only autoencoder loss,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.284013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.205359Z digest=sha256:c69df36de7501a6f8904594c9e863145a93d90cef3c673358f6a1c7178f2d8f8

Observation 7dc0f064-a8a4-444a-98f9-1b116e5e5a28 · outbound

This paper cites Controlvc: Zero-shot voice conversion with time-varying controls on pitch and speed,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Controlvc: Zero-shot voice conversion with time-varying controls on pitch and speed,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.269197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.327182Z digest=sha256:3e3d98a433b03a900ea25afa48b836b691b2d26952762417812fff7068009415

Observation 5b65b122-8ace-47f4-8fcd-ad16828b37ba · outbound

This paper cites Streamable Speech Representation Disentanglement and Multi-Level Prosody Modeling for Live One-Shot V oice Conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Streamable Speech Representation Disentanglement and Multi-Level Prosody Modeling for Live One-Shot V oice Conversion,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.253979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.387237Z digest=sha256:282bc8e63f190e0aeb5e856c5ceb33c5c34fa00f63257e0fd15291992e1425b7

Observation db8fc721-dc17-4452-8afd-bd97b4987118 · outbound

This paper cites Streamvc: Real-time low-latency voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Streamvc: Real-time low-latency voice conversion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.239188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.429437Z digest=sha256:b53899699e099a3013c7f3f93a90fdc60edb4636f061447bb36f3c795c315eb9

Observation 5fd170c3-41f0-43cd-92d2-d365d32067cf · outbound

This paper cites TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:20.468550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:20.468550Z digest=sha256:fae07d9e53781919af0bb209bf033bb02f6afcc88d8e61d2442512fc655a58ed

Observation a770f3f7-d1b2-4e12-ab4a-000c0cdf4d44 · outbound

This paper cites Isdrama: Immersive spatial drama generation through multimodal prompting,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Isdrama: Immersive spatial drama generation through multimodal prompting,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:20.561847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:20.561847Z digest=sha256:111fe5171d4f63202642e24a4323b4e802e3ecf1ff50a082087aa7ec755fe5ba

Observation 2cb7ec70-99cb-4f02-9f9f-5c5fcae37b5e · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion A comparison of discrete and soft speech units for improved voice conversion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.224119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.642241Z digest=sha256:5d6804c4d8261c3a6b0ab75776235eca356610975f81d877c94d41136bb89f7a

Observation d56cdefa-73c8-4ca9-8f2f-eabfa2678478 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:20.755757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:20.755757Z digest=sha256:df0112c91b93ab18fb9f4b24e8ec4bc79d133e9f971d40e742c910e3d07a654c

Observation 61d08d86-e121-47f0-8d73-727c8a3b5f26 · outbound

This paper cites Phonetic pos- teriorgrams for many-to-one voice conversion without parallel data training,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Phonetic pos- teriorgrams for many-to-one voice conversion without parallel data training,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.196974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.822747Z digest=sha256:e75b567c712894e27f4f021ddbe49dbf57f9146a0b3aaba8c12fb32252df83a8

Observation 7a52f3ad-e015-45dc-8824-89840f32ea6a · outbound

This paper cites Starganv2-vc: A diverse, unsuper- vised, non-parallel framework for natural-sounding voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Starganv2-vc: A diverse, unsuper- vised, non-parallel framework for natural-sounding voice conversion,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.180932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.892516Z digest=sha256:21a7e14cd3d920e0b1693b9385665129bf935bd92ebd5cdbb701bc9ef92f7c1e

Observation 91124d21-72f8-47bc-90d3-cf7e4999109a · outbound

This paper cites End-to-end streaming model for low-latency speech anonymization,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion End-to-end streaming model for low-latency speech anonymization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.166104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:20.956605Z digest=sha256:3aad4f108ddbde3172a494f2cc5d6ce700775c52cb8d43b6deaf2af5d97a3c5b

Observation 6a760ef3-0401-43bf-bd99-f3d878867ffd · outbound

This paper cites Contrastive predictive coding supported factorized variational autoen- coder for unsupervised learning of disentangled speech representations,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Contrastive predictive coding supported factorized variational autoen- coder for unsupervised learning of disentangled speech representations,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.151943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:21.020533Z digest=sha256:62ea0128f378368284242ba27767cf42d0ac18ee810f67c470b915a9d0689909

Observation 96a890f9-a376-4685-8dd6-17cd857f0e3c · outbound

This paper cites Neural analysis and synthesis: Reconstructing speech from self-supervised representations,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Neural analysis and synthesis: Reconstructing speech from self-supervised representations,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.136970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:21.081301Z digest=sha256:3dd226ebdee5fc5e7b108f49902f7df8c54b2c755bce3e38ad467b443a6658db

Observation 1b9e8804-0a17-4676-90be-d75e25940d71 · outbound

This paper cites Lm-vc: Zero-shot voice conversion via speech generation based on language models,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Lm-vc: Zero-shot voice conversion via speech generation based on language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:21.209242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:21.209242Z digest=sha256:09533b60e51bcab51af9f3146906d998d5424419249dfbfe9384dfd12640dc0d

Observation 8c907b39-3523-476b-88cc-fb4502ebad7b · outbound

This paper cites An investigation of streaming non-autoregressive sequence-to-sequence voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion An investigation of streaming non-autoregressive sequence-to-sequence voice conversion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.111657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:21.349497Z digest=sha256:e9e8b6c5bc57082304c3f72e7ced68281c9c748570f469a6a1be735289707860

Observation 7b4fa9f6-cead-40b6-99e0-3546858cac79 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.078539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:21.451568Z digest=sha256:efa1abc1a9ada2632f5b8fa2e43720d65e3e2ea92cd75ddf0d3494af90f2cd0d

Observation 20d40fe8-b908-477f-8a3b-8040792bb92c · outbound

This paper cites Non- autoregressive sequence-to-sequence voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Non- autoregressive sequence-to-sequence voice conversion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:28.987253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:21.521693Z digest=sha256:ce0202c6246beb19bf28c1e88c7ab9e96ace679eaa236f89f2bc92a58d5439ca

Observation 99af9cd1-284f-41ba-b65d-df1f8d8ce90e · outbound

This paper cites FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:21.646838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:21.646838Z digest=sha256:d8670f53e45cdfa688b6701cedb94d64bd36f74dea5421550b0aa6d2c0e9e6bb

Observation 413f6a08-89af-41ff-aa52-b939132cfaf1 · outbound

This paper cites Streaming voice conversion via intermediate bottleneck features and non-streaming teacher guidance,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Streaming voice conversion via intermediate bottleneck features and non-streaming teacher guidance,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:28.922925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:21.788288Z digest=sha256:e90b407ce6f3d7a6a2e829476ea11e52c81a5fe5dad1df4df5f5bf7a51c7c926

Observation d80a12b9-c892-4f3d-b002-b69ee35365d6 · outbound

This paper cites Dualvc 2: Dynamic masked convolution for unified streaming and non-streaming voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Dualvc 2: Dynamic masked convolution for unified streaming and non-streaming voice conversion,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:27.723679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:21.890854Z digest=sha256:e17ba01d1687db5652db8e53022d52682f40b4591814bf0a557fc688707df6c5

Observation d29c1678-9cc9-459f-bf71-ea0f94c461cd · outbound

This paper cites Alo-vc: Any-to-any low-latency one-shot voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Alo-vc: Any-to-any low-latency one-shot voice conversion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:26.410264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.006855Z digest=sha256:81346f5dd4287ae6a84b4883d6e1482e4302d32e5e668699349b71c26697a15a

Observation 29fa0b9f-0546-45cd-9807-a5059a1343d4 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:26.009254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.085174Z digest=sha256:3ef47c41d3c49c0cec1112fe9d683a287a2e38cd82c010aaeb45b45b5afed7f7

Observation 7f253c7c-e32b-4f9c-9cbe-6eff216ecd08 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:22.132763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:22.132763Z digest=sha256:f1d16fe86cfd1285a406f795bb09086d0a344d27149cdff32c75956793d13b62

Observation f02d83d8-498d-4482-b44d-365368ac1e63 · outbound

This paper cites Attentron: Few-shot text-to- speech utilizing attention-based variable-length embedding,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Attentron: Few-shot text-to- speech utilizing attention-based variable-length embedding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.846045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.184149Z digest=sha256:615c32aaf7a17b025ec5c34780f4bfe367d13a2e2b4fa916a2119ace9a968223

Observation 6bf52a16-2edb-4b8d-8599-c30401a74128 · outbound

This paper cites Normalization driven zero- shot multi-speaker speech synthesis.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Normalization driven zero- shot multi-speaker speech synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.663916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.238284Z digest=sha256:efc885a930e799300c8ff39945fd9665356bd58fbf338295c6aebf30c06e3477

Observation 6914e1ce-e1a9-4fd4-84a3-6b8419a2dfcc · outbound

This paper cites Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:57:23.740340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.279661Z digest=sha256:dc9b62ba5f736d1c38e4d65959828c45f8936a7407aaceada2d2171989734920

Observation 59926bab-b637-4466-ac8a-85b31fa7c54d · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.482486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.352742Z digest=sha256:048990c5fa037108ba08607462c8c5003870432e6c77b84abe19c2ff8d809c83

Observation 14f439df-7073-4c64-87c6-958de311e3f5 · outbound

This paper cites Styler: Style factor modeling with rapidity and robustness via speech decomposition for expressive and controllable neural text to speech,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Styler: Style factor modeling with rapidity and robustness via speech decomposition for expressive and controllable neural text to speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.383341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.405331Z digest=sha256:1e3ac11ffc812e49d2682d60ca37b70441fdc3273d1453b6002f30e46ad32e51

Observation e10b327d-5ade-4bca-85bb-6fc4cceee77a · outbound

This paper cites Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.247102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.447883Z digest=sha256:8256e70183393e07c5fb29f3036a31d8e9c35ad0bc19a7f077900737b9c0e6bb

Observation b9c96fa5-2edf-40a3-b1a4-501bdbc90c3b · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.119522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.539901Z digest=sha256:d41f2d7b6cc6c02ff6734d83ce2471a287b1f9aa3e3e6d3ef91f75bb19a70429

Observation 0fabef11-8d8b-4b56-867a-9e69465e9f56 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:22.591196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:22.591196Z digest=sha256:4bdf1a7e66bd9df6acec44389da43d5ce86ee5221476b2959c642457b18a71e3

Observation df94fc8e-f56e-4d3c-b282-fb44fed0ee30 · outbound

This paper cites Fastpitch: Parallel text-to-speech with pitch prediction,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Fastpitch: Parallel text-to-speech with pitch prediction,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.969094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.641193Z digest=sha256:bf19b46b604fafa698cfc1a043477ffd09eee41e340cb81fa7e27cc5a829b210

Observation aa4e9205-55d7-43d1-98be-f6f2e64cde40 · outbound

This paper cites Tcsinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Tcsinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.822684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.706455Z digest=sha256:e1d9f9c9fee81448dda20214c904f9dedf2f9df645efe8af67f15c5ee7a14079

Observation a45c2ba1-c090-4c1f-98f4-1dabd0e10053 · outbound

This paper cites Online clustered codebook,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Online clustered codebook,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.718540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.774621Z digest=sha256:db1ef4e09bb84747b3a83eae5763723b3701091c7ada9cbec8dc15ddeaf24dad

Observation 413b1fd2-3a26-447e-b60c-260114e88022 · outbound

This paper cites Versatile framework for song generation with prompt-based control,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Versatile framework for song generation with prompt-based control,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:22.843384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:22.843384Z digest=sha256:1ab5ec720cf6e44b093c324e8f6c5273f73b3f654b651da00d2763cc780029e3

Observation d6d7515e-12de-4b4a-9708-c365d02980b0 · outbound

This paper cites Neural discrete representa- tion learning,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Neural discrete representa- tion learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.601983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.872499Z digest=sha256:be663554491e402a6bb209ecf1c69024eec113d7792968ed98a40803d36ba4cc

Observation 192d2527-d4dd-4399-9079-c3b0aa3481c2 · outbound

This paper cites Stylesinger: Style transfer for out-of-domain singing voice synthesis,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Stylesinger: Style transfer for out-of-domain singing voice synthesis,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.472361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.907363Z digest=sha256:5d169bd7713eca08daf530131bd777a52f5efb77f183ccbc70a491471327d7ed

Observation 70e69050-c6de-4617-bf63-2846ad79f51a · outbound

This paper cites Attention is all you need,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Attention is all you need,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.366429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:22.946130Z digest=sha256:b76c0095dcf2d540bf3459718caa4ed3e116a409f6568960632b71717e692674

Observation a3c0a13f-26ed-4cf6-8595-a7cfd7a99d5f · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:23.040525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:23.040525Z digest=sha256:b3182394e00ce0b9b627fd6168d66c611a3c34c211453b4c516fbe4aa5f079b8

Observation 2d7addaf-f07e-44a2-94c2-cd3f7f332b14 · outbound

This paper cites Deconvolution and checkerboard artifacts,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Deconvolution and checkerboard artifacts,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:23.075340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:23.075340Z digest=sha256:2aee1fb32d1d13e1ed42e7d41d910a3d0c127c4e63c93a992931411663d09853

Observation cc8f04d0-4f09-4b09-b8e2-1ba40b78caf3 · outbound

This paper cites Least squares generative adversarial networks,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Least squares generative adversarial networks,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:23.098616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:23.098616Z digest=sha256:058bb38a2fe60ca8839d75a9996a8cf55229310c3f7f51248eb3be34cc5d5327

Observation e00da223-e34f-4250-bffe-f10ee2f37467 · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Libritts: A corpus derived from librispeech for text-to-speech,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.234647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:23.174327Z digest=sha256:b53f184c09d178a95e9a7ddac16d2b1658cbe6465a59b8ecd5921bb02574a1d1

Observation 29c8e372-8a71-449e-84b6-078ce53fbe70 · outbound

This paper cites Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.096294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:23.233471Z digest=sha256:c2991c2340fa5da2da1bef3d7b14fbcbd4cf834c60f58b18852e9547177b4bd8

Observation 8969748c-db12-4245-bde3-787651920458 · outbound

This paper cites Any-to-any generation via composable diffusion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Any-to-any generation via composable diffusion,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:23.957767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:23.261921Z digest=sha256:feabc71a5f7dde055c29027f9af6bfede1803e140bd16dd3093c5d49d86be70a

Observation 4fa7bbd9-077c-481f-a3e6-0f43e14ac6e6 · outbound

This paper cites Any-to-many voice conversion with location-relative sequence-to-sequence modeling,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Any-to-many voice conversion with location-relative sequence-to-sequence modeling,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:23.295591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:23.295591Z digest=sha256:c5c41bb8b7839f4d3ea696c8e917a2226482b32f13500f45a42f762520e06df6

Observation 5a24e83c-6473-4627-860b-cd2028b88c1c · outbound

This paper cites QuickVC: Any-to-many Voice Conversion Using Inverse Short-time Fourier Transform for Faster Conversion.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion QuickVC: Any-to-many Voice Conversion Using Inverse Short-time Fourier Transform for Faster Conversion

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:57:23.502029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:23.386903Z digest=sha256:820ca4b2a20887038896aefe11c646692166ca4a168174d3c97c108169dc8495

Pith citing papers

Observation 316d538f-1f07-4436-96e3-d160f469a179 · inbound

X-VC: Zero-shot Streaming Voice Conversion in Codec Space cites this paper.

X-VC: Zero-shot Streaming Voice Conversion in Codec Space Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:20:29.828049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:20:01.885544Z digest=sha256:312c1da9404f9e026444d750a61be4d9c88a4936bff60d372967a1f399a9b151

Observation 67293642-b383-4b5c-8fda-7708339eed94 · inbound

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer cites this paper.

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:16:12.321170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T21:17:48.886421Z digest=sha256:92d1def2b3d6e644ae24dfaca71ffe30d4ba1b0865902ddf6cf71cb732197fae