Pith. sign in

Paper Citation Record · LEDGER

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion

As of 5 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2601.18094.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.18094 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T11:35:58.050305Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:00:27.852836Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19c87387-2e29-4079-8009-eae1873d6d4a · outbound

This paper cites Streaming voice con- version via intermediate bottleneck features and non- streaming teacher guidance.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Streaming voice con- version via intermediate bottleneck features and non- streaming teacher guidance

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.271534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:63de71cdf83f6b64fb65d7f27e178aaf51a9b109c444ecc6979812fe4deb3f80

Observation 91d816af-a723-4002-afea-6fabc1640502 · outbound

This paper cites Yingmusic-svc: Real- world robust zero-shot singing voice conversion with flow- grpo and singing-specific inductive biases.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Yingmusic-svc: Real- world robust zero-shot singing voice conversion with flow- grpo and singing-specific inductive biases.Arxiv

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.178893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:3c467f7f3ef7ccc9c999a7d51d9fc903b3c1bb55245b10984df01e3bf3240dea

Observation 9f190ef8-61d3-433d-a157-c20f0831b08c · outbound

This paper cites Neural analysis and synthesis: Reconstructing speech from self-supervised representations.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Neural analysis and synthesis: Reconstructing speech from self-supervised representations

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.290104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:4a36a0bb4a9a8749f425d395b4219bf40e15c267c8beff1eafac9e46a10a43d6

Observation 8b30aaee-f30f-4754-8ea4-0e9033182326 · outbound

This paper cites an unresolved cited work.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-22T11:36:29.217619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:faba81eee4bb16b22a2a9a75b5826ed0fd0bcf55d77c7983ea0751514f0adbf9

Observation b498b656-2ec5-4a15-ad9b-f63a3b024ab6 · outbound

This paper cites The nus sung and spoken lyrics corpus: A quantitative comparison of singing and speech.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion The nus sung and spoken lyrics corpus: A quantitative comparison of singing and speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.202283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:ec9efc9614e5746c785c57aade5f8d23e3b1ad369d5f628cb782b00d10f7d538

Observation 88d8c067-597a-447c-85a9-736238fd5692 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Moshi: a speech-text foundation model for real-time dialogue

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.278134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:31652335f9e981404012331e5b46f3b5ceb9275ac1736fbb21a05e87b41bea39

Observation 5bb3d4b0-ccc1-4000-b6cb-903dd435bc0f · outbound

This paper cites Switch transformers: scaling to trillion parame- ter models with simple and efficient sparsity.Journal of Machine Learning Research, 23.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Switch transformers: scaling to trillion parame- ter models with simple and efficient sparsity.Journal of Machine Learning Research, 23

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.298051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:aa094bb80f932d695264c9354f0253e18b23a5d0ffa6b18ac4714bdf26843b6e

Observation 3998f120-7296-481b-9081-a875bb7c3e99 · outbound

This paper cites Zico Kolter, and Kaiming He.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Zico Kolter, and Kaiming He

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.195657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:f846210afad3b356ef1296149a4ac8c8a8b1d997971482cc06abe5e56ba97233

Observation f88b34bf-2489-41e1-aa3f-91ed6ca779bc · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Bigvgan: A universal neural vocoder with large-scale training.Arxiv

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.188972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:229a927b5ad198af5ad0d30819d4ffa67039347967d9ad28ad1f038332edda14

Observation 297ec85a-436d-4898-8281-49a3721bcb25 · outbound

This paper cites Emilia: A large-scale, extensive, multilingual, and diverse dataset for speech generation.Transactions on Audio, Speech and Language Processing, 33:4044–4054.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Emilia: A large-scale, extensive, multilingual, and diverse dataset for speech generation.Transactions on Audio, Speech and Language Processing, 33:4044–4054

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.185821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:9dce6a2b32c46ce191dc03676ba87d4db0daa397a854a36a4a734c812f66f302

Observation 1c320591-e7cc-4d2c-b6c1-cc75f5a69cc1 · outbound

This paper cites HuBERT: Self- supervised speech representation learning by masked pre- diction of hidden units.Transactions on Audio, Speech, and Language Processing, 29:3451–3460.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion HuBERT: Self- supervised speech representation learning by masked pre- diction of hidden units.Transactions on Audio, Speech, and Language Processing, 29:3451–3460

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.205561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:3c0be2cd5ba905af4fafa28ca89362133f50c24ae3278333f716ea4d030b3668

Observation cfecea9d-92bc-453f-9db6-89d94ecbdf42 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion LoRA: Low-rank adaptation of large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.192479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:b1947e96aa4361d2d83f23deec90f2786e7f3b50bd48bb751fb2311b8d07c567

Observation 29be1a8e-c65e-4da5-9f2c-aa131a215253 · outbound

This paper cites Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.232832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:63bf9b3b4083bd778f0cb15832c141cc0a939ffb10b0a95ebf07d05774b8e9c5

Observation 70f90b48-c7f0-4bf2-b541-b2e68a123495 · outbound

This paper cites The singing voice conversion challenge.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion The singing voice conversion challenge

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.268487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:9dcf9ff06a90b938a77f25d5f25fce21d495245e849eedcdd74e09c235d2358f

Observation 87851fa9-cddf-444d-b4b4-bf994d401481 · outbound

This paper cites DiTAR: Diffusion transformer autoregressive modeling for speech generation.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion DiTAR: Diffusion transformer autoregressive modeling for speech generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.199239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:7d21afdaa425e9d9e89f4dcb50cc0a16a95ff9aa668bfe83ae779b7641a870a0

Observation 227dce47-46fc-4752-9625-517637dab163 · outbound

This paper cites Ref-vc: Robust, expressive and fast zero-shot voice conversion with diffusion trans- formers.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Ref-vc: Robust, expressive and fast zero-shot voice conversion with diffusion trans- formers.Arxiv

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.214582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:301f7b1b9235d9198fa8883df1e6a0ea6a79466ebeab259aa38e6eb2c93e346e

Observation 31ad7fda-1618-49d6-903c-3f66cefd5260 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion mod- els.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion mod- els

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.262253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:39d91cabb4ab99d36ea4ed7ed7c93badadc9dff62780f6ac6f7b77f5e42947d1

Observation 88103a3d-4d45-4173-840b-2f6a1c657f1b · outbound

This paper cites Efficient multilingual asr finetuning via lora language ex- perts.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Efficient multilingual asr finetuning via lora language ex- perts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.249541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:771bc62eaa9dc6345c02bf35f638f27be44cdd2fd4dfb249f043c12b81426a90

Observation 48720ead-b5b8-4888-93f9-7e690c4a353e · outbound

This paper cites an unresolved cited work.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-22T11:36:29.211144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:82683269c96953ed0682f2e230e06e2135370ac5af77bd576c3b896e0662c49e

Observation 8db6212e-b3fc-4c1b-9f32-357467838f71 · outbound

This paper cites Transferring source style in non-parallel voice conversion.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Transferring source style in non-parallel voice conversion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.208336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:6bc7b2515a5d352dd41cce5e047febe94dfaecc3a05da2ae0b3a1191edbffb6d

Observation 481fda3a-022a-4a31-a607-7f7d42905c77 · outbound

This paper cites Learning the beauty in songs: Neural singing voice beautifier.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Learning the beauty in songs: Neural singing voice beautifier

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.182168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:96f202dc24fc09f29a0f8e22ea06820a42dce0201b241688f6da8937bc0c0750

Observation 45127fd4-855c-48d2-8891-cb7ac3ec736e · outbound

This paper cites Zero-shot voice conversion with diffusion transformers.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Zero-shot voice conversion with diffusion transformers.Arxiv

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.258872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:e9c0b8fc3d995641a862cbf616115155d0bf9c1f65b6698f60693ffa1f0923ba

Observation bd7e4782-d2d0-47c4-8dfd-529a7a296c17 · outbound

This paper cites Hdmole: Mixture of lora experts with hi- erarchical routing and dynamic thresholds for fine-tuning llm-based asr models.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Hdmole: Mixture of lora experts with hi- erarchical routing and dynamic thresholds for fine-tuning llm-based asr models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.255874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:d66f835ba2672b017f1fa9ebd5dfb4d3f842a5889636d116943e98b6121acb46

Observation 5dc1f525-bcf4-4776-991a-63a43de8933a · outbound

This paper cites Scalable diffusion models with transformers.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Scalable diffusion models with transformers.Arxiv

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.242735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:38421b76290d41c32cb7142572afb2a21708bd8a1aeece9e1885c4c6c93f2a5c

Observation 2b7133db-176d-4531-8ea3-5d03cfea5c38 · outbound

This paper cites Vibevoice technical report.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Vibevoice technical report.Arxiv

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.229863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:f22f48056292b8bb9dfa023dda8aa2a5f2f99af83a807fa657c401058841116b

Observation b0e87392-298a-4c8b-b905-679510ad55c8 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.265309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:9d53e4ec9466fa59750b4f4236b29ceb705af219db8ceece1b8ac3767a2d7d62

Observation e6c60e03-6bee-4aef-86bb-5d839e6f9ab2 · outbound

This paper cites Singing voice data scaling-up: An intro- duction to ace-opencpop and ace-kising.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Singing voice data scaling-up: An intro- duction to ace-opencpop and ace-kising.Arxiv

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.294575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:24dd874e2a23c28aa680a5344bb9164634a4d83543d7a66b358de0df382862d2

Observation 6d6c033f-3228-4f13-b64a-ff0bd6371b28 · outbound

This paper cites Li, Hao Wang, Shiyin Kang, and H.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Li, Hao Wang, Shiyin Kang, and H

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.235876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:394f8d725359b33f854b30f0b4ec6934ae6583d35654988d31c453e674edc2ea

Observation d3d40483-91aa-4c07-b1a5-855490497f68 · outbound

This paper cites Multimodal latent language modeling with next-token diffusion.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Multimodal latent language modeling with next-token diffusion.Arxiv

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.287788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:fc5acece840a36864aba7fbab7b3043c4e8f0e32d00b98ad30f9f65e9e904b5a

Observation 867c0439-611b-47dd-ae25-1ad9952d47dd · outbound

This paper cites Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis.Arxiv

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.239273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:469adb11db5d52e992d155708ab5950d8c04745ac5ba9e1999dbdaa1b1f9bc1c

Observation 32192538-672b-41e9-b944-09775968a27b · outbound

This paper cites Metis: A foundation speech generation model with masked generative pre-training.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Metis: A foundation speech generation model with masked generative pre-training.Arxiv

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.246299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:912af49839d2b2f9808b61cc1808d45e11600b6a7a2016e01603a5a696797a98

Observation 74b166b9-f8bb-4336-9097-cb386f322334 · outbound

This paper cites Moe-tts: Enhancing out-of-domain text understanding for description-based tts via mixture-of-experts.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Moe-tts: Enhancing out-of-domain text understanding for description-based tts via mixture-of-experts.Arxiv

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.252744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:aebcd81fdcb08d361d567baf21fe56e9bf8b28962edba6c902463d5af4d74f06

Observation d61cd101-c106-486f-b5f5-cbcdb990159a · outbound

This paper cites Uniaudio: An audio foundation model to- ward universal audio generation.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Uniaudio: An audio foundation model to- ward universal audio generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.284992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:ef8111a75370f172e68ed6f4e1b9c68c260594e4735beabd8a68a5ce88be38c8

Observation 48d0ee07-9e78-4e2c-88c5-b5cb466be1ab · outbound

This paper cites Llasa: Scal- ing train-time and inference-time compute for llama-based speech synthesis.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Llasa: Scal- ing train-time and inference-time compute for llama-based speech synthesis.Arxiv

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.281589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:2b1e1a54937620070a2ff6a847683fb2f90d26362e0a992240d5215532d74fef

Observation 802ef9e6-633e-41de-9018-be6165b51762 · outbound

This paper cites Megabyte: Predicting million-byte sequences with multi- scale transformers.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Megabyte: Predicting million-byte sequences with multi- scale transformers

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.223750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:5d2e806286d825009bd58aed43274bae586a053bbbd14a655960a131f5c85054

Observation 7fc96d1b-709d-462f-9f00-d01719115e12 · outbound

This paper cites Takin-VC: Expressive zero-shot voice conversion via adaptive hybrid content encoding and en- hanced timbre modeling.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Takin-VC: Expressive zero-shot voice conversion via adaptive hybrid content encoding and en- hanced timbre modeling

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.292338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:a2ba5950a918d30d0fa67ba17a25a0f0e181ddace75d492facc25c2b5b670229

Observation 6c6efad3-e0fc-4932-8e3c-7e772ddf3c82 · outbound

This paper cites SoundStream: An end-to-end neural audio codec.Trans- actions on Audio, Speech, and Language Processing, 30:495–507.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion SoundStream: An end-to-end neural audio codec.Trans- actions on Audio, Speech, and Language Processing, 30:495–507

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.226774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:d80e37afd66e5a61fd9c6a28d448d17b67dc35df87b0748cb078d4716963f6a5

Observation 3f3a50f3-1873-4b26-8585-e27c36a65171 · outbound

This paper cites M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.220666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:9e0cd105cea7b553331f53ed3ed10512bfccac2b1ec159199c7b5164861adefc

Observation ece65de7-8748-4a7a-8970-0fc1230d80fc · outbound

This paper cites Transfusion: Predict the next token and diffuse images with one multi-modal model.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Transfusion: Predict the next token and diffuse images with one multi-modal model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.274768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:21b3208e2da246362f8624829ae6f68d4aaf7799b4ecbd9d1ac75f79bf11db32

Pith citing papers

Observation a47df495-4001-4464-829e-102a64b06aad · inbound

MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection cites this paper.

MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T18:00:27.852836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:00:27.852836Z digest=sha256:759d037ae5babd2fdcdd5009b3adb2d611f3a32e11e03df29382089e1dd46d52