Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T11:35:58.050305Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2601.18094.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T11:35:58.050305Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:00:27.852836Z
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 19c87387-2e29-4079-8009-eae1873d6d4a · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Streaming voice con- version via intermediate bottleneck features and non- streaming teacher guidance
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 91d816af-a723-4002-afea-6fabc1640502 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Yingmusic-svc: Real- world robust zero-shot singing voice conversion with flow- grpo and singing-specific inductive biases.Arxiv
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f190ef8-61d3-433d-a157-c20f0831b08c · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Neural analysis and synthesis: Reconstructing speech from self-supervised representations
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8b30aaee-f30f-4754-8ea4-0e9033182326 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b498b656-2ec5-4a15-ad9b-f63a3b024ab6 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion The nus sung and spoken lyrics corpus: A quantitative comparison of singing and speech
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88d8c067-597a-447c-85a9-736238fd5692 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Moshi: a speech-text foundation model for real-time dialogue
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5bb3d4b0-ccc1-4000-b6cb-903dd435bc0f · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Switch transformers: scaling to trillion parame- ter models with simple and efficient sparsity.Journal of Machine Learning Research, 23
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3998f120-7296-481b-9081-a875bb7c3e99 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Zico Kolter, and Kaiming He
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f88b34bf-2489-41e1-aa3f-91ed6ca779bc · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Bigvgan: A universal neural vocoder with large-scale training.Arxiv
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 297ec85a-436d-4898-8281-49a3721bcb25 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Emilia: A large-scale, extensive, multilingual, and diverse dataset for speech generation.Transactions on Audio, Speech and Language Processing, 33:4044–4054
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c320591-e7cc-4d2c-b6c1-cc75f5a69cc1 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion HuBERT: Self- supervised speech representation learning by masked pre- diction of hidden units.Transactions on Audio, Speech, and Language Processing, 29:3451–3460
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfecea9d-92bc-453f-9db6-89d94ecbdf42 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion LoRA: Low-rank adaptation of large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29be1a8e-c65e-4da5-9f2c-aa131a215253 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70f90b48-c7f0-4bf2-b541-b2e68a123495 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion The singing voice conversion challenge
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87851fa9-cddf-444d-b4b4-bf994d401481 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion DiTAR: Diffusion transformer autoregressive modeling for speech generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 227dce47-46fc-4752-9625-517637dab163 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Ref-vc: Robust, expressive and fast zero-shot voice conversion with diffusion trans- formers.Arxiv
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 31ad7fda-1618-49d6-903c-3f66cefd5260 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion mod- els
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88103a3d-4d45-4173-840b-2f6a1c657f1b · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Efficient multilingual asr finetuning via lora language ex- perts
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48720ead-b5b8-4888-93f9-7e690c4a353e · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8db6212e-b3fc-4c1b-9f32-357467838f71 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Transferring source style in non-parallel voice conversion
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 481fda3a-022a-4a31-a607-7f7d42905c77 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Learning the beauty in songs: Neural singing voice beautifier
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45127fd4-855c-48d2-8891-cb7ac3ec736e · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Zero-shot voice conversion with diffusion transformers.Arxiv
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bd7e4782-d2d0-47c4-8dfd-529a7a296c17 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Hdmole: Mixture of lora experts with hi- erarchical routing and dynamic thresholds for fine-tuning llm-based asr models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5dc1f525-bcf4-4776-991a-63a43de8933a · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Scalable diffusion models with transformers.Arxiv
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2b7133db-176d-4531-8ea3-5d03cfea5c38 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Vibevoice technical report.Arxiv
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0e87392-298a-4c8b-b905-679510ad55c8 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e6c60e03-6bee-4aef-86bb-5d839e6f9ab2 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Singing voice data scaling-up: An intro- duction to ace-opencpop and ace-kising.Arxiv
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6d6c033f-3228-4f13-b64a-ff0bd6371b28 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Li, Hao Wang, Shiyin Kang, and H
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3d40483-91aa-4c07-b1a5-855490497f68 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Multimodal latent language modeling with next-token diffusion.Arxiv
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 867c0439-611b-47dd-ae25-1ad9952d47dd · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis.Arxiv
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 32192538-672b-41e9-b944-09775968a27b · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Metis: A foundation speech generation model with masked generative pre-training.Arxiv
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 74b166b9-f8bb-4336-9097-cb386f322334 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Moe-tts: Enhancing out-of-domain text understanding for description-based tts via mixture-of-experts.Arxiv
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d61cd101-c106-486f-b5f5-cbcdb990159a · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Uniaudio: An audio foundation model to- ward universal audio generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48d0ee07-9e78-4e2c-88c5-b5cb466be1ab · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Llasa: Scal- ing train-time and inference-time compute for llama-based speech synthesis.Arxiv
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 802ef9e6-633e-41de-9018-be6165b51762 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Megabyte: Predicting million-byte sequences with multi- scale transformers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7fc96d1b-709d-462f-9f00-d01719115e12 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Takin-VC: Expressive zero-shot voice conversion via adaptive hybrid content encoding and en- hanced timbre modeling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c6efad3-e0fc-4932-8e3c-7e772ddf3c82 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion SoundStream: An end-to-end neural audio codec.Trans- actions on Audio, Speech, and Language Processing, 30:495–507
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f3a50f3-1873-4b26-8585-e27c36a65171 · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ece65de7-8748-4a7a-8970-0fc1230d80fc · outbound
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Transfusion: Predict the next token and diffuse images with one multi-modal model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a47df495-4001-4464-829e-102a64b06aad · inbound
MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.