Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:42:51.295730Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2608.11590.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:42:51.295730Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:42:51.095849Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-16T00:42:51.460473Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 98a2c2dc-f6b3-4b3c-8074-158163818324 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b79a8942-3775-423d-b626-9630037b115d · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation By decomposing human voice into style, content, and prosody, CookV oice successfully unifying multiple voice generation task within a unified model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a86d69a5-5434-405c-898d-0a53177e79d3 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation be907e9c-246b-4088-b820-9c97f3426f48 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 334ab11e-101c-4624-9871-3381d6496da4 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Problem Formulation Human V oice Generation (HVG) is to generate waveform from different modality of control signal
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 69ae5454-0479-4350-b803-a2ce807068bc · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation This section will present the architecture of CookV oice in detail
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 74e7d7a9-1b77-43f9-b4be-a010e1e8d6c6 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 09750361-4f73-4943-9c79-bbb85e092703 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a8d7a2b9-4130-4d71-8e6e-d6c6991cb19f · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6e6461fa-f1dd-4461-80f2-26fc76b36fa5 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5f72e243-2ba4-4986-88e2-d1ebb6485937 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation First, due to re- source constraints, CookV oice has not yet been scaled up
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a37fe608-667b-48db-bff1-6fba55f0da50 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe8eb4b9-9524-4521-94c9-caab31cddeb3 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c686c9-89e8-42e3-85f1-9ec539400888 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Parastyletts: Toward effi- cient and robust paralinguistic style control for expressive text-to- speech generation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cf94d5fd-9273-42d3-bb02-07f2d22d7f1c · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe291923-ad50-4d6d-ad76-f99a0b6b1633 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Diffsinger: Singing voice synthesis via shallow diffusion mechanism,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173cc171-d93c-4ca4-baaf-67182a6f41ac · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Stylesinger: Style transfer for out-of- domain singing voice synthesis,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009b6443-3e49-41c7-b426-14c095605731 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Tcsinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 974fac2a-68cb-4534-85db-f9e026e5b2e9 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Vevo2: A unified and controllable framework for speech and singing voice generation,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 44a02b03-9794-4069-86ee-0e291628071a · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45d1b6a3-090a-4493-9669-e6bdd0ee010c · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Param- eta: Towards learning disentangled paralinguistic speak- ing styles representations from speech,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5d62aaf7-8d75-45a5-8d63-47ff93c4e04a · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Scalable diffusion models with transform- ers,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c36f22b-0f7b-4694-a4fc-aaa26f3db153 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Flow Matching for Generative Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ccf258c-424b-447b-9257-f5ed5b413d3c · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Fastspeech: Fast, robust and controllable text to speech,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12131af5-38dc-4fe5-a1d8-9bd92cdd06e6 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation M4singer: A multi-style, multi- singer and musical score provided mandarin singing corpus,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad2d683-fcdf-47c6-9e7e-6f8e61f1bac7 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Vevo: Controllable zero-shot voice imitation with self- supervised disentanglement,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d17a0e1f-4fa3-4448-8e36-28f637dd63a7 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd90a01-1665-4d68-b65a-076bb6eac9df · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Mpnet: Masked and permuted pre-training for language understanding,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d83b9c17-b824-4ef4-926e-2b3fe6922816 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Attention is all you need,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c328557-4944-40be-9b3c-1d8dd6b3f201 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0ca79004-7c70-42aa-a609-fc8b51d69561 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Tacotron: Towards End-to-End Speech Synthesis
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217af388-708a-4e4a-b532-f2b313c347d5 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Chinese mandarin female corpus,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b8010839-144b-4a12-9e34-729c08864903 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation The lj speech dataset,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276cbf09-14fd-4586-b6b0-33106e777b4a · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a38fdb-4e3e-4635-8f11-f0333b8acf07 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Crema-d: Crowd-sourced emotional multimodal actors dataset,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 020e2760-ffe3-464a-b42f-3870e233481b · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Common phone: A multilingual dataset for robust acoustic modelling,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ebf432da-1285-4d7a-be81-5392174b4368 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Genshin voice: A multi-lingual voice dataset from Genshin Impact,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ed5c18f1-fcf9-42c2-b65d-1ae6fa2b0bcc · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 930af073-8f3b-4301-90d2-e8ca61305640 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Robust speech recognition via large-scale weak supervision,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b12c8d-717f-4cf9-adfd-358121caedba · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b2f59396-0571-4e70-9455-1eb1fe06c189 · outbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Denoising diffusion probabilis- tic models,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 98a2c2dc-f6b3-4b3c-8074-158163818324 · inbound
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.