Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:10:40.920937Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2608.08362.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:10:40.920937Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:10:40.788931Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-12T00:10:41.540631Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d3906908-4e24-4528-b219-2ec9d33f18fe · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ca53c06-e6f2-47ae-b9b7-64d9808f0602 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f3eb38eb-8e72-44bd-ae26-ebe1c70f1d3e · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 440044ac-ccbf-4749-bcf6-10ac8c78b8d1 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6a494d1b-0a14-4601-98f0-8bc0bcb2ac01 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 33dac800-e9b4-4dcf-be69-a7637841b420 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis outpainting
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a0d2cb77-2aa9-4380-bf83-f557d79ce680 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 097e6d95-3fde-4a8f-96ae-2401dbab8ce6 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 39e83fd4-f028-4ba0-a7c9-4a397bfa4e5b · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation caf97688-a2fb-4f82-b633-172179f6c8d7 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis First, our experiments are mainly conducted on English speech, so the effectiveness of CTRLSPEECHfor multilingual or code-switching synthesis re- mains unexplored
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 92c5ee93-08e0-4bad-8be0-32c56eeae2c5 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis 2D- 16003984 through the Amazon-UT Austin HUB
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ee42a50d-1072-413f-9fe8-6934b39bcb0b · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis All authors remain fully responsible for the con- tent of this manuscript
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a7d6d96c-9730-46e6-bc04-c727ef0716c1 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Ditar: Diffusion transformer autoregressive modeling for speech generation,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 027dec4c-984b-4e54-97f3-e3d7d95a6796 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ffd9c1-4804-4199-b409-a07fe33fd398 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc0ab8f-49f9-4847-84ec-fe1532ef48a4 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6409aaac-722e-4444-bd63-62c7de4cd51a · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c7967d7-8218-47f1-a511-30027b4a84f4 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis V oicecraft-x: Unifying multilingual, voice- cloning speech synthesis and speech editing,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c025b68e-de5c-4b9f-a609-f03798ad1bf6 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Prompttts: Control- lable text-to-speech with text descriptions,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a09c230-013b-4930-8825-116294c725d3 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis PromptTTS 2: Describing and Generating Voices with Text Prompt
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084a1113-77e1-4aaf-836c-35857e15c28e · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Instructtts: Modelling expressive tts in discrete latent space with natural lan- guage style prompt,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b6a76d86-87f8-48ce-ac8e-37bd5dc32208 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02056422-fa70-46ed-b4bb-963fc2b1fe60 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c700938b-0062-403e-a8ce-f40cb29a6b4c · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Emovoice: Llm-based emotional text-to-speech model with freestyle text prompting,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e951b5e-af90-44a8-bb98-d41fec64bdc4 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Scaling rich style- prompted text-to-speech datasets,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 79554df0-aee5-4816-9128-5e0f94e067c7 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe29e582-b9c6-49c5-be71-2ed5e983fda7 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Vevo2: A unified and controllable frame- work for speech and singing voice generation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77015045-1b72-499d-b93f-68edfef23e56 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Mela-tts: Joint transformer-diffusion model with representation alignment for speech synthesis,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 737ebbf5-97a5-4e61-babd-a06700e970b1 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Streammel: Real-time zero-shot text-to-speech via interleaved continuous autoregressive model- ing,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation da98a471-810b-4dd6-bc07-a3d615f58348 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis High-fidelity audio compression with improved rvqgan,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff38436-b674-4f67-bf61-c21289a03c4c · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beed7ec7-66d0-4a8d-bd5d-29038c9034eb · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0755047a-85de-4b8a-bc6b-0994e4d246ff · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ff20da-d32e-4471-a84a-e20044149be6 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis World: a vocoder-based high-quality speech synthesis system for real-time applications,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 189bfeb3-c9ad-4c0e-8c7e-ef64b20f164f · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Fast and reliable f0 estimation method based on the period extraction of vocal fold vi- bration of singing voice and speech,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b72ed2be-5ade-4a44-acc6-a08a36515fb1 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Continuous-token diffu- sion for speaker-referenced tts in multimodal llms,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c808d725-d472-4221-8b5e-224e44c2a1e5 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b16eb8-10f0-4ab0-b52d-b78e3a58f35a · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd65d55-f795-4692-8566-41c299c3f5cd · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e17983-8473-41d0-9f8d-acb9088a8536 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis The lj speech dataset,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f85e8df-6a19-498c-bf75-d4205c79f7bd · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Semantic-vae: Semantic- alignment latent representation for better speech synthesis,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da8fe183-b02b-44f7-a098-089abdda655d · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Classifier-Free Diffusion Guidance
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 306d53bf-9ffc-46cf-bb8f-106d97ccf6a7 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Robust speech recognition via large-scale weak supervision,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d0016ae-58fc-4a76-b7b2-9bf848779e45 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ed6a26b-8b9c-4c3b-901a-7ed2ba1a11a2 · outbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Drawspeech: Expressive speech synthesis using prosodic sketches as control conditions,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6a494d1b-0a14-4601-98f0-8bc0bcb2ac01 · inbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.