Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:32:24.151140Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 7 inbound Pith citation observations for arXiv:2412.06602.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:32:24.151140Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:57.611094Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T12:24:39.704893Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d6878a65-541c-486d-a016-74a4b6ada6ab · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey • 2 points: It loosely follows the instruc- tions but misses key elements or timing in parts
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3a9db7a4-81d2-44aa-bbb3-6bad7270ceb7 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey • 2 points: Noticeably synthetic; some un- natural artifacts remain
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6571bb23-6e58-4ae1-851e-97d0c7b089ba · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey {transcript}
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bc464c2f-4265-4bd8-8681-969f8a033022 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey In Pro- ceedings of the 32nd ACM International Conference on Multimedia, pages 1255–1264
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ecc71c65-57fa-416a-ba47-2d1417464cdc · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59212e56-b630-4a0c-a627-112f10d8710f · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8d21e48-135a-49c3-974e-ee72316c30c3 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a48b56a-ea5f-44e5-b3c5-c91699a4fef8 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Advances in Neural Information Processing Systems, 36:53728– 53741
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bd28493a-eff6-428c-b292-10e6f4bb2c26 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey A Survey on Neural Speech Synthesis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b4e9803-615b-4b18-994a-37972b25e66e · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fecd5ed-31e5-42fb-a104-c5c451a205fd · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfd49210-b857-46ad-a415-62ea5a818b66 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ca8827d7-9fe1-4990-aa62-bd9ad8a7474a · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Speech vocoder is the last com- ponent that converts the intermediate acoustic fea- tures into a waveform that can be played back
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 909a04ca-16c9-4b60-94ea-d2a8d2d4ec1d · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
Reference 2001
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d82ed43d-ac82-46bb-84af-9682cc594194 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey A lower MCD value in- dicates a higher similarity between synthesized and reference speech, meaning better speech synthesis quality
Reference 2008
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 10ef6b29-1f0f-4fea-9e65-4069343e7774 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61ed7bf-e415-42a3-a621-40b685f5b6ef · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey speak calmly
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 33f32af4-c5f0-41bf-aa7b-08609e27b16f · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02937b2f-4361-4dd5-b3b2-0ff753baa0d2 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d23af91-f43b-42e5-8a7c-f30840e33a93 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0ae0045-a581-40fc-a133-ecb6925a0a34 · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a09387-1d28-43f3-a61f-51d1074a3f8d · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ae3806-5571-4a12-a529-2c9c05b018ec · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Hierarchical Control of Emotion Rendering in Speech Synthesis
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a279d26-5421-4784-a928-f7085b96354b · outbound
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e230f18e-81fa-4d4d-8444-6c259f26b889 · inbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49542191-d96c-4e3d-b266-6d82137c50fe · inbound
Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da47824a-1952-498a-937c-9ff0fa131f91 · inbound
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d5efeb-243f-41f0-8ad0-0e4b99600dbe · inbound
TokenChain: A Discrete Speech Chain via Semantic Token Modeling Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6e27a746-8fdd-4de7-b05f-cdd5558197e6 · inbound
Position: Towards Responsible Evaluation for Text-to-Speech Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c35a8c-70a6-4ef1-8039-8338f2332ce1 · inbound
CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 231ab4b1-d86d-4661-9316-3f6467049e24 · inbound
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.