Pith. sign in

Paper Citation Record · LEDGER

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

As of 13 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 7 inbound Pith citation observations for arXiv:2412.06602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06602 v3

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:32:24.151140Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:57.611094Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T12:24:39.704893Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6878a65-541c-486d-a016-74a4b6ada6ab · outbound

This paper cites • 2 points: It loosely follows the instruc- tions but misses key elements or timing in parts.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey • 2 points: It loosely follows the instruc- tions but misses key elements or timing in parts

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.452307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.141778Z digest=sha256:b4fd8eb6e7b82d83154587c03eff6fa3dba65f98bd2dc9a2dda4145cf6aa3e0c

Observation 3a9db7a4-81d2-44aa-bbb3-6bad7270ceb7 · outbound

This paper cites • 2 points: Noticeably synthetic; some un- natural artifacts remain.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey • 2 points: Noticeably synthetic; some un- natural artifacts remain

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.433638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.146371Z digest=sha256:20604ccc7c197dac2968337a26aca43a6ddd6b614d40f8e6eb11b120e63116b6

Observation 6571bb23-6e58-4ae1-851e-97d0c7b089ba · outbound

This paper cites {transcript}.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey {transcript}

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T19:32:24.415747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.151140Z digest=sha256:a4917b28647a9f4d2d33cf42004e59ebceed6ebfbff33fb840259db927352d1e

Observation bc464c2f-4265-4bd8-8681-969f8a033022 · outbound

This paper cites In Pro- ceedings of the 32nd ACM International Conference on Multimedia, pages 1255–1264.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey In Pro- ceedings of the 32nd ACM International Conference on Multimedia, pages 1255–1264

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.564296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.061207Z digest=sha256:1cf999736e055eebc4a537d8858203210f2e81f647ee9bd3b0302004ef0df5a4

Observation ecc71c65-57fa-416a-ba47-2d1417464cdc · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.075261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.075261Z digest=sha256:12561ac73c583de6701feb9b4561f9e8fe6caab5632921ca01b3bab1df4fab31

Observation 59212e56-b630-4a0c-a627-112f10d8710f · outbound

This paper cites an unresolved cited work.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:32:24.548472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.080857Z digest=sha256:37af856526069c2bbf0e01002197b26723cdf9021f062e6982a93ffdf6091de2

Observation e8d21e48-135a-49c3-974e-ee72316c30c3 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.085937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.085937Z digest=sha256:535b1a0ed1b7cecb94875b0b26db8b7609c5244bc21cd945a2b5ccca7f406bd6

Observation 1a48b56a-ea5f-44e5-b3c5-c91699a4fef8 · outbound

This paper cites Advances in Neural Information Processing Systems, 36:53728– 53741.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Advances in Neural Information Processing Systems, 36:53728– 53741

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.533789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.091663Z digest=sha256:a27373d6557d7bb02bdd3811e066226db9beb8f675da77a1957964827ca742f3

Observation bd28493a-eff6-428c-b292-10e6f4bb2c26 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey A Survey on Neural Speech Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.106056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.106056Z digest=sha256:2acdfcfb1b030a9a3b690e7497b462002a768ee612fdf3285f6314486f29c725

Observation 4b4e9803-615b-4b18-994a-37972b25e66e · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.110824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.110824Z digest=sha256:8a14c7684c933a99b578d55c5a7a9e6cabe57ab728dbf3da1337e7ddb6a27983

Observation 0fecd5ed-31e5-42fb-a104-c5c451a205fd · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.120019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.120019Z digest=sha256:39a49be4a97499833d4fdfef51925eaaf64ae6ae98fc302e6333740e26804b22

Observation bfd49210-b857-46ad-a415-62ea5a818b66 · outbound

This paper cites an unresolved cited work.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:32:24.502947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.128863Z digest=sha256:d6d6fdeedac14309696a64bd06c7c89694d3f38152d12f6b5dd5b92d892dea7c

Observation ca8827d7-9fe1-4990-aa62-bd9ad8a7474a · outbound

This paper cites Speech vocoder is the last com- ponent that converts the intermediate acoustic fea- tures into a waveform that can be played back.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Speech vocoder is the last com- ponent that converts the intermediate acoustic fea- tures into a waveform that can be played back

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.485242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.133064Z digest=sha256:b28a8ea593ddf8a0d27cf5c941c612903c2950c7dd691fe2ea9a153bdfd0cdef

Observation 909a04ca-16c9-4b60-94ea-d2a8d2d4ec1d · outbound

This paper cites Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.096350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.096350Z digest=sha256:fc5f2fce7002c944fcf23605e7204b07dd237216b13e61febf2c7d4a605cdc5f

Observation d82ed43d-ac82-46bb-84af-9682cc594194 · outbound

This paper cites A lower MCD value in- dicates a higher similarity between synthesized and reference speech, meaning better speech synthesis quality.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey A lower MCD value in- dicates a higher similarity between synthesized and reference speech, meaning better speech synthesis quality

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.469901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.136896Z digest=sha256:21f841c0305ddf78b2fa4d593829c55143511dea264b50fb7cd0331f3f394c43

Observation 10ef6b29-1f0f-4fea-9e65-4069343e7774 · outbound

This paper cites Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.042006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.042006Z digest=sha256:5dd89f9237bf2e38790358d70ce46738f637547b2ad178715c1d0fe85e9ee15e

Observation a61ed7bf-e415-42a3-a621-40b685f5b6ef · outbound

This paper cites speak calmly.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey speak calmly

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.518923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:32:24.124409Z digest=sha256:9af04b1fe3035970fcd7ea54ed286d017bbda16ed1a4bd39f7553ed6c27c8196

Observation 33f32af4-c5f0-41bf-aa7b-08609e27b16f · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.035818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.035818Z digest=sha256:48c4ef1d24149ae683b6eac5027ad1233d4afc69abae139d71a58bf328cc0ba1

Observation 02937b2f-4361-4dd5-b3b2-0ff753baa0d2 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.048348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.048348Z digest=sha256:897a923f240d342813745880d5e89a7f1b249d14e7ecd334c6025d5794848384

Observation 8d23af91-f43b-42e5-8a7c-f30840e33a93 · outbound

This paper cites SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.115348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.115348Z digest=sha256:a3192b0b3003e25951d4395260c200ff3105c4831434e85aa32a2ab9fa44ed4b

Observation b0ae0045-a581-40fc-a133-ecb6925a0a34 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.101443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.101443Z digest=sha256:8f3053fbe1952cd608069bbeff9340b4bc0414b156c1e8731f9777a84d6ca78a

Observation b9a09387-1d28-43f3-a61f-51d1074a3f8d · outbound

This paper cites SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.067231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.067231Z digest=sha256:e2bb26f55025a41353c7f4162549d24f4dec5abfc4a0ea7dddd8852ac99fe438

Observation 99ae3806-5571-4a12-a529-2c9c05b018ec · outbound

This paper cites Hierarchical Control of Emotion Rendering in Speech Synthesis.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Hierarchical Control of Emotion Rendering in Speech Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.055265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.055265Z digest=sha256:4bb32d0c930c7d19534f6e3e96bf27cb6e086355acfb6a38468d1afb5eba5559

Observation 8a279d26-5421-4784-a928-f7085b96354b · outbound

This paper cites SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.029868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.029868Z digest=sha256:a554d130dc9a17cc4a7bb61935a0926b2e00c27be9573cb482232c38a3419746

Pith citing papers

Observation e230f18e-81fa-4d4d-8444-6c259f26b889 · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:57.611094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:57.611094Z digest=sha256:d64e3c3d0d0fb05f0201ed5c2b8260d2827bec4c5f8d65563f1632fd19f79eda

Observation 49542191-d96c-4e3d-b266-6d82137c50fe · inbound

Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis cites this paper.

Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:25:27.548842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:25:27.548842Z digest=sha256:d52ade97a67c583b35aab35f8b50a5fe5b9f7c6f31b5443729e864da516f2561

Observation da47824a-1952-498a-937c-9ff0fa131f91 · inbound

Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models cites this paper.

Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:25:29.118241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:25:29.118241Z digest=sha256:4d5d1e2e81fc6a2d6db3df7b3bb5268163a3f3752504e28319529c51d7de08d1

Observation 80d5efeb-243f-41f0-8ad0-0e4b99600dbe · inbound

TokenChain: A Discrete Speech Chain via Semantic Token Modeling cites this paper.

TokenChain: A Discrete Speech Chain via Semantic Token Modeling Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:16:09.518750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T09:14:58.542628Z digest=sha256:8b94138cfa0aed624fb524764137e6da867cfce34625dfc07f83be4d588da038

Observation 6e27a746-8fdd-4de7-b05f-cdd5558197e6 · inbound

Position: Towards Responsible Evaluation for Text-to-Speech cites this paper.

Position: Towards Responsible Evaluation for Text-to-Speech Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T11:07:41.013784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:07:41.013784Z digest=sha256:8ae4c16dd546c45ad5276be4d8d41979547be43088c6628ec03aedf322baed61

Observation f2c35a8c-70a6-4ef1-8039-8338f2332ce1 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.009111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:7a1f2c624f45c5e7f84ff3fdda3be91d54596bfa034985b2b0596041bf6eb424

Observation 231ab4b1-d86d-4661-9316-3f6467049e24 · inbound

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations cites this paper.

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 5

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T12:24:39.706455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T12:20:43.383513Z digest=sha256:a9830ef412849320c9deb60234520984077c2eb6d0dcb90c0ee8c07b302fd063