Pith. sign in

Paper Citation Record · LEDGER

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

As of 21 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 7 inbound Pith citation observations for arXiv:2412.06602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06602 v3

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:32:24.151140Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:57.611094Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T12:24:39.704893Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6878a65-541c-486d-a016-74a4b6ada6ab · outbound

This paper cites • 2 points: It loosely follows the instruc- tions but misses key elements or timing in parts.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey • 2 points: It loosely follows the instruc- tions but misses key elements or timing in parts

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.452307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.141778Z digest=sha256:f49870ee983c293c43ff508274b6164d82db10f9d9d7111b19a299df9c8aa850

Observation 3a9db7a4-81d2-44aa-bbb3-6bad7270ceb7 · outbound

This paper cites • 2 points: Noticeably synthetic; some un- natural artifacts remain.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey • 2 points: Noticeably synthetic; some un- natural artifacts remain

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.433638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.146371Z digest=sha256:203c3f7001405ca436452134a9f197f8d3e266e3f929f3e0d7fb9212aa45053d

Observation 6571bb23-6e58-4ae1-851e-97d0c7b089ba · outbound

This paper cites {transcript}.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey {transcript}

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T19:32:24.415747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.151140Z digest=sha256:10e55179a88e902fd86d660467c58dc0652cf53f601f6c7f005443749e8e29f2

Observation bc464c2f-4265-4bd8-8681-969f8a033022 · outbound

This paper cites In Pro- ceedings of the 32nd ACM International Conference on Multimedia, pages 1255–1264.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey In Pro- ceedings of the 32nd ACM International Conference on Multimedia, pages 1255–1264

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.564296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.061207Z digest=sha256:22bbf9471155a52ac82c1d850af0b1a59e3c65e290b50bc84e53916cf43aadf9

Observation ecc71c65-57fa-416a-ba47-2d1417464cdc · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.075261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.075261Z digest=sha256:d70018454df0f49f5a11137968f46fc5f5d27bdbe6c37b6089d864ff489646a8

Observation 59212e56-b630-4a0c-a627-112f10d8710f · outbound

This paper cites an unresolved cited work.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:32:24.548472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.080857Z digest=sha256:ab89f35b6432df74f50106ccb33ab07c31c2e644cbf20d71a6f11f69baa53db4

Observation e8d21e48-135a-49c3-974e-ee72316c30c3 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.085937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.085937Z digest=sha256:d52f68fa24842fc1b24c801582de3f9fa43006093bf38ace54a6a32222c3929c

Observation 1a48b56a-ea5f-44e5-b3c5-c91699a4fef8 · outbound

This paper cites Advances in Neural Information Processing Systems, 36:53728– 53741.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Advances in Neural Information Processing Systems, 36:53728– 53741

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.533789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.091663Z digest=sha256:85ef21db453649dbc932378dda7bc636f4be2fc7f39fb2aea9c7ed196002b9d5

Observation bd28493a-eff6-428c-b292-10e6f4bb2c26 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey A Survey on Neural Speech Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.106056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.106056Z digest=sha256:6fea6086f08b85eafce8407b368e2004a51ce3387c561f503bd2b080cf089d1d

Observation 4b4e9803-615b-4b18-994a-37972b25e66e · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.110824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.110824Z digest=sha256:5424a7cebcdaa821cf76ba619c99090706c84dd8eef96f42e5e8185f84f8a47d

Observation 0fecd5ed-31e5-42fb-a104-c5c451a205fd · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.120019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.120019Z digest=sha256:d4fa61fca47954f8ad4d370d610ef6820410eb4825ce5498d377f2294d1a766d

Observation bfd49210-b857-46ad-a415-62ea5a818b66 · outbound

This paper cites an unresolved cited work.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:32:24.502947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.128863Z digest=sha256:e8235a881647c7821b75b7a74649f1948b44ece89930ca2168f9f3dd5560e380

Observation ca8827d7-9fe1-4990-aa62-bd9ad8a7474a · outbound

This paper cites Speech vocoder is the last com- ponent that converts the intermediate acoustic fea- tures into a waveform that can be played back.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Speech vocoder is the last com- ponent that converts the intermediate acoustic fea- tures into a waveform that can be played back

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.485242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.133064Z digest=sha256:56698574ff50b4656842e5cbf65bcf53e1b695c878fe8233994f2cc779858f54

Observation 909a04ca-16c9-4b60-94ea-d2a8d2d4ec1d · outbound

This paper cites Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.096350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.096350Z digest=sha256:2409cd853a55dc3f5f4bdfdd9e30780ba7e6bb8ac7b81221dddc9b3c90e7f1ca

Observation d82ed43d-ac82-46bb-84af-9682cc594194 · outbound

This paper cites A lower MCD value in- dicates a higher similarity between synthesized and reference speech, meaning better speech synthesis quality.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey A lower MCD value in- dicates a higher similarity between synthesized and reference speech, meaning better speech synthesis quality

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.469901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.136896Z digest=sha256:fb03364b987a3314e9a92331aa39e5e1b06954e71caf1cc84bd4b5f7be3f18fe

Observation 10ef6b29-1f0f-4fea-9e65-4069343e7774 · outbound

This paper cites Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.042006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.042006Z digest=sha256:2b39ac5cfba8988b6991ae870c48afa80b205736cb3b4154b23e40306cbbe649

Observation a61ed7bf-e415-42a3-a621-40b685f5b6ef · outbound

This paper cites speak calmly.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey speak calmly

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:32:24.518923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:32:24.124409Z digest=sha256:9b66d10da1ca64529a91922dc270b392a0bc57ee9a80f8f6bbcc4c47541bc414

Observation 33f32af4-c5f0-41bf-aa7b-08609e27b16f · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.035818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.035818Z digest=sha256:48c4ef1d24149ae683b6eac5027ad1233d4afc69abae139d71a58bf328cc0ba1

Observation 02937b2f-4361-4dd5-b3b2-0ff753baa0d2 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.048348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.048348Z digest=sha256:897a923f240d342813745880d5e89a7f1b249d14e7ecd334c6025d5794848384

Observation 8d23af91-f43b-42e5-8a7c-f30840e33a93 · outbound

This paper cites SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.115348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.115348Z digest=sha256:2740e3376a1bbb6a37f76a49cca7bc94edc02de62458fd1016ae9f035091d6c3

Observation b0ae0045-a581-40fc-a133-ecb6925a0a34 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.101443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.101443Z digest=sha256:d16924863cd981e49659039f93873cfaaf3640934ea92711f75f8b7288c77c9a

Observation b9a09387-1d28-43f3-a61f-51d1074a3f8d · outbound

This paper cites SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.067231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.067231Z digest=sha256:f25a6f0557b9bf2c8e8c92d93115bd19e466ad76d35c3c8a44b6f3c30d609945

Observation 99ae3806-5571-4a12-a529-2c9c05b018ec · outbound

This paper cites Hierarchical Control of Emotion Rendering in Speech Synthesis.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Hierarchical Control of Emotion Rendering in Speech Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.055265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.055265Z digest=sha256:b348eac28b41f2d1cdf2b3abbf8e9cb35ad8eb11087c730f44baefc54668300f

Observation 8a279d26-5421-4784-a928-f7085b96354b · outbound

This paper cites SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.029868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.029868Z digest=sha256:a554d130dc9a17cc4a7bb61935a0926b2e00c27be9573cb482232c38a3419746

Pith citing papers

Observation e230f18e-81fa-4d4d-8444-6c259f26b889 · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:57.611094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:57.611094Z digest=sha256:896f5d468b3f034035689cacbcddc4a3dcfda509c90751020647d08f7a3911b3

Observation 49542191-d96c-4e3d-b266-6d82137c50fe · inbound

Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis cites this paper.

Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:25:27.548842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:25:27.548842Z digest=sha256:8b35fcf73c3ec3ceb8dde7537e459f6c285249ad4065b3b927c74a4fcaa41f5a

Observation da47824a-1952-498a-937c-9ff0fa131f91 · inbound

Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models cites this paper.

Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:25:29.118241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:25:29.118241Z digest=sha256:66b53d7fce92f1a26dc4c761892ef1ffd7fc7e2f44a3de48ee53b3257f6c40d1

Observation 80d5efeb-243f-41f0-8ad0-0e4b99600dbe · inbound

TokenChain: A Discrete Speech Chain via Semantic Token Modeling cites this paper.

TokenChain: A Discrete Speech Chain via Semantic Token Modeling Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:16:09.518750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T09:14:58.542628Z digest=sha256:02820370d2d470ade2cea6500f896d9d0d47de9721fb25717fc7699a15916d8f

Observation 6e27a746-8fdd-4de7-b05f-cdd5558197e6 · inbound

Position: Towards Responsible Evaluation for Text-to-Speech cites this paper.

Position: Towards Responsible Evaluation for Text-to-Speech Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T11:07:41.013784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:07:41.013784Z digest=sha256:7bddb285697735d0974f528098b8e645abef891f069992898873e1fbfbdc28fc

Observation f2c35a8c-70a6-4ef1-8039-8338f2332ce1 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.009111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:0d475a345f472d2049c78109b471756308e0a30d5e80050dc7fdc3c1b34efe19

Observation 231ab4b1-d86d-4661-9316-3f6467049e24 · inbound

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations cites this paper.

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 5

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T12:24:39.706455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T12:20:43.383513Z digest=sha256:c5fca2621a66a059416e1f92888bded22a018291b902530fbee869124acc4761