Pith. sign in

Paper Citation Record · LEDGER

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2608.08362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08362 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:10:40.920937Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:10:40.788931Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:10:41.540631Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3906908-4e24-4528-b219-2ec9d33f18fe · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.893849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.774360Z digest=sha256:dd160a21c87d56dfda027d7175b328bf0da5c4ed3562895d639a29bba59df714

Observation 7ca53c06-e6f2-47ae-b9b7-64d9808f0602 · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.884413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.778165Z digest=sha256:dbcaf31a18f7d0a6382e5ba9a5e82ce2919c138a0c5296e61dc7b692c27f7cbf

Observation f3eb38eb-8e72-44bd-ae26-ebe1c70f1d3e · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.875334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.781218Z digest=sha256:ba32922beb012ae1a3919b9f5988b22d23511ac8df4e97ceb0567ccb73082658

Observation 440044ac-ccbf-4749-bcf6-10ac8c78b8d1 · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.866574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.785035Z digest=sha256:76318dd25e10ed1bcb516a36811381d24bedbe7fc5089f60b16efc5e5977b6e9

Observation 6a494d1b-0a14-4601-98f0-8bc0bcb2ac01 · outbound

This paper cites CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:10:41.544806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.788931Z digest=sha256:718f5ff16573101c8d7eba6771af014359c461a65760f74e3b83438e74c75cb3

Observation 33dac800-e9b4-4dcf-be69-a7637841b420 · outbound

This paper cites outpainting.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis outpainting

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.856686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.792815Z digest=sha256:43e08571f92066d2032e24243fa23e70b503cd97c7bcba55c101b524e22f5805

Observation a0d2cb77-2aa9-4380-bf83-f557d79ce680 · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.846877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.796541Z digest=sha256:d5f0625670175c92da3ffb60c2dc8b20d3daad2585eeb0b30210bc1045bbde11

Observation 097e6d95-3fde-4a8f-96ae-2401dbab8ce6 · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.837289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.800135Z digest=sha256:047eb01d2994e1d22c3d9e574d72bc544fb1fb5bd35a7f705e938a539b2c275a

Observation 39e83fd4-f028-4ba0-a7c9-4a397bfa4e5b · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.827861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.803689Z digest=sha256:cb087ca956e88a669b0de60ab7bbe7a21133edc4809babccdfb9bf9b311f1361

Observation caf97688-a2fb-4f82-b633-172179f6c8d7 · outbound

This paper cites First, our experiments are mainly conducted on English speech, so the effectiveness of CTRLSPEECHfor multilingual or code-switching synthesis re- mains unexplored.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis First, our experiments are mainly conducted on English speech, so the effectiveness of CTRLSPEECHfor multilingual or code-switching synthesis re- mains unexplored

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.818313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.806688Z digest=sha256:cc41b70fcfca68c56ee007929680a58750a9c3047cb2c90c392f32febfd92d9a

Observation 92c5ee93-08e0-4bad-8be0-32c56eeae2c5 · outbound

This paper cites 2D- 16003984 through the Amazon-UT Austin HUB.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis 2D- 16003984 through the Amazon-UT Austin HUB

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.807232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.809810Z digest=sha256:1d1a1848be7c52bc08dea8b376e262ee582b77e49d0aeabf1b4d1e2742ee2ed4

Observation ee42a50d-1072-413f-9fe8-6934b39bcb0b · outbound

This paper cites All authors remain fully responsible for the con- tent of this manuscript.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis All authors remain fully responsible for the con- tent of this manuscript

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.796694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.812951Z digest=sha256:3c1c5509148d4b101c2c5f74610e83d095cb4ee35ae41dd0b6cc5b8202bf25df

Observation a7d6d96c-9730-46e6-bc04-c727ef0716c1 · outbound

This paper cites Ditar: Diffusion transformer autoregressive modeling for speech generation,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Ditar: Diffusion transformer autoregressive modeling for speech generation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.816253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.816253Z digest=sha256:d318eaffa84811e11a0e29076312f630dcfeb6a27d617248f8cbabc150fc2d9d

Observation 027dec4c-984b-4e54-97f3-e3d7d95a6796 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.819409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.819409Z digest=sha256:d9137b25d56ac35513e65ce0da6a7ec2b8c860831b63c90fc6eef03746e6406a

Observation 24ffd9c1-4804-4199-b409-a07fe33fd398 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.823347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.823347Z digest=sha256:22ee8af17dfd5a442103be20c9063328a4839431b08caa14ddc43a77f4ea6b93

Observation 4dc0ab8f-49f9-4847-84ec-fe1532ef48a4 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.826976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.826976Z digest=sha256:2a312fa476238feef24858ee326e15b448817ba575aaa8739db376345eaad4d8

Observation 6409aaac-722e-4444-bd63-62c7de4cd51a · outbound

This paper cites F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.830784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.830784Z digest=sha256:0bc732d4994b39a6422eafb0bf2e34f32a55cf55ff80f0179ab602e509ccb5cc

Observation 1c7967d7-8218-47f1-a511-30027b4a84f4 · outbound

This paper cites V oicecraft-x: Unifying multilingual, voice- cloning speech synthesis and speech editing,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis V oicecraft-x: Unifying multilingual, voice- cloning speech synthesis and speech editing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.781590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.834489Z digest=sha256:0bab85ab1167b16d13036f980433c82d5e733c68321a0a376517ddb635f2138e

Observation c025b68e-de5c-4b9f-a609-f03798ad1bf6 · outbound

This paper cites Prompttts: Control- lable text-to-speech with text descriptions,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Prompttts: Control- lable text-to-speech with text descriptions,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.838301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.838301Z digest=sha256:116d1a8bcb1f9387dd930f6c0f012a89d4f6b472f3051de277bd4f6b7397cb64

Observation 4a09c230-013b-4930-8825-116294c725d3 · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.841801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.841801Z digest=sha256:6c70c1078eab458f56281fb332ca4aced2d03e01e6c7e93b13c0de042cd91094

Observation 084a1113-77e1-4aaf-836c-35857e15c28e · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural lan- guage style prompt,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Instructtts: Modelling expressive tts in discrete latent space with natural lan- guage style prompt,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.765506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.846245Z digest=sha256:8d707780fa1666d7bf19d9aa40a8d4ed839485e24ae8319cf8d8b6a7ccd2161d

Observation b6a76d86-87f8-48ce-ac8e-37bd5dc32208 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.849652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.849652Z digest=sha256:4fd6cf19282c4732b55d8b1d74c0f3e1b27b4dd49d76c556def429ca0a103dbc

Observation 02056422-fa70-46ed-b4bb-963fc2b1fe60 · outbound

This paper cites V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.755234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.853557Z digest=sha256:901ce629eb9153c63430aad7b119278fd0c723e2085736284e51b2481eb78f10

Observation c700938b-0062-403e-a8ce-f40cb29a6b4c · outbound

This paper cites Emovoice: Llm-based emotional text-to-speech model with freestyle text prompting,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Emovoice: Llm-based emotional text-to-speech model with freestyle text prompting,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.856918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.856918Z digest=sha256:182d79f96ca07bcfcaa89eb9719f3cea260c0dc2f70badcbe578c25212e23773

Observation 4e951b5e-af90-44a8-bb98-d41fec64bdc4 · outbound

This paper cites Scaling rich style- prompted text-to-speech datasets,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Scaling rich style- prompted text-to-speech datasets,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.739353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.859754Z digest=sha256:297c2c38f9deef21d4cf75097b6438e107b6aa69dba49c81d8b7200f78609724

Observation 79554df0-aee5-4816-9128-5e0f94e067c7 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.862468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.862468Z digest=sha256:e341f505c85c6d72b0cf9c23eea1b955d33fd96bb3bb31107ee45fc14455fb33

Observation fe29e582-b9c6-49c5-be71-2ed5e983fda7 · outbound

This paper cites Vevo2: A unified and controllable frame- work for speech and singing voice generation,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Vevo2: A unified and controllable frame- work for speech and singing voice generation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.865457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.865457Z digest=sha256:b603ce433b721c7875cd441ee1d46d09b287b1a35f3f80450d1b4868d3cb127d

Observation 77015045-1b72-499d-b93f-68edfef23e56 · outbound

This paper cites Mela-tts: Joint transformer-diffusion model with representation alignment for speech synthesis,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Mela-tts: Joint transformer-diffusion model with representation alignment for speech synthesis,

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:10:41.326818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.868011Z digest=sha256:2b9cf9ffa31c902b70d009b4642d858e6e58ec247624606c6fe907bdcfc101f9

Observation 737ebbf5-97a5-4e61-babd-a06700e970b1 · outbound

This paper cites Streammel: Real-time zero-shot text-to-speech via interleaved continuous autoregressive model- ing,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Streammel: Real-time zero-shot text-to-speech via interleaved continuous autoregressive model- ing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.728008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.871337Z digest=sha256:bae30f1421f39fe7ffdbbf3087e5ba5b8386afaf7958e35f73138acd2e6181fa

Observation da98a471-810b-4dd6-bc07-a3d615f58348 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis High-fidelity audio compression with improved rvqgan,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.873950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.873950Z digest=sha256:7099fbe0f78ebbbca8454f2bb912757e927a22b3cf03a548f1f6cb20393ce1b0

Observation 5ff38436-b674-4f67-bf61-c21289a03c4c · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.876765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.876765Z digest=sha256:d6b01bde6e90edca58eb988b34ccb1cce95ecf9927f404eabf0f3b737b491f85

Observation beed7ec7-66d0-4a8d-bd5d-29038c9034eb · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.880326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.880326Z digest=sha256:bd73ad2c050308cc2f729f49f9591acb8e3ea72843faa5b35b6a4440291ce4ca

Observation 0755047a-85de-4b8a-bc6b-0994e4d246ff · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.883298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.883298Z digest=sha256:4fd1351c2c03064e9bfa0cfa46d46d2ca4bc760b11f5a40cbc4a2fa277288c98

Observation c6ff20da-d32e-4471-a84a-e20044149be6 · outbound

This paper cites World: a vocoder-based high-quality speech synthesis system for real-time applications,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis World: a vocoder-based high-quality speech synthesis system for real-time applications,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.605365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.886358Z digest=sha256:d9e86f5686c0d2b20d4e82150b204d6066a9ff7c74e463a690959485b951ed97

Observation 189bfeb3-c9ad-4c0e-8c7e-ef64b20f164f · outbound

This paper cites Fast and reliable f0 estimation method based on the period extraction of vocal fold vi- bration of singing voice and speech,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Fast and reliable f0 estimation method based on the period extraction of vocal fold vi- bration of singing voice and speech,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.594371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.889367Z digest=sha256:9f7fba79755e4ab12c2ddb6af16e5683dfdef194d67b1cbe3bd55635bbf5657f

Observation b72ed2be-5ade-4a44-acc6-a08a36515fb1 · outbound

This paper cites Continuous-token diffu- sion for speaker-referenced tts in multimodal llms,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Continuous-token diffu- sion for speaker-referenced tts in multimodal llms,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.892028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.892028Z digest=sha256:6f62932ae8ea4c3d8d883263422a347ddbc2631dd976d477c5a045851843d880

Observation c808d725-d472-4221-8b5e-224e44c2a1e5 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.895164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.895164Z digest=sha256:789e5279cc79fd876042a15e8957c67ae27e792ba0943a48a696e588f9699010

Observation d8b16eb8-10f0-4ab0-b52d-b78e3a58f35a · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.898571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.898571Z digest=sha256:987417991a29fe52033f0571e46936b93913b41e193a4a32d428c113f6668a1a

Observation ddd65d55-f795-4692-8566-41c299c3f5cd · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.901548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.901548Z digest=sha256:a3adcebe8b4e6cc77a291d336844f3bc2634b3ea1e9d692fccaea75f6e6f35fb

Observation b8e17983-8473-41d0-9f8d-acb9088a8536 · outbound

This paper cites The lj speech dataset,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis The lj speech dataset,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.904905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.904905Z digest=sha256:5ddae4a237ceda92043e206a62d2ca935d756f2babcfa143b5c9152b778fe9ff

Observation 5f85e8df-6a19-498c-bf75-d4205c79f7bd · outbound

This paper cites Semantic-vae: Semantic- alignment latent representation for better speech synthesis,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Semantic-vae: Semantic- alignment latent representation for better speech synthesis,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.907818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.907818Z digest=sha256:c0872c2efb4ab7895f3dffbba4eb4cf90a576b16a5e5bc8ae1e08dbe764dd20f

Observation da8fe183-b02b-44f7-a098-089abdda655d · outbound

This paper cites Classifier-Free Diffusion Guidance.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Classifier-Free Diffusion Guidance

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.910596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.910596Z digest=sha256:b9dc268b56f8d936eacd238b5472c283d03ceb2b5f3f848752306bc3cdbadfab

Observation 306d53bf-9ffc-46cf-bb8f-106d97ccf6a7 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Robust speech recognition via large-scale weak supervision,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.913810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.913810Z digest=sha256:488b915ee655e6bdfb84ac279a8da3f70d448f4ddabe25f18e497451616ac05a

Observation 2d0016ae-58fc-4a76-b7b2-9bf848779e45 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.917178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.917178Z digest=sha256:b5b94160f8b19f6705e12a19519eedd2b2073d2975daf4bd0f6dd398ffd1eb7b

Observation 9ed6a26b-8b9c-4c3b-901a-7ed2ba1a11a2 · outbound

This paper cites Drawspeech: Expressive speech synthesis using prosodic sketches as control conditions,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Drawspeech: Expressive speech synthesis using prosodic sketches as control conditions,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.556446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.920937Z digest=sha256:f6b430d1300a88c75f554e8950f7c023aa7dca040be757482af8455018ad138b

Pith citing papers

Observation 6a494d1b-0a14-4601-98f0-8bc0bcb2ac01 · inbound

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis cites this paper.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:10:41.544806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:10:40.788931Z digest=sha256:718f5ff16573101c8d7eba6771af014359c461a65760f74e3b83438e74c75cb3