Pith. sign in

Paper Citation Record · LEDGER

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2608.08362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08362 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:10:40.920937Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:10:40.788931Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:10:41.540631Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3906908-4e24-4528-b219-2ec9d33f18fe · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.893849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.774360Z digest=sha256:b6cd0a9c6d89e6439f8c2fdf3868b5a23651b3197a4d0efebddfe6f684cbe844

Observation 7ca53c06-e6f2-47ae-b9b7-64d9808f0602 · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.884413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.778165Z digest=sha256:b2eabd5d6de1c770b0f17175dd1a17e9f82bddd0b03be48ece80d5a4afe91f53

Observation f3eb38eb-8e72-44bd-ae26-ebe1c70f1d3e · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.875334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.781218Z digest=sha256:8d8a4bea22892b4662a1c6c03a0d5f83e5aa9d0249e3975748b8de7f8b69cdf9

Observation 440044ac-ccbf-4749-bcf6-10ac8c78b8d1 · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.866574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.785035Z digest=sha256:bc38141cfe6ad3b1cb1fcbdcf8f497c83cd9eb6b043c63711977913fe1e2ff52

Observation 6a494d1b-0a14-4601-98f0-8bc0bcb2ac01 · outbound

This paper cites CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:10:41.544806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.788931Z digest=sha256:c8e9ee8294e397d8f754b5b5d3b69ed394eac9815cdf2f7167a9aa6d07c634db

Observation 33dac800-e9b4-4dcf-be69-a7637841b420 · outbound

This paper cites outpainting.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis outpainting

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.856686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.792815Z digest=sha256:b2479321505d976ff53d31750b823ab248de67decc5e549091b42f878765bd0e

Observation a0d2cb77-2aa9-4380-bf83-f557d79ce680 · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.846877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.796541Z digest=sha256:dc29b739c584b49e2959b4af9fae0ab40d350b37521ac1aaa22bed597cd1a564

Observation 097e6d95-3fde-4a8f-96ae-2401dbab8ce6 · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.837289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.800135Z digest=sha256:1838c0fa802af98e7a8bde3601547cc215547114e77408b08275c04ddf53b8a6

Observation 39e83fd4-f028-4ba0-a7c9-4a397bfa4e5b · outbound

This paper cites an unresolved cited work.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:10:41.827861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.803689Z digest=sha256:217be3ef53a441eb7c6433a7e9e18507967cd25d7f49808887e50e93ee9de2e2

Observation caf97688-a2fb-4f82-b633-172179f6c8d7 · outbound

This paper cites First, our experiments are mainly conducted on English speech, so the effectiveness of CTRLSPEECHfor multilingual or code-switching synthesis re- mains unexplored.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis First, our experiments are mainly conducted on English speech, so the effectiveness of CTRLSPEECHfor multilingual or code-switching synthesis re- mains unexplored

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.818313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.806688Z digest=sha256:168c95b1c989c3fd13e807f9cdf6a600446add73819224988546ec9270fe4af6

Observation 92c5ee93-08e0-4bad-8be0-32c56eeae2c5 · outbound

This paper cites 2D- 16003984 through the Amazon-UT Austin HUB.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis 2D- 16003984 through the Amazon-UT Austin HUB

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.807232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.809810Z digest=sha256:f5a4d9b38cf4a41dcd9e6f3638ef6b2f6f2cb4bb55cedbddbd5ebcc27502d7d8

Observation ee42a50d-1072-413f-9fe8-6934b39bcb0b · outbound

This paper cites All authors remain fully responsible for the con- tent of this manuscript.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis All authors remain fully responsible for the con- tent of this manuscript

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.796694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.812951Z digest=sha256:1c6c78c83f32e066625e5255b76e47cd1f1c0841df332da5afd8227f7039379e

Observation a7d6d96c-9730-46e6-bc04-c727ef0716c1 · outbound

This paper cites Ditar: Diffusion transformer autoregressive modeling for speech generation,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Ditar: Diffusion transformer autoregressive modeling for speech generation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.816253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.816253Z digest=sha256:a045450fe617cb661e70fd0f0de23735c4792c4b36c4a81fe2b8df101da869b6

Observation 027dec4c-984b-4e54-97f3-e3d7d95a6796 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.819409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.819409Z digest=sha256:9c2fd1ccad2f23f94f01c0034af47a06e91a8f6922819520007f004116c7d59c

Observation 24ffd9c1-4804-4199-b409-a07fe33fd398 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.823347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.823347Z digest=sha256:415c7153f5718f225c66ccdb94523777f1add4b5960957b1e062406a63b99f8f

Observation 4dc0ab8f-49f9-4847-84ec-fe1532ef48a4 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.826976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.826976Z digest=sha256:4228a957d8b12c085adcccc0b8edfe90c14288240f43a935355c47e95aae92fb

Observation 6409aaac-722e-4444-bd63-62c7de4cd51a · outbound

This paper cites F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.830784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.830784Z digest=sha256:fa0d20f2f68395022787d837b89a62d254d6a5ae312f7037c1029853a6c42fdf

Observation 1c7967d7-8218-47f1-a511-30027b4a84f4 · outbound

This paper cites V oicecraft-x: Unifying multilingual, voice- cloning speech synthesis and speech editing,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis V oicecraft-x: Unifying multilingual, voice- cloning speech synthesis and speech editing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.781590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.834489Z digest=sha256:e8374f75a170180353d57793915909003292f54fbcbb503b9d39f4508421e7ba

Observation c025b68e-de5c-4b9f-a609-f03798ad1bf6 · outbound

This paper cites Prompttts: Control- lable text-to-speech with text descriptions,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Prompttts: Control- lable text-to-speech with text descriptions,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.838301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.838301Z digest=sha256:5c4c305dd2dbe7ed86610c7f13cfc922822c474bbeaa6132d570980bf99bc3ce

Observation 4a09c230-013b-4930-8825-116294c725d3 · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.841801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.841801Z digest=sha256:be0bdba53fcf2ff81a3e580269806e2a81e801b7fcb1fa983f4e4f5d2a2a6e7d

Observation 084a1113-77e1-4aaf-836c-35857e15c28e · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural lan- guage style prompt,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Instructtts: Modelling expressive tts in discrete latent space with natural lan- guage style prompt,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.765506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.846245Z digest=sha256:c80d99c1931754bada82cb58c6f864c6e96e6e6972aba405f0188e880498b1d9

Observation b6a76d86-87f8-48ce-ac8e-37bd5dc32208 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.849652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.849652Z digest=sha256:b362c0257f80ff18dd21ef86bc2a21bae5370368e37c0ff7f234f72b8e5183a9

Observation 02056422-fa70-46ed-b4bb-963fc2b1fe60 · outbound

This paper cites V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.755234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.853557Z digest=sha256:460d01e24388d9a31834e98d729824d047e4e8fdc58b81b719718bf283d76380

Observation c700938b-0062-403e-a8ce-f40cb29a6b4c · outbound

This paper cites Emovoice: Llm-based emotional text-to-speech model with freestyle text prompting,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Emovoice: Llm-based emotional text-to-speech model with freestyle text prompting,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.856918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.856918Z digest=sha256:5b47b6417bb19db103044374f247ea55860e17cf35659ddee4402d9fbd4ae444

Observation 4e951b5e-af90-44a8-bb98-d41fec64bdc4 · outbound

This paper cites Scaling rich style- prompted text-to-speech datasets,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Scaling rich style- prompted text-to-speech datasets,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.739353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.859754Z digest=sha256:0be24b95493abcfad9687a65547dd3d7afd9837c829aff8a0faf069f0510c897

Observation 79554df0-aee5-4816-9128-5e0f94e067c7 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.862468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.862468Z digest=sha256:27741463dda60884135942e6cb4eab113c167e34771791552da7798dbac49570

Observation fe29e582-b9c6-49c5-be71-2ed5e983fda7 · outbound

This paper cites Vevo2: A unified and controllable frame- work for speech and singing voice generation,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Vevo2: A unified and controllable frame- work for speech and singing voice generation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.865457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.865457Z digest=sha256:d902dc591ffa5f94bee33dab6320ba8c99e2be2f28fcaf7e2cba6214114ecec4

Observation 77015045-1b72-499d-b93f-68edfef23e56 · outbound

This paper cites Mela-tts: Joint transformer-diffusion model with representation alignment for speech synthesis,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Mela-tts: Joint transformer-diffusion model with representation alignment for speech synthesis,

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:10:41.326818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.868011Z digest=sha256:e521465e8eb2073002d958a9ba0b3151c5ee816fd1347757b5cb4ff8c8bd1478

Observation 737ebbf5-97a5-4e61-babd-a06700e970b1 · outbound

This paper cites Streammel: Real-time zero-shot text-to-speech via interleaved continuous autoregressive model- ing,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Streammel: Real-time zero-shot text-to-speech via interleaved continuous autoregressive model- ing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.728008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.871337Z digest=sha256:25e3db2b3e10bfa14b1b5aec2ff13f2a67065dc73f88e147f8fad3bae26974ee

Observation da98a471-810b-4dd6-bc07-a3d615f58348 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis High-fidelity audio compression with improved rvqgan,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.873950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.873950Z digest=sha256:a033c87fb702fea02fffd15be4dd0d1ec2f26287752e30ca2c414ca9662fdd4d

Observation 5ff38436-b674-4f67-bf61-c21289a03c4c · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.876765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.876765Z digest=sha256:ad1c20e7b1dc7798a082cfcb7e31a119a8ff2ba4dd9979bebce980e059d0b431

Observation beed7ec7-66d0-4a8d-bd5d-29038c9034eb · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.880326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.880326Z digest=sha256:7a9639c8c093f18945794e4fb92fcc69dc6b8622c69daf94b33d72773d423ec6

Observation 0755047a-85de-4b8a-bc6b-0994e4d246ff · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.883298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.883298Z digest=sha256:a75143f82e2b5de294453c4f10c065f5b6a23d0f8d5c5d6e50cddcfdc7c0aa24

Observation c6ff20da-d32e-4471-a84a-e20044149be6 · outbound

This paper cites World: a vocoder-based high-quality speech synthesis system for real-time applications,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis World: a vocoder-based high-quality speech synthesis system for real-time applications,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.605365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.886358Z digest=sha256:b95b603040771be274fc958f237eb099acdbd8a2b382786643908c0214fb3ce5

Observation 189bfeb3-c9ad-4c0e-8c7e-ef64b20f164f · outbound

This paper cites Fast and reliable f0 estimation method based on the period extraction of vocal fold vi- bration of singing voice and speech,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Fast and reliable f0 estimation method based on the period extraction of vocal fold vi- bration of singing voice and speech,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.594371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.889367Z digest=sha256:6f3ba9bb81ec0e29dd79350453cb24e451ce62855bdced2b3944ce2b203d126e

Observation b72ed2be-5ade-4a44-acc6-a08a36515fb1 · outbound

This paper cites Continuous-token diffu- sion for speaker-referenced tts in multimodal llms,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Continuous-token diffu- sion for speaker-referenced tts in multimodal llms,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.892028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.892028Z digest=sha256:ede8f4b0234d82ab1d92dd52ad47302d7fc3f178244113dd84560b38a8d19e0b

Observation c808d725-d472-4221-8b5e-224e44c2a1e5 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.895164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.895164Z digest=sha256:bb91dbd7271a68a11d5e19651f1fc55aa5f00a1f3c4cdd74930e321e24b7225b

Observation d8b16eb8-10f0-4ab0-b52d-b78e3a58f35a · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.898571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.898571Z digest=sha256:6156c8a109fc2ceefcf171bb78ee3ca6dc99f3741017c0a2e774d58a2eb9f5e2

Observation ddd65d55-f795-4692-8566-41c299c3f5cd · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.901548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.901548Z digest=sha256:47e32334776e0315bb4a3a5cb57cd25c040b947d6930d4aa69fb44a6ec290116

Observation b8e17983-8473-41d0-9f8d-acb9088a8536 · outbound

This paper cites The lj speech dataset,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis The lj speech dataset,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.904905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.904905Z digest=sha256:306fd930e09721d13f7fa46abf06c3dd7bfcdd57c132c2b71d6e0af9ffe5dd9b

Observation 5f85e8df-6a19-498c-bf75-d4205c79f7bd · outbound

This paper cites Semantic-vae: Semantic- alignment latent representation for better speech synthesis,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Semantic-vae: Semantic- alignment latent representation for better speech synthesis,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.907818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.907818Z digest=sha256:58764947b194d506a609e21ed4511ac8f0af861c74ae1cd22b927a0e88529320

Observation da8fe183-b02b-44f7-a098-089abdda655d · outbound

This paper cites Classifier-Free Diffusion Guidance.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Classifier-Free Diffusion Guidance

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.910596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.910596Z digest=sha256:db10ac985756f3aba3b075bfa115b35b1b3d0b94c4ca979b8532c274708f9900

Observation 306d53bf-9ffc-46cf-bb8f-106d97ccf6a7 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Robust speech recognition via large-scale weak supervision,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.913810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.913810Z digest=sha256:7c3ed484b24957611fb240e457f1f6ab1580c5432507dc47255d489cad2f7063

Observation 2d0016ae-58fc-4a76-b7b2-9bf848779e45 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.917178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.917178Z digest=sha256:46869b27281d96465ec2dd8f6306b3aa44bfd403543106383e2bea91f56cfa16

Observation 9ed6a26b-8b9c-4c3b-901a-7ed2ba1a11a2 · outbound

This paper cites Drawspeech: Expressive speech synthesis using prosodic sketches as control conditions,.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis Drawspeech: Expressive speech synthesis using prosodic sketches as control conditions,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:10:41.556446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.920937Z digest=sha256:7bf5ce66b23e5b748054360a17ec3c71d463438beb098fcc41e638225536718b

Pith citing papers

Observation 6a494d1b-0a14-4601-98f0-8bc0bcb2ac01 · inbound

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis cites this paper.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:10:41.544806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:10:40.788931Z digest=sha256:c8e9ee8294e397d8f754b5b5d3b69ed394eac9815cdf2f7167a9aa6d07c634db