Pith. sign in

Paper Citation Record · LEDGER

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

As of 20 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2501.04256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04256 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:41:05.512701Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:19.134392Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation 7ed7a4e9-b270-46b6-9c38-b4da6b45fb67 · outbound

This paper cites Review: Prosodic patterns in english conversation,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Review: Prosodic patterns in english conversation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.729039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.223811Z digest=sha256:6c6b13218a24c7782b86bd5541f453cdb70429e8dc7a2972a85d0e46e8d7fc35

Observation f96bab2f-df9c-499a-9184-9bc16d7a53aa · outbound

This paper cites Supervised and unsu- pervised approaches for controlling narrow lexical focus in sequence- to-sequence speech synthesis,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Supervised and unsu- pervised approaches for controlling narrow lexical focus in sequence- to-sequence speech synthesis,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.706053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.236746Z digest=sha256:5bf9f78542497eb46f9de5ba6eb6d5bb344bef0f41d4bc9c78e2df191524ec1f

Observation b2cafa12-d255-46a0-a3ac-fa7e3d8fe149 · outbound

This paper cites A Survey on Neural Speech Synthesis.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions A Survey on Neural Speech Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.248498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.248498Z digest=sha256:be4b40cc11066341fe81a14714a4952c664b37dc75ea534332c00bba605e1020

Observation 7283c252-0281-4a8c-a366-db276c7200c3 · outbound

This paper cites NaturalSpeech: End-to-end text-to-speech synthesis with human-level quality,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions NaturalSpeech: End-to-end text-to-speech synthesis with human-level quality,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.683491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.254758Z digest=sha256:fd9ee55b4297bc0c7e1c4bccecf0bd365d45b1fdb5dfa168eef39019376457bc

Observation 4258bcaf-114f-4fd7-8e4f-ae7d127ca892 · outbound

This paper cites Emotion rendering for conversational speech synthesis with heterogeneous graph-based context modeling,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Emotion rendering for conversational speech synthesis with heterogeneous graph-based context modeling,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.664604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.260233Z digest=sha256:b54d603675da4795698b6da9a19ae15668e961fe526544f9f553c57020839a1e

Observation c2cb0723-18e9-4108-bb7e-81a96bde7300 · outbound

This paper cites Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.644394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.266983Z digest=sha256:6d8783f4c3a2f2cb0872e9397f7c3fa9d0733d972f11cb384c4a96019d24e3e0

Observation d3dadb15-99c4-42ff-807b-0203a4de419d · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.274200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.274200Z digest=sha256:0f1fba5ec60f701d2fe946281fc98f694271c61de89794654346544f9db0781e

Observation 93bbd24b-ed81-455b-8619-2ab0f40e3eb8 · outbound

This paper cites Towards multi-scale style control for expressive speech synthesis,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Towards multi-scale style control for expressive speech synthesis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.624385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.279726Z digest=sha256:7715bae3ed0e4c82ef88a28c08897e9374c0c6bed5954ca13c2bcafedf1bc598

Observation 7903e203-d90b-446d-b6a4-8b2a4e9e3ce8 · outbound

This paper cites Contrastive context-speech pretraining for expressive text-to-speech synthesis,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Contrastive context-speech pretraining for expressive text-to-speech synthesis,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.602743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.287278Z digest=sha256:5600302bf57fe5e0d9b60f0ef5ee50ca5dc12cf7d96370cc3c1fc45853600d0e

Observation 983c872f-43f9-4b2c-a6a1-4569aa63bad7 · outbound

This paper cites StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.292855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.292855Z digest=sha256:281ad3ba5c2abb4178fe49284886f733eef01f058ac2d9c3b2b1049affe18016

Observation ef1b36aa-1cd8-4351-9241-e36eafef24c4 · outbound

This paper cites NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.582016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.299029Z digest=sha256:4a92ead994d91ae6430de93f6afd687610b21a44d65711372619111f97b2d2ec

Observation 1c759169-0b6b-42be-bf31-d0ab2ceebe71 · outbound

This paper cites NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.562966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.304869Z digest=sha256:573a64af820550b769f5a54eb0c556727482501f0a498a7c8640c37aab8ec39e

Observation 9313f290-fd6a-4a42-b03f-88644450ded0 · outbound

This paper cites Principal style components: Expressive style control and cross-speaker transfer in neural tts,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Principal style components: Expressive style control and cross-speaker transfer in neural tts,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.541323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.312897Z digest=sha256:260069742cfd7ad7001ab14212076f5095471a524a5e9e07d9ee3db7698b11c8

Observation c22c75b0-5fe7-462d-892f-3dfede9a74fc · outbound

This paper cites Multi-speaker expressive speech synthesis via multiple factors decoupling,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Multi-speaker expressive speech synthesis via multiple factors decoupling,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.514266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.324855Z digest=sha256:6d54930f95a376909692e84334afdf37ab96629b22c35e7fae7e36926ae09315

Observation 372752d7-074a-4fb2-8346-55afcaac34c9 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.331488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.331488Z digest=sha256:918b959ded0c085520d186fac27c5767526923a2555abf2d64eb9d89149caa83

Observation e0854797-73ba-41d9-9c83-a67a0b53735a · outbound

This paper cites Controllable speaking styles using a large language model,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Controllable speaking styles using a large language model,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.492669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.339200Z digest=sha256:664900ab30e44495cc6073d58c0838b9e422c899f4656de3e4b78b6e767ea657

Observation 7a126917-fb4b-442e-9675-03a4fca3f1c2 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.346831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.346831Z digest=sha256:e91d06535b451b7275ad031a0763c6049609d1365c0545ea43db322517a22831

Observation c26423aa-3e62-4d57-ba67-c72e1909d8b6 · outbound

This paper cites InstructTTS: Modelling expressive tts in discrete latent space with natural language style prompt,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions InstructTTS: Modelling expressive tts in discrete latent space with natural language style prompt,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.473423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.353554Z digest=sha256:dc070fb77980792ff8ada683b18439df873ff51a5ae1a1e2bbea59914600046d

Observation 481a9617-b00e-4e79-9c8f-45c7d41fdc71 · outbound

This paper cites Controlling emotion in text-to-speech with natural language prompts,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Controlling emotion in text-to-speech with natural language prompts,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.453619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.358581Z digest=sha256:c28867a63ff582d19c7ed24ff1743093d5bef455f395d8f502734cc3f2bffa7d

Observation 476e4551-b98e-4449-b8b6-9b008cd9de12 · outbound

This paper cites PromptStyle: Controllable style transfer for text-to-speech with natural language descriptions,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions PromptStyle: Controllable style transfer for text-to-speech with natural language descriptions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.434421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.364723Z digest=sha256:79f015e78d805e2832ce98fe4937159a7386ba14ac94977ed4c9a5188e960aea

Observation 76573e23-412f-47db-8df6-1f4499b66754 · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Prompttts: Controllable text-to-speech with text descriptions,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.410604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.370739Z digest=sha256:f4effaee746f2594d3e8b2734ee0609e936fe6247f6e74a71a0cd4f3e674ae6c

Observation c2c9d3ff-2ea6-4468-8c90-6303649fe0a0 · outbound

This paper cites UniAudio: Towards universal audio generation with large language models,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions UniAudio: Towards universal audio generation with large language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.388153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.378985Z digest=sha256:53af6aebb4752ae3efb92b4d6214919589cde7caa11351634284406b6e534cac

Observation ed8b3ecd-9a45-4ff7-82b7-01d72fac5568 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.384890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.384890Z digest=sha256:e1b964aa6b2478c8b3975b697a63baa99966bc127bb778c7ca96b49c220fce66

Observation 638742f1-b16e-4155-a992-b3c8f94fff4d · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.397681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.397681Z digest=sha256:a2ea165f7d4e91bd3a06b96f9f660276971964cd343a4c9d42e3dd5ba8e1a0ee

Observation e08c34ec-3071-4c82-896e-f14bdbf5f07b · outbound

This paper cites Towards general-purpose text-instruction- guided voice conversion,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Towards general-purpose text-instruction- guided voice conversion,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.350652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.404258Z digest=sha256:1a05c00914b1073407212a793d13d01e060ad9b34efb78cd4bfaef8c50ab2b05

Observation 2b77f819-72de-432a-becf-d49126aaaa9a · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.316711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.410582Z digest=sha256:427b9fa1a16d15df267459869443470f89e09a63919f0044c2344b398090463a

Observation aa2b4d87-8672-4382-9656-7ac88541cc3d · outbound

This paper cites Communicating emotion: The role of prosodic features,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Communicating emotion: The role of prosodic features,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.293761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.417589Z digest=sha256:c7ff61e2d97c67f1bcf00acfc0a3375f46d55117be62905e2a3383d184e79ea1

Observation 7554618b-c9bc-463c-b809-0fcbadbacc9c · outbound

This paper cites Speech prosody enhances the neural processing of syntax,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Speech prosody enhances the neural processing of syntax,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.267189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.424812Z digest=sha256:7c82eaad9e94fc38b21be351403b5c71b67eb89089ebe93664db49064ef57405

Observation a5a2a896-3376-4c11-8e20-a9bbc32029fe · outbound

This paper cites Smoothing and differentiation of data by simplified least squares procedures.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Smoothing and differentiation of data by simplified least squares procedures

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.430583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.430583Z digest=sha256:74fa6561d0bf6ebbc438e1ad639a756accd9cf075cb0053ee80d9ac4ca37c503

Observation 1d6c3ab9-bc4b-441e-a2f6-1982750e480d · outbound

This paper cites Auto-encoding variational bayes,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Auto-encoding variational bayes,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.436238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.436238Z digest=sha256:ab652cb30f39eb7b89689f4752d04ea4d7663c81d720b76d5c958cad55047a0f

Observation 46f9ed2b-71af-4d09-82cd-56eea750f524 · outbound

This paper cites Fast- Speech: Fast, robust and controllable text to speech,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Fast- Speech: Fast, robust and controllable text to speech,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.050228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.442259Z digest=sha256:0734167d1b8b7205358f6e3136f119d24828b87a97a27ceb2430c91b2ba8b2d8

Observation 75cdbde0-ff5f-4f59-a8df-b966c3a4c97e · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions High- resolution image synthesis with latent diffusion models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:06.011224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.448117Z digest=sha256:5b70a213615a40e4cb988ab31749e0a8c09aec1387d7e82b57944d5d71288e81

Observation d061b293-93ce-45e5-b679-11e770161819 · outbound

This paper cites Denoising diffusion probabilistic models,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Denoising diffusion probabilistic models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:05.985105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.453438Z digest=sha256:d3a2ed8295f040704fe54165c2a321feffbc52c6617d5d5ac5becad0b6eafd4d

Observation b3aabb62-e90f-4bb5-8158-991c067ce2e8 · outbound

This paper cites AudioLDM: Text-to-audio generation with latent dif- fusion models,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions AudioLDM: Text-to-audio generation with latent dif- fusion models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:05.961826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.458663Z digest=sha256:7330387a2595f0731e84cc8ee52650673869c5141ca722accaa6c3863185a832

Observation 50dc9f95-e88f-45da-990f-05ff0f797a25 · outbound

This paper cites The lj speech dataset,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions The lj speech dataset,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.463927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.463927Z digest=sha256:874bf8537b4468c0a3242355d288c9819eb1646498b2828467faed7b5ed7989c

Observation ff91d98f-0b32-45a2-a232-83d0e1c67fe4 · outbound

This paper cites Attention is all you need,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Attention is all you need,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:05.470695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:05.470695Z digest=sha256:4814dddd18cac8dced0a99167a073cca6b25d6e4701d8f80109d8dce51fcfe6d

Observation 6ca0b6d7-7058-4eec-aaea-0c2072a50f4c · outbound

This paper cites AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:05.892570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.480700Z digest=sha256:832e6cf4c6da5faf27301dd7372633d150e413eb5a7b43b8d8fea9480c7268e8

Observation e447c0f9-651a-46fa-b0aa-b804df390b0d · outbound

This paper cites HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:05.863672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.487576Z digest=sha256:08f4f807a07eabc3d07aeb9cfe4796c89ba361bc7baa62f81ed5937a033f3c78

Observation bca90c41-3ba0-475a-b9d9-ef63d8bb8fe8 · outbound

This paper cites Adam: A method for stochastic optimization,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Adam: A method for stochastic optimization,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:05.842468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.493299Z digest=sha256:59c3e29fe1b81c54cadae459c9a3f3690a3600fe4872b7ed7cb5894b55154e81

Observation 05d7924e-c7c0-4a1e-b8de-2c60cdcff6d4 · outbound

This paper cites Decoupled weight decay regularization,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Decoupled weight decay regularization,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:05.811226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.498808Z digest=sha256:ed3e7c479248eee11b5053aec2c19d3a58d9207763716eb943c7e09e112e87ce

Observation 14cd95dc-0adb-4123-9c0c-2de2677903ae · outbound

This paper cites FastSpeech 2: Fast and high-quality end-to-end text to speech,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions FastSpeech 2: Fast and high-quality end-to-end text to speech,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:05.787324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.504092Z digest=sha256:3fdbca3424497052203e0c59bd3aafadf45f71dacb45b320fe6efd43c8d1e36b

Observation 7da0bfd6-2b54-4525-9bab-93cde587c03f · outbound

This paper cites Individual comparisons by ranking methods,.

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions Individual comparisons by ranking methods,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:41:05.760813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T21:41:05.512701Z digest=sha256:872bd3fe3ff73b6097de70caa9f1b958c25a3c4376820db794ec5c95eb1dbd2c

Pith citing papers

Observation ccb609e8-ddf5-4482-843d-273d1b8e377a · inbound

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models cites this paper.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.134392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.134392Z digest=sha256:bbf082c69b49cec0caf4189c4eb99d2af26259192d46d26bddd4904a29a06d7b

Observation 994c00c7-d34c-468b-92ae-d298e02b0200 · inbound

Improving French Synthetic Speech Quality via SSML Prosody Control cites this paper.

Improving French Synthetic Speech Quality via SSML Prosody Control DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T16:54:33.231805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-05T16:54:32.874696Z digest=sha256:1915fda1ad5e60e2ff50e7c08602be161cb3ed5f27dadadb0e9f6340fb643498