Pith. sign in

Paper Citation Record · LEDGER

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

As of 23 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.20945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20945 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:52.529754Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27adb6b3-b24d-4ae4-8078-4cbcf618dc2c · outbound

This paper cites Mega- tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Mega- tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.933915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.228144Z digest=sha256:7fd35ea2960e71da6807a5ac22c7ef1abb72f0aecbd02bf11c373f992678b873

Observation 49a3a70a-a4b9-4624-8064-993c0af597aa · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.265834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.265834Z digest=sha256:5cc0a45cdc0bf2b5cb555d8c0c6e773c9ea5216dcdfcb2698d379373c1d240ae

Observation 663c1c04-c3e4-47e7-9303-e95d307ed5f0 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.332197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.332197Z digest=sha256:de676412c3f2f46571a880ab3beb9f18bac4671ced79a4807a280c984041d4a0

Observation 80bebff5-0883-408a-bb4e-bb46906172b1 · outbound

This paper cites Imaginary voice: Face-styled diffusion model for text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Imaginary voice: Face-styled diffusion model for text-to-speech,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.768449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.379308Z digest=sha256:a5e01c030d7176ff072b5d4abef4884898bcd7e67c19dddbf0ed5ca0c4d35db9

Observation 3bae3fb5-4477-40bf-a466-d1206948ef89 · outbound

This paper cites SYNTHE-SEES: Face based text-to-speech for virtual speaker,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis SYNTHE-SEES: Face based text-to-speech for virtual speaker,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.560254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.454168Z digest=sha256:fbd3a6a8a626c7799b373e79faadc7af886897c7dbd1f9d30886f8147879d276

Observation 921cc365-a7bf-4ae9-aafd-b477a96c2b58 · outbound

This paper cites Face2Speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Face2Speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.395195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.528613Z digest=sha256:399a60aa4c27a63cafda37ad030b50cd373878f2e298d4125adcb5c55544a51f

Observation 3cf14a4c-5247-4d3f-a3fd-edced867f33d · outbound

This paper cites FVTTS : Face based voice synthesis for text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis FVTTS : Face based voice synthesis for text-to-speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.239168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.580539Z digest=sha256:f40427a72bf6dde609d6013fa4f6927f4beeb1cc0994f31c4768f4eb862aab32

Observation a53c088a-a93b-47ee-9a23-c88649a88f76 · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.088342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.633123Z digest=sha256:9412566a407b7d9102c5375437293e76ff2e9bc23b8fd73e3b33c0fa8c6c58f6

Observation ffd35756-5c7b-4282-9d92-0eb646be2e2d · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Prompttts: Controllable text-to-speech with text descriptions,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.942659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.701372Z digest=sha256:bd8c59614f02bbe11ec51421edfe99411c86fc545a720899aa418cc0b8173afb

Observation 3556fb19-c8c4-4690-a30d-2355895035fb · outbound

This paper cites PromptTTS++: Controlling speaker identity in prompt-based text-to- speech using natural language descriptions,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis PromptTTS++: Controlling speaker identity in prompt-based text-to- speech using natural language descriptions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.781073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.766499Z digest=sha256:e184e10d6d15281f7c9a9b76e36f752969f6fb8e95af7808757ac8c3765f8778

Observation 96b659af-c3d2-464f-83f5-8ed7af6292f7 · outbound

This paper cites UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.830879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.830879Z digest=sha256:2c78da02c3a7170a874898a4c5056937caaa8dd5fd337a05efc23a09dd4e8dfc

Observation 6de8182b-399d-45b8-bde7-dc17676d2a97 · outbound

This paper cites MM-TTS: Multi- modal prompt based style transfer for expressive text-to-speech synthe- sis,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MM-TTS: Multi- modal prompt based style transfer for expressive text-to-speech synthe- sis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.631249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.902999Z digest=sha256:d09599043f8993faf5f27dc1a6127718396676d7945e1b761916ec4efe495028

Observation 9b3f1fe2-a79f-49e2-ac7e-39310830e03f · outbound

This paper cites Gen- eralized end-to-end loss for speaker verification,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Gen- eralized end-to-end loss for speaker verification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.479581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:50.984883Z digest=sha256:1bd4ef4ddde1c48b8fecc8d4bd90f8494460b0266e334f31ff45b342be7efb62

Observation 42cd64a7-6d95-4511-bf4f-adf0c4a098f7 · outbound

This paper cites Additive margin softmax for face verification,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Additive margin softmax for face verification,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.350835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:51.073749Z digest=sha256:27eea0e1e1cd4b62f203ad56bc5e49875f9bda231814f82638715c8fc7c50dd6

Observation 513fb0e2-5d0e-4c9d-980a-02ad682a3cba · outbound

This paper cites Bridging the gap between object and image-level representations for open-vocabulary detection,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Bridging the gap between object and image-level representations for open-vocabulary detection,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.148216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:51.143577Z digest=sha256:78918463c1baa3464f1ac1e1a1015df9ca844c8cba59c80409efbfcdccedb97a

Observation 1a202008-871a-415c-9415-9faaffdcd3d8 · outbound

This paper cites Joint- teaching: Learning to refine knowledge for resource-constrained un- supervised cross-modal retrieval,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint- teaching: Learning to refine knowledge for resource-constrained un- supervised cross-modal retrieval,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.002813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:51.255727Z digest=sha256:f4703a2308fc82ae0ef632614574ce898e0c7c79ca6a28299683e96740b5eb22

Observation 8de811c5-5e1f-49dc-865e-18f221ead62d · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Representation Learning with Contrastive Predictive Coding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.359236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.359236Z digest=sha256:4ee9787dcb4f9cf4213b847f074541cfb831ad8743e1e6154607321631800dd8

Observation 283108dd-26b0-4501-ae07-6e229d58e159 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.818807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:51.427727Z digest=sha256:b6bb8968039144e553b5d43dd2d7c66fb9237109fd9d56d3f9d8a2f558dae9b6

Observation a29d89ed-059f-4cef-a988-bf5424033794 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LRS3-TED: a large-scale dataset for visual speech recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.542194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.542194Z digest=sha256:225172af54f79a078334af877507d0fe07f796961e49a9bd1ca3f17c17f44a7f

Observation 4f76282e-dc9d-4e21-a18a-023dd6db6970 · outbound

This paper cites Multi- caption text-to-face synthesis: Dataset and algorithm,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Multi- caption text-to-face synthesis: Dataset and algorithm,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.588907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:51.614760Z digest=sha256:e8f8bbd31533feb5a6b540e3ad2393804ec91713c21b05b5e48c810e669c658c

Observation 5eb4f155-d80c-43eb-b966-9739821a71fa · outbound

This paper cites LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.707283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.707283Z digest=sha256:590112188d1cddb45a55692aaab2650676177565a3cf5c43484513f8d92394f9

Observation d4b436d8-9b4e-4335-b459-f4b37136a926 · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.837350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.837350Z digest=sha256:4affc2601fc108422dc346db132d9df0192c9bf73f46f59f262ed0493709a198

Observation 1eb94adb-cfed-44c4-8ab4-2efea73e8fff · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.942109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.942109Z digest=sha256:97cbc481dfe658c06b9d9f6f42e1ffe0b39e0f327f4fb79be704a9b348e749aa

Observation 30bb92a6-34a9-4db0-8670-fc1c27a258f0 · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Facenet: A unified embedding for face recognition and clustering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.414214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:52.028232Z digest=sha256:c020c257c42cd50dfc4f84103eb1d645af96e09a866ad6dbe2b47c3e0743d9ab

Observation a1b04ad8-8725-496d-a7bc-e9b50f752036 · outbound

This paper cites Vggface2: A dataset for recognising faces across pose and age,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Vggface2: A dataset for recognising faces across pose and age,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.221017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:52.103593Z digest=sha256:7126501067976291d806e804a13a5cb68fb3b9ed255d148ef45a7c4b2cbe30ce

Observation 0398b027-d2a8-43c4-b82e-21d5e5777b9b · outbound

This paper cites Joint face detection and alignment using multitask cascaded convolutional networks,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint face detection and alignment using multitask cascaded convolutional networks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.040774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:52.210637Z digest=sha256:9ef36506aaf69d09d7a0f92b1dafd7ae587f0bd0ea8d725d084b7d5c4191e375

Observation 0859996d-b0fe-41e4-8386-aaebfc5501b1 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:52.286082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:52.286082Z digest=sha256:ed7ed471bb2b917fbcf6c7fcffa954c8243116ddefc54f55b6277f6100a0bb1a

Observation 5e0dcf08-eb5f-4c7c-b3d9-6e454ef07ed8 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis VoxCeleb2: Deep Speaker Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:52.418324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:52.418324Z digest=sha256:1ef7d5140e4ca8b8dca87a6b84ebe58bd92071724d045ac8b4c58b7854b961f8

Observation 385dc97e-106a-4cb1-b335-b4c575076666 · outbound

This paper cites Explor- ing the limits of transfer learning with a unified text-to-text transformer,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Explor- ing the limits of transfer learning with a unified text-to-text transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:52.825995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:43:52.529754Z digest=sha256:f027dcec99f9df59e017383b479df3969d16628de0a0bc08a44f48d1f0d9d370

Pith citing papers

No inbound Pith citation observations are available.