Pith. sign in

Paper Citation Record · LEDGER

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.20945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20945 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:52.529754Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27adb6b3-b24d-4ae4-8078-4cbcf618dc2c · outbound

This paper cites Mega- tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Mega- tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.933915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.228144Z digest=sha256:88be96c07cd03957fb8f39011626084fe203330ed5d74b6eb47c5482a62b2c8e

Observation 49a3a70a-a4b9-4624-8064-993c0af597aa · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.265834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.265834Z digest=sha256:4ce146645973876e1fe04bf1b800512d4fa56c6bc960d13436d1078238869dd7

Observation 663c1c04-c3e4-47e7-9303-e95d307ed5f0 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.332197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.332197Z digest=sha256:9d4cd522f955d1cdb6b2f3938a0cc19b7e8d935d3aca060c1784ec0e33796ad7

Observation 80bebff5-0883-408a-bb4e-bb46906172b1 · outbound

This paper cites Imaginary voice: Face-styled diffusion model for text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Imaginary voice: Face-styled diffusion model for text-to-speech,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.768449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.379308Z digest=sha256:0f3b34d0c811f6096fa880c540d0e530cec2d32be65f279c43406cbbef09a05e

Observation 3bae3fb5-4477-40bf-a466-d1206948ef89 · outbound

This paper cites SYNTHE-SEES: Face based text-to-speech for virtual speaker,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis SYNTHE-SEES: Face based text-to-speech for virtual speaker,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.560254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.454168Z digest=sha256:b398f33c028a4a599330a491bbdcb37e5e3345032a454842eff80e9f146edf8f

Observation 921cc365-a7bf-4ae9-aafd-b477a96c2b58 · outbound

This paper cites Face2Speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Face2Speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.395195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.528613Z digest=sha256:9c8c56213d2c9518b6b89addd00f8e7d4d7d29340ff8d5c480554c87176101b0

Observation 3cf14a4c-5247-4d3f-a3fd-edced867f33d · outbound

This paper cites FVTTS : Face based voice synthesis for text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis FVTTS : Face based voice synthesis for text-to-speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.239168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.580539Z digest=sha256:4bdda66d9b293ad5da18726e01e0df790f60135d80dd3dfcbfff05aa94ddb6c6

Observation a53c088a-a93b-47ee-9a23-c88649a88f76 · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.088342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.633123Z digest=sha256:7cdfd975805b9d40f42f0d254174389ba5609f7769bb1a2c758a6a6aa7bf5a5c

Observation ffd35756-5c7b-4282-9d92-0eb646be2e2d · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Prompttts: Controllable text-to-speech with text descriptions,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.942659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.701372Z digest=sha256:73075bc06f72e62028f7ccf7c64e5c0770ff353bef0904d15b53e368c1ad414c

Observation 3556fb19-c8c4-4690-a30d-2355895035fb · outbound

This paper cites PromptTTS++: Controlling speaker identity in prompt-based text-to- speech using natural language descriptions,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis PromptTTS++: Controlling speaker identity in prompt-based text-to- speech using natural language descriptions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.781073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.766499Z digest=sha256:f01bea535110be4c14f2ef2ed78a466f2d761ae56e38a3453829c214e550105a

Observation 96b659af-c3d2-464f-83f5-8ed7af6292f7 · outbound

This paper cites UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.830879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.830879Z digest=sha256:68f76ec268cc98be16da756467bf09af318e4a2cc467d5e2136fff0c610b85dd

Observation 6de8182b-399d-45b8-bde7-dc17676d2a97 · outbound

This paper cites MM-TTS: Multi- modal prompt based style transfer for expressive text-to-speech synthe- sis,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MM-TTS: Multi- modal prompt based style transfer for expressive text-to-speech synthe- sis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.631249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.902999Z digest=sha256:9a173fbf29bdd0626eefe20acad3df882acdc86c4324c8d6f4c71d81d061bcf6

Observation 9b3f1fe2-a79f-49e2-ac7e-39310830e03f · outbound

This paper cites Gen- eralized end-to-end loss for speaker verification,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Gen- eralized end-to-end loss for speaker verification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.479581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:50.984883Z digest=sha256:d7a1cf8e1c98d6f6b79bb6c22d485975f028c4c9e76b6e3dc37af989718c8888

Observation 42cd64a7-6d95-4511-bf4f-adf0c4a098f7 · outbound

This paper cites Additive margin softmax for face verification,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Additive margin softmax for face verification,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.350835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:51.073749Z digest=sha256:098a5d54151ee462dde2071fb9bd3b180a1c942cccff30267f5bf0cfaa2cab85

Observation 513fb0e2-5d0e-4c9d-980a-02ad682a3cba · outbound

This paper cites Bridging the gap between object and image-level representations for open-vocabulary detection,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Bridging the gap between object and image-level representations for open-vocabulary detection,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.148216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:51.143577Z digest=sha256:872ac0d6668a575ee7031f30100619b9ef9af023cf0d895cb4af2e277eb43f2c

Observation 1a202008-871a-415c-9415-9faaffdcd3d8 · outbound

This paper cites Joint- teaching: Learning to refine knowledge for resource-constrained un- supervised cross-modal retrieval,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint- teaching: Learning to refine knowledge for resource-constrained un- supervised cross-modal retrieval,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.002813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:51.255727Z digest=sha256:1a663a696391ef263518a4aa7b4bfe39f9d303754dc08d87776e37b479a62a58

Observation 8de811c5-5e1f-49dc-865e-18f221ead62d · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Representation Learning with Contrastive Predictive Coding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.359236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.359236Z digest=sha256:52248ec787545cf6d298226701709afc5d7aaa563b8920f823fa491a0a986256

Observation 283108dd-26b0-4501-ae07-6e229d58e159 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.818807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:51.427727Z digest=sha256:45855684cec4bf895d0fa64402841f331673d51bf9a797eb4de4c9077234d77d

Observation a29d89ed-059f-4cef-a988-bf5424033794 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LRS3-TED: a large-scale dataset for visual speech recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.542194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.542194Z digest=sha256:2a314f2c5fb03ce1a3ba2f331c54a16a7510e00653f5822928a95ad7f2a8f200

Observation 4f76282e-dc9d-4e21-a18a-023dd6db6970 · outbound

This paper cites Multi- caption text-to-face synthesis: Dataset and algorithm,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Multi- caption text-to-face synthesis: Dataset and algorithm,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.588907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:51.614760Z digest=sha256:1ff7b12c377ca296c2177ed02362d4e4d8a9e3d996fd4a12d476e6906b51df94

Observation 5eb4f155-d80c-43eb-b966-9739821a71fa · outbound

This paper cites LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.707283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.707283Z digest=sha256:91dff0c1f686464f67241efcfe515f07a1650542b0a486b74bd9eae573f0618c

Observation d4b436d8-9b4e-4335-b459-f4b37136a926 · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.837350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.837350Z digest=sha256:d1b7507c54a4da3d1099dcf41d77b4102d70c09948ec03b49952f3475bb027d6

Observation 1eb94adb-cfed-44c4-8ab4-2efea73e8fff · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.942109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.942109Z digest=sha256:ce6e0ca9c6cd66cc578d787df7004ef1b14fde00a1f34d94e178a2c683d84faf

Observation 30bb92a6-34a9-4db0-8670-fc1c27a258f0 · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Facenet: A unified embedding for face recognition and clustering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.414214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:52.028232Z digest=sha256:49418bc2eb779eb3cd44d0177635e6ce558a78879178508ff2c0cb349ee53208

Observation a1b04ad8-8725-496d-a7bc-e9b50f752036 · outbound

This paper cites Vggface2: A dataset for recognising faces across pose and age,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Vggface2: A dataset for recognising faces across pose and age,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.221017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:52.103593Z digest=sha256:5265a8e9d8d4c692a231f25ddbb17135efc0114808dec28c758fee86e747db44

Observation 0398b027-d2a8-43c4-b82e-21d5e5777b9b · outbound

This paper cites Joint face detection and alignment using multitask cascaded convolutional networks,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint face detection and alignment using multitask cascaded convolutional networks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.040774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:52.210637Z digest=sha256:cc233e77f8523fd66971fdff5c1cee077a76b95bac4cfadfb38f8d11e7aae0b9

Observation 0859996d-b0fe-41e4-8386-aaebfc5501b1 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:52.286082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:52.286082Z digest=sha256:c919093db6fe904b5dd778115ccdecb545eaf200c9acdc8f127e7a3834f2490b

Observation 5e0dcf08-eb5f-4c7c-b3d9-6e454ef07ed8 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis VoxCeleb2: Deep Speaker Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:52.418324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:52.418324Z digest=sha256:c496edffb52833cf409f7cce3547734063a70d988e7df3c172475e78417cbc89

Observation 385dc97e-106a-4cb1-b335-b4c575076666 · outbound

This paper cites Explor- ing the limits of transfer learning with a unified text-to-text transformer,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Explor- ing the limits of transfer learning with a unified text-to-text transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:52.825995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:43:52.529754Z digest=sha256:b912de0f687cf018c73cacbeb8881021f19e4023801ac5c381d81730e7d1efa9

Pith citing papers

No inbound Pith citation observations are available.