Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:52.529754Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.20945.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:52.529754Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 27adb6b3-b24d-4ae4-8078-4cbcf618dc2c · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Mega- tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49a3a70a-a4b9-4624-8064-993c0af597aa · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 663c1c04-c3e4-47e7-9303-e95d307ed5f0 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80bebff5-0883-408a-bb4e-bb46906172b1 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Imaginary voice: Face-styled diffusion model for text-to-speech,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bae3fb5-4477-40bf-a466-d1206948ef89 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis SYNTHE-SEES: Face based text-to-speech for virtual speaker,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 921cc365-a7bf-4ae9-aafd-b477a96c2b58 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Face2Speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cf14a4c-5247-4d3f-a3fd-edced867f33d · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis FVTTS : Face based voice synthesis for text-to-speech,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a53c088a-a93b-47ee-9a23-c88649a88f76 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffd35756-5c7b-4282-9d92-0eb646be2e2d · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Prompttts: Controllable text-to-speech with text descriptions,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3556fb19-c8c4-4690-a30d-2355895035fb · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis PromptTTS++: Controlling speaker identity in prompt-based text-to- speech using natural language descriptions,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96b659af-c3d2-464f-83f5-8ed7af6292f7 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de8182b-399d-45b8-bde7-dc17676d2a97 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MM-TTS: Multi- modal prompt based style transfer for expressive text-to-speech synthe- sis,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b3f1fe2-a79f-49e2-ac7e-39310830e03f · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Gen- eralized end-to-end loss for speaker verification,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42cd64a7-6d95-4511-bf4f-adf0c4a098f7 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Additive margin softmax for face verification,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 513fb0e2-5d0e-4c9d-980a-02ad682a3cba · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Bridging the gap between object and image-level representations for open-vocabulary detection,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a202008-871a-415c-9415-9faaffdcd3d8 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint- teaching: Learning to refine knowledge for resource-constrained un- supervised cross-modal retrieval,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8de811c5-5e1f-49dc-865e-18f221ead62d · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Representation Learning with Contrastive Predictive Coding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283108dd-26b0-4501-ae07-6e229d58e159 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a29d89ed-059f-4cef-a988-bf5424033794 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LRS3-TED: a large-scale dataset for visual speech recognition
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f76282e-dc9d-4e21-a18a-023dd6db6970 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Multi- caption text-to-face synthesis: Dataset and algorithm,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5eb4f155-d80c-43eb-b966-9739821a71fa · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b436d8-9b4e-4335-b459-f4b37136a926 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eb94adb-cfed-44c4-8ab4-2efea73e8fff · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30bb92a6-34a9-4db0-8670-fc1c27a258f0 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Facenet: A unified embedding for face recognition and clustering,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1b04ad8-8725-496d-a7bc-e9b50f752036 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Vggface2: A dataset for recognising faces across pose and age,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0398b027-d2a8-43c4-b82e-21d5e5777b9b · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint face detection and alignment using multitask cascaded convolutional networks,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0859996d-b0fe-41e4-8386-aaebfc5501b1 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e0dcf08-eb5f-4c7c-b3d9-6e454ef07ed8 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis VoxCeleb2: Deep Speaker Recognition
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 385dc97e-106a-4cb1-b335-b4c575076666 · outbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Explor- ing the limits of transfer learning with a unified text-to-text transformer,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.