Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2406.04904.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:04.860178Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T18:15:21.439861Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3b44d82e-5b4b-41e4-b5dc-22689577b166 · inbound
WavChat: A Survey of Spoken Dialogue Models XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001694fd-0561-48ab-b3b8-b86b4adfdc41 · inbound
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381e9395-9d0b-40f2-93d5-f6b78ea70f34 · inbound
SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99f0a98a-20ff-4fe6-ac2b-0712939f2c70 · inbound
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4a5a56-897e-4fa7-b9fc-06b5aa821aa5 · inbound
Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4addb919-8823-4a23-bd6d-f0e0d6cfc1ba · inbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b07bdf7-89e2-4fbc-bd3f-5ea2dc2abbf2 · inbound
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90fc0f9c-b398-4c03-9fd7-39c5f559cfe7 · inbound
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39e4809-08c4-4f34-9f4c-803701fbe69a · inbound
LoRP-TTS: Low-Rank Personalized Text-To-Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f878d8-1be9-4f38-b1e1-9f7b288489c2 · inbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb412036-97d5-4785-9018-d1fe2474c66e · inbound
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0525e029-dec9-4ae9-b6b4-495de409f2fa · inbound
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac6a78c-a3f6-4d43-b8ba-3408509a47ad · inbound
Tell me Habibi, is it Real or Fake? XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cf99313-2e9f-452c-ad98-1832a9f59aad · inbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 571a9844-a593-4402-ac6d-3aebec911138 · inbound
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 975c135b-5556-4155-bad1-0e7301bc06e9 · inbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d9afc30-17f2-4d57-80ad-22a25d830b56 · inbound
Speaking images. A novel framework for the automated self-description of artworks XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6edb5495-89ea-4980-8374-716e9984d724 · inbound
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edd82026-6223-44f0-bb48-c9e302ec281e · inbound
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b05cb50-b374-45c7-bd32-9380b858d038 · inbound
ClaritySpeech: Dementia Obfuscation in Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 799f044e-036b-4f20-a99e-2041da5c1500 · inbound
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cd6b2ae-2341-4b02-845f-a85ed67504cc · inbound
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f00805-43b9-4391-b1f9-c4d742d2f610 · inbound
XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c86899-5cac-4b9d-88c5-b7d52daf0ed0 · inbound
KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669adc67-2243-4bdd-85d2-9bca6c09596b · inbound
VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2e9eef-15e4-4f97-aadc-7ed3d18401ba · inbound
ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc7db468-af1d-45d7-aa7a-683ac46d9380 · inbound
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f25c780d-31be-4780-9fc6-e4680d44d235 · inbound
PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1fe1ca6c-7490-4a4b-b22b-9ab94179b8ef · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 00793784-0913-4347-8e8e-541819072528 · inbound
ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 33706689-396b-4f9d-8b80-a85ebcb489c5 · inbound
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 965f1998-4cd8-47e4-b757-8f6294a251cf · inbound
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6cc82a9c-52ba-42da-a254-31a51a9486ed · inbound
VoxCPM2 Technical Report XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7269ea7c-ab9d-4626-aff0-12e5934a871a · inbound
FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 441dfd27-e0cf-4caf-9bf7-99575e308264 · inbound
ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d5ee5b57-f32b-4296-ac7b-0c1f4e02ec94 · inbound
An Evaluation Framework for Text-to-Speech Voice Reconstruction XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ee8dd6a8-4006-4da5-a60b-885fb9a5e925 · inbound
Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b8ea5c51-58d3-474d-b08a-9fedb9e9d281 · inbound
Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7458003e-2b4e-4e18-b7b1-5c7fa97dd716 · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 469e6569-08bc-43f5-ba0c-b09e242ddf96 · inbound
Conversational Human Audio-visual Talking Dialogue Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7758f09-b868-4f3e-a232-1f9dd80f55e6 · inbound
BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8f49aece-a048-4a0f-9b5b-bca81cdea929 · inbound
FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0814c3d-3be4-4616-810c-4572890cc750 · inbound
FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598472a0-17c8-4bad-a142-30cdc42e8ed0 · inbound
What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd0d35ac-7503-41bd-bf97-9c3a0b5ea50a · inbound
Large Audio Language Models for Spoofing-Aware Speaker Verification XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ba440a-6caa-404a-8a3f-2ea1a06ce577 · inbound
Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605b69c7-9d09-4ff3-a2c6-dcfc54792be3 · inbound
Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d279d8-29c1-4fd5-afc8-1017556a9f8e · inbound
AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0c1e541-27dc-4abf-b71e-723a4dae2e4a · inbound
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 152
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 168de320-8cbc-4adb-9c74-5a0875765363 · inbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f6b3a3-7e86-4b8a-bf33-72248c77c095 · inbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.