Pith. sign in

Paper Citation Record · LEDGER

XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2406.04904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04904 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:04.860178Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T18:15:21.439861Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b44d82e-5b4b-41e4-b5dc-22689577b166 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.058757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.058757Z digest=sha256:48073c9b9e1b728aef7d45965b3d07f905a740661b5bf127aa8cc987eb3f6cbc

Observation 001694fd-0561-48ab-b3b8-b86b4adfdc41 · inbound

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation cites this paper.

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:53:11.569946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:53:11.569946Z digest=sha256:aca3460d9ec76cfecf00d935259da9ef0ead0229715da4dec099e3b4b915b96c

Observation 381e9395-9d0b-40f2-93d5-f6b78ea70f34 · inbound

SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters cites this paper.

SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T05:43:00.592537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:43:00.592537Z digest=sha256:e9c1c80b6ebdfc365b78d259c56df4e270c9d7bccdd7fde1479b5e548da89466

Observation 99f0a98a-20ff-4fe6-ac2b-0712939f2c70 · inbound

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation cites this paper.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.512970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.512970Z digest=sha256:473de1e1736cf5b82f6746884cc716b2c93f498b588cdc48d00b84929dbd5f8f

Observation 9a4a5a56-897e-4fa7-b9fc-06b5aa821aa5 · inbound

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis cites this paper.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.663955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.663955Z digest=sha256:983f8cdf25655161590e997806a807bc012d206112dd4fbe3d1f3b6a3096e944

Observation 4addb919-8823-4a23-bd6d-f0e0d6cfc1ba · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.427049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.427049Z digest=sha256:533303b62f930b37951268eb92e403ddd6d4635a7c01ca19868c4eecc7ccb994

Observation 4b07bdf7-89e2-4fbc-bd3f-5ea2dc2abbf2 · inbound

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System cites this paper.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.496354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.496354Z digest=sha256:18efdc00acba876f06f024a6c6796823367aa9fbbe2c881d7a508f55a38f9a1c

Observation 90fc0f9c-b398-4c03-9fd7-39c5f559cfe7 · inbound

ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts cites this paper.

ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T18:27:07.057268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:27:07.057268Z digest=sha256:56b63b947846274ce4294c8f0f504db69e656039e654c176a780868275204bde

Observation f39e4809-08c4-4f34-9f4c-803701fbe69a · inbound

LoRP-TTS: Low-Rank Personalized Text-To-Speech cites this paper.

LoRP-TTS: Low-Rank Personalized Text-To-Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:21:44.159764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:21:44.159764Z digest=sha256:7af2ea10ac23d48e16389ccb79b1af573781e082326b4c61245b76a1eeee983b

Observation 44f878d8-1be9-4f38-b1e1-9f7b288489c2 · inbound

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder cites this paper.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.232012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.232012Z digest=sha256:00b98cc51ffc72f1b229edca2e4d3ceb69b46a3a7528e2f1af9aebcb9c738652

Observation fb412036-97d5-4785-9018-d1fe2474c66e · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.665380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.665380Z digest=sha256:f3573e01b109de15393002c2c1fbf05312d58d797106143c608092eeee21d7f2

Observation 0525e029-dec9-4ae9-b6b4-495de409f2fa · inbound

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning cites this paper.

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:01.090834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:01.090834Z digest=sha256:5a26e3b63c1f13a3c0e10f14b7729c6634b1d8bfcdb8598e717d1ada606e8503

Observation fac6a78c-a3f6-4d43-b8ba-3408509a47ad · inbound

Tell me Habibi, is it Real or Fake? cites this paper.

Tell me Habibi, is it Real or Fake? XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:14.471696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:14.471696Z digest=sha256:afb305d11c2eed4afd86f5c684aa666a014fe029e2fce70d55e551d0bf04bb76

Observation 7cf99313-2e9f-452c-ad98-1832a9f59aad · inbound

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation cites this paper.

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:25.084551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:25.084551Z digest=sha256:5c2f67ac98158f78f81ddc1ffeb3e9d8223380ba80c31e4d42538113725e84dc

Observation 571a9844-a593-4402-ac6d-3aebec911138 · inbound

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation cites this paper.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.213609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.213609Z digest=sha256:e2714553fade1f1f075e76bb6f63f6045651a747313297b5bea7ea918199a4a0

Observation 975c135b-5556-4155-bad1-0e7301bc06e9 · inbound

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation cites this paper.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.422722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.422722Z digest=sha256:0e9fca9098c2f47ded3f00b02488f79beb95895771b9bfd94bbb9b448fd90e7f

Observation 8d9afc30-17f2-4d57-80ad-22a25d830b56 · inbound

Speaking images. A novel framework for the automated self-description of artworks cites this paper.

Speaking images. A novel framework for the automated self-description of artworks XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:51.188095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:51.188095Z digest=sha256:855148a533a8167b202ff2df5fd2abc544c3673e381a2451a0cf07e8c065799c

Observation 6edb5495-89ea-4980-8374-716e9984d724 · inbound

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges cites this paper.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:50.893642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:50.893642Z digest=sha256:c96825e7ab55a4457e35d2702ce104b4b29968e43e310f4379a0b3fb1428eb9d

Observation edd82026-6223-44f0-bb48-c9e302ec281e · inbound

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech cites this paper.

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:00:06.463235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:00:06.463235Z digest=sha256:944c3e1b2d6cc52979979ca0146aac6ae03409e9ba5b589b4280ce6ede22b2f9

Observation 2b05cb50-b374-45c7-bd32-9380b858d038 · inbound

ClaritySpeech: Dementia Obfuscation in Speech cites this paper.

ClaritySpeech: Dementia Obfuscation in Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:31.756606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:31.756606Z digest=sha256:ff38ebbd18ae510c3743397413293d598de314c800f1cdb492ce3d4de6ce3d70

Observation 799f044e-036b-4f20-a99e-2041da5c1500 · inbound

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations cites this paper.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.583767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.583767Z digest=sha256:e48224d13c4d32323d7ef31bcfcd5f1d2465e90aec67ae11f08341480a64a480

Observation 3cd6b2ae-2341-4b02-845f-a85ed67504cc · inbound

AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations cites this paper.

AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:45:47.353222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:45:47.353222Z digest=sha256:25af51d2b9c8565b3db0b982f8562a33c52fbc758a02c8bbf136c7290454fdc3

Observation e7f00805-43b9-4391-b1f9-c4d742d2f610 · inbound

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation cites this paper.

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:17:38.311772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:17:38.311772Z digest=sha256:58d3f0595ee8a67f560ed54cafe5620dca9eea9282f5d83b87ab1691cac2c23d

Observation 34c86899-5cac-4b9d-88c5-b7d52daf0ed0 · inbound

KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features cites this paper.

KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T22:17:25.394751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:17:25.394751Z digest=sha256:67eccb09c5901a5406ca74598a827decad8963de3b1f9a0d88a31be93125cf68

Observation 669adc67-2243-4bdd-85d2-9bca6c09596b · inbound

VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks cites this paper.

VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T15:51:48.942690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:51:48.942690Z digest=sha256:d2b8855aa3771c61c9be34a827e63cc49ff415d44d64dfe761a1efa30b132102

Observation aa2e9eef-15e4-4f97-aadc-7ed3d18401ba · inbound

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis cites this paper.

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:17:45.506100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:17:45.506100Z digest=sha256:037bc757744c475b97237c05d15dd0b3f6aa09d585f605d15e21671dfe38c98b

Observation fc7db468-af1d-45d7-aa7a-683ac46d9380 · inbound

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection cites this paper.

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:02:23.378403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T05:01:57.044869Z digest=sha256:8ef7c4be59ec1b5e4bc13f9605a845ecfa907c608db7f9bb8ca757783f9e34fe

Observation f25c780d-31be-4780-9fc6-e4680d44d235 · inbound

PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing cites this paper.

PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:00:58.518478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:21:43.379525Z digest=sha256:956942042dd17d0ea76cd4ada06768d024b2bdc7e9a5119376842841ea53492e

Observation 1fe1ca6c-7490-4a4b-b22b-9ab94179b8ef · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.105245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:d95c2d90a484927204fab0c08743f209399b21cf78bf50c12ef587a2babf338f

Observation 00793784-0913-4347-8e8e-541819072528 · inbound

ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks cites this paper.

ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.593634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:31:12.802897Z digest=sha256:f2c36bff5aceef99d705d6a79656c0018d411318ba1e6e9ebaa57579f4e15130

Observation 33706689-396b-4f9d-8b80-a85ebcb489c5 · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:16.067515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T04:35:58.032597Z digest=sha256:f71c697ca6487210a771a736a655700b26959ffbe9422998554e71210524cd40

Observation 965f1998-4cd8-47e4-b757-8f6294a251cf · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:13.922407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T01:43:48.555523Z digest=sha256:2ef1a9e3301be2fa7259e7a2c98fc1986201670c96c45a8f7560e8bc81f30793

Observation 6cc82a9c-52ba-42da-a254-31a51a9486ed · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.766782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:fb897f43a03f1a6e32a3ef6b2222b4c8ca1924433157b789fc9ed634a3d74fc8

Observation 7269ea7c-ab9d-4626-aff0-12e5934a871a · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.138017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:3676d6309d5fd5e81b65b767db81aa2473122e6ae47e14cd75319518fba4c02c

Observation 441dfd27-e0cf-4caf-9bf7-99575e308264 · inbound

ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots cites this paper.

ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:58:57.489402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T01:05:20.432379Z digest=sha256:9960da7a30da9039799a179a0c521a5f72043670b38771cf47d0b7462d8cd212

Observation d5ee5b57-f32b-4296-ac7b-0c1f4e02ec94 · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.713811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:545438d5cec5dd7abfa4affde3bf681bf8b6bbd1d8f6d47f40f7e84d85bc1edf

Observation ee8dd6a8-4006-4da5-a60b-885fb9a5e925 · inbound

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean cites this paper.

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:50.257927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:24:03.268025Z digest=sha256:09285094c3166a1548ed5a5ee48c2ab437754d9f58cbd42c075d48fb4b6f4bc9

Observation b8ea5c51-58d3-474d-b08a-9fedb9e9d281 · inbound

Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection cites this paper.

Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:55:50.944908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T04:28:58.482098Z digest=sha256:efed76b44616866d968cae4bda849ad693272a01455b44a163d9d0730493535a

Observation 7458003e-2b4e-4e18-b7b1-5c7fa97dd716 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.186169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:45fd4cdc4b24224a6cf62f8a9eef29aab8422f31960a6e1e72105e82a8bf3948

Observation 469e6569-08bc-43f5-ba0c-b09e242ddf96 · inbound

Conversational Human Audio-visual Talking Dialogue Generation cites this paper.

Conversational Human Audio-visual Talking Dialogue Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T06:59:35.258176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:59:35.258176Z digest=sha256:8db1b8b2f20a70db6f601d9407bf53eb1f46dc3a702e9ff97615feb54d9669cb

Observation e7758f09-b868-4f3e-a232-1f9dd80f55e6 · inbound

BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech cites this paper.

BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-08T18:15:21.442023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T18:09:22.207379Z digest=sha256:fe41d90961ea4455fe9f74bb6a693d7d289d26d20cf86f6a555f5340aeb3c744

Observation 8f49aece-a048-4a0f-9b5b-bca81cdea929 · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T02:22:47.820537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:22:47.820537Z digest=sha256:a4054c556d9a9ec404b0739e1e84862786dcdb5ce81cd1c8923a4003d3d697b9

Observation f0814c3d-3be4-4616-810c-4572890cc750 · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T07:39:22.737607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:39:22.737607Z digest=sha256:9902ffb564c074c18ead271704e69fe9a8bec838f3cf97491ddad54b1c2b5b9e

Observation 598472a0-17c8-4bad-a142-30cdc42e8ed0 · inbound

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection cites this paper.

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T14:47:58.592181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:47:58.592181Z digest=sha256:259f242d0e74df595023148484e2ada9fee20225e3a319ec2a6a891ac8c5e3a5

Observation fd0d35ac-7503-41bd-bf97-9c3a0b5ea50a · inbound

Large Audio Language Models for Spoofing-Aware Speaker Verification cites this paper.

Large Audio Language Models for Spoofing-Aware Speaker Verification XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:12:29.763464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:12:29.763464Z digest=sha256:5a113fb7824dcc84d451e15f72a18faf5dadff924cbd69db823da6990327da00

Observation a8ba440a-6caa-404a-8a3f-2ea1a06ce577 · inbound

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech cites this paper.

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T15:22:08.936770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:22:08.936770Z digest=sha256:ba0712b39d746a9438876c59ca1f2c0629b01a955b6c60ecff1fc7f6a9616db3

Observation 605b69c7-9d09-4ff3-a2c6-dcfc54792be3 · inbound

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry cites this paper.

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:44:55.937047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:44:55.937047Z digest=sha256:3b9ef104a316593ad664a59018e45bb6da8c4c305fb22ff46ee9b010468d88d2

Observation 93d279d8-29c1-4fd5-afc8-1017556a9f8e · inbound

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks cites this paper.

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:21.155761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:50:21.155761Z digest=sha256:617058a84c551c88b04275554122c6e6de95b0303f9d521377b835dbe09081d5

Observation e0c1e541-27dc-4abf-b71e-723a4dae2e4a · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:42.302908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:42.302908Z digest=sha256:6b30329a7f119a35e775508316a026fa9a6cdc6cce14369e68de871aef191e99

Observation 168de320-8cbc-4adb-9c74-5a0875765363 · inbound

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder cites this paper.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.601698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.601698Z digest=sha256:bc4fbde6c04c4a9a4b90043225012bf0c1a18ed0000500686abbdaa94b677504

Observation 24f6b3a3-7e86-4b8a-bf33-72248c77c095 · inbound

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization cites this paper.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.860178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.860178Z digest=sha256:034d6c2ced2ee212d47880b098b6737e7658cdb2b81e9246a18114cb8034c88c