Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:22:00.140173Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 41 inbound Pith citation observations for arXiv:2506.21619.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:22:00.140173Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:50.544402Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T12:19:49.633371Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ba6482a5-3d62-479f-b595-93150211c630 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fbb62bb-fe07-419b-8184-9c3d8349b7d2 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech C.; Vidler, J.; and Roedig, U
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 857060fd-8691-4b3e-a815-b652d61d99d7 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55ae0086-9ecf-43fc-b66e-0443261bdd61 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech o lge, E.; G \
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed11d40d-2d66-4e37-8525-c8ea8db6105e · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech T.; Rubanova, Y.; Bettencourt, J.; and Duvenaud, D
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c44ba08-afc8-4dbf-94ea-b407c5446dd0 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e47c365-9195-4e19-8118-c14e329e8df8 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7f49650-edd8-42f0-9b75-ed6c67c51a73 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3389ed9b-60d1-4229-a949-802ace936876 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed3fd843-b8d4-447d-8fe2-389ea4370182 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 743b8f20-f8f9-484d-9729-16dcc0f97530 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9b7afa-6e45-4de3-965f-74d7e921dc57 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eef22172-612c-406c-8739-4ca04512428d · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109b26b9-bb32-4ca1-a88b-3a458c154038 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 248496b5-51e7-4c45-834d-fdca8bf8d666 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; and Wang, H
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68dc8c7f-92f3-436d-8588-5b356d0e5ff6 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; Liu, Y.; Zhao, S.; and Kanda, N
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff85a550-8a16-46de-9218-987d9ba51989 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ecb148e-6917-4155-b0bc-9b5bbfe8e867 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1112caf4-2c01-4233-bab7-13bf90d6c3d1 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f04ef65d-a581-47f3-8108-7d8574e282b2 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b94280-cdb3-4f31-a909-433f45771f95 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e314fc3a-0310-4ba0-888d-305ab9815519 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ff03d1-ba1d-4f90-b261-66173c4e3946 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a83a6da2-1f0f-4950-84ee-9026ec8c28eb · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e58924-eafc-479b-b71f-e7919cd052b9 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50297084-4b08-47df-8773-dffb35428c2b · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb32930b-3b78-44ad-baa0-fadb9101c5b4 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fe15248-fe58-4e00-8cc3-3c21292f24be · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76fed29e-242d-4dfa-b6fc-41b2f8fc9140 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55540a81-a7ae-44cd-91a1-fd5e4d506dc2 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec264d95-654f-440a-9c15-2b2f25920f8a · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b651a00-37be-422f-bbfc-7bbf58c16c60 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; Han, C.; Raghavan, V.; Mischler, G.; and Mesgarani, N
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2bdd39a2-b877-43b9-ba76-26e40715669f · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4926be1-b5a6-4c69-9098-2f81b5e375b4 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Zero-shot Voice Conversion with Diffusion Transformers
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db8cd0e8-bcf4-4dc7-9249-62ea0776d5a3 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7051c0b-f641-4363-940d-13858c01e34d · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Finite Scalar Quantization: VQ-VAE Made Simple
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25004936-e1d2-498e-a7ea-46875215afb1 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42046312-291d-4ca5-8cb5-79b7fa1cdd94 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca11a80b-9add-4a97-88fa-d01d3087f14f · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51dd94e3-a1a6-48c6-9188-206b659b9008 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ecc25f5-0aeb-4093-ba3e-7273c48a84ff · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4753148e-6bf7-473f-8d88-4f3c673be2e8 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6bf2331-1b89-4039-938e-dc3b51fde4bb · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; Gonzalez, J.; and Escalera, S
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6c00cfd-550f-4e87-9953-321632221f2c · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d58f3e58-36fc-4fd0-8dcd-579b40551c1b · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech E.; Hinton, G
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 214db0ad-301f-4718-9fc1-d18dda973ad7 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3098054-309b-4bc4-bf24-f287f3b7d03d · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a39cf59c-84a4-4e32-b488-bc1979a61e5b · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6357baba-1f9d-446b-aecb-937f7b65bf4e · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech LLaMA: Open and Efficient Foundation Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c60b0e3d-ffb6-432d-806c-cc2f525b1e2e · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b861883c-b8da-4abf-abeb-f1d59ec23138 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech N.; Kaiser, L.; and Polosukhin, I
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cf52851-2267-4464-ac50-5ebe5ba8e95d · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04e4ef9c-3e63-4826-b1d8-4ba7a760e217 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87488ee5-d784-4d51-ba47-6cc934239f83 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Qwen3 Technical Report
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3532d1ea-3948-4fe4-a0b4-111f73970903 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abde2652-5844-45b5-bc1b-6b14c8aafb41 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b14c852-8f34-4761-bb46-b3461b79d073 · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63838768-dc09-4026-8baa-6769dd2aae4b · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; and Li, H
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5da1b1c-023a-4ea3-99d9-63d9d35a2d6f · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6cb74769-28fd-4cfb-baad-e0f37447b91e · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech , " * write output.state after.block = add.period write newline
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c4a7b8d-536e-41d3-aa1d-07323ea46f4d · outbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech write newline
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf142752-2af7-45bc-9cb5-7725da8b0756 · inbound
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e574c656-9807-4a67-97f2-0ac729782f23 · inbound
JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bf6e419-1013-4bed-8f91-02ff709e7272 · inbound
Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 878d1998-5126-475c-b0c5-ebb94ffe3517 · inbound
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d1be585-afb8-4162-8a6d-24ba3001dc1f · inbound
Sharp spectral estimates for free boundary problems arising in plasma physics IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b01ef42-a07b-480a-a239-4229b954d18f · inbound
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ef942c6-ff61-45d8-b44d-fa72ec1c2ae4 · inbound
CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 713eefad-13f9-4a93-b473-955a704b9278 · inbound
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03223873-a2f7-4fdd-b941-df791230ca24 · inbound
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a395226-b98f-488a-bfcc-795612a049bd · inbound
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec3507fd-42e6-4b9f-af66-555e92b2e64e · inbound
RTCFake: Speech Deepfake Detection in Real-Time Communication IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 598b00b6-b804-4a36-8a51-1ca6d28b3aee · inbound
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ba17a93-84ad-456e-98e6-97359ecd5833 · inbound
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 235e690a-a3f6-4cf0-b49c-a4865a81eac1 · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe258340-112e-4730-990b-7bc43fc0e81f · inbound
AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7b3091e-e029-4632-9041-6273c134b334 · inbound
AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c36b7f05-03d9-4e50-a73a-d1e5dccec9b1 · inbound
DeepSlide: From Artifacts to Presentation Delivery IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdc9418e-d138-41ec-8f3e-14063841f124 · inbound
From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5dc0fd2-ca5e-4249-aa34-398776f16c9a · inbound
AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 579f3555-339b-4a1c-9417-fdf45cda185a · inbound
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8144dd9c-78d0-4466-b216-679c9b414ff1 · inbound
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2830a0fb-9f2d-4eea-aede-ebac1bee36ff · inbound
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22e6c8c1-4c35-4e15-8038-3119b51f83a9 · inbound
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6815c925-6871-483f-bd92-24d4edcbc89c · inbound
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16f5240f-3f96-4f1c-bfe6-8bd5849f9f77 · inbound
UniVocal: Unified Speech-Singing Code-Switching Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61fcceb1-562f-404d-837d-ba92ee5d31c3 · inbound
Resonant Minds: Closed-Loop Social Avatars with Theory of Mind IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0be20842-bbf1-47f6-9ecb-4f2e9eb7446c · inbound
Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc9f9742-5fa6-4233-a3fb-1f7dd5b91dd0 · inbound
TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8581762d-eee4-4332-9ce6-51353a517a4c · inbound
SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c83c0c63-cb90-42d8-827d-dff5dd9adeb3 · inbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a8912fd-aeef-48a7-b060-41486a59264a · inbound
EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3fcbdf0b-04d4-4188-a95c-e24b4faa8ff1 · inbound
An Evaluation Framework for Text-to-Speech Voice Reconstruction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8dde18c9-d7bb-4c47-b60f-d6c59ab96167 · inbound
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04d45292-d23e-45b6-b3ed-a2e00b37f448 · inbound
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997da372-83b8-4418-a395-df4cb3399f5d · inbound
How to Leverage Synthetic Speech for LLM-Based ASR Systems? IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b46fa76-577d-4a5a-aaf2-9b813acae9fd · inbound
How to Leverage Synthetic Speech for LLM-Based ASR Systems? IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf0a7c94-810c-486b-80c0-50d42fcc5bd4 · inbound
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28686fc6-7cc7-4693-887f-911c13a40eba · inbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a132f17-e405-42be-b8b7-27ec31c2d703 · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0cdb54b-a9a6-4ab2-9ced-30c4f7c6082c · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2a5988c-5344-4936-b561-59216c21ed4d · inbound
Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.