Pith. sign in

Paper Citation Record · LEDGER

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

As of 9 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 30 inbound Pith citation observations for arXiv:2502.05512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05512 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:07:00.588636Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:07:00.481427Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6530d5cc-7b61-4af1-ad91-7f92bfd767cd · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.481427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.481427Z digest=sha256:7277524b677be391bd686b035f4b1e9b826f87a6d51d8891085d9e94280a11df

Observation a670a59a-cc8b-4b05-946f-0b0e678666c0 · outbound

This paper cites [BT], prompt text, text, [ET], [BA], prompt audio, audio, [EA].

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System [BT], prompt text, text, [ET], [BA], prompt audio, audio, [EA]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:00.919750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:00.486856Z digest=sha256:92b25db66be78cda7cba5880f714a679017b4d5f0aa1e65e66c7e4c41d44ee20

Observation f09f3234-2912-49b8-9f20-b1e6f601f349 · outbound

This paper cites Dataset All training data was collected from the internet, with an initial 120,000 hours of raw audio.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Dataset All training data was collected from the internet, with an initial 120,000 hours of raw audio

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T19:07:00.905617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:00.491564Z digest=sha256:237ad6fac695e3bd81b9dca8ca3079f98195f36bb451abe77b23f151be191166

Observation 4b07bdf7-89e2-4fbc-bd3f-5ea2dc2abbf2 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.496354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.496354Z digest=sha256:3ef66f7644c5e36c6243315e2836608709489faf9f5a8bbdcc773aa9df317e59

Observation 56fb5224-3e15-411f-9401-d2b9b29b2629 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.501225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.501225Z digest=sha256:fab45e635d4ddef31f207178705aac33ef91366961a2681258f66b0fa38acd11

Observation 99cdb013-4c37-4c63-aff5-ef9260ebb7f3 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.505811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.505811Z digest=sha256:dc5542711d1d2956bb130099a6870151acc0e74321725b50d9582cea1ff4e55a

Observation 0fadcd10-c646-4158-be36-c0dddd35c6bb · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.510661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.510661Z digest=sha256:159464fae5c28233972f4da7a17b15f8b1b1d8510f5e7d15900ea8f70ebc10bb

Observation 5ad645aa-8a29-424e-865c-dd703843400a · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.515212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.515212Z digest=sha256:dae9ffd4cee775dac2c3703fe897584b16caa6f676fed9a0e2d57a79a8327eb6

Observation 9ca92841-67a1-4390-9ec0-f00df24db369 · outbound

This paper cites Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.520391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.520391Z digest=sha256:eebf79fc40008ef1889046155dbc23f4bf9d7a04681f909cb9f73f172263e4d5

Observation f8bede2c-be27-4dc2-85dd-36125b0de7dd · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.524666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.524666Z digest=sha256:3826d568a469b28d5cfe0eed2c67bcc6b2e962e8664ca711ceafbc9841b016dc

Observation 89e29809-4d7c-40d5-a481-5b6ae6b1667b · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.528669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.528669Z digest=sha256:6b3a1344a63448e5d410c8cc26dcc7aa85be78b9b43fa6c64dcfc08ec3540d4f

Observation 9e1ef766-eea4-4b45-86f0-abf236a47554 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.532973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.532973Z digest=sha256:4f17709e69a4e38b00837818b78105ebfa8afa65a8c573211e71431cabe98c35

Observation 54a85830-836c-4f55-ac96-7fdf24b2aad4 · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.537933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.537933Z digest=sha256:18eae344b8e658ccf81eeb18e6b6d597554a0a6957874b78ff5e1698f2a95c6e

Observation ecf8f860-9809-4c0e-8dfd-c0805f94c1a2 · outbound

This paper cites Better speech synthesis through scaling.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Better speech synthesis through scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.542381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.542381Z digest=sha256:d38ebb6be2214164eafa8e79778591a203d99e8abdd1c99b0fea901585c85a61

Observation 98c78b57-5ca5-488d-8648-a662f0e598d4 · outbound

This paper cites Neural discrete represen- tation learning,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Neural discrete represen- tation learning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.546792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.546792Z digest=sha256:1a534bc0c2fade3d1ef16488ee9cf2b6e123c40c6b01c42cf1d218dc206899ea

Observation 21ab49b3-9225-4430-bf1b-39e4d32a42d1 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Finite Scalar Quantization: VQ-VAE Made Simple

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.550792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.550792Z digest=sha256:c134dc4956f332a7081fd328c32222bd8fc0273868220b1aaddc5b6d3ea6ec47

Observation 579bf133-2af6-4ed5-8767-e6b373bc5b1c · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.555072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.555072Z digest=sha256:510a8e02b73cf235c945745f3fc1b3225ee500924868ff7caf43c28c44ef5ce8

Observation 8b9e2605-4eb6-4c05-9646-c89d1cd2044c · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow match- ing,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Matcha-tts: A fast tts architecture with conditional flow match- ing,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.559444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.559444Z digest=sha256:40ed64c496b39b1f13b864831b67605550c8f16e2b21345bd44da4f01c995832

Observation 906e5a61-7860-47a1-bd52-349f57a0e54e · outbound

This paper cites Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.563570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.563570Z digest=sha256:1fe60aa0b157754495481552b9c645804201f4a51cee3da5b8cd7b13f29ca51a

Observation 4b29838f-8228-4c82-b65f-c4478b6af569 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Lib- rispeech: an asr corpus based on public domain audio books,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.567811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.567811Z digest=sha256:ea684b8cd95b5da29c151fa36eefcabc6ed3d185cd4fa4d1af4707298b4c8d5f

Observation 51618b6c-1f70-47c8-9318-c6eeb7266fcc · outbound

This paper cites Aishell-1: An open- source mandarin speech corpus and a speech recognition base- line,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Aishell-1: An open- source mandarin speech corpus and a speech recognition base- line,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.571915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.571915Z digest=sha256:30382bc7b6f0d06e4b996998dd8028cb55d15aa73414c0039bafffb32f639449

Observation 278b1ffd-223a-4ed5-b5d0-e5c554b3a460 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Common Voice: A Massively-Multilingual Speech Corpus

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.575954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.575954Z digest=sha256:6a57539940871e73aa683615d2420cd5e3a6fbe786f1d51c7c7d3cfcbb35fd43

Observation 97b06b7c-850a-4ebe-9f0e-e9a21b7871f5 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.580468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.580468Z digest=sha256:76e49674598f61df0169975ccb3467dd8c3a049b0e872cd67771b91f0134a5d0

Observation 27a95564-2160-45b0-8af3-8ffe9a76837a · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Robust speech recognition via large-scale weak supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.584577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.584577Z digest=sha256:01bd9182b1dc8df0094b330113381054b456a9ed0e42929e811f8d84643ceba6

Observation 56c59340-6fa5-4555-87f6-1aabf4fc68a6 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.588636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.588636Z digest=sha256:290128b06131daf697d010a961717f0040ace3326acdb3a5859c60964203777b

Pith citing papers

Observation 6530d5cc-7b61-4af1-ad91-7f92bfd767cd · inbound

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System cites this paper.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.481427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.481427Z digest=sha256:7277524b677be391bd686b035f4b1e9b826f87a6d51d8891085d9e94280a11df

Observation cceffa41-b9b1-4038-bf68-672cd9a46730 · inbound

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information cites this paper.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.979998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.979998Z digest=sha256:f2f45c612718c073de85dbf5f8827af7d00236f8b6cda27d4f5980a114ae0c5d

Observation ca5794d7-3d49-4f58-bef4-4742bc912a27 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.569998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:a56702885b2f66e3325b421562980195e387e554730eae9068e7e0618ba935a8

Observation 03ab161a-d36d-4a93-9e79-0c4f7662f124 · inbound

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning cites this paper.

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:01.135750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:01.135750Z digest=sha256:54274765ba8ecb4dd26e28f5a7f84e6d03ab98c0a44bcdd9e1393cc8b418344b

Observation 743b8f20-f8f9-484d-9729-16dcc0f97530 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.901086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.901086Z digest=sha256:92fc2cfd7738744d9ba770a9d54f9a3a0a78b8bf85bb32dca58ccda765676f6f

Observation 6df3652a-8560-48d3-9cff-ff5f41003c0b · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:32:03.696915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:9d87570cc3f7152885aa0cc0174d4f5149aad31c59fd71f2dd294316bb744367

Observation b005b13d-8d31-46e1-aaf2-2a23300bdd76 · inbound

Robust Residual Finite Scalar Quantization for Neural Compression cites this paper.

Robust Residual Finite Scalar Quantization for Neural Compression IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T18:21:55.877250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:21:55.877250Z digest=sha256:02da290e41de191c40225274989f5645265d7c0a7a26138de5c1a8fc046f2cc4

Observation f6b11d61-fc62-4741-9e45-9c2fd67d14da · inbound

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot cites this paper.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.299724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.299724Z digest=sha256:a8248cd5b9befa074f5fc570f7c1339396b381602dfede6a7ce8d1acaf02d7ba

Observation 0ee0c70b-a7db-4912-9cc0-47f9039061f6 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.502401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.502401Z digest=sha256:de3040e3f84e516fe907d8e477d8b60fea9399dcc7f690ab830db796333df6d4

Observation c0d2a60b-8c49-41a8-b9ca-d0d7f9072168 · inbound

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection cites this paper.

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:02:23.325753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:01:57.044869Z digest=sha256:e547e03418f4b4077f3a500f9e50b486face663a1e7294a95c52add7e45aaa6e

Observation a6cc9382-8537-4204-9c75-ad97e33a7f16 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.762622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:a8b1fbf12a0e935c8131b56b4f72b99fdee20a1d5bf7f0bcd144eadee0f75bd5

Observation 9b8586ec-f497-4039-87a1-671dcfc07bf8 · inbound

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan cites this paper.

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.716219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:40:10.657590Z digest=sha256:e0708f3eea5766c82f08c9067b607e95efe4254c4b7d9f148aeccec9a460eea0

Observation 7f9524ac-1a88-4870-b2e1-c130847c7aeb · inbound

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models cites this paper.

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:29:55.951508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T10:29:51.313835Z digest=sha256:63d8bd0487d84be70c6aaeb024f70e42b28877e5df377414230d7b61053ed015

Observation 066c6c75-1c36-4075-ab14-f14e9810b178 · inbound

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition cites this paper.

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:07.162050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:22:08.670559Z digest=sha256:00dcad52e710e628aeab30760bd45c2d9cfbe242034a2d37fb9977fe2a8e4671

Observation 32632292-6130-4474-9201-86b204be205b · inbound

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing cites this paper.

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:02.483890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:03:15.572657Z digest=sha256:4903808fe8356663ea4ffca713402d8e6e20f9c79b001d899dc851e9f46c109c

Observation 9253335f-202f-4fdd-ba32-02f22861ae7f · inbound

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing cites this paper.

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.232303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:59:34.622096Z digest=sha256:4812d6e957b864f6cf7c24de7bd1907a071015178737a0301e956cbce821e7b8

Observation 6ede54ab-6ac8-4d72-911d-eb83bc9a9429 · inbound

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing cites this paper.

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:29:37.697055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:29:37.697055Z digest=sha256:7afbab57ed74ef03456203cc3ad6afecbb1958a2a146fcb602ba716153a0f530

Observation 5c0fc56d-4fe0-47e3-a0b1-fb609c5fe0c9 · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:04:47.235530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:710713b5deb2a2f803696b739ccee959776d3fc1b4b4485f2fc35c69cd4b6460

Observation a21a2f5c-3be3-4b7a-b59b-38e9beef61e2 · inbound

RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System cites this paper.

RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:10.281890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T19:59:00.255786Z digest=sha256:f2ae0d7d48eb1be71ed5e3a947290f7b14f50700a7023986ebac50c718680b5a

Observation 7f83fb1f-d3f9-4bf7-a807-81baadf1f855 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.600587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:b94a1e89942c7ab0f0127485b167fad4c041c3e6a7549050d026dc9b13d521e4

Observation 3cc031fc-1a1c-4bbb-be1f-c66b39d9a931 · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.270673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:2002061b4545c516794121ed3e89001a3704023096bfe1828f532c5b788ed8ee

Observation ac48d7d0-7235-4016-ac85-5d345ec6882a · inbound

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis cites this paper.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.890163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:d5b434815346def9c41c0a6033ace1e030dbbe632d762d83b52e2242135a406a

Observation c8a2368c-09d1-4a4a-85c6-16870403ee90 · inbound

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation cites this paper.

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.683613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:30:11.718647Z digest=sha256:fd2ca4bf306ccf3aa2561d7e308392cf2153c651ec4d292a97755cc60c667192

Observation 09e111f7-6c4a-4bf2-a6a8-1ced99766cb2 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.601926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:a7ccbf073d3583bd7cadedf99f1f5d6b44babd965f6e3bd9cf39e98e55dd0664

Observation 8380a461-bd96-48c6-baaa-409dd38bda18 · inbound

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy cites this paper.

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:06:53.393587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T04:37:26.452974Z digest=sha256:aa5fbd5b40ebe4dee3cd57723104637b37cb901defb6e9fad31ba22ed8691064

Observation f4b6d5e9-3a04-4460-bec4-761851ae30b0 · inbound

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction cites this paper.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.949342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:9b227f7dab898da15cc99bd918747a0e52b1a164ae2a75550e42c503b3943d37

Observation 0e620d67-2f50-448b-8493-788856d9f13d · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.194761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:414b289da81fba8a3d3a3c2b076f4ab00ab5b947fc7dbe20abb1d9c648b7e8f2

Observation 472c3733-0573-4870-bc53-b39520c86dba · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 137

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.283406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:fae3ca9fd604f9b99eddc3fbfdebfe03aee2d51ea9764d4a9502837674501eb0

Observation 82b3df18-4735-4230-91db-c7ca95c4d4ea · inbound

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling cites this paper.

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T06:27:04.866415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:27:04.866415Z digest=sha256:660396a05513978389c1d0261682b81137d65b1e67efccb411c3b129785785f3

Observation 148d8915-cd7c-42ea-b5af-27826fc4e2d2 · inbound

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System cites this paper.

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:44.391170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:44.391170Z digest=sha256:865a3fd0b0ca2bf53c2e0c4b8d385f75f2c351734e58bd6b8fa08f220470944b