Pith. sign in

Paper Citation Record · LEDGER

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 41 inbound Pith citation observations for arXiv:2506.21619.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21619 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:22:00.140173Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:50.544402Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:19:49.633371Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba6482a5-3d62-479f-b595-93150211c630 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:51.474748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:51.474748Z digest=sha256:f0e0baf1064d18ebdd8b189087cf89c7d1e44f77b326b68e14637174c2b185c2

Observation 1fbb62bb-fe07-419b-8184-9c3d8349b7d2 · outbound

This paper cites C.; Vidler, J.; and Roedig, U.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech C.; Vidler, J.; and Roedig, U

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:09.568747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:51.714747Z digest=sha256:afac6f3b5a119b13489a6dd67096235b39e2224432e740aa1abd1b037d891a5c

Observation 857060fd-8691-4b3e-a815-b652d61d99d7 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:09.343035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:51.891560Z digest=sha256:d581b06e4623b38030ce7a175af78356bbf859ba4b12163069a0e53936cb3386

Observation 55ae0086-9ecf-43fc-b66e-0443261bdd61 · outbound

This paper cites o lge, E.; G \.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech o lge, E.; G \

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:09.141449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:51.984742Z digest=sha256:94b110312de57fab48301072bcf89250ebf25c724f8b51bc309a6e21d0b173f6

Observation ed11d40d-2d66-4e37-8525-c8ea8db6105e · outbound

This paper cites T.; Rubanova, Y.; Bettencourt, J.; and Duvenaud, D.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech T.; Rubanova, Y.; Bettencourt, J.; and Duvenaud, D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:08.907713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.090822Z digest=sha256:cf689bda9af2cddfd1efbd3c04e27dc96496c7b1e3fad6f132940eac9e62196f

Observation 7c44ba08-afc8-4dbf-94ea-b407c5446dd0 · outbound

This paper cites Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.191734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.191734Z digest=sha256:b74b50d773cff2c2d94196713532d237233c55c35afb844fd02de9ca598110a1

Observation 1e47c365-9195-4e19-8118-c14e329e8df8 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:08.664334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.374989Z digest=sha256:36755e36eb28995482c4f916841b8c27561e84e4b7c0bd9def0f8c05e635b3ac

Observation c7f49650-edd8-42f0-9b75-ed6c67c51a73 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.491309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.491309Z digest=sha256:0fe878a7ccca86efe49f3a7bb653a9f9a82552456096c474cac48adbfced66cf

Observation 3389ed9b-60d1-4229-a949-802ace936876 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:08.443414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.584839Z digest=sha256:725e896bd40bd97acd5398a1470791e04334087634ab6897bdd1935210c8da54

Observation ed3fd843-b8d4-447d-8fe2-389ea4370182 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:08.224525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.694748Z digest=sha256:938440c526ff5b9f27b5e271111939a74cb78b43e66349fa1bedfa81f29489d3

Observation 743b8f20-f8f9-484d-9729-16dcc0f97530 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.901086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.901086Z digest=sha256:a467de33f4ebe23396699e2d91e822855dd34eac1df7eca321826fc30263d9dc

Observation 2e9b7afa-6e45-4de3-965f-74d7e921dc57 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:07.977427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.044748Z digest=sha256:53a7c4aed40c1255cee734eab811e705d5f0cdce25cabad6d97de39216505c88

Observation eef22172-612c-406c-8739-4ca04512428d · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:53.174876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:53.174876Z digest=sha256:bd75a6f8d0febc8b99f30105ce46f21053c9e0b066175d32e326b833b266ab8c

Observation 109b26b9-bb32-4ca1-a88b-3a458c154038 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:53.286323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:53.286323Z digest=sha256:221af78d6c564f53cbf8a174f2091b98c2f9c3932802a7c0e1eeb8cf71bc4097

Observation 248496b5-51e7-4c45-834d-fdca8bf8d666 · outbound

This paper cites A.; and Wang, H.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; and Wang, H

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:07.662270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.404945Z digest=sha256:e0bf7097bc7bd2b703d7767f59d5ddc41793e4c68f8e061eea245c2189151bbb

Observation 68dc8c7f-92f3-436d-8588-5b356d0e5ff6 · outbound

This paper cites E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; Liu, Y.; Zhao, S.; and Kanda, N.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; Liu, Y.; Zhao, S.; and Kanda, N

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:07.397033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.514918Z digest=sha256:b6497e77fb7e10014e65f5089ef1f08bb5092b990e3624b8efce0a44956210fb

Observation ff85a550-8a16-46de-9218-987d9ba51989 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:07.073937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.654824Z digest=sha256:04e01726d3a0ba0106c0b1f9930e96d51d45090f60185551d7365f730c5d69f0

Observation 3ecb148e-6917-4155-b0bc-9b5bbfe8e867 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:06.782697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.794860Z digest=sha256:8906c29a62f0e5df160e25776326ec2a359ac1e9720e5ffe7c7d025ba181b434

Observation 1112caf4-2c01-4233-bab7-13bf90d6c3d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:53.970330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:53.970330Z digest=sha256:0f457c9a1d970c6c8592f7a43e1d31b82a132496c8c768d43a5789aa12baa7d0

Observation f04ef65d-a581-47f3-8108-7d8574e282b2 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:54.069634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:54.069634Z digest=sha256:0b5af87cd4a1e6ce39092c8421b7498a76ec15d6357837ef13f54adbc7c03f64

Observation c8b94280-cdb3-4f31-a909-433f45771f95 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:06.449439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:54.324773Z digest=sha256:6462ba92e29f268987ceb5c86beec66e2c65639e61d5145baf14f7efd069262e

Observation e314fc3a-0310-4ba0-888d-305ab9815519 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:54.509932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:54.509932Z digest=sha256:6e336407676e531a8529ea8ae1b2b87e46c39272eff4fe2eb2f4f5c5e18c8dd6

Observation 15ff03d1-ba1d-4f90-b261-66173c4e3946 · outbound

This paper cites ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:22:01.640762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:54.594770Z digest=sha256:197054a285ffffa1a19294c75aa90f9fb66d2aa3f70f25db622f0979b8f8f2c8

Observation a83a6da2-1f0f-4950-84ee-9026ec8c28eb · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.093152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.093152Z digest=sha256:84aa00b5b9378103acf6eaff42c841d3ef84d0faffb906fdb3da39854bd11fac

Observation a2e58924-eafc-479b-b71f-e7919cd052b9 · outbound

This paper cites SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.224747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.224747Z digest=sha256:601a2444a1d6ef6a4f65629ecb9582e725c03a494231c4db80de7a78be41fe8b

Observation 50297084-4b08-47df-8773-dffb35428c2b · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:06.135488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:55.490696Z digest=sha256:7b5e9f2f18ec57df6e15b0469df5b3458baf98a05104341b6c210f6778b84a8a

Observation eb32930b-3b78-44ad-baa0-fadb9101c5b4 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:05.968350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:55.624763Z digest=sha256:cf00714617cefec2d7321cabe9dd072bc766b2094083dc5cf792545b20585900

Observation 0fe15248-fe58-4e00-8cc3-3c21292f24be · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.719910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.719910Z digest=sha256:5ff6d398bc3e64987a64804f87a6888a55ae9bcf65e76312d6a31e5e623bcef2

Observation 76fed29e-242d-4dfa-b6fc-41b2f8fc9140 · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.843335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.843335Z digest=sha256:304aed2c5449d2465bd1d3fc24eba30385c73b47562204b95e7507a281b466d2

Observation 55540a81-a7ae-44cd-91a1-fd5e4d506dc2 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:05.796837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:55.959394Z digest=sha256:38d44099ec2d2cb33166806dfcfa15cda113b5a2dc1f976f54a14f6837e0e65f

Observation ec264d95-654f-440a-9c15-2b2f25920f8a · outbound

This paper cites FleSpeech: Flexibly Controllable Speech Generation with Various Prompts.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.065145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.065145Z digest=sha256:b60b001a8fb65dc80a06879cf99725ec85420f9732a2d9fa4c4ee94caaa7d63b

Observation 2b651a00-37be-422f-bbfc-7bbf58c16c60 · outbound

This paper cites A.; Han, C.; Raghavan, V.; Mischler, G.; and Mesgarani, N.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; Han, C.; Raghavan, V.; Mischler, G.; and Mesgarani, N

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:05.560496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:56.158020Z digest=sha256:b3902057e5902fc5e73aac6651aee52baa0cdd351910a8a850d9b168954a9eff

Observation 2bdd39a2-b877-43b9-ba76-26e40715669f · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:05.244739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:56.230762Z digest=sha256:6b7b720023c839e002cbc4c085ff6644d52425195f2221de3b2e750951986d88

Observation d4926be1-b5a6-4c69-9098-2f81b5e375b4 · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Zero-shot Voice Conversion with Diffusion Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.314992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.314992Z digest=sha256:cc90542fb12f49cc808bc94c9623e030a5dcc96c7d90ca352621ca16c8a617da

Observation db8cd0e8-bcf4-4dc7-9249-62ea0776d5a3 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:04.820065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:56.474829Z digest=sha256:2de2b3981cf48aff815ebc216fdbadd98d3357affdc18e408b3ceb576c8fb9e6

Observation d7051c0b-f641-4363-940d-13858c01e34d · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Finite Scalar Quantization: VQ-VAE Made Simple

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.557191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.557191Z digest=sha256:82961d7096691db5e5fde638b0325c47fb6d241e19c51109f00c76217e453156

Observation 25004936-e1d2-498e-a7ea-46875215afb1 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.630048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.630048Z digest=sha256:37aa68f907ddba456f40fec733ac244523155a951d3b8d98ab96bf39f15fe1aa

Observation 42046312-291d-4ca5-8cb5-79b7fa1cdd94 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.720288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.720288Z digest=sha256:56afbf66fe55e4a6871ddde15a3448ca716097a2def62d2ea61417505e1d1ccb

Observation ca11a80b-9add-4a97-88fa-d01d3087f14f · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:04.538604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.085294Z digest=sha256:89c364a6bb190c8d8af766783fccbd6ad2fec1cd2da2094a4beb9e4ef696edbb

Observation 51dd94e3-a1a6-48c6-9188-206b659b9008 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:04.386285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.212866Z digest=sha256:ea52022bc78ff07a01f894537b3b9e7492048ee5b574ed824579a4fa811cdf1e

Observation 9ecc25f5-0aeb-4093-ba3e-7273c48a84ff · outbound

This paper cites W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:57.367656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:57.367656Z digest=sha256:0e216e576bb16c459f45a3a739f2f731cb670eb7f0c318905d0af43418209c90

Observation 4753148e-6bf7-473f-8d88-4f3c673be2e8 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:04.191326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.484383Z digest=sha256:20a492bb7260e48aaf4b6e97e89464c68f219504e80005e23cfec99a253c36ca

Observation b6bf2331-1b89-4039-938e-dc3b51fde4bb · outbound

This paper cites A.; Gonzalez, J.; and Escalera, S.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; Gonzalez, J.; and Escalera, S

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:04.028061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.624749Z digest=sha256:3f97c0d432339aa943a06b6c96cb02a58c5f1f6c64e78502f53b4e40419b033a

Observation f6c00cfd-550f-4e87-9953-321632221f2c · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:57.740164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:57.740164Z digest=sha256:8ff117350d6504299ba232883f438305678fe46e432a1a0dfb979470b1487a21

Observation d58f3e58-36fc-4fd0-8dcd-579b40551c1b · outbound

This paper cites E.; Hinton, G.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech E.; Hinton, G

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:03.851801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.864744Z digest=sha256:6d21ba50ca093757bd21253f4d854ddbbc1eff86be8b84ef5eb94de2ac56b5b5

Observation 214db0ad-301f-4718-9fc1-d18dda973ad7 · outbound

This paper cites DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:22:01.078010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.044747Z digest=sha256:b1f28a6f21f6a3a412a345d3b59bb7263f59dc7339cfdd8ff2251c72d2a8445c

Observation e3098054-309b-4bc4-bf24-f287f3b7d03d · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:03.634747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.179341Z digest=sha256:153b9e975d40672c360a850b1966851b6efff1bdf36bbdf5f58ee4251b39503d

Observation a39cf59c-84a4-4e32-b488-bc1979a61e5b · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:03.332892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.359642Z digest=sha256:9ef801d0da3167ce21677d5052846f3d1990036416502af6f6c9b28736929b71

Observation 6357baba-1f9d-446b-aecb-937f7b65bf4e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech LLaMA: Open and Efficient Foundation Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:58.542592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:58.542592Z digest=sha256:831746995dc36c5d5f55c7dafbbdcd271b76637e3b03bb661b78823c70bec0b2

Observation c60b0e3d-ffb6-432d-806c-cc2f525b1e2e · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:03.046429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.696284Z digest=sha256:82d38ceb00eac7e9529a89ad5f3eaa25b10daf90c0bf3cb4eaeb3e7312363a3e

Observation b861883c-b8da-4abf-abeb-f1d59ec23138 · outbound

This paper cites N.; Kaiser, L.; and Polosukhin, I.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech N.; Kaiser, L.; and Polosukhin, I

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:02.851166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.874889Z digest=sha256:9886dbd5ace78aa9dd7f1856d7b2a964520d19ed5aae4d0e7081fb9a8810fa80

Observation 7cf52851-2267-4464-ac50-5ebe5ba8e95d · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.042807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.042807Z digest=sha256:57069db5f62756701967131c4962c2b270880c834b94fc9c031792b8ec4ed367

Observation 04e4ef9c-3e63-4826-b1d8-4ba7a760e217 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.231724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.231724Z digest=sha256:fed1fc61277ab26dcac237f18e9c03d15cbbb49532ccc882b36d578e32f940b5

Observation 87488ee5-d784-4d51-ba47-6cc934239f83 · outbound

This paper cites Qwen3 Technical Report.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Qwen3 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.316444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.316444Z digest=sha256:a6d2d5169591d7c5230296f777a11d7af67f5e9dca2519dc7a5f35c0da8a993d

Observation 3532d1ea-3948-4fe4-a0b4-111f73970903 · outbound

This paper cites SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.444748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.444748Z digest=sha256:85c581bcf430f6a2e2db52e174c23594ad648010f6a8c747d10a2fed5065b726

Observation abde2652-5844-45b5-bc1b-6b14c8aafb41 · outbound

This paper cites Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.575812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.575812Z digest=sha256:479d3e2c9606db7ca016cff9348d0a64775504d7837b3ca7f3eb0d13741fc2de

Observation 4b14c852-8f34-4761-bb46-b3461b79d073 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:02.663861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:59.683308Z digest=sha256:85d21920e59a825a6e05b6bdbdf60895d7a35ec7de9bc2171d8ce5db59a09ec9

Observation 63838768-dc09-4026-8baa-6769dd2aae4b · outbound

This paper cites W.; and Li, H.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; and Li, H

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:02.432658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:59.768998Z digest=sha256:786629a7629231954f3ffc871b11b2e4cb8c08b11f26d5140d78d1ffbe7f116a

Observation d5da1b1c-023a-4ea3-99d9-63d9d35a2d6f · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:02.151155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:21:59.914789Z digest=sha256:b0fd75efa761ecd9a84272ebd767d1f281cbed9504ce84f61999a742b63e9619

Observation 6cb74769-28fd-4cfb-baad-e0f37447b91e · outbound

This paper cites , " * write output.state after.block = add.period write newline.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech , " * write output.state after.block = add.period write newline

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:22:00.026278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:22:00.026278Z digest=sha256:4e76841ebff5845c03500a95aa7afdde5a25a33d4e7d8cc087862c96af5c182c

Observation 4c4a7b8d-536e-41d3-aa1d-07323ea46f4d · outbound

This paper cites write newline.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:22:00.140173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:22:00.140173Z digest=sha256:b399acb7cd85426f53be691b50f70ea2f5f6837b3f65bad2806e7b43cf7fd609

Pith citing papers

Observation cf142752-2af7-45bc-9cb5-7725da8b0756 · inbound

TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis cites this paper.

TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-21T15:34:15.080474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T15:33:08.885667Z digest=sha256:f2974715b77b438b442c260fbf574998c7a1bd5a484e5df9bdcfc74ed9fa7d1f

Observation e574c656-9807-4a67-97f2-0ac729782f23 · inbound

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion cites this paper.

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:57:42.883160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:55:32.044506Z digest=sha256:e361d00d3845ac7258ee6c634c9c9c057846a89de47d2296061a1b750a1a0752

Observation 3bf6e419-1013-4bed-8f91-02ff709e7272 · inbound

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios cites this paper.

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T22:51:21.928324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:51:21.928324Z digest=sha256:e24d1b11baa2c3d09ff2a5deca16858d32eea0e7bd05f3ff344a31c8614986d8

Observation 878d1998-5126-475c-b0c5-ebb94ffe3517 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.931802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:7e729b2fd7cdc58fb2a46ef399cc8e864dcf37481728550cefe7b87c190ef04f

Observation 5d1be585-afb8-4162-8a6d-24ba3001dc1f · inbound

Sharp spectral estimates for free boundary problems arising in plasma physics cites this paper.

Sharp spectral estimates for free boundary problems arising in plasma physics IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T19:56:33.793167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:56:33.793167Z digest=sha256:42fafec3665a8951f9b95c0ef7ff75ba9d25a83adcd8032a373b9b63345e279a

Observation 3b01ef42-a07b-480a-a239-4229b954d18f · inbound

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection cites this paper.

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.994347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:55:09.141906Z digest=sha256:5c0182d2e4b2023d592c11a054b6b5f8530d65d8008b88dd7dc8f2b121682574

Observation 0ef942c6-ff61-45d8-b44d-fa72ec1c2ae4 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.305794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:e0acf6f6b233649d804fef30dbfb3879695f3b0f8c9bf9c99d042a118043a578

Observation 713eefad-13f9-4a93-b473-955a704b9278 · inbound

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing cites this paper.

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:51:00.754781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:46:58.010112Z digest=sha256:3cdf68b2668a35d3d961d112f68e07c05ded6678240d269d3a69532bb7bf0603

Observation 03223873-a2f7-4fdd-b941-df791230ca24 · inbound

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation cites this paper.

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.500305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:36:07.090627Z digest=sha256:81d7a449e0711bcb8e3dd5ba471bbdb94837167ef7f9dd1654b2a8aeb2c50420

Observation 8a395226-b98f-488a-bfcc-795612a049bd · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:10.007737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:af0c34299477a64672273f412cbaac4d3690d995f36fb446852a449b43a26ec9

Observation ec3507fd-42e6-4b9f-af66-555e92b2e64e · inbound

RTCFake: Speech Deepfake Detection in Real-Time Communication cites this paper.

RTCFake: Speech Deepfake Detection in Real-Time Communication IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:26:17.854987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T05:29:45.895667Z digest=sha256:88cc50e9b0b8853c4510f1d276d8fb57de06fb227c3f0ff873ddf3f291c2144a

Observation 598b00b6-b804-4a36-8a51-1ca6d28b3aee · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.102735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:3ee2c0c8546d5655be60c4725f3f9c692c069006e4088e7c77f83216436b6811

Observation 3ba17a93-84ad-456e-98e6-97359ecd5833 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.806423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.806423Z digest=sha256:af12f8d94ce0e9b426ea8ee1bf167cf9c0f22eca86c1501bd90b281474b491d4

Observation 235e690a-a3f6-4cf0-b49c-a4865a81eac1 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.070996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:6c72036d8c30d1617e8160bacd7cb6194ba443d520079f4f37b01179c807db34

Observation fe258340-112e-4730-990b-7bc43fc0e81f · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.280453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:06:54.972846Z digest=sha256:e0fee432c42489cdf1d69c458eb2cb17a22967fcd28609916dcaaf5e99fe9dc6

Observation c7b3091e-e029-4632-9041-6273c134b334 · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:29:52.767862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T08:28:21.790110Z digest=sha256:f8e94cd7a4864597519fc84984e0d37b80e38a2a8af484db76ef762a8f60293f

Observation c36b7f05-03d9-4e50-a73a-d1e5dccec9b1 · inbound

DeepSlide: From Artifacts to Presentation Delivery cites this paper.

DeepSlide: From Artifacts to Presentation Delivery IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:03:10.135599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T18:03:02.253422Z digest=sha256:e8ad0c01c648c0683fccaccb3cc10fd00173f262ae671720d5a4381caef6f37e

Observation cdc9418e-d138-41ec-8f3e-14063841f124 · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.031428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:04e5ee58366fea114b16c59e1d2416806a5d827685381fbba7399c88072b0627

Observation a5dc0fd2-ca5e-4249-aa34-398776f16c9a · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.232955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:4b11e29464676f9fb4f1a533e8fdb7462022337b6882a56f7389bf5b4eae6e01

Observation 579f3555-339b-4a1c-9417-fdf45cda185a · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:00:59.040259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T02:56:06.910758Z digest=sha256:c81ce492267f337294ef2d4105e1ae06efb72e9797be80915c5169469668b3cc

Observation 8144dd9c-78d0-4466-b216-679c9b414ff1 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.802184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.802184Z digest=sha256:c4cd0e4d6f323d7157e5d2a0823733fd46f2f69c02aaa0f5b340ff99b214a645

Observation 2830a0fb-9f2d-4eea-aede-ebac1bee36ff · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.082455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:be1717eca6419760176c07c8da69c217541336ea018be927b0a87a1f79f2821e

Observation 22e6c8c1-4c35-4e15-8038-3119b51f83a9 · inbound

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech cites this paper.

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:48.523055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:35:06.352720Z digest=sha256:67d3f3950a0abee182f9697a3fb2cd78600ac2360f09f9568e410f9ad0953a6e

Observation 6815c925-6871-483f-bd92-24d4edcbc89c · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 107

Resolution
malformed identifier
arxiv_id, observed 2026-07-01T20:46:13.835687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:f2a204c8d8d2ae27960e5aef52ffa989875608ea16776b36c5df39f0f6581bb2

Observation 16f5240f-3f96-4f1c-bfe6-8bd5849f9f77 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:46:24.634388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:210216a84b82a0deeec224864a6832dd0e092187a5e8cf8e1f2e3db71b94da6a

Observation 61fcceb1-562f-404d-837d-ba92ee5d31c3 · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.224385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:52a1412a97852ab816fdf4d2ba673ad60bb9b95d366e3f1db777fc202c02eb04

Observation 0be20842-bbf1-47f6-9ecb-4f2e9eb7446c · inbound

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models cites this paper.

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:20.002939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:16:07.007759Z digest=sha256:a96fc5f406448bab7ba842bc57e3dc674ac93fff4ea9b819f585eb0005452adc

Observation cc9f9742-5fa6-4233-a3fb-1f7dd5b91dd0 · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:07:35.857767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:262dfbb6d768342bbc200d7f651d435a0408e7082d116dd9891383a3144f90cc

Observation 8581762d-eee4-4332-9ce6-51353a517a4c · inbound

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations cites this paper.

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.371843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T08:38:51.371500Z digest=sha256:b90ba45a0e018b57dae060eb8a2a3e88a99cf6c9b7a949b6ea55668f5f683daf

Observation c83c0c63-cb90-42d8-827d-dff5dd9adeb3 · inbound

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction cites this paper.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.958514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:f88980ffc5cc37d21a102231cd905f45efebec5132691dcaba5892f67cab6585

Observation 7a8912fd-aeef-48a7-b060-41486a59264a · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.572183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:e7cb8bdca1ed5df3c03f4be36f7a180553af7118fc5dd6a678423cc6668a02ef

Observation 3fcbdf0b-04d4-4188-a95c-e24b4faa8ff1 · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.701979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:ddaf87fb7d8e579e639815a03c2df1d7412142debe37afe0ac9b2d87507fc893

Observation 8dde18c9-d7bb-4c47-b60f-d6c59ab96167 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:19:49.634822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:72046c8ecd641fd361b62f97c77078be19d324ebc82bcaf43d8593fd323b7190

Observation 04d45292-d23e-45b6-b3ed-a2e00b37f448 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:5a969c3aed0e90d6a591d472e5055b301dfb76baa4e88ecaf344fcfd7367a1de

Observation 997da372-83b8-4418-a395-df4cb3399f5d · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:54:40.709861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:35:03.660671Z digest=sha256:080496d53174fd9c0fb3d1dd14fd48b00d2b550be9b86b8c3b0cdb4b1c943a60

Observation 5b46fa76-577d-4a5a-aaf2-9b813acae9fd · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T11:11:08.878850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:11:08.878850Z digest=sha256:ddca0c03b6c679ea3767aafaafe9b9a7c15b4f183115c13786cb28e99f7223bb

Observation cf0a7c94-810c-486b-80c0-50d42fcc5bd4 · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.182730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:d9385e6c54ecb1c778ff86752d208c753979825cde90850451338b3bae6dc009

Observation 28686fc6-7cc7-4693-887f-911c13a40eba · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.726590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.726590Z digest=sha256:63810ccf6e97c2f87d89a9267a688fe9fb2a82aa80c23d429591301250905867

Observation 8a132f17-e405-42be-b8b7-27ec31c2d703 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:31.531580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:31.531580Z digest=sha256:53aa8b00d524d274bdb88808e0d5ea22164fe7547e5d3888269859730ad1736e

Observation a0cdb54b-a9a6-4ab2-9ced-30c4f7c6082c · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:50.544402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:50.544402Z digest=sha256:f8b50158b4a66fd2ffaed91c9de47a9043e7ba2a422c6641a1adc1a44acef2f9

Observation f2a5988c-5344-4936-b561-59216c21ed4d · inbound

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models cites this paper.

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:39.932052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:39.932052Z digest=sha256:bad6107903e81ab79eb2041ff2efdd14b43451d20f7053ad12c087281159dedc