Pith. sign in

Paper Citation Record · LEDGER

Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 81 inbound Pith citation observations for arXiv:2503.01710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.01710 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 81 of 81 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:05:11.478547Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca3a1a53-ced5-4735-89a8-7469ee6554d1 · inbound

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget cites this paper.

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T06:05:11.478547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:05:11.478547Z digest=sha256:091909f6757f9603c7ca1599ec1a8d4acd952c9b481fc79de88b44391ce145c6

Observation 4c99b361-cddf-44b0-b87b-f9819dd76fb3 · inbound

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing cites this paper.

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:03.293855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:03.293855Z digest=sha256:868c45495eadd96f88c04bbdd82e0d1cce65850d058e45b7e748bd94f63a9abc

Observation 320ec6cd-47ca-4755-b7f4-245507619faa · inbound

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese cites this paper.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.866945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.866945Z digest=sha256:aae424144b0f98f7a176d20a6c736cc42a0b7ca146d5c0949e97df1caf5f0e59

Observation 2a745d27-1788-4f6c-9bd8-47e1b4265e34 · inbound

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information cites this paper.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.592637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.592637Z digest=sha256:d43305d364bdc93af13acd75e5434f7d75b76579348c11ef0b7e936929557542

Observation 658331e9-759f-4387-83f8-826b084a0d92 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:2ceaa8d195e35638c434565c83359d56996127c9e98111647a482abae1fbe2e8

Observation 083c2582-08da-4a78-b51a-176982c32b9b · inbound

Speaking images. A novel framework for the automated self-description of artworks cites this paper.

Speaking images. A novel framework for the automated self-description of artworks Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:52.259140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:52.259140Z digest=sha256:6945e5c69fddaf764ec84ce2236f666cde1d696b399cb204ba45d06a663c0424

Observation 45102f47-a413-4735-b698-7e1c2858eaa9 · inbound

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching cites this paper.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.314055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.314055Z digest=sha256:c2465b53eca405441b3fc796eaa98cd214a36f6591a26fc5c88cdfca7a0b9fc1

Observation e5f9d465-6cd7-435c-bc08-08a9f234aa88 · inbound

Optimizing Multilingual Text-To-Speech with Accents & Emotions cites this paper.

Optimizing Multilingual Text-To-Speech with Accents & Emotions Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:29.495278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:29.495278Z digest=sha256:8d64977eed27785fbb532e72db7f1f80542156e52adb65793a7a770a8046cb20

Observation 7cf52851-2267-4464-ac50-5ebe5ba8e95d · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.042807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.042807Z digest=sha256:4054c27202cd025193e66fc17bb25ecd8b5710372f8d53bd5e15a66f6cea25f3

Observation bf87162d-6f74-49b5-b76f-603227d7de29 · inbound

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability cites this paper.

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:06:18.858838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:06:18.858838Z digest=sha256:f9a5d64cef6cd9978e97c118db761c35680975c6008b14eb5060310c0b9361d9

Observation 5a06dcc2-5cfc-46c5-a2c9-3a91297b2948 · inbound

ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition cites this paper.

ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:31.298864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:23:31.298864Z digest=sha256:68bbefcd4be6ebadba7ed265ad4c72b98457584e07987544c3620b7f2d6bb04d

Observation 89afd59f-6147-4c85-95e2-9fa326083994 · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:32:03.669368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:7a49ba98336ad57a232b49c472f0af7a4fed20d87309c545c359eff4dfc1b6bd

Observation e4ebd70e-7347-4e0b-897d-2364e7a2b092 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.918034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.918034Z digest=sha256:999c1634f5024bf701341abebeba5f6559f3a838ef14a391cee950816097625c

Observation a925d1f3-84b8-455a-9fad-945bdf949598 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:1a3c53214e8099c36c8b5dbf53509b487c43d72f7524b9bfed66d99c20a5bf52

Observation 0c5d9294-2c5c-4955-820b-fd37b48e7a1b · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.167239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.167239Z digest=sha256:b41d7d569c396aa935eab3e50a71b67d09c3835f0826aab43ac544d37234d070

Observation 9accb710-3a08-439c-b5ad-e7e7fbbba6c0 · inbound

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts cites this paper.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.369330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.369330Z digest=sha256:0bbdbcbfd989f4208d295d67946719d4e5d957217bf11c735d4a8a4e0319e5db

Observation 61f5c358-f70e-48f2-9779-a8badee88a3c · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.823558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.823558Z digest=sha256:fed43ccd1953048fd82841a6c2d30f6fc1b6f6e8ca1c83417c692ec6d66eb26e

Observation c68d5cec-178c-4b7f-87a1-08a6cf00b825 · inbound

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot cites this paper.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.305208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.305208Z digest=sha256:6a00e39f3d103caef940be674bc6d9583c1a3b64934c8eab35854b2eb215c96b

Observation 03362d43-70ad-431d-86e9-4539288b7823 · inbound

Audio Deepfake Verification cites this paper.

Audio Deepfake Verification Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:12:07.934329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:12:07.934329Z digest=sha256:86f9ed1468ea9378f446abb31e855e1400726847dc7dab9c2fcd2245873a0a45

Observation af68322a-adf3-429a-9b59-fdc25d09a134 · inbound

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching cites this paper.

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:21.114170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:51:21.114170Z digest=sha256:83606217bf9bb47ebad0066f930935cceece739607c1c34cd3f6ada3243f5e28

Observation 4ff6b021-b5f3-451e-894c-08844f99ef49 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.596167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.596167Z digest=sha256:389010e955a6669f3ba59f62d2a59b1594fa09b69f0e8965b1a55d0953a978da

Observation 825cea1f-649e-4269-ac2d-da8e597cbc63 · inbound

Speaker Anonymisation for Speech-based Suicide Risk Detection cites this paper.

Speaker Anonymisation for Speech-based Suicide Risk Detection Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:48:32.047386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:48:32.047386Z digest=sha256:f9e97a2950b6dbb30e387129b5d86e71bc2425811888430a61ed8643718a91df

Observation 06f231c7-7dd5-4e17-92c0-188fee700bf8 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.515375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.515375Z digest=sha256:b177063e26e5b39cde0bee44fc652dd0ac809db4175d04137e416de925b9515f

Observation 2a8b220f-d17b-4d31-b1b5-ea219981fda6 · inbound

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement cites this paper.

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T08:29:17.918593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:29:17.918593Z digest=sha256:5bb544600f187aba08790a901e0111960af5f8efb6e532d72638a2240590411e

Observation 815c3a28-de4e-4fe9-94b1-972d24b9bfd3 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:84a5fb5dac1957f617571fc0c053397e70186677633e096684d4c354e6a9b645

Observation d5fc9a45-8b02-43b6-939b-541e33fa06e6 · inbound

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization cites this paper.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:b3789280a0d6f14814b3042b1c74c10cb5ebc394527b70d943a087151ce3d549

Observation 0ee2ca03-4cb0-4c66-ba8b-feabbae8ded2 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:0193d7b1ed061a369fcde537a05d7274a3659917055dbd95ed4116dda38120d8

Observation d3203e74-6950-41b6-b3f9-9b6ad7aae489 · inbound

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models cites this paper.

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T10:29:51.313835Z digest=sha256:c31a3c3fcea915d46e0fa45042fcabf33f527296a0fdefc98d8c2c253da9a67e

Observation 032baf92-2d34-4419-a84a-9db2a127cc2c · inbound

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing cites this paper.

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T16:03:15.572657Z digest=sha256:a82abd7286970958244e7aa9b8618b8b12f1a3755967de53dcf5766ba096b0fe

Observation 9e41184c-0120-46dc-823d-c196fdc6e7cb · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:a3303a0c8094b4e63d58630a65a16a335725465a76aa23567b1a14f2761e971c

Observation 5bb3292c-5288-40fd-be9f-aa3f4257990a · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:43c0002a72c7bbf1c729a82fed1a062d6918a04f951f92786ab3b3097475e271

Observation 7759d619-2cfb-40c0-9a1f-c3b129291c2a · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:3fc978fd6129d694b57a76ba2b312c6db159a9f8f74cf701caa059db3c786922

Observation e1e6834f-1a4a-4a34-9992-c3c0b4df4f7c · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:db00b22db9535fdbac6ddd3c1e847bf3017b859c8bc6f028ba15a73fd5e822a1

Observation cc57d162-1654-46ea-aa7e-a8702aab346b · inbound

RTCFake: Speech Deepfake Detection in Real-Time Communication cites this paper.

RTCFake: Speech Deepfake Detection in Real-Time Communication Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T05:29:45.895667Z digest=sha256:cf888444c37e0caf8aa69f33348abcecaa57ab6a6cf54c6eaaff6dba283cc237

Observation 6c12d119-f452-48b7-828a-6f0fc7a4ea25 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:d6a432ac692cbe60cf948396aa0f0606cb390e418b3b7bc6ed612a8e53b935ff

Observation efb8d1e8-9160-4f13-8a1a-d8e873d46580 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.920158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.920158Z digest=sha256:1098e5c287aa260e3d01e6ab72d7cb521c922dbda5f70c4ef5d3d8f984bff1eb

Observation 81a62402-3d51-45c6-98cf-bb2906a96b7f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:b62da3d90dd21ea9587fccc9c685e74d8570b9124a2563a5cf75c584c6436894

Observation 6977a6bd-2144-48d5-99e7-34103be00a11 · inbound

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation cites this paper.

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:11:50.733585Z digest=sha256:ddbcc7d3baf4dcc2d602ba06b4eb8dcf529b868bb9f8efa786dcf9b572491a4c

Observation e007d887-c74f-4f1b-919e-9e9e6e87eee2 · inbound

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech cites this paper.

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:04:44.050889Z digest=sha256:3816a164114dfd6e2f88fb87ff8a0df3da1d0b41979201712e8c1219b42cb0fe

Observation 0b14dfbb-e19b-4af2-a375-5ab8589f6bc3 · inbound

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue cites this paper.

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:08:55.753359Z digest=sha256:a4f928744a5553b9d18f60bb1c6909e586741a955ea0c2486146ec4988f444d1

Observation 6ae81e12-bb5b-46de-aff7-f569e95e1a1b · inbound

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cites this paper.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T01:04:54.506749Z digest=sha256:c6d41f9225bf9645b8a6e25d45a8419aeef5a96bcddfe4eecd171508dfbdb2fa

Observation 9607944b-5c3b-41ae-b04c-c10646d4901c · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T19:02:43.603286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:33efe6716ce49100ee42daa86df3014efe750ff1205326d1403eb570489193ad

Observation a98bcf48-2a04-4ac6-8f94-fb52873bf9d3 · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:19:03.226500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:7eda0cf4fddc899225a0b1dfbed028388770f58ff775ca52958c7d70d0bfdf6e

Observation 29e3e707-dd00-4e28-b5dd-054cff31428d · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-22T03:00:59.013899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T02:56:06.910758Z digest=sha256:bdcb73672a705356b1810ea5f5a2d4dfdb1dbe017d2fe1809728c0ce5b28b1da

Observation 1ccef922-c3a6-4f84-9aa0-7095ea3f6202 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.566575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.566575Z digest=sha256:6d99b19f8c7e291d8ae1b2e932d1a1eaaa7b94598bb548e0967e8be1c16f50df

Observation 85d2052b-70ad-48c4-94d6-25e41c559dfa · inbound

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio cites this paper.

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T23:14:01.868617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T23:08:06.858906Z digest=sha256:454e9cc100a6c54d586b373122dadf1705598874df98e82e8c18b8b50a30e618

Observation b7f74d6d-11cc-4c4a-a803-fb5f05c44891 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:26:13.152245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:270f8aaae91c0e28dbc86fa2d02ba4d92281eb20ffcdffeca5569d337e7b39e2

Observation 8665f702-ea46-4599-849c-b89776716389 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.744079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6d9816d6f9377b8b48ac6cae97f638875ebf63e56bffa9377fbddba37e03a059

Observation 3174fc14-969f-4e16-b1f8-dda8ee890e65 · inbound

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding cites this paper.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-28T05:21:39.804685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:1f56e365e458bdf758b9c5a29099286a2d20764c5673ef4b389fbd379b3d8992

Observation 1a1f96d6-7f35-4392-9505-788a2d154bae · inbound

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech cites this paper.

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T15:27:04.877027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T23:53:55.385445Z digest=sha256:37502da6cffd0ae97c80a0d5374039ad06e82a45577b4f3fecee9e800490cb15

Observation e8aeb4cb-e157-4f1e-952e-49575cdc4a3d · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:07:35.861159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:805659ffb54476184c6c74e38ab91fa2c64a42c683bf7a264be60a707c0d8889

Observation 4c5aa081-7ff1-4e12-b4b4-28eb23a8961a · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:36.146102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:67fb93d2d03b7bef6c48af7b364a13e13cf102ff65859c82f04e61552b5db6c2

Observation 0491a8e6-a336-4ea9-82c2-9999197fb4a0 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.903098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:2e34065aa63c980f26d0495c2130e553134e4c3ddd4728b62a79cdb2234b0ce2

Observation aa152316-2677-43d6-9aa4-a45bfef35ba0 · inbound

SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion cites this paper.

SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T07:39:38.407351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T13:19:43.545346Z digest=sha256:3151a805829be72f1c8d060f3b90c1591cf300016be7e20d690d6ec9d80d671b

Observation 587579d2-b9a5-4802-a375-8e9153cda095 · inbound

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity cites this paper.

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T07:39:39.524651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T12:58:27.702138Z digest=sha256:d3bbd211da7a9836142a56433ec748aa8d1d43d32ebc8b0b788bbbba85cab8ad

Observation 86040e40-24ad-4087-9471-978b9f711cf6 · inbound

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead cites this paper.

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:41.337402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T11:42:44.035902Z digest=sha256:c7c059c1733cf931bfa24161febf83462c859161fba7a75426477435e2fcd7d7

Observation f2ca7e14-86e9-45ab-9112-d512c0572d1c · inbound

AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation cites this paper.

AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:41.882848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T11:37:39.960212Z digest=sha256:147f9ff4999a77e5c528c3ee6f9d73b85d4a8d588e04564d0c1f3971d5696ecc

Observation b6ca7271-a549-4947-87c6-96b4d5e84e2f · inbound

ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech cites this paper.

ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:42.158604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T11:33:32.067771Z digest=sha256:d92bc3593022562abd3524585e21b5479b14f08013afc5280c7fbe527cb69d02

Observation d3c53df9-fa03-4088-a98d-fc2409abcfd6 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:19:49.663900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:1a5d7ef84453a07fbee0082f1c13f5d73c89dadf8f45180470cdf48a3e78b4bc

Observation 863ace9b-822c-4cb7-b5f6-4a8a1a89481e · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:a9989785adfd6560ccdc518273373ebab8f9faa8d1799932357577ee8cc22c73

Observation dc2309cc-a33c-4d8a-854a-4a958ba64342 · inbound

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations cites this paper.

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:20:07.170409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T20:21:47.681999Z digest=sha256:bcf724e35f39e7b13e675573d9756a66e335707861807e32a72b15f1831d7d08

Observation 526c3869-1c27-4d3b-a04b-08bda1d6c4c1 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 209

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T11:45:47.168565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:f5a64f270050ac1d4696bc040aacc6e0fd32733f9e99a12b2b9b4b3ce55a6ad6

Observation bd6ce112-e622-4ab7-8cbf-dba419703643 · inbound

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech cites this paper.

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T21:24:36.925360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:24:36.925360Z digest=sha256:4dc9e917ecb1ee8eee6a856355164ed00caa6ea7972ff6df51070219f97afac8

Observation d9825cfb-f65c-4be8-9c6a-9eae8e64d9fa · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 243

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.654037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:2d01ce29b0b420d5ddb7d591ccd67f63438385a01698b41263d7567e8d270c1a

Observation 308b3215-00ce-4f70-b0cf-f6b0318882bd · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 243

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:b9c385c9c993bbe563d1f54420dcb9befbcc2dd61214ef0561a91cc12e202475

Observation a83b9699-f724-4176-8de1-21615976fa81 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:1fd0282f88340c58725a240a6c1a4c9dfdb34e9adc8905ee36b2ac17571c32f0

Observation 7ed4bf9d-46f4-47da-b1d1-88d7a861b82c · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T02:22:47.820537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:22:47.820537Z digest=sha256:bfc9448c3a0b6458055221bced413385eec091875b360e7ca29e5942d0e49b30

Observation efa9c8c8-f4b2-4418-9237-48cb0f51b532 · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T07:39:23.645758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:39:23.645758Z digest=sha256:65b66de5cfb4d67c32e01ac5b59d39674c13672c1845794f007494b0029a8ba2

Observation 69ce9690-b3f7-48be-9344-2ff340b19901 · inbound

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs cites this paper.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.564412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.564412Z digest=sha256:49c72e069f4cb95706bd85af7a750e4812dce8057cf1ceea0c211bdd4549c4f0

Observation 20c6fb94-ab8a-4374-abb5-e373528bc0a0 · inbound

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm cites this paper.

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T23:35:21.089160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:35:21.089160Z digest=sha256:a635daf6915df612fe84617f723e09fb8e9dc20ac1a0ed9814f41b87e1c3af23

Observation 65821001-88bc-4a4f-8301-65f88dc8ba45 · inbound

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning cites this paper.

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-30T11:37:00.186862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:37:00.186862Z digest=sha256:b63568ea92896ba1c3ef639d0cad7f7000e675483ac2d4deb988831708a30297

Observation 3c693153-dad8-4601-af31-ad398c47d5ec · inbound

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning cites this paper.

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T09:59:39.136823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:59:39.136823Z digest=sha256:a3f06688e7e2de4042ed17686796b85b9d24b0564d00fa52c4d954e75ea7b66b

Observation 4ca58211-2071-4b36-82ff-2a6c4b26cb4b · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:55.557807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:55.557807Z digest=sha256:5b22c62bc71136257f76ec503a7ad73363056cc4505fe3763cc4968ab69e68df

Observation 94e003f1-8bd2-432f-a719-819c7391d48c · inbound

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech cites this paper.

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T15:22:08.955751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:22:08.955751Z digest=sha256:57cfab5997f5cb1ae86c8ac46811c892d47735b504900979966b1afc2eb7bcd5

Observation 8f826f4d-6309-454f-a685-1ce916139be0 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:29.149236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:29.149236Z digest=sha256:e48fdb8d55ce93f7270d263fb91b94a7680ae0f41b3f00cfe888007a0fa5ef43

Observation c9cabb25-a778-4980-9293-5f01de63f73c · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.455337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.455337Z digest=sha256:4c6046f4c6bf2e374c6fbe745935fd0a76e47e9fa4f9f255c8cd9042995bef3b

Observation 9efab4d2-cd97-49f1-94fa-edeb1e746723 · inbound

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation cites this paper.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.539324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.539324Z digest=sha256:54f9549404f9096abb75a9d6edf786b46fc2b5d9d9f63d074e20193e1d77c872

Observation 7c966cb8-0a3c-4ff9-953c-929f7103eef2 · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:42.137255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:42.137255Z digest=sha256:823294d4dc55a1d20b6394ef57245c3998ed2c6e237c45e242ce8f2dad4d585f

Observation 2000ce91-14ae-4b79-9ae6-bb88bd612383 · inbound

Luna-TTS Family Technical Report cites this paper.

Luna-TTS Family Technical Report Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.975736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.975736Z digest=sha256:f4771001c12d1ddc89bcfdd3ba5f936692fb91bd287f3b967614b30b85511698

Observation a1c13e3b-44d7-4ed7-934c-2568ce3b2d8d · inbound

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization cites this paper.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.986772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.986772Z digest=sha256:2e6066e5d44dd779f20c4d1b64b275a1e9ef79cea7bef82b8d54457ca777ae07

Observation 1f4fc372-c793-4fc4-96fd-5ad5fa39434e · inbound

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec cites this paper.

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:22:15.051545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:22:15.051545Z digest=sha256:9bb60d8fceeb135ecc3de3d71ed099419ead8a2eb6b9be1cc48b5be6b8ad631a