Pith. sign in

Paper Citation Record · LEDGER

Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2503.01710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.01710 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:08.167239Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 658331e9-759f-4387-83f8-826b084a0d92 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:d1db0a0bab5ae920c4c926b493e3fc4ccb379218e38c5c489f8678c0f85eb455

Observation 89afd59f-6147-4c85-95e2-9fa326083994 · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:32:03.669368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:1a4eae31f01511cfc384ce926579c5ac282fb202f13d51d26a6806242e2007da

Observation a925d1f3-84b8-455a-9fad-945bdf949598 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:0828bd12944c53a4a13532dbd06bcb8b1c9a29a0fd1d5b8e22faaeacb67cb9cb

Observation 0c5d9294-2c5c-4955-820b-fd37b48e7a1b · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.167239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.167239Z digest=sha256:f8736e6df7d78fdb6ed216e19fc3e3f7300fd20de65726cfcfc115fa17dc31e0

Observation 9accb710-3a08-439c-b5ad-e7e7fbbba6c0 · inbound

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts cites this paper.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.369330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.369330Z digest=sha256:05620f660f42365002294798fcca7c93a9872262834d3307df99053c3c74b4ac

Observation 61f5c358-f70e-48f2-9779-a8badee88a3c · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.823558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.823558Z digest=sha256:cb82dcd3844210e695afb8a02c50073a6e28f0a21142223044802ced47896a47

Observation c68d5cec-178c-4b7f-87a1-08a6cf00b825 · inbound

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot cites this paper.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.305208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.305208Z digest=sha256:1ecdba0576818ef5fafa25381a5a703db7c8f2042ec0ba29a7e764f2fcc4a0bb

Observation af68322a-adf3-429a-9b59-fdc25d09a134 · inbound

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching cites this paper.

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:21.114170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:51:21.114170Z digest=sha256:8af62c146f9db385392b86a2f44625eb9126a1d600b0e2536fa2aeb507886236

Observation 4ff6b021-b5f3-451e-894c-08844f99ef49 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.596167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.596167Z digest=sha256:5cf7e127361efca89eb5a1230ed2292f0d3b362253484ea01f05f19c400d6a44

Observation 06f231c7-7dd5-4e17-92c0-188fee700bf8 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.515375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.515375Z digest=sha256:33944fe45068cfac5e3d1c3b8ba9a6e0d45387b5e60a0c5a19b8a9d9b4f7697c

Observation 2a8b220f-d17b-4d31-b1b5-ea219981fda6 · inbound

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement cites this paper.

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T08:29:17.918593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:29:17.918593Z digest=sha256:fbbb4b22df2b987b5be9e2d3a623c073f7c78b54df02750fc7d04ef83b83a1ac

Observation 815c3a28-de4e-4fe9-94b1-972d24b9bfd3 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:acf79e62a1392f56e64061c0073b9e96449727aa8357b3750b3376a9b3734dfa

Observation d5fc9a45-8b02-43b6-939b-541e33fa06e6 · inbound

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization cites this paper.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:981268ed56b5dbb729734b5c8ea78d45b9991a967c526b9555c4e7efdcca5f7a

Observation 0ee2ca03-4cb0-4c66-ba8b-feabbae8ded2 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:fff253900946e08fac22f6ba7d4d63f69abb30b8e0d98d0ab55d02587e89c6d5

Observation d3203e74-6950-41b6-b3f9-9b6ad7aae489 · inbound

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models cites this paper.

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T10:29:51.313835Z digest=sha256:f44e344919299ea7f4327945d1d2c72627c6169a714be696587418d0dc41bbda

Observation 032baf92-2d34-4419-a84a-9db2a127cc2c · inbound

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing cites this paper.

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T16:03:15.572657Z digest=sha256:d790f5c218b5b214e01e2cb0eea660fd8bddfc9f0d850cad946158b1cc116021

Observation 9e41184c-0120-46dc-823d-c196fdc6e7cb · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:b280d7af5d54222479a2269aace187d02c1cf381f5a96e43643911c0883b43f5

Observation 5bb3292c-5288-40fd-be9f-aa3f4257990a · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:6d9227f51357dabb4e64b50a5438f7ee4a2c952fd597a49cdd2172a1d857485d

Observation 7759d619-2cfb-40c0-9a1f-c3b129291c2a · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:c0ee0347ff87080b6fddcf1d8b8e09495f6c8c1bcd032d40b8c4636c84636309

Observation e1e6834f-1a4a-4a34-9992-c3c0b4df4f7c · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:c33cef095a7652316b2a2a03a29003bf0d20447fcfb7c43d752851872dba5aca

Observation cc57d162-1654-46ea-aa7e-a8702aab346b · inbound

RTCFake: Speech Deepfake Detection in Real-Time Communication cites this paper.

RTCFake: Speech Deepfake Detection in Real-Time Communication Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T05:29:45.895667Z digest=sha256:5025ea908e1c2fa3c510d8b1a00cf9c54bc391fbdb6c14a17cf877e4b1913268

Observation 6c12d119-f452-48b7-828a-6f0fc7a4ea25 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:2a23e35e50c8ad8ba43089f4bc24751cfc226d4569cd47c5b65bdd8b9e65d59d

Observation efb8d1e8-9160-4f13-8a1a-d8e873d46580 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.920158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.920158Z digest=sha256:9c72ce964611588660ef59f75e6be8fb250c71024ca8706fa189e5853ab0bf74

Observation 81a62402-3d51-45c6-98cf-bb2906a96b7f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:3c6bce9c62563588d6641343d47a1369d04b47883b7304f94ffadbaa1c287fd1

Observation 6977a6bd-2144-48d5-99e7-34103be00a11 · inbound

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation cites this paper.

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:11:50.733585Z digest=sha256:ac5199678c60883892531701a4f9bb1d0418492c2d3ff93677b769ed0a99c2d5

Observation e007d887-c74f-4f1b-919e-9e9e6e87eee2 · inbound

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech cites this paper.

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:04:44.050889Z digest=sha256:8dd7225b717e61a444de49bb865802e56e01ac5910ac74cfc4a4215a4c339e0b

Observation 0b14dfbb-e19b-4af2-a375-5ab8589f6bc3 · inbound

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue cites this paper.

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:08:55.753359Z digest=sha256:18f1c3c562f66326da449e01ab615fa30b19fb0370f3ceb1b59e4456a6aa2467

Observation 6ae81e12-bb5b-46de-aff7-f569e95e1a1b · inbound

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cites this paper.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T01:04:54.506749Z digest=sha256:29cab53d91683951945d2a690c54797f5e856c57835948f48c5daf5fd61cc4d0

Observation 9607944b-5c3b-41ae-b04c-c10646d4901c · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T19:02:43.603286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:adc874500231a36d90d59e6274236e030a9fa3e7d7c1a0fdb3532a51fdf3774d

Observation a98bcf48-2a04-4ac6-8f94-fb52873bf9d3 · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:19:03.226500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:16964bae7a51021fdf1f4cdcbae579a66ce05f99cbf62dc58f1f26753d650ea3

Observation 29e3e707-dd00-4e28-b5dd-054cff31428d · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-22T03:00:59.013899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T02:56:06.910758Z digest=sha256:8e66799b1267b7e7a8a9bb447404bf0e04005c676deb2110dbafe250f165159d

Observation 1ccef922-c3a6-4f84-9aa0-7095ea3f6202 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.566575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.566575Z digest=sha256:40c80b54f4ba1dcd450272e8ddc65d8be8cfc7354a75e499d1f37ba5326e4ab4

Observation 85d2052b-70ad-48c4-94d6-25e41c559dfa · inbound

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio cites this paper.

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T23:14:01.868617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T23:08:06.858906Z digest=sha256:6ba8148c21668434e214df44aa5ff90a1378c9f47eaaf8100e72c708d9cda668

Observation b7f74d6d-11cc-4c4a-a803-fb5f05c44891 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:26:13.152245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:df1df2df4f9e3a825f5373ff9c2a68ce745d149a9e64a80fe1f26d71264cc5dd

Observation 8665f702-ea46-4599-849c-b89776716389 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.744079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:2afba8a3cbd0aea7c228f139f145f562bc464de02a72d39faa9ad7cb66b4e57b

Observation 3174fc14-969f-4e16-b1f8-dda8ee890e65 · inbound

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding cites this paper.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-28T05:21:39.804685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:b3d6fc81ea331fa028414f836b19c134ae52401181294d49cf0263c7f94cc500

Observation 1a1f96d6-7f35-4392-9505-788a2d154bae · inbound

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech cites this paper.

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T15:27:04.877027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T23:53:55.385445Z digest=sha256:6cceaed376a0ed4890993888eea3e2712d977000e699d9bc8e6fd3666aee0a37

Observation e8aeb4cb-e157-4f1e-952e-49575cdc4a3d · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:07:35.861159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:2570567d884128846eab6e05fa1400b3d6681673016d064c7138b88882f40dca

Observation 4c5aa081-7ff1-4e12-b4b4-28eb23a8961a · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:36.146102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:890a0a1cee09b49968b0c61b6f191bb82478092c4eca471838aae49061bf4d64

Observation 0491a8e6-a336-4ea9-82c2-9999197fb4a0 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.903098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:c478b07443e1f96c260d4d310ae51784f5b71fd54920d58fde6f494b8f2b03d3

Observation aa152316-2677-43d6-9aa4-a45bfef35ba0 · inbound

SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion cites this paper.

SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T07:39:38.407351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T13:19:43.545346Z digest=sha256:9c2d5a53bafd894ebdda0a8597e87208ecf0d8d4ec9e3c9782b486a98f46a09d

Observation 587579d2-b9a5-4802-a375-8e9153cda095 · inbound

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity cites this paper.

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T07:39:39.524651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T12:58:27.702138Z digest=sha256:be1a8a2abede7a56faed3858edd96ff86b1bcc68ae321d7d5358eecba61c84a9

Observation 86040e40-24ad-4087-9471-978b9f711cf6 · inbound

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead cites this paper.

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:41.337402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:42:44.035902Z digest=sha256:69bdd7ac067b1ed7ba7cd3019e876c271096515c309802446cedf374dd40825b

Observation f2ca7e14-86e9-45ab-9112-d512c0572d1c · inbound

AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation cites this paper.

AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:41.882848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:37:39.960212Z digest=sha256:2ab96b307caa0cfea9de75f88138324929fd2f47ca84af011139ee89121b8bd5

Observation b6ca7271-a549-4947-87c6-96b4d5e84e2f · inbound

ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech cites this paper.

ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:42.158604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:33:32.067771Z digest=sha256:24f1adfb61b138bfa8ae0e0c082e4be8f1bf494693b6ba41b9f53661b232acd6

Observation d3c53df9-fa03-4088-a98d-fc2409abcfd6 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:19:49.663900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:c8800acfb9e39e5bbd8cbfaaffe55d893bd098b0e23e6205fdb90a8b0e8745b4

Observation 863ace9b-822c-4cb7-b5f6-4a8a1a89481e · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:4dc7dd29bf14b6ef16e4e7102c4c2c994a714f93fd52f752ecf9b83086d48cc3

Observation dc2309cc-a33c-4d8a-854a-4a958ba64342 · inbound

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations cites this paper.

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:20:07.170409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T20:21:47.681999Z digest=sha256:d34bfaebd30b2c630254e11e60e789d587fbf15aaa64bde1d3561612f78bb962

Observation 526c3869-1c27-4d3b-a04b-08bda1d6c4c1 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 209

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T11:45:47.168565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:1715cbd2f609582d34024d0e680e4c9d8b4e22e7dfd3301f0a8cd2e0a437e37b

Observation bd6ce112-e622-4ab7-8cbf-dba419703643 · inbound

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech cites this paper.

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T21:24:36.925360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:24:36.925360Z digest=sha256:6bcd0b85ac222ceaa1691a30f8d4c619f5020fb9b9981132c75800fe329aa105

Observation d9825cfb-f65c-4be8-9c6a-9eae8e64d9fa · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 243

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.654037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:aa2ebec9b51ff4c52abd92630b223fc02ed5d8ec6682c58230d4e4c34c3fd965

Observation 308b3215-00ce-4f70-b0cf-f6b0318882bd · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 243

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:aca0c2a2f7919ee513338958c706d489da90776c35c9b3317c50610596f8a0c2

Observation a83b9699-f724-4176-8de1-21615976fa81 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:87a298baf382fc2a8b7eeafc4a8584576ef38206155c003806908d5ff7519a0f

Observation 7ed4bf9d-46f4-47da-b1d1-88d7a861b82c · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T02:22:47.820537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:22:47.820537Z digest=sha256:602cb83368b4df3e71baa19e80ecd07091e992fdd0c1de893ed76dad117d7038

Observation efa9c8c8-f4b2-4418-9237-48cb0f51b532 · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T07:39:23.645758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:39:23.645758Z digest=sha256:0831cb7301f3b6bd4eb8613abc66103a252cb3d61484e1c7809ababcdd58730a

Observation 69ce9690-b3f7-48be-9344-2ff340b19901 · inbound

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs cites this paper.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.564412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.564412Z digest=sha256:2d591641efc780ff7df02bdddacd5b46f3a6154c1f2f356b136f18fadf3656b2

Observation 20c6fb94-ab8a-4374-abb5-e373528bc0a0 · inbound

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm cites this paper.

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T23:35:21.089160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:35:21.089160Z digest=sha256:7cd6541dd8a90463a2ca117fbc228bdb717d969aba5080962c3643e48f9680c4

Observation 65821001-88bc-4a4f-8301-65f88dc8ba45 · inbound

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning cites this paper.

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-30T11:37:00.186862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:37:00.186862Z digest=sha256:0aaa4c7e0e9102567732ad222768956559cb51849dd118b60bced53d5cb28a7a

Observation 3c693153-dad8-4601-af31-ad398c47d5ec · inbound

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning cites this paper.

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T09:59:39.136823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:59:39.136823Z digest=sha256:913de9fde9fa13fc1172ddb14a5e674c9d1caefbaa2c64f9ba70da4d9aed268b

Observation 4ca58211-2071-4b36-82ff-2a6c4b26cb4b · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:55.557807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:55.557807Z digest=sha256:3ed8119fbccdf500691ebc117d13bf6fda68743b58924a2fc0f3b539dfac55c9

Observation 8f826f4d-6309-454f-a685-1ce916139be0 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:29.149236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:29.149236Z digest=sha256:ea0a51067d314f0407f67f3f1b2bcc1e524e9115477ea3a556b6399470c29c10