Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2503.01710.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:08.167239Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 658331e9-759f-4387-83f8-826b084a0d92 · inbound
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 89afd59f-6147-4c85-95e2-9fa326083994 · inbound
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a925d1f3-84b8-455a-9fad-945bdf949598 · inbound
Step-Audio 2 Technical Report Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0c5d9294-2c5c-4955-820b-fd37b48e7a1b · inbound
Adaptive Duration Model for Text Speech Alignment Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9accb710-3a08-439c-b5ad-e7e7fbbba6c0 · inbound
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f5c358-f70e-48f2-9779-a8badee88a3c · inbound
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c68d5cec-178c-4b7f-87a1-08a6cf00b825 · inbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af68322a-adf3-429a-9b59-fdc25d09a134 · inbound
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ff6b021-b5f3-451e-894c-08844f99ef49 · inbound
DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f231c7-7dd5-4e17-92c0-188fee700bf8 · inbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8b220f-d17b-4d31-b1b5-ea219981fda6 · inbound
UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 815c3a28-de4e-4fe9-94b1-972d24b9bfd3 · inbound
Qwen3-TTS Technical Report Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5fc9a45-8b02-43b6-939b-541e33fa06e6 · inbound
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ee2ca03-4cb0-4c66-ba8b-feabbae8ded2 · inbound
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d3203e74-6950-41b6-b3f9-9b6ad7aae489 · inbound
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 032baf92-2d34-4419-a84a-9db2a127cc2c · inbound
ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e41184c-0120-46dc-823d-c196fdc6e7cb · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5bb3292c-5288-40fd-be9f-aa3f4257990a · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7759d619-2cfb-40c0-9a1f-c3b129291c2a · inbound
NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1e6834f-1a4a-4a34-9992-c3c0b4df4f7c · inbound
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc57d162-1654-46ea-aa7e-a8702aab346b · inbound
RTCFake: Speech Deepfake Detection in Real-Time Communication Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c12d119-f452-48b7-828a-6f0fc7a4ea25 · inbound
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation efb8d1e8-9160-4f13-8a1a-d8e873d46580 · inbound
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a62402-3d51-45c6-98cf-bb2906a96b7f · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6977a6bd-2144-48d5-99e7-34103be00a11 · inbound
Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e007d887-c74f-4f1b-919e-9e9e6e87eee2 · inbound
Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b14dfbb-e19b-4af2-a375-5ab8589f6bc3 · inbound
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6ae81e12-bb5b-46de-aff7-f569e95e1a1b · inbound
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9607944b-5c3b-41ae-b04c-c10646d4901c · inbound
SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a98bcf48-2a04-4ac6-8f94-fb52873bf9d3 · inbound
AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29e3e707-dd00-4e28-b5dd-054cff31428d · inbound
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ccef922-c3a6-4f84-9aa0-7095ea3f6202 · inbound
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d2052b-70ad-48c4-94d6-25e41c559dfa · inbound
Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7f74d6d-11cc-4c4a-a803-fb5f05c44891 · inbound
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8665f702-ea46-4599-849c-b89776716389 · inbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3174fc14-969f-4e16-b1f8-dda8ee890e65 · inbound
CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a1f96d6-7f35-4392-9505-788a2d154bae · inbound
GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8aeb4cb-e157-4f1e-952e-49575cdc4a3d · inbound
TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4c5aa081-7ff1-4e12-b4b4-28eb23a8961a · inbound
FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0491a8e6-a336-4ea9-82c2-9999197fb4a0 · inbound
End-to-End Training for Discrete Token LLM based TTS System Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa152316-2677-43d6-9aa4-a45bfef35ba0 · inbound
SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 587579d2-b9a5-4802-a375-8e9153cda095 · inbound
LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86040e40-24ad-4087-9471-978b9f711cf6 · inbound
Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2ca7e14-86e9-45ab-9112-d512c0572d1c · inbound
AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b6ca7271-a549-4947-87c6-96b4d5e84e2f · inbound
ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d3c53df9-fa03-4088-a98d-fc2409abcfd6 · inbound
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 863ace9b-822c-4cb7-b5f6-4a8a1a89481e · inbound
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc2309cc-a33c-4d8a-854a-4a958ba64342 · inbound
CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 526c3869-1c27-4d3b-a04b-08bda1d6c4c1 · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 209
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd6ce112-e622-4ab7-8cbf-dba419703643 · inbound
DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9825cfb-f65c-4be8-9c6a-9eae8e64d9fa · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 243
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 308b3215-00ce-4f70-b0cf-f6b0318882bd · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 243
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a83b9699-f724-4176-8de1-21615976fa81 · inbound
ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed4bf9d-46f4-47da-b1d1-88d7a861b82c · inbound
FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa9c8c8-f4b2-4418-9237-48cb0f51b532 · inbound
FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ce9690-b3f7-48be-9344-2ff340b19901 · inbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c6fb94-ab8a-4374-abb5-e373528bc0a0 · inbound
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65821001-88bc-4a4f-8301-65f88dc8ba45 · inbound
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c693153-dad8-4601-af31-ad398c47d5ec · inbound
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca58211-2071-4b36-82ff-2a6c4b26cb4b · inbound
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f826f4d-6309-454f-a685-1ce916139be0 · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.