Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:00:45.443144Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 22 inbound Pith citation observations for arXiv:2509.02020.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:00:45.443144Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:33:04.314045Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:10:07.264049Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cefde6a1-a184-456f-8e52-c5e09ce4e093 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 489b64a4-604d-404a-9fa5-6e6923f7132c · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d251310-1f84-4d67-a078-5a8b47f82b65 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2eacaa-c0cd-46c9-8700-f55b55ee16de · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29256767-0f50-4d5c-ba91-349d19f274cc · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6b11d61-fc62-4741-9e45-9c2fd67d14da · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c68d5cec-178c-4b7f-87a1-08a6cf00b825 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc062fb2-9d13-477c-897c-c469a86beab9 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 251bcca0-1cdb-4a25-87c9-fafb3c76c7ec · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc97a2fe-4c5b-4154-80eb-a4cf23513cda · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ec05d8-8df7-4746-a823-d2fe5760ef14 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20543e3d-db0e-4037-83ad-ebdca0906f54 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Podagent: A comprehensive framework for podcast generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0dbdac0b-c7a9-4d3c-8c3e-6e38adc875df · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Covomix: Advancing zero-shot speech generation for human-like multi-talker conversations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b146c7fd-0612-4ee7-b127-281212aa782f · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Covomix2: Advancing zero-shot dialogue generation with fully non-autoregressive flow matching
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48fddc6-1147-43b0-9c2b-560188e4da60 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot MoonCast: High-Quality Zero-Shot Podcast Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14c934d0-4c7f-485a-a657-0f88943a0a9d · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df16ea93-8aef-4689-a1b5-51eecc291ec9 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Parakeet, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 22b0f101-4787-4608-86c1-c0a5ae3ee09d · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f467db39-b9b1-492f-b7a2-8bc0d7bee920 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Text to spoken dialogue generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73c02f5b-98d9-445e-a231-b9897de33e5f · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot VibeVoice Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0afad6cd-6aff-4af9-9e06-252673b779b2 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Crossing the uncanny valley of conversa- tional voice., 2025
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ef406995-7bd4-458a-9f11-208424da88b2 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Codec does matter: Exploring the semantic shortcoming of codec for audio language model
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 889ef837-5acf-4ae4-bc89-3030145ccde1 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b986fc34-f565-4f28-9bb0-aabbcd94cffc · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25c7f6dc-5a26-403a-a0e8-eeb2b8cad26f · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Moshi: a speech-text foundation model for real-time dialogue
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f63e429c-afad-40d9-9618-54c74be991ff · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Robust speech recognition via large-scale weak supervision
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab7c535d-0bda-4896-b7da-61b089fcafac · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Soundstream: An end-to-end neural audio codec
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb76bd30-7352-4638-bae8-87eb1a383b24 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a6632e9-dec5-4d29-b7cc-3d67b369c452 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Scaling transformers for low-bitrate high-quality speech coding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c2fd124-eeea-489c-ac8e-3a0dbc21d53a · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Simple and controllable music generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 36f5adc5-97d5-4776-8b37-15868874c33f · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 78740931-f52f-485a-baba-662d9142bc05 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d8e8d91-3e33-49eb-8da4-a442370d71d5 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3d3c93-977f-44cf-b01c-dd3aba6139e1 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Qwen3 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2a2da1-0729-4284-94be-73e437d7e9ab · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Powerset multi-class cross entropy loss for neural speaker diarization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 349115e2-7dc5-49bb-82e8-d2889652fb37 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ca17f8e5-94ab-4327-a234-057193db9ad1 · outbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4bcdddd7-3fe6-47af-9e1d-29abea65a4a3 · inbound
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676e98a2-d3c0-4c1a-aa9a-15964a6a82ae · inbound
Qwen3-TTS Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85a899a5-1326-4855-9f0c-b55acd8dfa31 · inbound
CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation def0da8e-14cc-441f-90dd-899d9e8f6df3 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3e9c3660-7e7d-4837-85e6-e1db2e1f4755 · inbound
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea201a97-bad9-4520-a255-b746a50ce4f5 · inbound
Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c59f4b74-288e-485e-b9f3-8ef84eb59645 · inbound
From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b07cea6b-7b1f-499d-a707-763bc2d30046 · inbound
SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 347fa992-b235-4df3-881b-b67243fca0b3 · inbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23b15910-3c75-4b15-9600-4ec0c072da49 · inbound
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7403fb74-abe3-4dee-85de-5ee6503270fd · inbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e414a990-d9d6-4970-9391-d82056e56a09 · inbound
VoxCPM2 Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4af40350-8541-4072-9114-7128ce4677d3 · inbound
dots.tts Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 67ed1d33-45d0-4f92-b587-64bb4b7dd29f · inbound
TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8976db81-22b6-4a89-b588-c4602aa08a20 · inbound
FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3993a0d3-5e96-4bdc-b983-13a855139985 · inbound
EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 698cc9de-f8e5-476a-a516-bc9dbdae1d65 · inbound
Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dfe95108-84a6-49cc-bf55-78468548ad92 · inbound
HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 156c0386-b4dd-449f-940e-2a7df766a339 · inbound
DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8c5536-3544-4e4e-9205-72d53a0d1476 · inbound
ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b706bb34-4e39-4f81-b778-6bc3753d72b1 · inbound
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c8f1cfb-c94d-4215-ad4d-940288e57853 · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.