Pith. sign in

Paper Citation Record · LEDGER

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

As of 6 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 22 inbound Pith citation observations for arXiv:2509.02020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02020 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:00:45.443144Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:33:04.314045Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:10:07.264049Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved21
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cefde6a1-a184-456f-8e52-c5e09ce4e093 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.274842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.274842Z digest=sha256:156f45c5bc0be8454740258cbb00e3dac6b14f87ba3ef3ac2b7f2e7ecc8aaab9

Observation 489b64a4-604d-404a-9fa5-6e6923f7132c · outbound

This paper cites FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.279735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.279735Z digest=sha256:eb125e2a40a90021275f12e6106913c1643dc569283bacc90b10c91e4065fa15

Observation 5d251310-1f84-4d67-a078-5a8b47f82b65 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.284776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.284776Z digest=sha256:98eb95b60f04e9d37fb1d75a9f36acfb8c3b04ec3fc19d5f6abb66a6b2294277

Observation 7d2eacaa-c0cd-46c9-8700-f55b55ee16de · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.289852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.289852Z digest=sha256:5b9ee947658d686809740d5d77e7c94451010f90e7165e9664c7eda62798b86d

Observation 29256767-0f50-4d5c-ba91-349d19f274cc · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.294951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.294951Z digest=sha256:2a9718b34601bd9a5c666c6f561db943f10a15a6c18375f27403d19232298a10

Observation f6b11d61-fc62-4741-9e45-9c2fd67d14da · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.299724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.299724Z digest=sha256:8622285fea783093f60cfd48082be2ecccccd53af2471fbd8993f2f4e63acae4

Observation c68d5cec-178c-4b7f-87a1-08a6cf00b825 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.305208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.305208Z digest=sha256:1ecdba0576818ef5fafa25381a5a703db7c8f2042ec0ba29a7e764f2fcc4a0bb

Observation cc062fb2-9d13-477c-897c-c469a86beab9 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.309565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.309565Z digest=sha256:fe6b12824d56e6629041ffab5ee2e8d8bfbdf6a143146ff762140c3aae8cd5dc

Observation 251bcca0-1cdb-4a25-87c9-fafb3c76c7ec · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.129243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.314513Z digest=sha256:359e2f9edf7df8d75f23c8dc4253d04444b6ed0f9b3e3b031c5ed3835870b3a4

Observation cc97a2fe-4c5b-4154-80eb-a4cf23513cda · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.318658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.318658Z digest=sha256:50e42ec423fc85aa5f3b6db6731e4e96f01ba8ee42b880b4eb16a613e7dc8cd5

Observation 82ec05d8-8df7-4746-a823-d2fe5760ef14 · outbound

This paper cites AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.323277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.323277Z digest=sha256:b98df3139c2f58a0fa71c2cb9ece57c9e701b7dabeb1fe69249b24e7a8c711e7

Observation 20543e3d-db0e-4037-83ad-ebdca0906f54 · outbound

This paper cites Podagent: A comprehensive framework for podcast generation.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Podagent: A comprehensive framework for podcast generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.114580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.327810Z digest=sha256:979d26d14f28c20bcdee087f7b8d2168e0ed568e085c78c46a7b997ac1f6c2c4

Observation 0dbdac0b-c7a9-4d3c-8c3e-6e38adc875df · outbound

This paper cites Covomix: Advancing zero-shot speech generation for human-like multi-talker conversations.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Covomix: Advancing zero-shot speech generation for human-like multi-talker conversations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.099725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.332221Z digest=sha256:605d8d70f11e6ada70d78e17955b5502e09f3ead73109bd8958ab2874630360c

Observation b146c7fd-0612-4ee7-b127-281212aa782f · outbound

This paper cites Covomix2: Advancing zero-shot dialogue generation with fully non-autoregressive flow matching.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Covomix2: Advancing zero-shot dialogue generation with fully non-autoregressive flow matching

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.336483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.336483Z digest=sha256:bc3ca11ef45aa24f3b84050b874ec9ffee028a3748e7814ec2f18148ad22ecc8

Observation a48fddc6-1147-43b0-9c2b-560188e4da60 · outbound

This paper cites MoonCast: High-Quality Zero-Shot Podcast Generation.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot MoonCast: High-Quality Zero-Shot Podcast Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.340930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.340930Z digest=sha256:c593811e429d54913a2551ebf3bcf2a53e4e53e7ec35c0f639699d57e3100610

Observation 14c934d0-4c7f-485a-a657-0f88943a0a9d · outbound

This paper cites an unresolved cited work.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:00:46.085072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.345227Z digest=sha256:0955d5c0efd840ea49fe900545f7bb121eb4892dc1bb1f9b8edecfdcf25025c3

Observation df16ea93-8aef-4689-a1b5-51eecc291ec9 · outbound

This paper cites Parakeet, 2024.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Parakeet, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.055182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.354761Z digest=sha256:1ccb8b0ced78cc2ab0122564ef12de2ab365a263140e61e680c22569fd79c14d

Observation 22b0f101-4787-4608-86c1-c0a5ae3ee09d · outbound

This paper cites ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.359674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.359674Z digest=sha256:4276908b952340b8d13a57ff4c4b424b3c5065be80c282d3162a10f6db2e8ada

Observation f467db39-b9b1-492f-b7a2-8bc0d7bee920 · outbound

This paper cites Text to spoken dialogue generation.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Text to spoken dialogue generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.040932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.364350Z digest=sha256:31e5be85abf15bab27c43280a9a8ae24a8fed04d6b85b8ac3b71469564f47bfe

Observation 73c02f5b-98d9-445e-a231-b9897de33e5f · outbound

This paper cites VibeVoice Technical Report.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot VibeVoice Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.368477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.368477Z digest=sha256:11e0e86e0872d4e511457f9eb9022870f70cd9b866e4cb7874f1939b172c3773

Observation 0afad6cd-6aff-4af9-9e06-252673b779b2 · outbound

This paper cites Crossing the uncanny valley of conversa- tional voice., 2025.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Crossing the uncanny valley of conversa- tional voice., 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.025912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.372784Z digest=sha256:1dc10216c41693e744fb7e249d9f67df5a58d17f24fc5e6e19e5a93e1b3ecea1

Observation ef406995-7bd4-458a-9f11-208424da88b2 · outbound

This paper cites Codec does matter: Exploring the semantic shortcoming of codec for audio language model.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Codec does matter: Exploring the semantic shortcoming of codec for audio language model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.010433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.377062Z digest=sha256:9f515ace72635b0df99420d17e7b4d4ada326b646f1251d5b14c1a7712fd9ba0

Observation 889ef837-5acf-4ae4-bc89-3030145ccde1 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.381451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.381451Z digest=sha256:96a1450617e6590a9f6245869f97aa959c5e2585055f4894ff4814cc1a46397f

Observation b986fc34-f565-4f28-9bb0-aabbcd94cffc · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.385843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.385843Z digest=sha256:bad00b82fbc13fc53878b1e798a9711f1bfa3009938ecb937ddfc579ce4d0ec7

Observation 25c7f6dc-5a26-403a-a0e8-eeb2b8cad26f · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Moshi: a speech-text foundation model for real-time dialogue

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.391077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.391077Z digest=sha256:51ca681d6da13e5464807bb06acefb40beef38ff262f5af69904d00f5645fa14

Observation f63e429c-afad-40d9-9618-54c74be991ff · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Robust speech recognition via large-scale weak supervision

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.994370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.395542Z digest=sha256:893b8874974aa1e46f8047abe3c951426cd37d32a22f6bb030fdb912265da261

Observation ab7c535d-0bda-4896-b7da-61b089fcafac · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Soundstream: An end-to-end neural audio codec

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.979068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.399918Z digest=sha256:2cc27792d074d4add162502adb963eaa9a4e7ba20617fad92ef4067746510bdc

Observation cb76bd30-7352-4638-bae8-87eb1a383b24 · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.962367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.404142Z digest=sha256:3002d5c7d294d176eee5d8328dabf896fb0f39705823596c9e3c8ff8e84e3560

Observation 4a6632e9-dec5-4d29-b7cc-3d67b369c452 · outbound

This paper cites Scaling transformers for low-bitrate high-quality speech coding.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Scaling transformers for low-bitrate high-quality speech coding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.944951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.409065Z digest=sha256:345599fabb2533dce8510c3ef521b01c2ead8b5f6eece4db9c29b02624aa894e

Observation 5c2fd124-eeea-489c-ac8e-3a0dbc21d53a · outbound

This paper cites Simple and controllable music generation.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Simple and controllable music generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.929684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.413833Z digest=sha256:67fe698f8ba964e08d76ef6f2435cffb3c4567877ba6605652f713c732669577

Observation 36f5adc5-97d5-4776-8b37-15868874c33f · outbound

This paper cites Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.914030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.418098Z digest=sha256:275b4ffb2eedb38a3a7819637fdd7dcf1465e3b2d5e374f71300544cf914f9ad

Observation 78740931-f52f-485a-baba-662d9142bc05 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.422814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.422814Z digest=sha256:921b9e41bda288c91f00c64e5551f5c3cc90bbd6aa3b5f771ed468d57469bde4

Observation 6d8e8d91-3e33-49eb-8da4-a442370d71d5 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.427743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.427743Z digest=sha256:2db20adf6cb82bd5c8846ee0cbfec4be167b626c0af100e8a8d8f16ffceb3365

Observation bf3d3c93-977f-44cf-b01c-dd3aba6139e1 · outbound

This paper cites Qwen3 Technical Report.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Qwen3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.432306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.432306Z digest=sha256:c92fd810d60bc0db684cd4a1bcaa932edc27227750d988d9dc61bf77bdecd333

Observation eb2a2da1-0729-4284-94be-73e437d7e9ab · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Powerset multi-class cross entropy loss for neural speaker diarization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.898617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.437900Z digest=sha256:690d372c64e5e0b1bed9a2de3b8c23199c9577e7e2b6f7f5be740b98da5583dc

Observation 349115e2-7dc5-49bb-82e8-d2889652fb37 · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.882801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.443144Z digest=sha256:9b3b78f1e3f58c8eda93024acb8524d4adcfa45542c1a6e7ed2c9adf49d7358b

Observation ca17f8e5-94ab-4327-a234-057193db9ad1 · outbound

This paper cites an unresolved cited work.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Unresolved cited work

Reference 2025

Resolution
parse uncertain
raw_fallback, observed 2026-08-05T12:00:46.070208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T12:00:45.349698Z digest=sha256:f742dc3f44227a783e53f2f43ca8c553d93bf9128f070d7b608f225e8c5c1bec

Pith citing papers

Observation 4bcdddd7-3fe6-47af-9e1d-29abea65a4a3 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.314045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.314045Z digest=sha256:66ccdce36f09236770b727a83b22844d9de4f3af8b6c0b7ee1dead6babb5d3c4

Observation 676e98a2-d3c0-4c1a-aa9a-15964a6a82ae · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.169130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:1f687613853cb352a46e7e823e45340e14b7fd5d8e4f02585be7902c16d4bb55

Observation 85a899a5-1326-4855-9f0c-b55acd8dfa31 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.066357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:84d4c99eb4817a54791ac2941fab13d3cc11bcf542295d94fb2de85767b078af

Observation def0da8e-14cc-441f-90dd-899d9e8f6df3 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.022266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:ecf5f59768cf5e15c35c4c1730685bff9553de2212cf17fe3d8abd72b0f67921

Observation 3e9c3660-7e7d-4837-85e6-e1db2e1f4755 · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:06.900141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:5663a5c11ce057df17ceff8dd30affa788ab0f179e94276c7f20a999438cc36f

Observation ea201a97-bad9-4520-a255-b746a50ce4f5 · inbound

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech cites this paper.

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:18.903145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:04:44.050889Z digest=sha256:10c4eb4f36f41b20621e5d5d0cc4179a3b257ce869847466b55bcab86adca002

Observation c59f4b74-288e-485e-b9f3-8ef84eb59645 · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.046486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:7b148d67b7b3850943d98ceaabde17a3a4be85e2fda1a71c4a19e768ae68dbf4

Observation b07cea6b-7b1f-499d-a707-763bc2d30046 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.571927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:4416ddd3a13202f799ace50d52c84c9e3a1e2b7c98be06a50578d8e1ea0c5bfd

Observation 347fa992-b235-4df3-881b-b67243fca0b3 · inbound

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis cites this paper.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.902062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:bca7ed09ecc37d5bbf5b6da5c75a06b88bc15ff5e10ffdc5f5892018b3f55901

Observation 23b15910-3c75-4b15-9600-4ec0c072da49 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.142910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:ccb6f8484845b6d77cab2922092651f06638637e2d0504a75e4719bae82bb134

Observation 7403fb74-abe3-4dee-85de-5ee6503270fd · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.764041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0136bfe124c8b1fcd371c4586ab18cc2802cf4423b33eb522a606001a8a3b117

Observation e414a990-d9d6-4970-9391-d82056e56a09 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.789461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:fdcf7373fd572a8a74aa2091f38bd9d684186b5442d4e8b359fd6e37f665196d

Observation 4af40350-8541-4072-9114-7128ce4677d3 · inbound

dots.tts Technical Report cites this paper.

dots.tts Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.057397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:c58bf67a258c46a2f06e762f2e978fdef3c930b2b67c8e8cd1019907093143d3

Observation 67ed1d33-45d0-4f92-b587-64bb4b7dd29f · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:07:35.862568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:57f24e1b11d236ab68d43ab0c942dbf9ec32e46b4761b323d0da40bbba79c71d

Observation 8976db81-22b6-4a89-b588-c4602aa08a20 · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.155827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:aa75b36dbadc1ba5eef4dda84d03549e280bd38f708c8b4113f652bdfe90fbef

Observation 3993a0d3-5e96-4bdc-b983-13a855139985 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.569340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:dfd32e36a9f365b779412c5b7dd2368ad89a99eb3c712ba4532aa965344a4745

Observation 698cc9de-f8e5-476a-a516-bc9dbdae1d65 · inbound

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis cites this paper.

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:10:07.265869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T20:43:29.118351Z digest=sha256:eeffbe9f3bd2d146888b55ad200ad0591c0a45b04265a93d9b9ccf39231df105

Observation dfe95108-84a6-49cc-bf55-78468548ad92 · inbound

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech cites this paper.

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:57.629251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T02:05:46.290721Z digest=sha256:9143343feb4240f2d981a9402b371965a83c85b70ae35e53891e0f24d9eaf0ce

Observation 156c0386-b4dd-449f-940e-2a7df766a339 · inbound

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech cites this paper.

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T21:24:36.925360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:24:36.925360Z digest=sha256:8268f42e472fa00bf9f4d7f71a8a1843bf8850608a1ffdf7e52fe87b431004f8

Observation 5c8c5536-3544-4e4e-9205-72d53a0d1476 · inbound

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching cites this paper.

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:31:31.608149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:31:31.608149Z digest=sha256:4d1a65fc23040c3259550b0ee583c8171815297d79064c19e48f04fc5f4934f6

Observation b706bb34-4e39-4f81-b778-6bc3753d72b1 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:56.089198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:56.089198Z digest=sha256:ade129c15760d2c9c6c8cdcf2ea61fbe86087e6ceee5931ab9426d59b0d63a54

Observation 9c8f1cfb-c94d-4215-ad4d-940288e57853 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:29.493316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:29.493316Z digest=sha256:c155b7974fdd352a1d698cd3328613788e009f309580451591ec541f6659bf0c