Pith. sign in

Paper Citation Record · LEDGER

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

As of 21 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 25 inbound Pith citation observations for arXiv:2509.02020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02020 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:00:45.443144Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:34:47.830741Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:10:07.264049Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved21
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cefde6a1-a184-456f-8e52-c5e09ce4e093 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.274842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.274842Z digest=sha256:670a2fd81a5681906e897f534015b1b75a1fb53dc582356116a4be502d5c48be

Observation 489b64a4-604d-404a-9fa5-6e6923f7132c · outbound

This paper cites FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.279735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.279735Z digest=sha256:b2c1a5e5e5e021ce5ad3832e5b1ae43d2755538cdce198d86f484f6682f13de0

Observation 5d251310-1f84-4d67-a078-5a8b47f82b65 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.284776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.284776Z digest=sha256:2cea6345aa83029576c4066ac74692a4b2dace7132d3ec7a17814723af3eb074

Observation 7d2eacaa-c0cd-46c9-8700-f55b55ee16de · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.289852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.289852Z digest=sha256:c456b9929eb1b3e1a9e497bd15430f22a60b41dc6ae3daa1d591df4134844f77

Observation 29256767-0f50-4d5c-ba91-349d19f274cc · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.294951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.294951Z digest=sha256:f4249476ae7fb0b11e3aa78245a205698ccd2ee62417d50a8cffa19098f9a460

Observation f6b11d61-fc62-4741-9e45-9c2fd67d14da · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.299724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.299724Z digest=sha256:a8f241357b1988359cf28d40068f016810ec38e247dc520661057880fd8d7b87

Observation c68d5cec-178c-4b7f-87a1-08a6cf00b825 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.305208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.305208Z digest=sha256:6a00e39f3d103caef940be674bc6d9583c1a3b64934c8eab35854b2eb215c96b

Observation cc062fb2-9d13-477c-897c-c469a86beab9 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.309565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.309565Z digest=sha256:4620e2c7078c3c019bbafcbbd4337d13555f106f0440c52e65bbd51a016c3ea6

Observation 251bcca0-1cdb-4a25-87c9-fafb3c76c7ec · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.129243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.314513Z digest=sha256:9f0a675ed4cc425ea82baa11e0722b0eb4424f19d094c70a7e590987cf44e555

Observation cc97a2fe-4c5b-4154-80eb-a4cf23513cda · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.318658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.318658Z digest=sha256:a8f6e2d181a83611f045ffd5d903b7392508a5e554f1727a7100a7fc82172259

Observation 82ec05d8-8df7-4746-a823-d2fe5760ef14 · outbound

This paper cites AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.323277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.323277Z digest=sha256:a2bf183665d4337b5e27b76f3e84b65cb5053b04ddad9d9f308424aa28c4f046

Observation 20543e3d-db0e-4037-83ad-ebdca0906f54 · outbound

This paper cites Podagent: A comprehensive framework for podcast generation.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Podagent: A comprehensive framework for podcast generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.114580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.327810Z digest=sha256:9e8dc1eeb2df2b4429c71548d6280b34ebffe4091598c20deee9be6d1b9b37ae

Observation 0dbdac0b-c7a9-4d3c-8c3e-6e38adc875df · outbound

This paper cites Covomix: Advancing zero-shot speech generation for human-like multi-talker conversations.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Covomix: Advancing zero-shot speech generation for human-like multi-talker conversations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.099725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.332221Z digest=sha256:15be89cec0a5298082e26435a944502e2be6d9587a0f139ef2f2bfdb1f0c96f6

Observation b146c7fd-0612-4ee7-b127-281212aa782f · outbound

This paper cites Covomix2: Advancing zero-shot dialogue generation with fully non-autoregressive flow matching.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Covomix2: Advancing zero-shot dialogue generation with fully non-autoregressive flow matching

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.336483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.336483Z digest=sha256:34c6a851e44dc9b530003e274aeba170d38d245aa2729f2a3c94c4f360358d1a

Observation a48fddc6-1147-43b0-9c2b-560188e4da60 · outbound

This paper cites MoonCast: High-Quality Zero-Shot Podcast Generation.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot MoonCast: High-Quality Zero-Shot Podcast Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.340930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.340930Z digest=sha256:e89db65875de3016bacd7e524a3e73fc1c40cff2c68a3bfa4329080c617a47dc

Observation 14c934d0-4c7f-485a-a657-0f88943a0a9d · outbound

This paper cites an unresolved cited work.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:00:46.085072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.345227Z digest=sha256:16c37e3b20f364a97805e9ef938607a1fa8239805e78a7c8c27284736b614af9

Observation df16ea93-8aef-4689-a1b5-51eecc291ec9 · outbound

This paper cites Parakeet, 2024.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Parakeet, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.055182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.354761Z digest=sha256:9158541c688caa10052688dc75d85ae4be4ca11126ecd43dd40a10eac3d415eb

Observation 22b0f101-4787-4608-86c1-c0a5ae3ee09d · outbound

This paper cites ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.359674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.359674Z digest=sha256:dba3a4b2d7034b55a690cf4559e9bc19ed25c8895c145e74ae2904dbf007708d

Observation f467db39-b9b1-492f-b7a2-8bc0d7bee920 · outbound

This paper cites Text to spoken dialogue generation.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Text to spoken dialogue generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.040932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.364350Z digest=sha256:f193cadafab6c5b2bce95f8683f85e269f739953394b1fb9bb8471d6f08cee05

Observation 73c02f5b-98d9-445e-a231-b9897de33e5f · outbound

This paper cites VibeVoice Technical Report.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot VibeVoice Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.368477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.368477Z digest=sha256:07d426b5f7e5ef9ba8f073d9edb315418beacec8dab419514e120ce1831e1b5a

Observation 0afad6cd-6aff-4af9-9e06-252673b779b2 · outbound

This paper cites Crossing the uncanny valley of conversa- tional voice., 2025.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Crossing the uncanny valley of conversa- tional voice., 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.025912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.372784Z digest=sha256:9464c0f86d29fbd097a045c4ce110cc41570c9fcc09690306dfd9da93d601e56

Observation ef406995-7bd4-458a-9f11-208424da88b2 · outbound

This paper cites Codec does matter: Exploring the semantic shortcoming of codec for audio language model.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Codec does matter: Exploring the semantic shortcoming of codec for audio language model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:46.010433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.377062Z digest=sha256:fd9cb23ef319369fc4a3e60c6d24a3ea3aa39dd436480c969ba409b3435da6d8

Observation 889ef837-5acf-4ae4-bc89-3030145ccde1 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.381451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.381451Z digest=sha256:d1f5ae1978c9fd7615c46b92b2a3a42983883e80a4ad074bc1d9f474f70d7a5b

Observation b986fc34-f565-4f28-9bb0-aabbcd94cffc · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.385843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.385843Z digest=sha256:a9e15734a7fac46129144b5e906548424f274561a28b5f697d79e021bb298c21

Observation 25c7f6dc-5a26-403a-a0e8-eeb2b8cad26f · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Moshi: a speech-text foundation model for real-time dialogue

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.391077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.391077Z digest=sha256:8669492cde0104f93b017cb961c49f9075496fe400ae96527369e5857acad748

Observation f63e429c-afad-40d9-9618-54c74be991ff · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Robust speech recognition via large-scale weak supervision

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.994370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.395542Z digest=sha256:251e33681173989d7232d8b4e02faa7179823c62e90141f5d19828c45fa706e0

Observation ab7c535d-0bda-4896-b7da-61b089fcafac · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Soundstream: An end-to-end neural audio codec

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.979068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.399918Z digest=sha256:034195fd059ed1d4aa16f52d2816618860d2839b61b340d86d8ada8997a4066f

Observation cb76bd30-7352-4638-bae8-87eb1a383b24 · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.962367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.404142Z digest=sha256:7673795a9ea9ba53d8a4f42df0a3482a742e8081d554b19a3dffa97cea5a1d11

Observation 4a6632e9-dec5-4d29-b7cc-3d67b369c452 · outbound

This paper cites Scaling transformers for low-bitrate high-quality speech coding.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Scaling transformers for low-bitrate high-quality speech coding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.944951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.409065Z digest=sha256:422c63da4c0003f120c3b2c406b0b6afaf6e146b4f565fb1023c5092191c14bc

Observation 5c2fd124-eeea-489c-ac8e-3a0dbc21d53a · outbound

This paper cites Simple and controllable music generation.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Simple and controllable music generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.929684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.413833Z digest=sha256:9edbdea4e4a66e940cb4adb9840d0707d9569add2eb50ab7c8e8b8dca4db00a4

Observation 36f5adc5-97d5-4776-8b37-15868874c33f · outbound

This paper cites Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.914030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.418098Z digest=sha256:458cac732af4cf66a3d3cb479bea19fb9993471c30023eea4b7444effa737336

Observation 78740931-f52f-485a-baba-662d9142bc05 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.422814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.422814Z digest=sha256:8d5a9ef0f43eed26c03078dcf1c05e708827d3abe680c0ed3941016dc200affb

Observation 6d8e8d91-3e33-49eb-8da4-a442370d71d5 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.427743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.427743Z digest=sha256:fe7ec5b3a55cb6c854b970e0f5a2053f2eed9b79d27b72a62b5eb7d9ceef5984

Observation bf3d3c93-977f-44cf-b01c-dd3aba6139e1 · outbound

This paper cites Qwen3 Technical Report.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Qwen3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.432306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.432306Z digest=sha256:88d41cf6bad17a76d26eaf6778dac4bf1ea6948342fbbf0d94a565aa71682b03

Observation eb2a2da1-0729-4284-94be-73e437d7e9ab · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Powerset multi-class cross entropy loss for neural speaker diarization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.898617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.437900Z digest=sha256:5878d24b2c93690d542ce2f96df874f8759d5f27d936600b5aad8c429388caef

Observation 349115e2-7dc5-49bb-82e8-d2889652fb37 · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:45.882801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.443144Z digest=sha256:6c41559a1e3d1154874cdc8a247fc74cf96b2ea1ef6e386407c7a57d154569a6

Observation ca17f8e5-94ab-4327-a234-057193db9ad1 · outbound

This paper cites an unresolved cited work.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Unresolved cited work

Reference 2025

Resolution
parse uncertain
raw_fallback, observed 2026-08-05T12:00:46.070208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T12:00:45.349698Z digest=sha256:378bb82526b20db74e65658634aea28e4e78201b410f55140ea826ea3efc6939

Pith citing papers

Observation 4bcdddd7-3fe6-47af-9e1d-29abea65a4a3 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.314045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.314045Z digest=sha256:81359af492c2557b8a74c082cbb73049cd9cc0d7f8636106fd40fe4bfad93de9

Observation 676e98a2-d3c0-4c1a-aa9a-15964a6a82ae · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.169130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:5b261c597166e955425b386dce0bfb1a1c29259d1d6a1a54d1800b9bc35a4fbf

Observation 85a899a5-1326-4855-9f0c-b55acd8dfa31 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.066357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:ad5f89025ca139031047675c75421d6ff18cf81412e8d8bfd305d3b80039c088

Observation def0da8e-14cc-441f-90dd-899d9e8f6df3 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.022266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:307e26357092d7b739b4656c0bc7b8af1576b2b02c1de4ae9c4edf090f0d80c7

Observation 3e9c3660-7e7d-4837-85e6-e1db2e1f4755 · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:06.900141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:906563ed2a3f3aa2aa440bcb1f5dbc9bcaac8e8c1778c4627d27f130a30ac266

Observation ea201a97-bad9-4520-a255-b746a50ce4f5 · inbound

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech cites this paper.

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:18.903145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:04:44.050889Z digest=sha256:8a7f541c71c83b68cc5da33fd34c0878c420eb19dc145d1be937195ab8b38ba6

Observation c59f4b74-288e-485e-b9f3-8ef84eb59645 · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.046486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:bf073c07842b498f5d99a199cef39415a33827923b3c5fe06bb4b5bc0fd6f7cb

Observation b07cea6b-7b1f-499d-a707-763bc2d30046 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.571927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:124ff625d10268dade20899d0c60403c101c6449a7d590c538043c84d2522a34

Observation 347fa992-b235-4df3-881b-b67243fca0b3 · inbound

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis cites this paper.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.902062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:203bd29acdfc9f6665e999dd4e0d23d3b50d12740ef4bd25cbbe60e2682cd5a3

Observation 23b15910-3c75-4b15-9600-4ec0c072da49 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.142910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:acb241623151ddf1506f2f59af5f4070fe01f68ede0a455952db4a6b665770b6

Observation 7403fb74-abe3-4dee-85de-5ee6503270fd · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.764041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:41d33224faf4ac73d10705d98d587d3d660db8db199ee72773620f8e1de99ec9

Observation e414a990-d9d6-4970-9391-d82056e56a09 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.789461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:5ad658a7055a6505edba87a5397e5125c04cd7ae39e0a96570379fc2887cbe76

Observation 4af40350-8541-4072-9114-7128ce4677d3 · inbound

dots.tts Technical Report cites this paper.

dots.tts Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.057397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:5c37a18812068dee49a91390940da440dea2ffd0641661e345ba81a43cc6348d

Observation 67ed1d33-45d0-4f92-b587-64bb4b7dd29f · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:07:35.862568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:7788439ebd48154d5f16d9b067f8463704cd827b957eee3ad921a8084f1eb73a

Observation 8976db81-22b6-4a89-b588-c4602aa08a20 · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.155827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:17ea45a02e08ae78d36e74a19b0a24ca96ad4845792326b2c4ea07f6f7ac46f2

Observation 3993a0d3-5e96-4bdc-b983-13a855139985 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.569340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:fab317e5d1a48b2142f1f3fa0968d0313d176d00ba8ff8188961d0e8bc2161a0

Observation 698cc9de-f8e5-476a-a516-bc9dbdae1d65 · inbound

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis cites this paper.

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:10:07.265869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-25T20:43:29.118351Z digest=sha256:9c499d3b933c7d60856012adb1d6b41f91146fb87eea5c814daf689b6ec8e068

Observation dfe95108-84a6-49cc-bf55-78468548ad92 · inbound

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech cites this paper.

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:57.629251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T02:05:46.290721Z digest=sha256:b74e14cc491a44e8f1f4a69beab5e267d3b4933e7142a490af516e3f21806721

Observation 156c0386-b4dd-449f-940e-2a7df766a339 · inbound

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech cites this paper.

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T21:24:36.925360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:24:36.925360Z digest=sha256:db8d6ee8811b368f99d4e5f3edc174c68fa462c678fae9c5103c67df953664dd

Observation 5c8c5536-3544-4e4e-9205-72d53a0d1476 · inbound

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching cites this paper.

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:31:31.608149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:31:31.608149Z digest=sha256:2975704ad236a1277d11985834cc7201e09517fcc6b3e0c44ca2afd2ef33e7ef

Observation b706bb34-4e39-4f81-b778-6bc3753d72b1 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:56.089198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:56.089198Z digest=sha256:bb782ec666827ed3d0dbc14a164a6dd5afec7e8a1263c417d270d0da204daafc

Observation 9c8f1cfb-c94d-4215-ad4d-940288e57853 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:29.493316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:29.493316Z digest=sha256:532efa4f8721362ed4ce5cbbf19d962f9401e028d0df81bfc5891af6f1834913

Observation fc61158c-106d-4638-89d6-178270575d88 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.608557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.608557Z digest=sha256:cf4d855e21b65c7ee0dce6b435874314c0d246b93e39fa105054051e8d6fae23

Observation 89fb8f6b-3ee3-48eb-91b6-9efa46182242 · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:41.980552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:41.980552Z digest=sha256:fe9b092e89bb4cfbd93df20c8b5cdfb5b82973b68516137486fa55a9847dcea8

Observation f6a00cd0-c484-42da-ae3f-df317a6ab822 · inbound

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents cites this paper.

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:47.830741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:47.830741Z digest=sha256:9474f910f024b5787ed0c27d89ef07f2813c59b06df414b4648e84d5ac3dac3b