Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:05.011278Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.11737.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:05.011278Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6299754f-6941-43b3-9c65-1897d925b67e · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f6b3a3-7e86-4b8a-bf33-72248c77c095 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef19d6b0-298f-446a-838c-2426a92e094b · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7db203f-2dcf-4328-b109-506bf25fb32d · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Ds-codec: Dual-stage training with mirror-to-nonmirror architecture switching for speech codec
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 775d7f71-93cd-4c9b-aacd-a5b1b1044671 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e56797-04fd-4565-8829-026f3014147c · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d0725086-5c4b-4806-88ac-ad3c29b57b46 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Neural codec language models are zero-shot text to speech synthesizers.IEEE Transactions on Audio, Speech and Language Processing, 33:705–718, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3a9f138-c9ae-41a3-b09a-77a8b5b633de · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97897886-1e82-4a23-8a54-9682a6b4b788 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf54ab6f-ceb9-43f5-8639-0c71bc5ee143 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0f3479-6dfa-4191-9acf-269cb5071608 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff943c9-eaae-4729-be9c-2d189e7ef9cc · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47121a70-ff15-4021-a5c8-1efac8aa7538 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization High Fidelity Neural Audio Compression
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde3c9fd-c867-4eae-8885-5230e107130d · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee9a78e7-0c36-4084-b44b-9198b2f3fd84 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 848bc857-ef5c-43a4-a6b0-e6b640ba068f · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Conformer: Convolution-augmented Transformer for Speech Recognition
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68c1bcb-43ad-440c-b9ef-a1e1f4946ff5 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76cc8bc-1f3a-4b96-927c-a71265d63cf7 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Didispeech: A large scale mandarin speech corpus
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2cd431-891c-4a07-ad23-0bb3de3965c2 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf71c01-bf65-4aa7-b665-7701c3e67bc2 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM transactions on audio, speech, and language processing, 29:3451–3460, 2021
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3657d78c-4938-47de-9af2-cfa2f82cc64d · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Ditar: Diffusion transformer autoregressive modeling for speech generation.arXiv preprint arXiv:2502.03930, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38fea359-7336-4eca-8aec-8ec2fa84dc1d · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb41eb7-2020-4432-89de-a3686b7633d0 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Libriheavy: A 50,000 hours asr corpus with punctuation casing and context
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 201224c3-6514-4c19-a9a5-45de04100265 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization High- fidelity audio compression with improved rvqgan.Advances in Neural Information Processing Systems, 36:27980–27993, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f71aab3-e39d-4989-b3f4-31a948676316 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e64899a-7835-46d9-bb90-30dbda7a87e4 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Zero-shot Voice Conversion with Diffusion Transformers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b115cb4-af6c-4869-9f42-bc6710c32c45 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea442264-10cd-41a2-98d6-e8cf3df62d6f · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a8bac5-196d-4068-8128-52d64fa6ed3c · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Librispeech-pc: Benchmark for evaluation of punctuation and capitaliza- tion capabilities of end-to-end asr models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ac5d2e-5115-4339-aa2e-12d74942f7f7 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Autoregressive speech synthesis without vector quantization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3caaa9a8-a762-44b4-abb2-2e4f28dba4d6 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Librispeech: an asr corpus based on public domain audio books
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d67925-d54f-4a3d-99e7-02fc5d60a9ff · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Scalable diffusion models with transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbfe518d-49fe-447c-aead-9e343fd9bef5 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Robust speech recognition via large-scale weak supervision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bef0a40f-7795-4b00-ac2a-a536a6ff8af7 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c13e3b-44d7-4ed7-934c-2568ce3b2d8d · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae40bb28-6d73-43a9-a72c-fa9cccee1085 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81311671-a9d0-4270-9e88-5af927fd25c1 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8019c36-794d-4b8e-9aed-58afb3dd4281 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e209853-2f7b-4c21-9c50-9e32064c4973 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization X-VC: Zero-shot Streaming Voice Conversion in Codec Space
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17a4ed74-2ab2-4784-b52c-030b5cb0a95a · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e29db32b-1c62-41d3-8dc8-db4e941dadf6 · outbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning.arXiv preprint arXiv:2509.24650, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.