Pith. sign in

Paper Citation Record · LEDGER

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.11737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11737 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:05.011278Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6299754f-6941-43b3-9c65-1897d925b67e · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.855832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.855832Z digest=sha256:071e6c457c586df92eee044167c742f465cb2a48e5465de5f39517d0fa6b62e0

Observation 24f6b3a3-7e86-4b8a-bf33-72248c77c095 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.860178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.860178Z digest=sha256:61e85ea8a30a11deb59dbc8707126b968f93d7f1768831518df7c03d3278aea0

Observation ef19d6b0-298f-446a-838c-2426a92e094b · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.864550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.864550Z digest=sha256:8011bf077d1c3c6e97b102de29c85ed354a8edd74dd4c2c18e5323581671ff43

Observation c7db203f-2dcf-4328-b109-506bf25fb32d · outbound

This paper cites Ds-codec: Dual-stage training with mirror-to-nonmirror architecture switching for speech codec.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Ds-codec: Dual-stage training with mirror-to-nonmirror architecture switching for speech codec

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.631902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:04.868619Z digest=sha256:c14bbb2c2c1177433b72e4abafa937c41c5a3cde28b59460eb7c8e3b4c344126

Observation 775d7f71-93cd-4c9b-aacd-a5b1b1044671 · outbound

This paper cites SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.872459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.872459Z digest=sha256:b4dbbc88773917c08700dc4c44066bccd9e63123c1c421ac307e3e1e54500b80

Observation 26e56797-04fd-4565-8829-026f3014147c · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.616681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:04.876165Z digest=sha256:7051174d854e17d33b6391fdde61483a6c03d2826910dc7f9959475a4c4f0769

Observation d0725086-5c4b-4806-88ac-ad3c29b57b46 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers.IEEE Transactions on Audio, Speech and Language Processing, 33:705–718, 2025.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Neural codec language models are zero-shot text to speech synthesizers.IEEE Transactions on Audio, Speech and Language Processing, 33:705–718, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.880497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.880497Z digest=sha256:ce4613b3e7c41a9adec11ea79d98e291df4790e15c6cf275db1d1ffeb15e0593

Observation f3a9f138-c9ae-41a3-b09a-77a8b5b633de · outbound

This paper cites F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.884261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.884261Z digest=sha256:0ea58f16b112079ad80de3a3cbc9231ecc5ebc735da81e8c75a06de92b30dfed

Observation 97897886-1e82-4a23-8a54-9682a6b4b788 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.888080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.888080Z digest=sha256:60d74a47984454a8349cfed65d235949748b083dfab3446b1b77b16f04041777

Observation cf54ab6f-ceb9-43f5-8639-0c71bc5ee143 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.891909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.891909Z digest=sha256:a7b2f2b15328ccd08336d9c939c690eea727c3290c6b074ffd7ec79179a99654

Observation 4f0f3479-6dfa-4191-9acf-269cb5071608 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.896325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.896325Z digest=sha256:48229e0d65205cb95095723b3435937fd56530dc13526670cb86909bda69d945

Observation dff943c9-eaae-4729-be9c-2d189e7ef9cc · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.900278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.900278Z digest=sha256:dda513cc6a4c5d263ebd60d65568dee85578799fc45f7b9d71520d02d820c9da

Observation 47121a70-ff15-4021-a5c8-1efac8aa7538 · outbound

This paper cites High Fidelity Neural Audio Compression.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization High Fidelity Neural Audio Compression

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.903997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.903997Z digest=sha256:d808923d77d33cd145bad20c9c014d4b9838bc924f9051f46d2c19f2aad65c0f

Observation cde3c9fd-c867-4eae-8885-5230e107130d · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.907580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.907580Z digest=sha256:a02b052f76eeb51bb15b96967dabacef543c61ea80817c75154727e0b9235c8c

Observation ee9a78e7-0c36-4084-b44b-9198b2f3fd84 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.911143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.911143Z digest=sha256:2bbef63232e3bd647e367d819a7b20bdf444157602d63a32449f9aef55e9930b

Observation 848bc857-ef5c-43a4-a6b0-e6b640ba068f · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.915043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.915043Z digest=sha256:c9ca224822aff4beb5042a93896a327509c1b67bdc4eb61447b1dce275aafb4f

Observation d68c1bcb-43ad-440c-b9ef-a1e1f4946ff5 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.919126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.919126Z digest=sha256:bb493fae2678132be788b8eb84f20a2740d883fb98aec7e2b3056829e814ba02

Observation f76cc8bc-1f3a-4b96-927c-a71265d63cf7 · outbound

This paper cites Didispeech: A large scale mandarin speech corpus.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Didispeech: A large scale mandarin speech corpus

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.923020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.923020Z digest=sha256:c53f0684efffbc0219700c604e9c4d6a81dc37606755ce0731c679750a7818b8

Observation 2f2cd431-891c-4a07-ad23-0bb3de3965c2 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.926907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.926907Z digest=sha256:8129f796f639baae9dc26aeae55e022efe0a57cc1131f0e028f46e6738066569

Observation faf71c01-bf65-4aa7-b665-7701c3e67bc2 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM transactions on audio, speech, and language processing, 29:3451–3460, 2021.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM transactions on audio, speech, and language processing, 29:3451–3460, 2021

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.549427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:04.931008Z digest=sha256:3dcb896cd7777e74ac9677eaf2abdc60063f51e70006d941b44c879bb4b942fb

Observation 3657d78c-4938-47de-9af2-cfa2f82cc64d · outbound

This paper cites Ditar: Diffusion transformer autoregressive modeling for speech generation.arXiv preprint arXiv:2502.03930, 2025.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Ditar: Diffusion transformer autoregressive modeling for speech generation.arXiv preprint arXiv:2502.03930, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.935124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.935124Z digest=sha256:624b03183f9b8c89a3cd64a8e730f3f9b56deb8f0524897755687556e06d255c

Observation 38fea359-7336-4eca-8aec-8ec2fa84dc1d · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.938708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.938708Z digest=sha256:c7ed5d05850900b12facde7bb69b331423336cbe29e4aecf4651a205237d2be0

Observation 7eb41eb7-2020-4432-89de-a3686b7633d0 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Libriheavy: A 50,000 hours asr corpus with punctuation casing and context

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.534863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:04.942649Z digest=sha256:0f861287d9379f576d879712314057ad6af99e77d9f5f9c3ae19a32df1a889e1

Observation 201224c3-6514-4c19-a9a5-45de04100265 · outbound

This paper cites High- fidelity audio compression with improved rvqgan.Advances in Neural Information Processing Systems, 36:27980–27993, 2023.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization High- fidelity audio compression with improved rvqgan.Advances in Neural Information Processing Systems, 36:27980–27993, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.521990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:04.946072Z digest=sha256:c6ea9fe83df15aad405563480fe0f03856cedf4f305f3e5e83e85c58ca535f52

Observation 0f71aab3-e39d-4989-b3f4-31a948676316 · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.949485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.949485Z digest=sha256:675279c97776b300e0a292bfeeb0c928da827636595f080e361a3003d8e109b5

Observation 6e64899a-7835-46d9-bb90-30dbda7a87e4 · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Zero-shot Voice Conversion with Diffusion Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.953206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.953206Z digest=sha256:b3581b0521223600abcbff4599134214e31ce13bd1ac93f882b67440162b38e7

Observation 4b115cb4-af6c-4869-9f42-bc6710c32c45 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.957117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.957117Z digest=sha256:e91b5eb29988890754ab7ebb55977031f17aa09317c15f10c0dd761807851b54

Observation ea442264-10cd-41a2-98d6-e8cf3df62d6f · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.960658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.960658Z digest=sha256:689ed10ac166acba19e32a44add9b0091e247d783c091db109d430718cb96d22

Observation c0a8bac5-196d-4068-8128-52d64fa6ed3c · outbound

This paper cites Librispeech-pc: Benchmark for evaluation of punctuation and capitaliza- tion capabilities of end-to-end asr models.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Librispeech-pc: Benchmark for evaluation of punctuation and capitaliza- tion capabilities of end-to-end asr models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.964166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.964166Z digest=sha256:51912b2223c5151b3cac2f871251abcb58d800239c827eff3a208d3bf406587c

Observation 90ac5d2e-5115-4339-aa2e-12d74942f7f7 · outbound

This paper cites Autoregressive speech synthesis without vector quantization.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Autoregressive speech synthesis without vector quantization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.967650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.967650Z digest=sha256:8286b43a97ebb2e152d054174494104b98cdca2ab10139d43c6ed4fe7df2b558

Observation 3caaa9a8-a762-44b4-abb2-2e4f28dba4d6 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Librispeech: an asr corpus based on public domain audio books

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.971209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.971209Z digest=sha256:f6e685b5da92455acfbf433a0c95e4b34a4d9c7e76bbfa75099477240abd6273

Observation a9d67925-d54f-4a3d-99e7-02fc5d60a9ff · outbound

This paper cites Scalable diffusion models with transformers.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Scalable diffusion models with transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.975118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.975118Z digest=sha256:fd0cf8354bd807903e89525a52a6e4f64f5159cac800f659db3184a8757c28b9

Observation dbfe518d-49fe-447c-aead-9e343fd9bef5 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Robust speech recognition via large-scale weak supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.978650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.978650Z digest=sha256:75c8afc1cc26cc0ef444586cd752caaa4f9b7290e399006d1f33acae09537e85

Observation bef0a40f-7795-4b00-ac2a-a536a6ff8af7 · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.982405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.982405Z digest=sha256:203c7035f18071f6aa9ad31069cd941c4581f2b9829d841323260b09981d0616

Observation a1c13e3b-44d7-4ed7-934c-2568ce3b2d8d · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.986772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.986772Z digest=sha256:5f6aa3f032206e8befc5f274b04865586af1f1253b2bc9ddb3019c1991fbe7d3

Observation ae40bb28-6d73-43a9-a72c-fa9cccee1085 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.990844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.990844Z digest=sha256:d428eeca7e1b01617d4e5e0427308ec7100db3ce2074f42a761c4d143aa00085

Observation 81311671-a9d0-4270-9e88-5af927fd25c1 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.994189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.994189Z digest=sha256:06bbe947555f8d5c8a20b3c3635a0953892fc4b541f93d0366df7510e61b886c

Observation c8019c36-794d-4b8e-9aed-58afb3dd4281 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.998310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.998310Z digest=sha256:7306049245894c0b09412ef924db17d407784ac6156a208f50489dbc39b856d3

Observation 3e209853-2f7b-4c21-9c50-9e32064c4973 · outbound

This paper cites X-VC: Zero-shot Streaming Voice Conversion in Codec Space.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization X-VC: Zero-shot Streaming Voice Conversion in Codec Space

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:05.002453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:05.002453Z digest=sha256:72d43fb309b64e2948555f2a827fb1f2f58b5e569df7756f758f18bcd445fe21

Observation 17a4ed74-2ab2-4784-b52c-030b5cb0a95a · outbound

This paper cites Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:05.007627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:05.007627Z digest=sha256:29f9802a7b97dbb5c32c495bf89a76e0faac9bfa15c1e8d880202353f69cb3c4

Observation e29db32b-1c62-41d3-8dc8-db4e941dadf6 · outbound

This paper cites V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning.arXiv preprint arXiv:2509.24650, 2025.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning.arXiv preprint arXiv:2509.24650, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:05.011278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:05.011278Z digest=sha256:244889e17870fc5d0b5fdcc251daa5cf4f134960c6e41f1d0d2e17fd366160c0

Pith citing papers

No inbound Pith citation observations are available.