Pith. sign in

Paper Citation Record · LEDGER

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

As of 9 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 16 inbound Pith citation observations for arXiv:2506.13053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13053 v3

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:45:17.467537Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:39:23.947687Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:34.896349Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46dd6f12-388d-47f2-b4d5-1b58801c4773 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Neural codec language models are zero-shot text to speech synthesizers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:24.192871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:13.214738Z digest=sha256:bbd5d23e56c99e67898bed1c8b76a5ad6d76481f7cf4c14f2e2d3a57f38e8f44

Observation 2cbfe618-9751-4ea9-90ad-f4aac7513486 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.288815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.288815Z digest=sha256:e44ec466d7ea16e4d461b9f0ff012efc262094269b964d1d81a71e59b6911cb1

Observation d4840c43-7c3b-4257-980b-e308b9d53508 · outbound

This paper cites E2 tts: Embarrassingly easy fully non- autoregressive zero-shot tts,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching E2 tts: Embarrassingly easy fully non- autoregressive zero-shot tts,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:24.038565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:13.380779Z digest=sha256:7cce9c069111a54506587391cae0c33cee39084c96c5d69fb00f0ec706cfa585

Observation 19e19724-5546-4746-8948-d7be6e8427d4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.443312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.443312Z digest=sha256:18e3642a0e02c14f285c363f12f9d97c35d11f13c2963461048fe70bd769ccb9

Observation c4ee4589-42be-40e8-b106-1220f2f921ad · outbound

This paper cites MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.893891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:13.540878Z digest=sha256:53e1bc8e0ac82b242efbeb6b3ecd5f6a41dd153688999a903fd06cae1f4491d4

Observation 47614f9a-f3b0-4357-b723-32f725028c9f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.608161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.608161Z digest=sha256:36c6d6f1d2afc4604561390217c650e0076286c198395e56a85cb4dc5eb6db08

Observation 811b01f7-0773-4610-a07c-7bf31ea74375 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.675822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.675822Z digest=sha256:71d57640de238220299d537b378f02fa61b3fb749da68f48686ec2e9dae18f8d

Observation 48f1086b-b0ef-4d23-aeb1-644adc8484d4 · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Libritts: A corpus derived from librispeech for text-to-speech,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.776702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:13.755949Z digest=sha256:2d7286091f9b4be427a25e286f2a460f7c537b99221c47215fb4b374669ff1db

Observation 1a7c5959-0e17-4458-b945-020459b91c96 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.847783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.847783Z digest=sha256:6d2724e33325273ec32f62022e714863f88cc871fcd0a2e6139e9dcb264f845e

Observation e190dbb3-1fbf-4a1f-9e04-e730369071e3 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.630165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:13.919856Z digest=sha256:2075e473c6a83e14820b25a4261be15e47c9bf0a213ebea2669b7818f819cafb

Observation c4b312cb-f7a2-492f-b2c2-3ea22dbe1a6a · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero- shot speech and singing synthesizers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Naturalspeech 2: Latent diffusion models are natural and zero- shot speech and singing synthesizers,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.479251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:13.988814Z digest=sha256:37e197be0096fa5a9adf88811608a7cbac8fd39b96038e090c241a7dbc2b89b5

Observation 0648c2e9-ec80-40e7-b40c-854a16fed321 · outbound

This paper cites Flow matching for generative modeling,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow matching for generative modeling,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.333080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:14.064823Z digest=sha256:2b2a5eb4e6484c05d4bfcd9b807691a46032403bdefe52bb933f8fc881977b70

Observation c726225b-2ddf-465f-9cf5-1cc4fdbf401a · outbound

This paper cites Sf- speech: Straightened flow for zero-shot voice clone,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Sf- speech: Straightened flow for zero-shot voice clone,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.200164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:14.143790Z digest=sha256:04674dd3ff56f5ce6f3390318299850710a463ea2805786a9fbdcc9adcd3ff68

Observation f7b3f264-2b8f-494d-a920-a01aff7949ae · outbound

This paper cites P-flow: A fast and data-efficient zero-shot tts through speech prompting,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching P-flow: A fast and data-efficient zero-shot tts through speech prompting,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.080631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:14.242517Z digest=sha256:1282d337b6b7e0b22428fea2115f4e9b1f9b99fe193d5703066767b85df7e66a

Observation 8e3f0334-850f-41da-9b83-dbe729364cf0 · outbound

This paper cites Attention is all you need,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Attention is all you need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.325866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.325866Z digest=sha256:de58f556bfc6c43ffcb888e358567bf96c3f785d60c1fb1cd0a99bae8e12eee4

Observation 1e0b8f32-b682-4c29-ba91-3741fc249547 · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Zipformer: A faster and better encoder for automatic speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.897794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:14.417592Z digest=sha256:22cc48b023af6a97c924548a7782c03a4dea1e0048e5a196d28f6493f9ca4db5

Observation 2110306a-b124-47bc-be6c-8bd87d96343e · outbound

This paper cites Classifier-free diffusion guidance,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Classifier-free diffusion guidance,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.760185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:14.504825Z digest=sha256:03839d57d4cabf6b1f12dbf03152508bd3e879a0799e9cfe6da326f919688054

Observation 170f12b1-dcbb-4192-bdf7-bbe89d575b1c · outbound

This paper cites Freeu: Free lunch in diffusion u-net,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Freeu: Free lunch in diffusion u-net,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.605205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:14.572199Z digest=sha256:e99671090d23143e0246dd0994e2ab84b74ad42405d0162a9f68e8f273248ec5

Observation 4a2bb12a-be56-4f7b-9cf2-5cc4415addbf · outbound

This paper cites U-dits: Downsample tokens in u-shaped diffusion transformers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching U-dits: Downsample tokens in u-shaped diffusion transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.210548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:14.658013Z digest=sha256:21771f54219008fc11ae92cde0bb8ecbdbab299ecd49af524d77500ace9c28db

Observation 1fad3369-6cbe-4ed7-be2d-5ae2ffe0c242 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Fastspeech: Fast, robust and controllable text to speech,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.731869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.731869Z digest=sha256:ba4715c5e22b33ff70529e4f53e46bf775b679017cab2035c7bb75a1924a243d

Observation 94ec5498-4dba-42f3-ae84-7ae06ad8cc7c · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Conformer: Convolution-augmented transformer for speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.047617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:14.832424Z digest=sha256:f0fdd815c842ff45aad4e4d5bca00370c6c6f4f786a1bbd74bc316a04fe594c2

Observation b9194bd8-2d98-4212-8bfe-fd2d1567bdf4 · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Glow-tts: A generative flow for text-to-speech via monotonic alignment search,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.930725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.930725Z digest=sha256:a045de6c1c4fc7915e837ebf5403ccf764366e737e8ea3e99b600ace51e2e0b1

Observation 64bcbf28-fe5b-4eff-a317-75f4b2b627ca · outbound

This paper cites Flow-tts: A non-autoregressive network for text to speech based on flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow-tts: A non-autoregressive network for text to speech based on flow,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.877248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.033007Z digest=sha256:092e6518ca901e0c721e8cb9201171784274b0a51d801cc8b73c4f9570fe88ad

Observation a84aa383-803f-4188-93d9-7e3c6359d2ab · outbound

This paper cites Simple- speech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Simple- speech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.707431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.151321Z digest=sha256:97a4c27838c62abb59d23de2e0bc01728e83d453b79be738a3c285e8e4d56b22

Observation e5c11d3b-3dc1-4568-b894-f926fe03ec62 · outbound

This paper cites DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.441890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.249852Z digest=sha256:5e114ed64466b3360fd27529ad46176403cf702bfe4c4ec34c1bf7d84b482a88

Observation 9954c653-8243-48f0-8f7c-bee76df3b892 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Convnext v2: Co-designing and scaling convnets with masked autoencoders,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:15.372934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:15.372934Z digest=sha256:e16ddaa4de33935ad55b8ed192281b3e668964e34bdbeb2d1b5df782d8efaf8b

Observation 38ae1325-a8b2-453f-a1ba-4bc4e81cc555 · outbound

This paper cites On distillation of guided diffusion models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching On distillation of guided diffusion models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.229679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.438075Z digest=sha256:0d3eaab6f6d529e2e9dc0ab3371150af3c23c4bfcf98d957a95a16cc1be69b6e

Observation 8a9aeb93-b8ac-45a2-9e03-00166a481c25 · outbound

This paper cites Tacotron: Towards end-to- end speech synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Tacotron: Towards end-to- end speech synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.032849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.495205Z digest=sha256:b31edc5afc3f189eb9afe15babd7cb4bd430d598cd5c1865cba4d69a16c0d7df

Observation c94e5b87-c82e-415e-b725-0232d8a39a93 · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:15.573777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:15.573777Z digest=sha256:2fcd967ea58e92573df1bbd20fa93fdd3e5c96bf990c045f20f6a311c922fd73

Observation 40a0df60-3838-450b-ae44-fdff04008690 · outbound

This paper cites Revisiting Over-Smoothness in Text to Speech.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Revisiting Over-Smoothness in Text to Speech

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:45:17.675731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.639463Z digest=sha256:8e3e471a2da53c11d9134f28c7dd72f9dd2922ecd3f9e399ff97995c00f1fb87

Observation e61555de-6d6c-4486-9be1-8c877b4ac97c · outbound

This paper cites Grad- tts: A diffusion probabilistic model for text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Grad- tts: A diffusion probabilistic model for text-to-speech,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.775881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.700325Z digest=sha256:4711cbb98e0dfe5994a5dad0e85fc087d6edf3ca602c4e57254ff1a16d996a7c

Observation 52d9923c-d046-424b-b4af-68dee322933b · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.589017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.776492Z digest=sha256:8f9de7e15ad7c8c22237d13418365e3f04065bfcd27c278742421c661746f024

Observation e5a4718e-8a2d-4a26-be94-da876beb5e8e · outbound

This paper cites Consistency models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Consistency models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.391621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.839694Z digest=sha256:df5f895f98de59658812f04f36d9ad0b4bcebd4744f16d603db6b0d019c7185b

Observation b152212b-1d78-43b0-bb85-10d3a4ece214 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow straight and fast: Learning to generate and transfer data with rectified flow,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.190716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:15.926474Z digest=sha256:07b61f956c73b42fb15652962c768005d80aea60d588ac9884981249e2878ae1

Observation f858a989-c3cc-448b-a3f4-7e1ede8df9f3 · outbound

This paper cites Comospeech: One-step speech and singing voice synthesis via consistency model,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Comospeech: One-step speech and singing voice synthesis via consistency model,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.931419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:16.004306Z digest=sha256:34b426e2cf4cff1abdd971ed3fa7c6c219f2924c077135c87866311c01cdfb7d

Observation e52ae0af-4104-4462-99cb-d18419bc7542 · outbound

This paper cites Reflow- tts: A rectified flow model for high-fidelity text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Reflow- tts: A rectified flow model for high-fidelity text-to-speech,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.730383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:16.052933Z digest=sha256:b54e6516bee48526e3e812cecb340f0c3d179f00f053876404f1fad0a901886d

Observation da6a2ad3-611e-4da0-9196-71b71c128ca1 · outbound

This paper cites V oiceflow: Efficient text- to-speech with rectified flow matching,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V oiceflow: Efficient text- to-speech with rectified flow matching,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.567457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:16.144964Z digest=sha256:f4c5b040e333cb498b5da8a3017f3bd220cf5177169b5c8ad38fdd32b7817e8f

Observation 6e52ee7d-d6cf-49f8-a13f-9e2be46f25db · outbound

This paper cites Flashspeech: Efficient zero-shot speech synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flashspeech: Efficient zero-shot speech synthesis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.354922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:16.189257Z digest=sha256:58007a51ba080010d1b229c2609918ce355aa5fb1b1509fd253deeb0456d1077

Observation fd3bdbbb-af80-4d3d-beef-ca010a89d28b · outbound

This paper cites Slimspeech: Lightweight and efficient text-to-speech with slim rectified flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Slimspeech: Lightweight and efficient text-to-speech with slim rectified flow,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.136763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:16.251832Z digest=sha256:da6a58c4ae224c1b0190a05a7671832ae55667f2a9720a90f8ee7c01000a3d51

Observation 92523cae-b210-41cc-ae14-2d9e03326ba3 · outbound

This paper cites Lightspeech: Lightweight and fast text to speech with neural architecture search,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Lightspeech: Lightweight and fast text to speech with neural architecture search,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.922989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:16.341634Z digest=sha256:8114b47042bc35f0d83c0d66fc69dc723db41bca3df8445abdc1b3546453a5a5

Observation ceb84728-7a54-48fa-8c99-e7d41459232b · outbound

This paper cites Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.724371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:16.410344Z digest=sha256:40e2da2c4b779a54b520c054a5b8cc97bc14f10d22fe676060c0bada23462b80

Observation a08e7b59-6a11-42a9-b2db-21d3c47cefe3 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Common voice: A massively-multilingual speech corpus,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.561055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:16.499791Z digest=sha256:1d9b75a13ab22f94411d1558da14eea9474de98004528e9dd6b8d955ab906b75

Observation 4bf2354c-ea56-44df-9c1f-b1f82f44ebda · outbound

This paper cites Didispeech: A large scale mandarin speech corpus,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Didispeech: A large scale mandarin speech corpus,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.414015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:16.572166Z digest=sha256:3623bbfd508e68c9587408c1ed28816d5721fdfa378bf4fd686dd247a52742d7

Observation b922f287-09de-418f-95d3-a394b95e554c · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.635875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.635875Z digest=sha256:226d6749577069776d916adec027945b2ee5bc181dee5388c7b93cf93dfe27ce

Observation 66d38499-0d1c-4dba-9295-38bca924f20c · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Robust speech recognition via large-scale weak supervision,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.715337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.715337Z digest=sha256:1f7c8c103fdf7b4b9ff5c1e7df0442ea0103262d043676920acdd24a1b52605d

Observation 28641d09-7533-4f30-9ff5-06211cc5fb0e · outbound

This paper cites Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.775147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.775147Z digest=sha256:e5df0ead2550e6dcff36de047790230c1d556a44730e594a4de3fe89ae6e7e22

Observation 87dc4532-6b2a-4054-85d9-5caaf380e904 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.855648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.855648Z digest=sha256:179588aa873df21b91dd8e93de3d9ba70a4672b498d5f687dd5783e539f037bc

Observation bf4eb6b6-3f38-412d-a71e-47c23e715ee4 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.920932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.920932Z digest=sha256:bfc4cd01e0546ab3c3e50bd1f547a2ae5b17df06fee26a88bf95ba4b3aa28cf8

Observation 523f1dfa-d805-481b-a534-f640a4ae254a · outbound

This paper cites Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.031709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.031709Z digest=sha256:a3bda89fb78d3ae9d7d1e01881a1fc0089033f079f67b90807351beaeb4dbe0c

Observation 31e12026-c525-468c-a783-f51f3b05bb96 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.197912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:17.111078Z digest=sha256:c0947ff73bfc2d1bf99ed30f88dd8aa831e53cbeec68d3487eae40f042c590b1

Observation 263e86b0-e18e-426c-8237-e6dcd3010d8e · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.180878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.180878Z digest=sha256:038c63df969fa2ac02b8cdf8347200d88209ff82d92dc4fb238a9cf705d93778

Observation 80c5d825-271b-4e00-9722-bc71a0a9a635 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.249552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.249552Z digest=sha256:1f525218909be144a2712875d7e7aba1350124c69faa51c1b9b813878b34c212

Observation 45102f47-a413-4735-b698-7e1c2858eaa9 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.314055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.314055Z digest=sha256:a4872e6a10cc23d4f4f8f54b526892f5f3fca308f63ef96a9c93ea5f1e5ab00f

Observation 11b0aff7-79ae-4cc5-9a05-3e065b821938 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.058741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:17.387188Z digest=sha256:4bbae5190c60b474dc227ee9f69659e4d640bd2343b9e212658a2831a42971a9

Observation df6b7d1e-4c64-4752-ab52-b75bfd80e94d · outbound

This paper cites Amphion: an open-source audio, music, and speech generation toolkit,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Amphion: an open-source audio, music, and speech generation toolkit,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:17.878084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:45:17.467537Z digest=sha256:4349b3c33e4573b646cc0f3b7dfe32ca204898c5c3556e75f493efd1c7de4ea5

Pith citing papers

Observation 1f256cb3-22a5-4e2f-be92-73228fa1566d · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:32:03.673585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:a85c7990c5a9cbcbc67bf163c5a3e7501383c24d8f406179d7815513cf9a050c

Observation ab2e20d5-7d2d-4ad0-b454-c474d228d099 · inbound

Universal Speech Content Factorization cites this paper.

Universal Speech Content Factorization ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-15T12:21:48.698333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:21:48.698333Z digest=sha256:6ce8e74e30804ed8ca29a21e08ce1281395273297d8f7cc15bd70efa33e6720a

Observation ef4f0b28-14ef-48b4-aad7-45836011a635 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.708672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:ef69f46dce453b1137bd974d74ec916fd5c10a27e618079eca0efd30845261c3

Observation a7b0fc38-7d06-4288-9379-6e9f4c2ce3d3 · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.048637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:12314e74450f0f113c4c41ed92cd6b4a6cddfbcebe3db120c6bd4aec1ca47b06

Observation 93f45acc-75fe-4625-954f-7155ac151c43 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:36:45.127244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:28:14.734682Z digest=sha256:7fb0ba425edf5bf61a44f84fcbd225364976b4b6feb60d6643edf724726e7c00

Observation a148795a-e901-425f-921a-39a80ed99b83 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.821926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T23:26:46.077894Z digest=sha256:3598e1a1031f381e46ba03bc528b3c0a661c8cc4df24d25c6189fae19bde9a40

Observation fce86089-99f4-4aab-9876-524fa05de4cd · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.049565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:c7b84e5cc6dedca4633e29165002351a22c67f568e67a36bf5350ee7c867e582

Observation 456e7331-2449-4af2-a6c5-fa6a9ce4ec35 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.093666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:52aa882bb0435bdaa9ca406a6f8ac73fdd13b1270eb535296b9d62a123e6deb5

Observation 35be436f-b14d-462d-9f24-c9059434f62d · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.787929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:aa151714ebea45a7e2ac0f4c6eddb3f0932babadbda89aab1dc307ca7651662a

Observation d30edfa3-c12d-4673-9a20-d17c3e8d3216 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:47:19.794924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:cf2f1e57eefe8f76c4ae8001281ad70cab1e22c47b91fc45971aec3bb5d794cb

Observation d1c78e43-7078-41ac-acf0-683a37a18d96 · inbound

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation cites this paper.

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:19.602737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:14:08.243894Z digest=sha256:68081532428dcd3f21d83a16aeb84b792277cbc800f553c801c356e2448f8d28

Observation a83fb087-1b43-4394-b06b-ed907f6a3bd5 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.898147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:964982ffad36191e39ac4fc87da8b197f99dcbbb949d928493c6fdbc047a6811

Observation 7a947619-156c-491b-a449-a4279363051c · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:402c7bbcca3882c29f2566dc486a41cc49311a233ae79e3d8857b6d21d03b5f2

Observation 0297f999-eccb-423d-876a-5296308efd40 · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T02:22:47.820537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:22:47.820537Z digest=sha256:6eff157f0e5064070f261444911bc5f6325a1f9235d645a9f43f190700f0e366

Observation e07bb246-f5cc-41a3-9d83-cc58e8eec40d · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T07:39:23.947687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:39:23.947687Z digest=sha256:3b6375136aa5836e3beb8c36ea836e664749f573e3b032e831cf5a7382012ed6

Observation 51c0f031-9d9e-4949-9112-f3042e5cf35e · inbound

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model cites this paper.

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T22:26:14.049693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:26:14.049693Z digest=sha256:db5de197f8446aa9c39975ccacc643b563f128b154642c93426b88d6dd9c995d