Pith. sign in

Paper Citation Record · LEDGER

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 16 inbound Pith citation observations for arXiv:2506.13053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13053 v3

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:45:17.467537Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:39:23.947687Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:34.896349Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46dd6f12-388d-47f2-b4d5-1b58801c4773 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Neural codec language models are zero-shot text to speech synthesizers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:24.192871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:13.214738Z digest=sha256:1d690240aa3fa5e7220a9e8eeab6599e74131c556b4c0c5beda9287735b646e6

Observation 2cbfe618-9751-4ea9-90ad-f4aac7513486 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.288815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.288815Z digest=sha256:895379adc223edf1698f52f749983c78f17e363ee4f02059b19e4ec8d91d679c

Observation d4840c43-7c3b-4257-980b-e308b9d53508 · outbound

This paper cites E2 tts: Embarrassingly easy fully non- autoregressive zero-shot tts,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching E2 tts: Embarrassingly easy fully non- autoregressive zero-shot tts,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:24.038565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:13.380779Z digest=sha256:9b08ed05c47c3f569539ec01e740af996fc8314b7bf91e998a9c5806292ab8fe

Observation 19e19724-5546-4746-8948-d7be6e8427d4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.443312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.443312Z digest=sha256:156d4ec63e05d2268ff43ccc736ed88f88399c392c09f27006a2592292f577ab

Observation c4ee4589-42be-40e8-b106-1220f2f921ad · outbound

This paper cites MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.893891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:13.540878Z digest=sha256:eda7a8d50fa0895d4f026057a311800f8edd34deab8979fea1efea22c2c65286

Observation 47614f9a-f3b0-4357-b723-32f725028c9f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.608161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.608161Z digest=sha256:35d2a359cd9d559d20a2a0202f17c4382637d98233242a4ea8631a4afc8133b5

Observation 811b01f7-0773-4610-a07c-7bf31ea74375 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.675822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.675822Z digest=sha256:8b938ccd30c4510e2c498ba1113143d2fbea9fa3124031e237e6efcfb3a2fcda

Observation 48f1086b-b0ef-4d23-aeb1-644adc8484d4 · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Libritts: A corpus derived from librispeech for text-to-speech,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.776702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:13.755949Z digest=sha256:b50739835a89f5a5a567200e48064f0333dbbc1ec62ad514a0d34fdd832fe7c0

Observation 1a7c5959-0e17-4458-b945-020459b91c96 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.847783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.847783Z digest=sha256:4c0fe002d9f193a33bbdfb78f31363826b7faa5213702d285a12a079029f8dfe

Observation e190dbb3-1fbf-4a1f-9e04-e730369071e3 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.630165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:13.919856Z digest=sha256:18753cb3fa864ccb92bc34b0d8d44079839dea818ef143babb733fa58c94041a

Observation c4b312cb-f7a2-492f-b2c2-3ea22dbe1a6a · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero- shot speech and singing synthesizers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Naturalspeech 2: Latent diffusion models are natural and zero- shot speech and singing synthesizers,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.479251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:13.988814Z digest=sha256:30196110dce18bf8590a8dc23545140cea37c2fe0b4726fe1712bacbed1d1473

Observation 0648c2e9-ec80-40e7-b40c-854a16fed321 · outbound

This paper cites Flow matching for generative modeling,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow matching for generative modeling,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.333080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:14.064823Z digest=sha256:c26868428b357ff4ac6bbaa26de2f28dbfecf488adfdbe53bb03b299bd9b4a84

Observation c726225b-2ddf-465f-9cf5-1cc4fdbf401a · outbound

This paper cites Sf- speech: Straightened flow for zero-shot voice clone,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Sf- speech: Straightened flow for zero-shot voice clone,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.200164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:14.143790Z digest=sha256:26a9e883c7bd9f8f264fe35a5f522051be118204777c7f1b1be8899a157d257c

Observation f7b3f264-2b8f-494d-a920-a01aff7949ae · outbound

This paper cites P-flow: A fast and data-efficient zero-shot tts through speech prompting,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching P-flow: A fast and data-efficient zero-shot tts through speech prompting,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.080631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:14.242517Z digest=sha256:776498616d6ed1648f8a2e32f67e6cfe55a432835e4d151fe5043f84c9220e4b

Observation 8e3f0334-850f-41da-9b83-dbe729364cf0 · outbound

This paper cites Attention is all you need,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Attention is all you need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.325866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.325866Z digest=sha256:b2bdcac1ac732e231aad677db8e47c49cfb49a1f7b7143eaae9b1ee18708bd65

Observation 1e0b8f32-b682-4c29-ba91-3741fc249547 · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Zipformer: A faster and better encoder for automatic speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.897794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:14.417592Z digest=sha256:2f33e01f9d8586e52e33ad82661dae413f84011a64aaf15df245562201829d17

Observation 2110306a-b124-47bc-be6c-8bd87d96343e · outbound

This paper cites Classifier-free diffusion guidance,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Classifier-free diffusion guidance,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.760185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:14.504825Z digest=sha256:915e0208840f48f393e329d4ba06164d7d4cadab0976c8c33cba26376d3bc45d

Observation 170f12b1-dcbb-4192-bdf7-bbe89d575b1c · outbound

This paper cites Freeu: Free lunch in diffusion u-net,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Freeu: Free lunch in diffusion u-net,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.605205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:14.572199Z digest=sha256:277a332f10c93e1567499e1fe3e599d5527484ed7d8912984b58a1caedb11897

Observation 4a2bb12a-be56-4f7b-9cf2-5cc4415addbf · outbound

This paper cites U-dits: Downsample tokens in u-shaped diffusion transformers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching U-dits: Downsample tokens in u-shaped diffusion transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.210548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:14.658013Z digest=sha256:abed4cfdd6d9b60151685e94875bf15ec39561b86b56c5ae4c16e9cc70b4e495

Observation 1fad3369-6cbe-4ed7-be2d-5ae2ffe0c242 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Fastspeech: Fast, robust and controllable text to speech,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.731869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.731869Z digest=sha256:88a2469f376cab9b6727ed1954ddc3241477979b4e266b8d2e72d6496e50fe0c

Observation 94ec5498-4dba-42f3-ae84-7ae06ad8cc7c · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Conformer: Convolution-augmented transformer for speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.047617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:14.832424Z digest=sha256:60b903cbc840fcdb1a441a0de18d80d832692c9b9a281cf375a9450be28d8909

Observation b9194bd8-2d98-4212-8bfe-fd2d1567bdf4 · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Glow-tts: A generative flow for text-to-speech via monotonic alignment search,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.930725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.930725Z digest=sha256:47cb8a610d181b28a672b2a534e592d6442a62ca9dc0c60ee6800ba57aa197da

Observation 64bcbf28-fe5b-4eff-a317-75f4b2b627ca · outbound

This paper cites Flow-tts: A non-autoregressive network for text to speech based on flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow-tts: A non-autoregressive network for text to speech based on flow,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.877248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.033007Z digest=sha256:5aa373d9abe82bb96d52e4d5e0cf27cdb65c1bf26fe1b32b16beeabf3cd414e9

Observation a84aa383-803f-4188-93d9-7e3c6359d2ab · outbound

This paper cites Simple- speech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Simple- speech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.707431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.151321Z digest=sha256:a85614bbb25961c1b2bc91d37939f34dbd218b240b5226bbab12c933e4349406

Observation e5c11d3b-3dc1-4568-b894-f926fe03ec62 · outbound

This paper cites DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.441890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.249852Z digest=sha256:567982bb4313a97a2ef63e483d4cfb2f1cb21e2be8911d87cd07154643997f65

Observation 9954c653-8243-48f0-8f7c-bee76df3b892 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Convnext v2: Co-designing and scaling convnets with masked autoencoders,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:15.372934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:15.372934Z digest=sha256:0b539f7a11af9943871a0dda2ce470c1c06764f2524b372021c7e98ee82b21b2

Observation 38ae1325-a8b2-453f-a1ba-4bc4e81cc555 · outbound

This paper cites On distillation of guided diffusion models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching On distillation of guided diffusion models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.229679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.438075Z digest=sha256:9e80fa05545b7bf9c5a32f7f57870a6e99a79bbdd12b97c5dfdd46a78c2ae34d

Observation 8a9aeb93-b8ac-45a2-9e03-00166a481c25 · outbound

This paper cites Tacotron: Towards end-to- end speech synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Tacotron: Towards end-to- end speech synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.032849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.495205Z digest=sha256:3222578ea25af37d5214c24f5073d7cedb9e717721889e732c01cf315d99b573

Observation c94e5b87-c82e-415e-b725-0232d8a39a93 · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:15.573777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:15.573777Z digest=sha256:7e0eefbc4ff35736e382758ef0b7bbef8777345613f151b3dfccf47f5136cf8c

Observation 40a0df60-3838-450b-ae44-fdff04008690 · outbound

This paper cites Revisiting Over-Smoothness in Text to Speech.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Revisiting Over-Smoothness in Text to Speech

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:45:17.675731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.639463Z digest=sha256:277c1991b14ada99178c9a1907549c5bcebf3a2294aed380123a64e6a8bc632b

Observation e61555de-6d6c-4486-9be1-8c877b4ac97c · outbound

This paper cites Grad- tts: A diffusion probabilistic model for text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Grad- tts: A diffusion probabilistic model for text-to-speech,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.775881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.700325Z digest=sha256:6f38c1de35fe8a1ba26271b43d7d9157671a3311c54ffb31fd0aefc881dfdda7

Observation 52d9923c-d046-424b-b4af-68dee322933b · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.589017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.776492Z digest=sha256:5fe0feab065f7266d10354f605b0236ffd8006100a3dc724055b31df45a8a62f

Observation e5a4718e-8a2d-4a26-be94-da876beb5e8e · outbound

This paper cites Consistency models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Consistency models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.391621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.839694Z digest=sha256:68d5ff9a38d3b969167eb3cbd224737347140bb5dba046aede36a454a89db9a8

Observation b152212b-1d78-43b0-bb85-10d3a4ece214 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow straight and fast: Learning to generate and transfer data with rectified flow,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.190716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:15.926474Z digest=sha256:b6d83874196856286f7f90c3a7bfa5e89d6fb5a5e4b70d651a348a69762ce63c

Observation f858a989-c3cc-448b-a3f4-7e1ede8df9f3 · outbound

This paper cites Comospeech: One-step speech and singing voice synthesis via consistency model,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Comospeech: One-step speech and singing voice synthesis via consistency model,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.931419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:16.004306Z digest=sha256:ac0d9cf3a7ee839dd86895ee9521c13d75fd9dda9f31af12a9b2fe4c9c8edd97

Observation e52ae0af-4104-4462-99cb-d18419bc7542 · outbound

This paper cites Reflow- tts: A rectified flow model for high-fidelity text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Reflow- tts: A rectified flow model for high-fidelity text-to-speech,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.730383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:16.052933Z digest=sha256:76ea82d3a372d7b30af7f449048aff6ef3b56da42707dc3e84b881cd548baae4

Observation da6a2ad3-611e-4da0-9196-71b71c128ca1 · outbound

This paper cites V oiceflow: Efficient text- to-speech with rectified flow matching,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V oiceflow: Efficient text- to-speech with rectified flow matching,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.567457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:16.144964Z digest=sha256:e9624eb5a050aaaddc632a9aa3150310a60b4e693ec9be860a686724bc42bd26

Observation 6e52ee7d-d6cf-49f8-a13f-9e2be46f25db · outbound

This paper cites Flashspeech: Efficient zero-shot speech synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flashspeech: Efficient zero-shot speech synthesis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.354922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:16.189257Z digest=sha256:f84d416abe4f11e116f5af9ef70dc183fc013571d367bb1da823b1ed66ac0ed5

Observation fd3bdbbb-af80-4d3d-beef-ca010a89d28b · outbound

This paper cites Slimspeech: Lightweight and efficient text-to-speech with slim rectified flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Slimspeech: Lightweight and efficient text-to-speech with slim rectified flow,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.136763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:16.251832Z digest=sha256:099428d3d0754c63ff026cef988214eecdfa13a25f609194ba9defb29ffae6ad

Observation 92523cae-b210-41cc-ae14-2d9e03326ba3 · outbound

This paper cites Lightspeech: Lightweight and fast text to speech with neural architecture search,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Lightspeech: Lightweight and fast text to speech with neural architecture search,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.922989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:16.341634Z digest=sha256:67361b3ccf238a98bb226d30f3bb7b212454865eacdcb0deda8cdf3d59cf7661

Observation ceb84728-7a54-48fa-8c99-e7d41459232b · outbound

This paper cites Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.724371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:16.410344Z digest=sha256:7a0e5a1426387f254d8cee91e10325bbb96cd18e3e522128b3792bbe090531eb

Observation a08e7b59-6a11-42a9-b2db-21d3c47cefe3 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Common voice: A massively-multilingual speech corpus,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.561055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:16.499791Z digest=sha256:eeb275c8801a1fc91ef940a60637478936ee8b8b9c71828322fbb49983da5450

Observation 4bf2354c-ea56-44df-9c1f-b1f82f44ebda · outbound

This paper cites Didispeech: A large scale mandarin speech corpus,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Didispeech: A large scale mandarin speech corpus,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.414015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:16.572166Z digest=sha256:bbdbff9de7e23706f4aed70b9aba4e8f8d2c96e1a7829db3ae6adb0fcb180fb9

Observation b922f287-09de-418f-95d3-a394b95e554c · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.635875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.635875Z digest=sha256:b830163b3f9cff0f66359a60b6f45194bc38588ac4d7f736e7365c7cb7adc8c4

Observation 66d38499-0d1c-4dba-9295-38bca924f20c · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Robust speech recognition via large-scale weak supervision,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.715337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.715337Z digest=sha256:59f1fa9621187755515651560f9da2f19a06b43f34c6645fd1be253086ee6c54

Observation 28641d09-7533-4f30-9ff5-06211cc5fb0e · outbound

This paper cites Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.775147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.775147Z digest=sha256:03603a70a720c4106b896f553f4a58e448010a6e41180672794565175bb956c3

Observation 87dc4532-6b2a-4054-85d9-5caaf380e904 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.855648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.855648Z digest=sha256:345f419e648dbec58703e1ce72bf41ec528290f0d1772252384841a6b3494746

Observation bf4eb6b6-3f38-412d-a71e-47c23e715ee4 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.920932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.920932Z digest=sha256:e479afbef00f89ede9176d121dba3d4d36c8aad0136bd9053d6494aba70142f8

Observation 523f1dfa-d805-481b-a534-f640a4ae254a · outbound

This paper cites Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.031709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.031709Z digest=sha256:5ca4756ec1e9d8cd02c46dda377c0456d87b82cb64ac6526956bb6e027ee6806

Observation 31e12026-c525-468c-a783-f51f3b05bb96 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.197912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:17.111078Z digest=sha256:3ed6862cba62da64e6440dfa3c5ee87da0c37008e68b2e21ec717e046c7629a9

Observation 263e86b0-e18e-426c-8237-e6dcd3010d8e · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.180878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.180878Z digest=sha256:0cc7d3262704d0fd5ed38af638dbf98fef26688b158aba9b8cf07f4639c817ef

Observation 80c5d825-271b-4e00-9722-bc71a0a9a635 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.249552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.249552Z digest=sha256:951931ef4304c23507c17a2ec3e33a40100b8c79322ba8e0f9b9eb831523a26f

Observation 45102f47-a413-4735-b698-7e1c2858eaa9 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.314055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.314055Z digest=sha256:c978de53d5836927274002f5e10df2bea62424ef5dc7616bbe6504a749b3081e

Observation 11b0aff7-79ae-4cc5-9a05-3e065b821938 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.058741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:17.387188Z digest=sha256:8e354080129d2f1fd3ebc0f2598f39bd8b30f4776c8b24fcad3b62adfeb05997

Observation df6b7d1e-4c64-4752-ab52-b75bfd80e94d · outbound

This paper cites Amphion: an open-source audio, music, and speech generation toolkit,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Amphion: an open-source audio, music, and speech generation toolkit,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:17.878084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:45:17.467537Z digest=sha256:517bb0556a52c4b8a6fc6c795a86c9967b5136cc622da20f034a87e2fc1e4d7a

Pith citing papers

Observation 1f256cb3-22a5-4e2f-be92-73228fa1566d · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:32:03.673585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:f2cbc433a38d38810a0f5ebf14bb4b9fc9161f7437c704009cf008eafeabd6be

Observation ab2e20d5-7d2d-4ad0-b454-c474d228d099 · inbound

Universal Speech Content Factorization cites this paper.

Universal Speech Content Factorization ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-15T12:21:48.698333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:21:48.698333Z digest=sha256:42f62b203eda377ad129cdd600aabf2f926e3e94a3e1b8c78cb4b2100dabb7ad

Observation ef4f0b28-14ef-48b4-aad7-45836011a635 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.708672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:147109afa53b9f3396cd4d4f3df85158e8a87349933d209812f8b4f123e84036

Observation a7b0fc38-7d06-4288-9379-6e9f4c2ce3d3 · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.048637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:6d8483f492ce5d2bd677db0bb3cd8aee2f41f2cee619ddf490a019e2ff233a9a

Observation 93f45acc-75fe-4625-954f-7155ac151c43 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:36:45.127244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:28:14.734682Z digest=sha256:9c2cfb9f3a5afeab8b5f15b2b12b6f61a1c00f89f74412ea291f9c7c1051b47f

Observation a148795a-e901-425f-921a-39a80ed99b83 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.821926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T23:26:46.077894Z digest=sha256:3f0e582837629ea78e9b8138cece54069437e30c92df86bc094f841d31d0108c

Observation fce86089-99f4-4aab-9876-524fa05de4cd · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.049565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:73330e6d96d9c5f9cf35ab22fa91ad1f945fe2c4fae370d0097d010036328777

Observation 456e7331-2449-4af2-a6c5-fa6a9ce4ec35 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.093666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:6513b41dd19cb7a715d61c1bd2b6e763e135a5fabc31f841e6331db98a6e0f19

Observation 35be436f-b14d-462d-9f24-c9059434f62d · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.787929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:75096e7957dac30e7c9d3541e1bf0ad72eaa743b2d551c6d9a15174d66e8caf6

Observation d30edfa3-c12d-4673-9a20-d17c3e8d3216 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:47:19.794924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:e3e15a134246fd077a15dfdc6b330582511e526e3d66e0ed8d73c265de344c60

Observation d1c78e43-7078-41ac-acf0-683a37a18d96 · inbound

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation cites this paper.

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:19.602737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:14:08.243894Z digest=sha256:1afaf7fd1b602bfdd9ae3880b75e6e5f6c03d42fa69998fd2778cfc0d705b558

Observation a83fb087-1b43-4394-b06b-ed907f6a3bd5 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.898147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:cb59c5ab3a33915d552f0a0e2ed6d14c968ddf654e836ff9b4b089e82fed3d31

Observation 7a947619-156c-491b-a449-a4279363051c · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:73c1e70291289afe80bc2ecf481e6d5401b84b76d07be642e4ea3a8c74e7a384

Observation 0297f999-eccb-423d-876a-5296308efd40 · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T02:22:47.820537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:22:47.820537Z digest=sha256:00d059b4b925f5d660d919d6af239ca45fb716a4249273f33e7878c46258d084

Observation e07bb246-f5cc-41a3-9d83-cc58e8eec40d · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T07:39:23.947687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:39:23.947687Z digest=sha256:f906e987543b05b57d28f567dd025df278dd246e0dab7c1cf6a7ec9a1b8d067a

Observation 51c0f031-9d9e-4949-9112-f3042e5cf35e · inbound

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model cites this paper.

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T22:26:14.049693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:26:14.049693Z digest=sha256:8661b367cc689ec1ad9cde344b35af66f75bc48b76b9cc6ead981d21628fefcf