Pith. sign in

Paper Citation Record · LEDGER

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

As of 6 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 1 inbound Pith citation observation for arXiv:2606.03455.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03455 v1

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T08:18:42.002083Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T08:35:47.752843Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 109 outbound references displayed

  • verified exact48
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ce5ec35-c2ca-49f5-a2e9-a3785716a6e5 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.811706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:fe5509ae5e0d691ef0ffdc48c6ef9f3ef36a24ff6873765cf93c8640e83c1b56

Observation df0517b9-3171-4a31-ba52-fe886fb22ab1 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:c0b4995b1c9c0ba75b46f2efa3b1d26e96b97dd5baaa7d2e32fdba997578fb3d

Observation c40da833-8482-4eeb-87e7-d065caf608f1 · outbound

This paper cites DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:9873961406729cfa6a4e3e63caf13a928505e5100aaa7cdb6cb398be2ba6973e

Observation f20ae24c-d64b-47e4-828b-bd08addbe3da · outbound

This paper cites WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis.Interspeech 2021, pages 3765–3769, 2021.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis.Interspeech 2021, pages 3765–3769, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ed2ee6f12d644aeebd25ede6f8bd5e4adceaac86197203354bfd800ec65e1d2c

Observation 16dedabf-4f7f-4a58-8162-5d6a50496af0 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.798400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6ee5a8614f50e420489874cb1dc42d02941376d5c4f9faba3d8f8150fa4c1b17

Observation 6218d65a-65a5-4d6e-b949-71cc7d3294f7 · outbound

This paper cites PixelFlow: Pixel-Space Generative Models with Flow.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixelFlow: Pixel-Space Generative Models with Flow

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.840803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f74feb662d9b8d313785fe2fb8344671673d9fd66dfbc48119e6b0fef4a83172

Observation 45f9bd5b-b056-4fff-a1b4-fe1d47b823ae · outbound

This paper cites On the Importance of Noise Scheduling for Diffusion Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling On the Importance of Noise Scheduling for Diffusion Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.809182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:55b83b2ca92bafb4e14540821951269a93dd2abfb3c8b1b5daf413aaa8f7606f

Observation 019f0ce0-9655-4450-b990-02b02b103f7d · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a8298065c0086f344f9c4f54874738c77f61521f1b74b1282223e9c074e4997c

Observation 30a31182-3299-4d74-be02-638bed1425c0 · outbound

This paper cites Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:2f455b0ba14bfc340135923741713f53a39669a219afa8cbe79e9bca7b588253

Observation db452e27-0a93-4d7b-91a2-7f95867d71bd · outbound

This paper cites arXiv preprint arXiv:2511.18822 (2025).

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling arXiv preprint arXiv:2511.18822 (2025)

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.729433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:3b7ce7d0f65892931e3d1c95921f60a0888ee908cdb174c0f97de4054aea1c20

Observation 4da41692-3df3-48e2-8bfd-ad1512f0dbe0 · outbound

This paper cites High Fidelity Neural Audio Compression.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling High Fidelity Neural Audio Compression

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.747814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:734e02a1f6ce2e99f917fa18d9c837310ea66ba148258ec9563ca0cb8fb30f9c

Observation 900dcf1e-ac6a-4e35-a7f5-d1b23f6b59ea · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.Advances in neural information processing systems, 34:8780–8794, 2021.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Diffusion Models Beat GANs on Image Synthesis.Advances in neural information processing systems, 34:8780–8794, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0e2c7a1775fd79989f266e0384141e156a8d25e334bd90d28b4b63e1085529f3

Observation df753d9d-693f-4cb8-80c7-091b7c29b7d0 · outbound

This paper cites End-to-End Adversarial Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling End-to-End Adversarial Text-to-Speech

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.837427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b4a919b4109284605276e89f2d24c5090d48715a36443fb358c4dbe38896fe54

Observation 84fd148e-3bfb-4b2a-b1d6-0feb44b79df1 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.752712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:d12b15507260f32e1057d26db954eccfaf56bdbf6a1ae511bd5eb482ebe69bc2

Observation 9ce8533d-e777-4ec8-aff8-fa9b56f7776e · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.729854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:cf4ff5903c9dacbfcd3bd16de562058472b1b0e5544f071a22a4570d6ffb6917

Observation 5eeff31d-1b2d-4086-9d7f-865bcf215b6a · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.732631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:faeb9779b7a4eb4f7bd858f9b3aa130a604c31fb4d537a7f66c9573252ebf7eb

Observation e4f26ebf-927d-405b-8dbb-12243196a13f · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:009a036997ccf2aa539560e25ea5a4d5b85dba5c8dae69bede4417bcbd676c2c

Observation cf141115-ac01-435e-baae-872024d51c15 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:758c97d8db0a3eec021c8d47d52f00a3f95664cac84cc06579f318a069525cca

Observation be4c75f3-b7e2-4970-9f3d-1e313c63015b · outbound

This paper cites E3 TTS: Easy End-to-End Diffusion-Based Text To Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling E3 TTS: Easy End-to-End Diffusion-Based Text To Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4ceb9f064a74d5c3ed4b667258440771fc34a3daac1792a8990e8a9f60ea2107

Observation c853f0d5-65bb-4044-a290-c6ee6eca0a88 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0f0e84241bc2f53379d4c48e3c574ef6421dfd73d859b9afdc4c1c968b72f01a

Observation 60db7b3b-d281-49f3-a971-56871a97da40 · outbound

This paper cites Moss-tts technical report,.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Moss-tts technical report,

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.777402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:bba46a4e1a7839adbf9fd1362c9c3b010b8aad435c3ecded2dbffc9856f7a216

Observation 4476e729-e5e4-4ced-aa61-c1d5762c8250 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.768674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:84326b59da158621f85412e86ac2cccad090436748937bd60588ba67e79fc409

Observation a267b415-7ff9-4590-80c3-78217b090f24 · outbound

This paper cites Didispeech: A Large Scale Mandarin Speech Corpus.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Didispeech: A Large Scale Mandarin Speech Corpus

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:c6cbd16bbe0efe4ce53a6785120f49aa374fae34a4b51601e0d2575146f175b6

Observation d1544d29-2011-4933-b183-1c149f16ce6c · outbound

This paper cites VoiceFlow: Efficient Text-To-Speech with Rectified Flow Matching.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VoiceFlow: Efficient Text-To-Speech with Rectified Flow Matching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:96d3fb036571ceddcef1fce0f25f5fd87ee461d1c2b69ab908ae694a6c199e97

Observation 1eb896f0-e2f4-47a9-896e-a36b0eeb781b · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.771613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f5dcec5c93b9f1a07bf5ce7823b59f4b66b0e2b7936861a9fb120150e16f3bd9

Observation 399552db-214b-40ed-8347-5df86bc09bc4 · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:904bb67ad1408a38127374546bb5bc665928d90b4dcbdbf541891c552f4c1548

Observation 2d58560a-0cb9-41a2-b162-8449f87fc371 · outbound

This paper cites Classifier-Free Diffusion Guidance.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Classifier-Free Diffusion Guidance

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.773456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ec0133c7598421a398c3950390fd46f1d261fabd72b2a6e9b4af512ab583c0f6

Observation e703ebeb-3f54-431f-9045-fe7dde22a6e1 · outbound

This paper cites Denoising Diffusion Probabilistic Models.Advances in neural information processing systems, 33:6840–6851, 2020.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Denoising Diffusion Probabilistic Models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4ff222f7c76ec9f8ae864b5cc0699cc3b3223ff84742358c29d2e0e3a017f235

Observation cd112712-f9d2-4ed8-acfe-95e02c0fe6cd · outbound

This paper cites simple diffusion: End-to-end diffusion for high resolution images.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling simple diffusion: End-to-end diffusion for high resolution images

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:af3fb2fd43823319ca35caf64aa916f3cd35d7032372ba1f775ca9e1b07b15f3

Observation dcddd738-1fa2-46fe-b487-f6ee85e8a883 · outbound

This paper cites Simpler Diffusion: 1.5 FID on ImageNet512 with pixel-space diffusion.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Simpler Diffusion: 1.5 FID on ImageNet512 with pixel-space diffusion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:84cd9495f2e890bbdaefb6a64d332b337b024cfc0a82de6e0b4bdcd9dc12454d

Observation 754f9d42-b7bf-4419-a134-6074274b2ad2 · outbound

This paper cites Qwen3-TTS Technical Report.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Qwen3-TTS Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.774185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b13cc7fb376019221c0a207b78ab228458e6f2374c9cdb9b6fab181f6211ea41

Observation 3c35d135-943c-4627-bae8-2c7ecbb8b3f7 · outbound

This paper cites FlowTS: Time Series Generation via Rectified Flow.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FlowTS: Time Series Generation via Rectified Flow

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.761621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b087ba15ba1fc6c19cc23df767e1d97b5d826228a9ab12f5c67dbab8d2a998e7

Observation 78fd2b98-7bc3-4e83-8b5b-5362f05080cd · outbound

This paper cites FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f5e7073cdfdbcc83d8becebc69f39a49938ae0f37d492c8ba3328ca73d9571de

Observation 6ff2e288-0835-4dfe-8332-a924ed469f2d · outbound

This paper cites ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-Speech

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:d23c21d1e7bcead8b81962371754adaed9aa652c2084e0d430fc44e4758b2966

Observation adc4592b-7fb3-4b4a-8ed7-95fea4218205 · outbound

This paper cites The lj speech dataset.https://keithito.com/LJ-Speech-Dataset/, 2017.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling The lj speech dataset.https://keithito.com/LJ-Speech-Dataset/, 2017

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:2bd18d558562972e14aab7369c59c71ac1bb93e76fdff7a9993f37b0bc32a661

Observation 1548522f-cae0-4606-8915-4cccf0898484 · outbound

This paper cites Diff-TTS: A Denoising Diffusion Model for Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.812312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0c9d913daf1b5698d4813fc4ef13e1de358298f386e47558c4d6c757008bf0d0

Observation 01b87c9d-b6ed-4cd4-a947-0efb9eb284f1 · outbound

This paper cites DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:fa3fd7ae93b1a017f450d8dacd6884f8bb501f2ebfe23ef6f9d7539f8d42f281

Observation 768e4f15-4a57-45e2-adb1-22e3139c3de8 · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.706809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0b025ba6f0b78eafe63dc68cd0bd35fa179f9e1df1134a53905c758acba7a94b

Observation aa02f98f-03b1-4561-8d4f-4f7b341c4665 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.706591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:fdf6a32bedccd57dfcd75c2c9c82eef37e669137c28e045a58b9c11ce392b58c

Observation b4eac9d8-23ab-4f62-a438-960e06501417 · outbound

This paper cites Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:e3016706a2093ef317314cfa24860927f33fc7c0d296e208153ddf84dd2daa29

Observation ea53c29d-d671-4eec-a403-b7b6e2ef4afa · outbound

This paper cites Auto-Encoding Variational Bayes.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Auto-Encoding Variational Bayes

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.700787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:13a89429dc70c347eee93830c61a6ba4fb667a43708cb09bdae432c0b94ac765

Observation 2ee8c748-cc88-4787-a8be-21c472cebfa0 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.Advancesin neural information processing systems, 33:17022–17033, 2020.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.Advancesin neural information processing systems, 33:17022–17033, 2020

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6ed7541f7e4e8e3e8557b340393505055d85bf0a54532c043f69e6b86425805d

Observation dca840f8-4fa0-41bc-97ce-cd422fa19f48 · outbound

This paper cites DiffWave: A Versatile Diffusion Model for Audio Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiffWave: A Versatile Diffusion Model for Audio Synthesis

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.829738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a42960a53c5f1bc6940430d1aaecfe8babaeb980b8c0bd347bfc811e9aa9c4b4

Observation e60b09b7-313d-46b8-a2ef-b546b3e1633c · outbound

This paper cites High-Fidelity Audio Compression with Improved RVQGAN.Advancesin Neural Information Processing Systems, 36:27980–27993, 2023.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling High-Fidelity Audio Compression with Improved RVQGAN.Advancesin Neural Information Processing Systems, 36:27980–27993, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a92573e7aacc9492579c08e7d4644abb364e5a20dba2c4135e9d4dc18108b9b6

Observation e5a1bf20-4e87-4599-a4e6-88cd7a38d279 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.768082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a27cd54f372baff6d0db936def8156a1ace0d3f0fefa861bd31b22243ad769b9

Observation 93762c56-0f5f-4aaf-b7f8-003f12fb0d8f · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:705f581ef2cd84bf6ca8a612893e524432447a61499eadead2f2c7d00ba657d9

Observation 0fc37142-53f8-46bb-83f4-d1143d9721ee · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b433e984ab7142f7670625e338d190a159bfb810e148d29b3ebfb2c1dd43166f

Observation f58b15e8-af53-49b9-9393-c734f69dbe9a · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.735273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ffb11c8198eb9f706bdb4bcecf9420db05a7141fcac81bb591eded230ae8fc43

Observation f277ce7e-50df-4401-9254-c44410190db8 · outbound

This paper cites Back to Basics: Let Denoising Generative Models Denoise.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Back to Basics: Let Denoising Generative Models Denoise

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.761116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a3d8842dd863862c6529a9d33bcc567627db08628e0de564c01af6aec14f3fb0

Observation 44f790ef-2204-43c1-9598-ef8a19186413 · outbound

This paper cites JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.840425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0fc79a895b3f6bb022056f061689f6e55d53353de093c99bd9b34749e1d09fb6

Observation d93cfb6a-c520-4951-8abc-63c44a4623ad · outbound

This paper cites Flow Matching for Generative Modeling.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Flow Matching for Generative Modeling

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.855011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:d27947982800668a2d655df8ed73174e5fb54ee622aedf2756c2ed0a367f6a07

Observation 9d1f613a-e5f6-4e7a-93fc-7039f1464f42 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.832283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:9d382ef6c738a19cb1248ae8f122121e65e24c9bfa8b480f4cd2a2a470858069

Observation 5d82fe48-e325-4ff3-9ad9-036537040823 · outbound

This paper cites Decoupled Weight Decay Regularization.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Decoupled Weight Decay Regularization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:e4b0abfffc31a106870fea9ede743fda4f1901fb9ec7e875a0bd1dff26b5cd7d

Observation 7f3c258a-852a-4242-a01d-cc80ea235884 · outbound

This paper cites DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.826742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:784194e25c271433a29dbc4b18a1df1b5f3ed113a91914939b7f2cd0e2fe66f8

Observation ffcf786b-c72f-4b11-a3fa-2347c54ae5f0 · outbound

This paper cites PixelGen: Improving Pixel Diffusion with Perceptual Supervision.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixelGen: Improving Pixel Diffusion with Perceptual Supervision

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.845806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4cc7f35a3fa9e5232b685f86c644fa03727a6ce161aba228b4cf9efd6c8d096f

Observation f31c0db0-3543-4d95-a6b7-28cb58225332 · outbound

This paper cites Matcha-TTS: A Fast TTS ArchitecturewithConditionalFlowMatching.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Matcha-TTS: A Fast TTS ArchitecturewithConditionalFlowMatching

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:8eab0c2127750d245a52afb66575b4e1aca8f06749c4072926b804a910c160df

Observation 8cdafe16-651d-49e0-a531-ad43a798df9f · outbound

This paper cites LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Model.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:16757170dabd5a200dc47dda7e1ad6a9184db1272136f5670adf527e13dd704c

Observation ba7daf40-9071-45a4-80c3-516cde68a34b · outbound

This paper cites Improved Denoising Diffusion Probabilistic Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Improved Denoising Diffusion Probabilistic Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6d146605e78f605653f2057ceee7c191474fd700ab718545089aebb31907f973

Observation f1c769f9-a12d-437b-bae7-50e3205c6852 · outbound

This paper cites Semantic-vae: Semantic-alignment latent representation for better speech synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Semantic-vae: Semantic-alignment latent representation for better speech synthesis

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.837808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4ca83627a708339992a1812c9fe102617b761d4d4115c85e704a0fd1331b67d4

Observation 0a9c96eb-7ed9-43d8-aa76-5bd61c4eefa5 · outbound

This paper cites Parallel WaveNet: Fast High-Fidelity Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Parallel WaveNet: Fast High-Fidelity Speech Synthesis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:bcdf9147205e4d08bae638d544eb54bf5ebe7be423cd81b856ff0de07dc1e33f

Observation ccc58cfe-27a9-4a9c-b488-67a14d863b52 · outbound

This paper cites Scalable Diffusion Models with Transformers.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Scalable Diffusion Models with Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:252bd36f7d1d3f1bdea4f2712b1ab9cf918b92cd7b0525e28b1dd1e91edd22e8

Observation 1288d50c-9117-447c-8a9e-df7df0f944e9 · outbound

This paper cites VOICECRAFT: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VOICECRAFT: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a3115b8eb8358ec4c0a2220122c981b97e88a63187caa6569b272f5055aa2136

Observation 0ea07226-6a13-4d74-bdd8-483523798cf9 · outbound

This paper cites VibeVoice: Expressive Podcast Generation with Next-Token Diffusion.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VibeVoice: Expressive Podcast Generation with Next-Token Diffusion

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f7b93fe3d711447be7055cee33c601c16c98409f0a91ec5a867ea602f865577e

Observation 61467cab-e5f6-43e0-8cf1-f7820e9b05e1 · outbound

This paper cites ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.832450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:901ca406d4a76032487d9363fff96d1e72aedb8e989c98823d171e6500871a1f

Observation 8d87da6c-ac09-46e4-84f8-6fc00f5f085e · outbound

This paper cites Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:017ea896fe379125db983d940f84f757abf93f460efcb9efbca6edde3aeba4d7

Observation b353a749-a4b4-4144-924d-9f88eb68b7c0 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Robust Speech Recognition via Large-Scale Weak Supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b52be376ab0f3601fe13ce5860278d4c67e49f23f61307510f64bb79a8888a68

Observation 3be336fa-3726-4214-87ef-5c60f15dba00 · outbound

This paper cites FastSpeech: Fast, Robust and Controllable Text to Speech.Advancesin neural information processing systems, 32, 2019.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FastSpeech: Fast, Robust and Controllable Text to Speech.Advancesin neural information processing systems, 32, 2019

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f9d157468f3fae87306030d52c204700d6fb03b0fd768f213fbf1f490ca11690

Observation 588b3f10-0b4d-43f6-9633-d1741589302c · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:24b237102c65d1ea5f1658135e3ad39f9f529ac4abd05e1302077177367c786e

Observation 387105e9-6335-4635-981d-487ee250f320 · outbound

This paper cites U-Net: Convolutional Networks for Biomedical Image Segmentation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling U-Net: Convolutional Networks for Biomedical Image Segmentation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4454aa85c5ac90751986b2141464d2b8355de89a73e560fb9c01cc1a87d20308

Observation 89d42bfb-bd2b-47e9-9828-db643118c589 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:5b104c0ac114c971679a3895256157a16fa13c2d229d786e2d592e64d8b7c454

Observation e2124967-1ff3-4a58-99ea-ec40a2394c72 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:78997c9a86cd1f4d69d39e121fea7ce4c85ce464ccbab0684d01ca795aac9cfb

Observation 311e960d-e47a-4882-9ac5-9415515192ac · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.843462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:348b563af993f83273ae394880317c753c3bc351efc93b8e24436bd409e6ebc1

Observation ee1ef2a2-d58c-46b0-a5ef-149e2bfa0a28 · outbound

This paper cites MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.817490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a8c7365547bb5a2d061f8182cf68d56da8d600d010833a9ead74948efad3c93c

Observation ff3ca7c0-f631-4154-b5a5-a89225372a12 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:9d8d12abd19198db11ec819116127a851e7f02540079aa7c9607494561287355

Observation 22236b6a-a76d-43f3-ad87-a72d169c297e · outbound

This paper cites Generative Modeling by Estimating Gradients of the Data Distribution.Advances in neural information processing systems, 32, 2019.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Generative Modeling by Estimating Gradients of the Data Distribution.Advances in neural information processing systems, 32, 2019

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:91944c9725ca6b6d605e076d9ddaa86655a916d907b9abb61ca198347bc3755a

Observation adbbee38-7e5a-4d54-bd29-5caa7eace8a8 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Score-Based Generative Modeling through Stochastic Differential Equations

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.814716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:985e6370d42991a1b6fd774ef66e99f87b1476b18dfff2a98508ce0ed7aa7727

Observation 5d0d0f5c-d8f8-4536-b140-43772a28ffbb · outbound

This paper cites RoFormer: Enhanced transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling RoFormer: Enhanced transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:3ca121ba15bcce8b7f833898e41c84047b5099ec8a28fc78b1bfa3e218754648

Observation 47596943-c5bb-40eb-a453-69a0a968f8d3 · outbound

This paper cites F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.779566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:80a9caff0226a25f12b6bc705816933d8304a9df789f7d89bb58049aadd5e564

Observation de6a525a-0689-434a-89d3-d92d240bb87a · outbound

This paper cites STFT Spectral Loss for Training a Neural Speech Waveform Model.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling STFT Spectral Loss for Training a Neural Speech Waveform Model

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a6d1bbb6248aea0b599243902a9fe50c777671cbf4fc80d554296c20a97e020b

Observation 5e94efc7-e50f-458e-9a74-c7d54a720bc2 · outbound

This paper cites NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality.IEEE Transactionson Pattern Analysis and Machine Intelligence, 46(6):4234–4245, 2024.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality.IEEE Transactionson Pattern Analysis and Machine Intelligence, 46(6):4234–4245, 2024

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6461e86ccc548728141c12fabe84ecd0ba304d7ebc041f5256681c4189bb1978

Observation e90ce36f-5056-4f40-a957-93a7fe4a7fd4 · outbound

This paper cites Relay Diffusion: Unifying diffusion process across resolutions for image synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Relay Diffusion: Unifying diffusion process across resolutions for image synthesis

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:8d77024d9327ff9924cc8cab714e42776b907e64d29cdff47bbf5fb7833edf26

Observation c90408f6-ba47-4573-8ab5-066a2e66222d · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling WaveNet: A Generative Model for Raw Audio

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.856964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:610f6f112f6da95e639d3709b7d3f8bcaf8da9f83f4bfa52bf26c027fa90083a

Observation 613fc9b5-8e3d-446c-9ab3-4b836eb18555 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.813970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:61a242dc5a871f3580650d15085f831b99409a4eb9e08e3d98cb4102c11c2660

Observation 550fec21-15f1-4126-bb4e-0a2ef6a59556 · outbound

This paper cites PixNerd: Pixel Neural Field Diffusion.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixNerd: Pixel Neural Field Diffusion

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.848770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6f6a72807b0bfc3b98e9222342a768744a470afdc85f975334eb640cb7457cba

Observation 31ea4afe-d4b8-4ede-a07e-8bee0a173d3f · outbound

This paper cites M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis.arXiv preprint arXiv:2512.04720, 2025.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis.arXiv preprint arXiv:2512.04720, 2025

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.782209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b986cf01a479c4aedeba3642ff6d53ea97077ac13acf9f5a50652c1cf4214433

Observation 8665f702-ea46-4599-849c-b89776716389 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.744079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:2afba8a3cbd0aea7c228f139f145f562bc464de02a72d39faa9ad7cb66b4e57b

Observation 2ee16ea4-22a0-4f0d-8c81-d09cd0a93592 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.738763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:47d70a5c60a6aa099c97d149c81f44bd1be5cc638be9218f9bd2d3a7e6c49b51

Observation 8a7b4e12-1723-4df8-8762-75ccc3d3aaf0 · outbound

This paper cites Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:aa449ce631f0761383831d702ddd7c6a3bba2c50c1be3acbb9296f6565ea4d4e

Observation 69a98e8a-b3aa-49cf-8309-6eced1dc7dc0 · outbound

This paper cites ConvNeXt V2: Co-Designing and Scaling ConvNets With Masked Autoencoders.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ConvNeXt V2: Co-Designing and Scaling ConvNets With Masked Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:93f06eed2f01d1c47bd338e811e8907e04e96f5871fa50debc771a1e787e1678

Observation 7a39acb0-45ea-42ea-afc2-6d1f97b3b7f3 · outbound

This paper cites TS3-Codec: Transformer-Based Simple Streaming Single Codec.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling TS3-Codec: Transformer-Based Simple Streaming Single Codec

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:7044cf156b8cc3eb2d26351eb95d6af87ce5ec853acd1d40f8d6eb1640a93fca

Observation 7403fb74-abe3-4dee-85de-5ee6503270fd · outbound

This paper cites FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.764041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0136bfe124c8b1fcd371c4586ab18cc2802cf4423b33eb522a606001a8a3b117

Observation ba13b69e-2615-4a35-acbc-eb134b1f42a9 · outbound

This paper cites Longcat-audiodit: High-fidelity diffusion text-to-speech in the waveform latent space.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Longcat-audiodit: High-fidelity diffusion text-to-speech in the waveform latent space

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.753373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:12024d77a9e5b94effcb94a4a9fa7e0e31a65e00ee6d7f6c13d3316e64d9f1d3

Observation 5f5e3a2c-dde7-499b-9c45-54bcf64c956c · outbound

This paper cites an unresolved cited work.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:2aa924ae09f448e26a1fe86e88b348c89f48f26b9152ea87c4ec0e88647db489

Observation 08f36623-25f8-496c-a33c-5567b5ac41d7 · outbound

This paper cites Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4d3f2ad084ea0fcb6805d59b6bcb3167038b7490f38f0e8e96381a5def767b43

Observation b09f98b4-38eb-4f9f-9700-9e4c990730bf · outbound

This paper cites Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation.arXiv preprint arXiv:2512.23278, 2025.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation.arXiv preprint arXiv:2512.23278, 2025

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.750020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:8ec7edd92da934f88d4af0f108baeebc8c6196737d15e01d94a9d6c152d59d73

Observation c53abe59-20e1-4186-af2d-831e5e1d587d · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.803185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:98cea15be72f6f5e36ad17ea351d8b9ed2f7ae82ca4e8fc295cd950530c765a7

Observation e32fd3f1-d01d-4f9e-b79a-365fa6e14786 · outbound

This paper cites PixelDiT: Pixel Diffusion Transformers for Image Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixelDiT: Pixel Diffusion Transformers for Image Generation

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.779918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:cc5537f9c1e6321933e8e56307dbd0efc8c587c2959b9ec0b5693cfdd117909d

Observation 1f0c0946-cdc3-45b3-9578-841c2d31ba8f · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 495–507, 2021.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling SoundStream: An End-to-End Neural Audio Codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 495–507, 2021

Reference 98

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:5c9f92c3224e9ef96f21ac72a92a32f0e863fbad9349c414dc61bb7df6341c20

Observation fcd78baa-5689-41d6-aa07-625d33992412 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.759013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:1af24c101c8bd1ce84ae8d50aeca5704bfc0dc1f62515de4b758e954e8592c6d

Observation 005df705-46a5-4519-af3e-ae2234d98c48 · outbound

This paper cites Normalizing Flows are Capable Generative Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Normalizing Flows are Capable Generative Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.794214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b9cb30f3b118345b08b491240b0ee4a54bc0a3545772cde6837ab3163ecd5f64

Pith citing papers

Observation 58bad382-20b4-4c5d-9025-9cca6096266b · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:47.752843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:47.752843Z digest=sha256:a88f5eb4c946b768a8e4573fe7af2c30004bb3d349a32630523cd9954158df0e