Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T08:18:42.002083Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 1 inbound Pith citation observation for arXiv:2606.03455.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T08:18:42.002083Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T08:35:47.752843Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 109 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2ce5ec35-c2ca-49f5-a2e9-a3785716a6e5 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df0517b9-3171-4a31-ba52-fe886fb22ab1 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Common Voice: A Massively-Multilingual Speech Corpus
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c40da833-8482-4eeb-87e7-d065caf608f1 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f20ae24c-d64b-47e4-828b-bd08addbe3da · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis.Interspeech 2021, pages 3765–3769, 2021
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16dedabf-4f7f-4a58-8162-5d6a50496af0 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6218d65a-65a5-4d6e-b949-71cc7d3294f7 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixelFlow: Pixel-Space Generative Models with Flow
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 45f9bd5b-b056-4fff-a1b4-fe1d47b823ae · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling On the Importance of Noise Scheduling for Diffusion Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 019f0ce0-9655-4450-b990-02b02b103f7d · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a31182-3299-4d74-be02-638bed1425c0 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db452e27-0a93-4d7b-91a2-7f95867d71bd · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling arXiv preprint arXiv:2511.18822 (2025)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4da41692-3df3-48e2-8bfd-ad1512f0dbe0 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling High Fidelity Neural Audio Compression
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 900dcf1e-ac6a-4e35-a7f5-d1b23f6b59ea · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Diffusion Models Beat GANs on Image Synthesis.Advances in neural information processing systems, 34:8780–8794, 2021
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df753d9d-693f-4cb8-80c7-091b7c29b7d0 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling End-to-End Adversarial Text-to-Speech
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 84fd148e-3bfb-4b2a-b1d6-0feb44b79df1 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ce8533d-e777-4ec8-aff8-fa9b56f7776e · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5eeff31d-1b2d-4086-9d7f-865bcf215b6a · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e4f26ebf-927d-405b-8dbb-12243196a13f · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf141115-ac01-435e-baae-872024d51c15 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4c75f3-b7e2-4970-9f3d-1e313c63015b · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling E3 TTS: Easy End-to-End Diffusion-Based Text To Speech
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c853f0d5-65bb-4044-a290-c6ee6eca0a88 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60db7b3b-d281-49f3-a971-56871a97da40 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Moss-tts technical report,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4476e729-e5e4-4ced-aa61-c1d5762c8250 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a267b415-7ff9-4590-80c3-78217b090f24 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Didispeech: A Large Scale Mandarin Speech Corpus
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1544d29-2011-4933-b183-1c149f16ce6c · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VoiceFlow: Efficient Text-To-Speech with Rectified Flow Matching
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eb896f0-e2f4-47a9-896e-a36b0eeb781b · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 399552db-214b-40ed-8347-5df86bc09bc4 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d58560a-0cb9-41a2-b162-8449f87fc371 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Classifier-Free Diffusion Guidance
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e703ebeb-3f54-431f-9045-fe7dde22a6e1 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Denoising Diffusion Probabilistic Models.Advances in neural information processing systems, 33:6840–6851, 2020
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd112712-f9d2-4ed8-acfe-95e02c0fe6cd · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling simple diffusion: End-to-end diffusion for high resolution images
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcddd738-1fa2-46fe-b487-f6ee85e8a883 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Simpler Diffusion: 1.5 FID on ImageNet512 with pixel-space diffusion
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 754f9d42-b7bf-4419-a134-6074274b2ad2 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Qwen3-TTS Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c35d135-943c-4627-bae8-2c7ecbb8b3f7 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FlowTS: Time Series Generation via Rectified Flow
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 78fd2b98-7bc3-4e83-8b5b-5362f05080cd · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ff2e288-0835-4dfe-8332-a924ed469f2d · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-Speech
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc4592b-7fb3-4b4a-8ed7-95fea4218205 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling The lj speech dataset.https://keithito.com/LJ-Speech-Dataset/, 2017
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1548522f-cae0-4606-8915-4cccf0898484 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01b87c9d-b6ed-4cd4-a947-0efb9eb284f1 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768e4f15-4a57-45e2-adb1-22e3139c3de8 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa02f98f-03b1-4561-8d4f-4f7b341c4665 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4eac9d8-23ab-4f62-a438-960e06501417 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea53c29d-d671-4eec-a403-b7b6e2ef4afa · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Auto-Encoding Variational Bayes
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2ee8c748-cc88-4787-a8be-21c472cebfa0 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.Advancesin neural information processing systems, 33:17022–17033, 2020
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca840f8-4fa0-41bc-97ce-cd422fa19f48 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiffWave: A Versatile Diffusion Model for Audio Synthesis
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e60b09b7-313d-46b8-a2ef-b546b3e1633c · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling High-Fidelity Audio Compression with Improved RVQGAN.Advancesin Neural Information Processing Systems, 36:27980–27993, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a1bf20-4e87-4599-a4e6-88cd7a38d279 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93762c56-0f5f-4aaf-b7f8-003f12fb0d8f · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc37142-53f8-46bb-83f4-d1143d9721ee · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f58b15e8-af53-49b9-9393-c734f69dbe9a · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f277ce7e-50df-4401-9254-c44410190db8 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Back to Basics: Let Denoising Generative Models Denoise
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44f790ef-2204-43c1-9598-ef8a19186413 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d93cfb6a-c520-4951-8abc-63c44a4623ad · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Flow Matching for Generative Modeling
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d1f613a-e5f6-4e7a-93fc-7039f1464f42 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d82fe48-e325-4ff3-9ad9-036537040823 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Decoupled Weight Decay Regularization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f3c258a-852a-4242-a01d-cc80ea235884 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ffcf786b-c72f-4b11-a3fa-2347c54ae5f0 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixelGen: Improving Pixel Diffusion with Perceptual Supervision
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f31c0db0-3543-4d95-a6b7-28cb58225332 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Matcha-TTS: A Fast TTS ArchitecturewithConditionalFlowMatching
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cdafe16-651d-49e0-a531-ad43a798df9f · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Model
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba7daf40-9071-45a4-80c3-516cde68a34b · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Improved Denoising Diffusion Probabilistic Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1c769f9-a12d-437b-bae7-50e3205c6852 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Semantic-vae: Semantic-alignment latent representation for better speech synthesis
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0a9c96eb-7ed9-43d8-aa76-5bd61c4eefa5 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Parallel WaveNet: Fast High-Fidelity Speech Synthesis
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc58cfe-27a9-4a9c-b488-67a14d863b52 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Scalable Diffusion Models with Transformers
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1288d50c-9117-447c-8a9e-df7df0f944e9 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VOICECRAFT: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea07226-6a13-4d74-bdd8-483523798cf9 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VibeVoice: Expressive Podcast Generation with Next-Token Diffusion
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61467cab-e5f6-43e0-8cf1-f7820e9b05e1 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d87da6c-ac09-46e4-84f8-6fc00f5f085e · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b353a749-a4b4-4144-924d-9f88eb68b7c0 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Robust Speech Recognition via Large-Scale Weak Supervision
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be336fa-3726-4214-87ef-5c60f15dba00 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FastSpeech: Fast, Robust and Controllable Text to Speech.Advancesin neural information processing systems, 32, 2019
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 588b3f10-0b4d-43f6-9633-d1741589302c · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387105e9-6335-4635-981d-487ee250f320 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling U-Net: Convolutional Networks for Biomedical Image Segmentation
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d42bfb-bd2b-47e9-9828-db643118c589 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2124967-1ff3-4a58-99ea-ec40a2394c72 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 311e960d-e47a-4882-9ac5-9415515192ac · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee1ef2a2-d58c-46b0-a5ef-149e2bfa0a28 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ff3ca7c0-f631-4154-b5a5-a89225372a12 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22236b6a-a76d-43f3-ad87-a72d169c297e · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Generative Modeling by Estimating Gradients of the Data Distribution.Advances in neural information processing systems, 32, 2019
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adbbee38-7e5a-4d54-bd29-5caa7eace8a8 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Score-Based Generative Modeling through Stochastic Differential Equations
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d0d0f5c-d8f8-4536-b140-43772a28ffbb · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling RoFormer: Enhanced transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47596943-c5bb-40eb-a453-69a0a968f8d3 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation de6a525a-0689-434a-89d3-d92d240bb87a · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling STFT Spectral Loss for Training a Neural Speech Waveform Model
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e94efc7-e50f-458e-9a74-c7d54a720bc2 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality.IEEE Transactionson Pattern Analysis and Machine Intelligence, 46(6):4234–4245, 2024
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e90ce36f-5056-4f40-a957-93a7fe4a7fd4 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Relay Diffusion: Unifying diffusion process across resolutions for image synthesis
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90408f6-ba47-4573-8ab5-066a2e66222d · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling WaveNet: A Generative Model for Raw Audio
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 613fc9b5-8e3d-446c-9ab3-4b836eb18555 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 550fec21-15f1-4126-bb4e-0a2ef6a59556 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixNerd: Pixel Neural Field Diffusion
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31ea4afe-d4b8-4ede-a07e-8bee0a173d3f · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis.arXiv preprint arXiv:2512.04720, 2025
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8665f702-ea46-4599-849c-b89776716389 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2ee16ea4-22a0-4f0d-8c81-d09cd0a93592 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a7b4e12-1723-4df8-8762-75ccc3d3aaf0 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69a98e8a-b3aa-49cf-8309-6eced1dc7dc0 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ConvNeXt V2: Co-Designing and Scaling ConvNets With Masked Autoencoders
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a39acb0-45ea-42ea-afc2-6d1f97b3b7f3 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling TS3-Codec: Transformer-Based Simple Streaming Single Codec
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7403fb74-abe3-4dee-85de-5ee6503270fd · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ba13b69e-2615-4a35-acbc-eb134b1f42a9 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Longcat-audiodit: High-fidelity diffusion text-to-speech in the waveform latent space
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5f5e3a2c-dde7-499b-9c45-54bcf64c956c · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Unresolved cited work
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f36623-25f8-496c-a33c-5567b5ac41d7 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b09f98b4-38eb-4f9f-9700-9e4c990730bf · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation.arXiv preprint arXiv:2512.23278, 2025
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c53abe59-20e1-4186-af2d-831e5e1d587d · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e32fd3f1-d01d-4f9e-b79a-365fa6e14786 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixelDiT: Pixel Diffusion Transformers for Image Generation
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f0c0946-cdc3-45b3-9578-841c2d31ba8f · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling SoundStream: An End-to-End Neural Audio Codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 495–507, 2021
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd78baa-5689-41d6-aa07-625d33992412 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 005df705-46a5-4519-af3e-ae2234d98c48 · outbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Normalizing Flows are Capable Generative Models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 58bad382-20b4-4c5d-9025-9cca6096266b · inbound
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.