Pith. sign in

Paper Citation Record · LEDGER

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

As of 23 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 1 inbound Pith citation observation for arXiv:2606.03455.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03455 v1

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T08:18:42.002083Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T08:35:47.752843Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 109 outbound references displayed

  • verified exact48
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ce5ec35-c2ca-49f5-a2e9-a3785716a6e5 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.811706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b1b5953df82e94b05c5d7e8875210f95bee429a8ab30c16253d5ecd4f45dc15c

Observation df0517b9-3171-4a31-ba52-fe886fb22ab1 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:c0b4995b1c9c0ba75b46f2efa3b1d26e96b97dd5baaa7d2e32fdba997578fb3d

Observation c40da833-8482-4eeb-87e7-d065caf608f1 · outbound

This paper cites DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:9873961406729cfa6a4e3e63caf13a928505e5100aaa7cdb6cb398be2ba6973e

Observation f20ae24c-d64b-47e4-828b-bd08addbe3da · outbound

This paper cites WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis.Interspeech 2021, pages 3765–3769, 2021.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis.Interspeech 2021, pages 3765–3769, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ed2ee6f12d644aeebd25ede6f8bd5e4adceaac86197203354bfd800ec65e1d2c

Observation 16dedabf-4f7f-4a58-8162-5d6a50496af0 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.798400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ce425cdfa44f7f7aee0ac449b4aa4d67935e431055fe7c4771c840ebf459fdf7

Observation 6218d65a-65a5-4d6e-b949-71cc7d3294f7 · outbound

This paper cites PixelFlow: Pixel-Space Generative Models with Flow.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixelFlow: Pixel-Space Generative Models with Flow

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.840803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:fef9bd738a76b76cc342618afb28e7ca270d512f8d92bf0f6f0362a1519b95b2

Observation 45f9bd5b-b056-4fff-a1b4-fe1d47b823ae · outbound

This paper cites On the Importance of Noise Scheduling for Diffusion Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling On the Importance of Noise Scheduling for Diffusion Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.809182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:40ccb21dd8df9646485c9b2386de4a9672cd19210b5f730cf629e3707b61b305

Observation 019f0ce0-9655-4450-b990-02b02b103f7d · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a8298065c0086f344f9c4f54874738c77f61521f1b74b1282223e9c074e4997c

Observation 30a31182-3299-4d74-be02-638bed1425c0 · outbound

This paper cites Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:2f455b0ba14bfc340135923741713f53a39669a219afa8cbe79e9bca7b588253

Observation db452e27-0a93-4d7b-91a2-7f95867d71bd · outbound

This paper cites arXiv preprint arXiv:2511.18822 (2025).

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling arXiv preprint arXiv:2511.18822 (2025)

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.729433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a7cbdb5e41be6a9984b7d16f93d8f6f4d782b13945f3b9d13a7fdf7d61ae4da0

Observation 4da41692-3df3-48e2-8bfd-ad1512f0dbe0 · outbound

This paper cites High Fidelity Neural Audio Compression.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling High Fidelity Neural Audio Compression

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.747814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:5086efc97e1b7360b23f62f3bf831538fd361a2e84653821d24bf49fc4adb5ef

Observation 900dcf1e-ac6a-4e35-a7f5-d1b23f6b59ea · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.Advances in neural information processing systems, 34:8780–8794, 2021.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Diffusion Models Beat GANs on Image Synthesis.Advances in neural information processing systems, 34:8780–8794, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0e2c7a1775fd79989f266e0384141e156a8d25e334bd90d28b4b63e1085529f3

Observation df753d9d-693f-4cb8-80c7-091b7c29b7d0 · outbound

This paper cites End-to-End Adversarial Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling End-to-End Adversarial Text-to-Speech

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.837427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:12de23bbe94154df022118a76577bd9f71c872612ec9dbadcaafbf5f0dbf3b01

Observation 84fd148e-3bfb-4b2a-b1d6-0feb44b79df1 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.752712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:eb170ce4c68b8c8661f891428afae152470965bf430ffd3d1ac72cd24d5762be

Observation 9ce8533d-e777-4ec8-aff8-fa9b56f7776e · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.729854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f9f058077aaa10e6de0a9043ee88a71f619b7686af6e7ab46f1b2be206bda21c

Observation 5eeff31d-1b2d-4086-9d7f-865bcf215b6a · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.732631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:eb0f4fb3ce70f127af4ff92f114f28688459c19882663f46f20ef24d57a8d95e

Observation e4f26ebf-927d-405b-8dbb-12243196a13f · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:009a036997ccf2aa539560e25ea5a4d5b85dba5c8dae69bede4417bcbd676c2c

Observation cf141115-ac01-435e-baae-872024d51c15 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:758c97d8db0a3eec021c8d47d52f00a3f95664cac84cc06579f318a069525cca

Observation be4c75f3-b7e2-4970-9f3d-1e313c63015b · outbound

This paper cites E3 TTS: Easy End-to-End Diffusion-Based Text To Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling E3 TTS: Easy End-to-End Diffusion-Based Text To Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4ceb9f064a74d5c3ed4b667258440771fc34a3daac1792a8990e8a9f60ea2107

Observation c853f0d5-65bb-4044-a290-c6ee6eca0a88 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0f0e84241bc2f53379d4c48e3c574ef6421dfd73d859b9afdc4c1c968b72f01a

Observation 60db7b3b-d281-49f3-a971-56871a97da40 · outbound

This paper cites Moss-tts technical report,.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Moss-tts technical report,

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.777402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:eae457637c61c218fbe991adf8bb3c983410dfd049f9d7f98b08b74c783c4829

Observation 4476e729-e5e4-4ced-aa61-c1d5762c8250 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.768674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:405ea3426161be22e2942c4d85b5b6092df889301fc6789ca399fd6b3a2264e9

Observation a267b415-7ff9-4590-80c3-78217b090f24 · outbound

This paper cites Didispeech: A Large Scale Mandarin Speech Corpus.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Didispeech: A Large Scale Mandarin Speech Corpus

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:c6cbd16bbe0efe4ce53a6785120f49aa374fae34a4b51601e0d2575146f175b6

Observation d1544d29-2011-4933-b183-1c149f16ce6c · outbound

This paper cites VoiceFlow: Efficient Text-To-Speech with Rectified Flow Matching.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VoiceFlow: Efficient Text-To-Speech with Rectified Flow Matching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:96d3fb036571ceddcef1fce0f25f5fd87ee461d1c2b69ab908ae694a6c199e97

Observation 1eb896f0-e2f4-47a9-896e-a36b0eeb781b · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.771613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:8ec88e2d84bdee44c8657880b9d48b270005c4216dda6669d7386b79f92f1447

Observation 399552db-214b-40ed-8347-5df86bc09bc4 · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:904bb67ad1408a38127374546bb5bc665928d90b4dcbdbf541891c552f4c1548

Observation 2d58560a-0cb9-41a2-b162-8449f87fc371 · outbound

This paper cites Classifier-Free Diffusion Guidance.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Classifier-Free Diffusion Guidance

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.773456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:d43f24cfaee8dbc1eeb0aaa813d50077b568b089e47d565c5a7537f24350b246

Observation e703ebeb-3f54-431f-9045-fe7dde22a6e1 · outbound

This paper cites Denoising Diffusion Probabilistic Models.Advances in neural information processing systems, 33:6840–6851, 2020.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Denoising Diffusion Probabilistic Models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4ff222f7c76ec9f8ae864b5cc0699cc3b3223ff84742358c29d2e0e3a017f235

Observation cd112712-f9d2-4ed8-acfe-95e02c0fe6cd · outbound

This paper cites simple diffusion: End-to-end diffusion for high resolution images.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling simple diffusion: End-to-end diffusion for high resolution images

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:af3fb2fd43823319ca35caf64aa916f3cd35d7032372ba1f775ca9e1b07b15f3

Observation dcddd738-1fa2-46fe-b487-f6ee85e8a883 · outbound

This paper cites Simpler Diffusion: 1.5 FID on ImageNet512 with pixel-space diffusion.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Simpler Diffusion: 1.5 FID on ImageNet512 with pixel-space diffusion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:84cd9495f2e890bbdaefb6a64d332b337b024cfc0a82de6e0b4bdcd9dc12454d

Observation 754f9d42-b7bf-4419-a134-6074274b2ad2 · outbound

This paper cites Qwen3-TTS Technical Report.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Qwen3-TTS Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.774185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:904e62a479fe5c97540fdb84db76c6098b6a5fc85dfe372319d9c2c29dabc727

Observation 3c35d135-943c-4627-bae8-2c7ecbb8b3f7 · outbound

This paper cites FlowTS: Time Series Generation via Rectified Flow.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FlowTS: Time Series Generation via Rectified Flow

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.761621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:7a8614a33c80438a7018c88e5459640c8d4045e891ec626045ee270ecca2ec69

Observation 78fd2b98-7bc3-4e83-8b5b-5362f05080cd · outbound

This paper cites FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f5e7073cdfdbcc83d8becebc69f39a49938ae0f37d492c8ba3328ca73d9571de

Observation 6ff2e288-0835-4dfe-8332-a924ed469f2d · outbound

This paper cites ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-Speech

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:d23c21d1e7bcead8b81962371754adaed9aa652c2084e0d430fc44e4758b2966

Observation adc4592b-7fb3-4b4a-8ed7-95fea4218205 · outbound

This paper cites The lj speech dataset.https://keithito.com/LJ-Speech-Dataset/, 2017.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling The lj speech dataset.https://keithito.com/LJ-Speech-Dataset/, 2017

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:2bd18d558562972e14aab7369c59c71ac1bb93e76fdff7a9993f37b0bc32a661

Observation 1548522f-cae0-4606-8915-4cccf0898484 · outbound

This paper cites Diff-TTS: A Denoising Diffusion Model for Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.812312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:8c5ba2ff8a45bf023fd87cd0b590f7adfad00038454e08beb7a228f7fb6c7b28

Observation 01b87c9d-b6ed-4cd4-a947-0efb9eb284f1 · outbound

This paper cites DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:fa3fd7ae93b1a017f450d8dacd6884f8bb501f2ebfe23ef6f9d7539f8d42f281

Observation 768e4f15-4a57-45e2-adb1-22e3139c3de8 · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.706809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:606377919c6099aeccccd44a0a315d196b8e62e6992ab755e859ede54ef143e3

Observation aa02f98f-03b1-4561-8d4f-4f7b341c4665 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.706591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f35d0afb3a15eb287afcc128c91f190223df09e3de3e45f35c163e02bd66fd48

Observation b4eac9d8-23ab-4f62-a438-960e06501417 · outbound

This paper cites Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:e3016706a2093ef317314cfa24860927f33fc7c0d296e208153ddf84dd2daa29

Observation ea53c29d-d671-4eec-a403-b7b6e2ef4afa · outbound

This paper cites Auto-Encoding Variational Bayes.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Auto-Encoding Variational Bayes

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.700787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:3b0d0cdeca00bf15f6f0288d3a589ba725bc9c7114b65da3dc5ddd856b5173a3

Observation 2ee8c748-cc88-4787-a8be-21c472cebfa0 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.Advancesin neural information processing systems, 33:17022–17033, 2020.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.Advancesin neural information processing systems, 33:17022–17033, 2020

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6ed7541f7e4e8e3e8557b340393505055d85bf0a54532c043f69e6b86425805d

Observation dca840f8-4fa0-41bc-97ce-cd422fa19f48 · outbound

This paper cites DiffWave: A Versatile Diffusion Model for Audio Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiffWave: A Versatile Diffusion Model for Audio Synthesis

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.829738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0bc3f9e44301aac6d1534a1999310a33a06d1b1d43373cc449a696344bc169e5

Observation e60b09b7-313d-46b8-a2ef-b546b3e1633c · outbound

This paper cites High-Fidelity Audio Compression with Improved RVQGAN.Advancesin Neural Information Processing Systems, 36:27980–27993, 2023.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling High-Fidelity Audio Compression with Improved RVQGAN.Advancesin Neural Information Processing Systems, 36:27980–27993, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a92573e7aacc9492579c08e7d4644abb364e5a20dba2c4135e9d4dc18108b9b6

Observation e5a1bf20-4e87-4599-a4e6-88cd7a38d279 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.768082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:3dc0db2fadefaf6840df4a1ad53be8696a76077c829347250c7e5a4a67cf09f8

Observation 93762c56-0f5f-4aaf-b7f8-003f12fb0d8f · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:705f581ef2cd84bf6ca8a612893e524432447a61499eadead2f2c7d00ba657d9

Observation 0fc37142-53f8-46bb-83f4-d1143d9721ee · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b433e984ab7142f7670625e338d190a159bfb810e148d29b3ebfb2c1dd43166f

Observation f58b15e8-af53-49b9-9393-c734f69dbe9a · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.735273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:034f8ac9ae6b80a62d5d304c188872de4a68dd7b47be4aef090a1be393f2511d

Observation f277ce7e-50df-4401-9254-c44410190db8 · outbound

This paper cites Back to Basics: Let Denoising Generative Models Denoise.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Back to Basics: Let Denoising Generative Models Denoise

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.761116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0f40d812c5f4296cf3c6ee559d7c9b38364a002707f105cd27b0208162bf0418

Observation 44f790ef-2204-43c1-9598-ef8a19186413 · outbound

This paper cites JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.840425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:112092e19a7a2cf816b422f5514dcba4eda675e4d072f5ca3c58fcd48c388499

Observation d93cfb6a-c520-4951-8abc-63c44a4623ad · outbound

This paper cites Flow Matching for Generative Modeling.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Flow Matching for Generative Modeling

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.855011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:07aa1560b77d8fadbab5b27ad3cc9a6cce2a87322833bc0baf625b43d8d46178

Observation 9d1f613a-e5f6-4e7a-93fc-7039f1464f42 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.832283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:38796946cae08ad7f5c785c3b13e5c2c1b46d7ed581c80155f525332a7ff1859

Observation 5d82fe48-e325-4ff3-9ad9-036537040823 · outbound

This paper cites Decoupled Weight Decay Regularization.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Decoupled Weight Decay Regularization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:e4b0abfffc31a106870fea9ede743fda4f1901fb9ec7e875a0bd1dff26b5cd7d

Observation 7f3c258a-852a-4242-a01d-cc80ea235884 · outbound

This paper cites DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.826742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:5bb2d181afd7e2e7d3c257e6a4e58a8cb281feea220db52ebd32fb593d50ff16

Observation ffcf786b-c72f-4b11-a3fa-2347c54ae5f0 · outbound

This paper cites PixelGen: Improving Pixel Diffusion with Perceptual Supervision.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixelGen: Improving Pixel Diffusion with Perceptual Supervision

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.845806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4d1382bfb7320ac638a80687f5b183a67814c7887cbc07efa7bb6d463bff3874

Observation f31c0db0-3543-4d95-a6b7-28cb58225332 · outbound

This paper cites Matcha-TTS: A Fast TTS ArchitecturewithConditionalFlowMatching.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Matcha-TTS: A Fast TTS ArchitecturewithConditionalFlowMatching

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:8eab0c2127750d245a52afb66575b4e1aca8f06749c4072926b804a910c160df

Observation 8cdafe16-651d-49e0-a531-ad43a798df9f · outbound

This paper cites LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Model.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:16757170dabd5a200dc47dda7e1ad6a9184db1272136f5670adf527e13dd704c

Observation ba7daf40-9071-45a4-80c3-516cde68a34b · outbound

This paper cites Improved Denoising Diffusion Probabilistic Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Improved Denoising Diffusion Probabilistic Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6d146605e78f605653f2057ceee7c191474fd700ab718545089aebb31907f973

Observation f1c769f9-a12d-437b-bae7-50e3205c6852 · outbound

This paper cites Semantic-vae: Semantic-alignment latent representation for better speech synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Semantic-vae: Semantic-alignment latent representation for better speech synthesis

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.837808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6524c6bc75920bc888fa003a3b489cecf59b8e37ccc7e7b997450d673aacd180

Observation 0a9c96eb-7ed9-43d8-aa76-5bd61c4eefa5 · outbound

This paper cites Parallel WaveNet: Fast High-Fidelity Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Parallel WaveNet: Fast High-Fidelity Speech Synthesis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:bcdf9147205e4d08bae638d544eb54bf5ebe7be423cd81b856ff0de07dc1e33f

Observation ccc58cfe-27a9-4a9c-b488-67a14d863b52 · outbound

This paper cites Scalable Diffusion Models with Transformers.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Scalable Diffusion Models with Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:252bd36f7d1d3f1bdea4f2712b1ab9cf918b92cd7b0525e28b1dd1e91edd22e8

Observation 1288d50c-9117-447c-8a9e-df7df0f944e9 · outbound

This paper cites VOICECRAFT: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VOICECRAFT: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a3115b8eb8358ec4c0a2220122c981b97e88a63187caa6569b272f5055aa2136

Observation 0ea07226-6a13-4d74-bdd8-483523798cf9 · outbound

This paper cites VibeVoice: Expressive Podcast Generation with Next-Token Diffusion.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VibeVoice: Expressive Podcast Generation with Next-Token Diffusion

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f7b93fe3d711447be7055cee33c601c16c98409f0a91ec5a867ea602f865577e

Observation 61467cab-e5f6-43e0-8cf1-f7820e9b05e1 · outbound

This paper cites ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.832450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:362aa8dce6cbd69a7a52e0b9bb2bfae77328b7ec3b4af437216f24393bd57bb4

Observation 8d87da6c-ac09-46e4-84f8-6fc00f5f085e · outbound

This paper cites Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:017ea896fe379125db983d940f84f757abf93f460efcb9efbca6edde3aeba4d7

Observation b353a749-a4b4-4144-924d-9f88eb68b7c0 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Robust Speech Recognition via Large-Scale Weak Supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b52be376ab0f3601fe13ce5860278d4c67e49f23f61307510f64bb79a8888a68

Observation 3be336fa-3726-4214-87ef-5c60f15dba00 · outbound

This paper cites FastSpeech: Fast, Robust and Controllable Text to Speech.Advancesin neural information processing systems, 32, 2019.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FastSpeech: Fast, Robust and Controllable Text to Speech.Advancesin neural information processing systems, 32, 2019

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f9d157468f3fae87306030d52c204700d6fb03b0fd768f213fbf1f490ca11690

Observation 588b3f10-0b4d-43f6-9633-d1741589302c · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:24b237102c65d1ea5f1658135e3ad39f9f529ac4abd05e1302077177367c786e

Observation 387105e9-6335-4635-981d-487ee250f320 · outbound

This paper cites U-Net: Convolutional Networks for Biomedical Image Segmentation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling U-Net: Convolutional Networks for Biomedical Image Segmentation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4454aa85c5ac90751986b2141464d2b8355de89a73e560fb9c01cc1a87d20308

Observation 89d42bfb-bd2b-47e9-9828-db643118c589 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:5b104c0ac114c971679a3895256157a16fa13c2d229d786e2d592e64d8b7c454

Observation e2124967-1ff3-4a58-99ea-ec40a2394c72 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:78997c9a86cd1f4d69d39e121fea7ce4c85ce464ccbab0684d01ca795aac9cfb

Observation 311e960d-e47a-4882-9ac5-9415515192ac · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.843462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:91ad350109fa0848555db265cc7c12cf5eb445ded5ac89c12696f1355ae47634

Observation ee1ef2a2-d58c-46b0-a5ef-149e2bfa0a28 · outbound

This paper cites MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.817490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:775d24dbe456978debcb75c7399a7a717d925773b316bef4ec869314c63a090d

Observation ff3ca7c0-f631-4154-b5a5-a89225372a12 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:9d8d12abd19198db11ec819116127a851e7f02540079aa7c9607494561287355

Observation 22236b6a-a76d-43f3-ad87-a72d169c297e · outbound

This paper cites Generative Modeling by Estimating Gradients of the Data Distribution.Advances in neural information processing systems, 32, 2019.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Generative Modeling by Estimating Gradients of the Data Distribution.Advances in neural information processing systems, 32, 2019

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:91944c9725ca6b6d605e076d9ddaa86655a916d907b9abb61ca198347bc3755a

Observation adbbee38-7e5a-4d54-bd29-5caa7eace8a8 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Score-Based Generative Modeling through Stochastic Differential Equations

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.814716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:9d470c751483c00025ab46691e171293c74e49e1ecbb94f6fd97a32badbd8f1c

Observation 5d0d0f5c-d8f8-4536-b140-43772a28ffbb · outbound

This paper cites RoFormer: Enhanced transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling RoFormer: Enhanced transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:3ca121ba15bcce8b7f833898e41c84047b5099ec8a28fc78b1bfa3e218754648

Observation 47596943-c5bb-40eb-a453-69a0a968f8d3 · outbound

This paper cites F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.779566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ff301f365e65e7f45cf2f0520fa7d7ef9bec19195813470b6b4774ea00f2354b

Observation de6a525a-0689-434a-89d3-d92d240bb87a · outbound

This paper cites STFT Spectral Loss for Training a Neural Speech Waveform Model.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling STFT Spectral Loss for Training a Neural Speech Waveform Model

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a6d1bbb6248aea0b599243902a9fe50c777671cbf4fc80d554296c20a97e020b

Observation 5e94efc7-e50f-458e-9a74-c7d54a720bc2 · outbound

This paper cites NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality.IEEE Transactionson Pattern Analysis and Machine Intelligence, 46(6):4234–4245, 2024.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality.IEEE Transactionson Pattern Analysis and Machine Intelligence, 46(6):4234–4245, 2024

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6461e86ccc548728141c12fabe84ecd0ba304d7ebc041f5256681c4189bb1978

Observation e90ce36f-5056-4f40-a957-93a7fe4a7fd4 · outbound

This paper cites Relay Diffusion: Unifying diffusion process across resolutions for image synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Relay Diffusion: Unifying diffusion process across resolutions for image synthesis

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:8d77024d9327ff9924cc8cab714e42776b907e64d29cdff47bbf5fb7833edf26

Observation c90408f6-ba47-4573-8ab5-066a2e66222d · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling WaveNet: A Generative Model for Raw Audio

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.856964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:48d4f3983b625f8e5690b4001dc62ef970e402b08a0dec318e12fc751de37b40

Observation 613fc9b5-8e3d-446c-9ab3-4b836eb18555 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.813970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:fd7c313ac2232551981fcf3e88c84a7274e36092e9f7d2e9dce61283d6c1a6f6

Observation 550fec21-15f1-4126-bb4e-0a2ef6a59556 · outbound

This paper cites PixNerd: Pixel Neural Field Diffusion.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixNerd: Pixel Neural Field Diffusion

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.848770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:71f5d210ea3e614908bb774b55c5cb5ee440b0061b995ef4c7fb78c053f44a0a

Observation 31ea4afe-d4b8-4ede-a07e-8bee0a173d3f · outbound

This paper cites M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis.arXiv preprint arXiv:2512.04720, 2025.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis.arXiv preprint arXiv:2512.04720, 2025

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.782209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:0577d1eae4566bbd3d8e248a23e691e3caab9835f63516ad310f0c6404ae9760

Observation 8665f702-ea46-4599-849c-b89776716389 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.744079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:a235c9bc2ccb986394ac772481c250313d95dcfd8599f2fa556e06c7d91c7cbb

Observation 2ee16ea4-22a0-4f0d-8c81-d09cd0a93592 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.738763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ceb6abd647f72ddbdcb76c185aabcde4ded3e781f0683e6fdfe44d402e5cf8a9

Observation 8a7b4e12-1723-4df8-8762-75ccc3d3aaf0 · outbound

This paper cites Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:aa449ce631f0761383831d702ddd7c6a3bba2c50c1be3acbb9296f6565ea4d4e

Observation 69a98e8a-b3aa-49cf-8309-6eced1dc7dc0 · outbound

This paper cites ConvNeXt V2: Co-Designing and Scaling ConvNets With Masked Autoencoders.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ConvNeXt V2: Co-Designing and Scaling ConvNets With Masked Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:93f06eed2f01d1c47bd338e811e8907e04e96f5871fa50debc771a1e787e1678

Observation 7a39acb0-45ea-42ea-afc2-6d1f97b3b7f3 · outbound

This paper cites TS3-Codec: Transformer-Based Simple Streaming Single Codec.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling TS3-Codec: Transformer-Based Simple Streaming Single Codec

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:7044cf156b8cc3eb2d26351eb95d6af87ce5ec853acd1d40f8d6eb1640a93fca

Observation 7403fb74-abe3-4dee-85de-5ee6503270fd · outbound

This paper cites FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.764041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:318b06dd40ced3b3b0395625bdb83e24e38ce5e83af0b1aa16c7b5071f831670

Observation ba13b69e-2615-4a35-acbc-eb134b1f42a9 · outbound

This paper cites Longcat-audiodit: High-fidelity diffusion text-to-speech in the waveform latent space.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Longcat-audiodit: High-fidelity diffusion text-to-speech in the waveform latent space

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.753373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:d94a8aac851369de419639cf3ee0ca14aed74c6a589ee8d540acb2231f501e79

Observation 5f5e3a2c-dde7-499b-9c45-54bcf64c956c · outbound

This paper cites an unresolved cited work.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:2aa924ae09f448e26a1fe86e88b348c89f48f26b9152ea87c4ec0e88647db489

Observation 08f36623-25f8-496c-a33c-5567b5ac41d7 · outbound

This paper cites Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:4d3f2ad084ea0fcb6805d59b6bcb3167038b7490f38f0e8e96381a5def767b43

Observation b09f98b4-38eb-4f9f-9700-9e4c990730bf · outbound

This paper cites Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation.arXiv preprint arXiv:2512.23278, 2025.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation.arXiv preprint arXiv:2512.23278, 2025

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.750020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:b02d762591e0dfc38dbfa501ddb9a14f83e53e51bfcad0645db4bc6ed3e791cd

Observation c53abe59-20e1-4186-af2d-831e5e1d587d · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.803185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:1101f77842f2e60c206797073a2d231ba27e2e6368a76c6d5c39e15011a8de30

Observation e32fd3f1-d01d-4f9e-b79a-365fa6e14786 · outbound

This paper cites PixelDiT: Pixel Diffusion Transformers for Image Generation.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling PixelDiT: Pixel Diffusion Transformers for Image Generation

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.779918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ba201c09581d87e557657fb7df7f60b7e050119f11f6894151dfae435437e102

Observation 1f0c0946-cdc3-45b3-9578-841c2d31ba8f · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 495–507, 2021.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling SoundStream: An End-to-End Neural Audio Codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 495–507, 2021

Reference 98

Resolution
unresolved
no resolver link, observed 2026-06-28T08:18:42.002083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:5c9f92c3224e9ef96f21ac72a92a32f0e863fbad9349c414dc61bb7df6341c20

Observation fcd78baa-5689-41d6-aa07-625d33992412 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.759013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f7d3941b3235c37469220bcab2f49ee948b56ca8171f75027c12963f73948470

Observation 005df705-46a5-4519-af3e-ae2234d98c48 · outbound

This paper cites Normalizing Flows are Capable Generative Models.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Normalizing Flows are Capable Generative Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.794214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:e348921324cebb3bf071eab3c04062b2ee993f5b482be64f54b1520e9dd95eed

Pith citing papers

Observation 58bad382-20b4-4c5d-9025-9cca6096266b · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:47.752843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:47.752843Z digest=sha256:1c36950500d100a10f278f0f0a8d556b3222148fed8ff578da2b864aad29996d