Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2406.18009.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:37:49.949120Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T10:37:01.584595Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f9b72dbf-3421-4be2-88ea-9ebd7141c765 · inbound
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a4d268d2-e593-4402-90d5-d7ba4783c17d · inbound
Zero-shot Voice Conversion with Diffusion Transformers E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3021b48b-e7f1-4a94-9036-9dd1c853f500 · inbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 978dfa2a-4ff1-432e-8ff7-204bc309cb87 · inbound
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ad64020b-cff5-4485-9922-e18c42653483 · inbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da9754b8-7a3f-49e2-ae31-f88bc3a5ac2e · inbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb4d767c-8237-4b66-8ebf-ce7fb7a4ca7f · inbound
Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc9ddd71-93f7-40da-94d6-6c2f9d161c84 · inbound
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56111811-5b5e-4c57-af0d-eb834e641a8d · inbound
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6b9603-3588-43e7-a355-04d4422fced4 · inbound
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 459caa60-3a6d-49c9-8acb-b7b8450965e5 · inbound
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76430cf7-6798-4e45-8f26-f2dd8121b684 · inbound
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c2e5d6-3b02-4eb5-adc9-9501f633fef7 · inbound
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6a85a2f9-0209-44fd-b840-7bb86963c6e9 · inbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572e330a-4104-4d2b-b35e-f211c1787355 · inbound
Next Tokens Denoising for Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b85f4d6d-3bfa-4a61-a653-8b689b6c30b2 · inbound
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0dac444-bc90-46c0-a043-27fcab1f0dd6 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1674fc9b-f029-4588-b108-cfe923adc7ce · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation facc5bea-8701-4760-9733-38eeacb80f97 · inbound
SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8c0d737c-63f1-414e-8a75-f261cdb0332b · inbound
DETECT-3B-Omni is Agnostic of Content and Demographics E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1df5631a-c75a-49ce-ad31-292479edaa5e · inbound
Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eb22d84c-56d5-40f1-a14d-1a95d0c6e2be · inbound
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e904c399-19c5-43ef-9edd-a6d6ce8a82ef · inbound
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e6decfd-b6b7-4b0d-9cd1-16704eff6bf9 · inbound
A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.