Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2406.05370.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:58:40.546152Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T00:04:22.411558Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 251b1e65-964d-4550-9ac4-6742880289fb · inbound
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f83f5ca-41ac-4880-8515-d288a2c21a99 · inbound
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91e21a39-50cf-492f-a160-8d31e37592d9 · inbound
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59084775-ccb8-4363-aad0-0ed7a15b0c55 · inbound
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6446f42-6a4d-4da2-8133-f9a06b7a46a7 · inbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f0abd3-8987-4043-b332-5830629f06d2 · inbound
Voice Adaptation for Swiss German VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bf7acdd-960c-4c94-9766-173297f31e72 · inbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 170af19d-7e84-4a5d-bd61-fb515caa88a5 · inbound
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3105a6d-23ac-44bc-adf1-d7dbaa5fd501 · inbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65de2375-289d-467f-a1fe-5fcfac7c0ea8 · inbound
Speaking images. A novel framework for the automated self-description of artworks VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c04f6727-a3cd-412a-8844-60e880dc20d2 · inbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd381a56-5b71-410b-ad15-d4d22877f2c0 · inbound
Dataset of News Articles with Provenance Metadata for Media Relevance Assessment VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 564ecf6b-fd82-4b41-b60e-b9bf15fb71bc · inbound
A Variational Framework for Improving Naturalness in Generative Spoken Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a995078-6c79-4198-9ada-73e1eedc2b32 · inbound
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8be1b90-e219-4d50-826e-65bebd9db008 · inbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c650765-46aa-4e7f-868f-ebc596c3d331 · inbound
Next Tokens Denoising for Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53ce4747-da48-485a-b0f5-1a133c8439b7 · inbound
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986fe835-d4af-4a6d-886d-bc6c77efe5ff · inbound
DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dbb91b5-f808-4ded-a2db-5a1266f0324c · inbound
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2b2830-f9b6-4d17-90a5-b4d3fc598206 · inbound
DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaedd33e-6e22-4792-acae-86dc149fecf8 · inbound
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a66343d-5dab-48a9-871d-109ba2cbdaee · inbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35c49ad4-9c77-458f-b191-99ac9d25e756 · inbound
Position: Towards Responsible Evaluation for Text-to-Speech VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21fdd2b4-abf1-44b8-a7b8-baf9025ebf8e · inbound
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0b2c8d6-4306-4cb9-89aa-27d3603f54a1 · inbound
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e7816cc-9440-42fb-8945-f208a6c3d830 · inbound
CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e4a4aa1-4677-401c-a4c4-1a60ac95ac1f · inbound
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f2bb8ca-5d3d-4fda-a860-ad6d9be317d6 · inbound
ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bb668d8-ef79-4305-acb6-567cefcafa54 · inbound
AST: Adaptive, Seamless, and Training-Free Precise Speech Editing VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85f745ae-e1e7-4840-ac6d-8a4fa94e34ae · inbound
AST: Adaptive, Seamless, and Training-Free Precise Speech Editing VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e50250-0aaa-4e0c-90dd-762047a61968 · inbound
SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 195876d7-3f1f-44ab-8cad-3a89abb7c570 · inbound
Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba4299f2-3240-4434-810f-20192360e422 · inbound
Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77c95389-f79c-4cd9-95ca-df1846225053 · inbound
Can We Hear from Events? Generating Speech from Event Camera VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3030e54-10c4-45dd-9755-9e24c05f2cd1 · inbound
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b24562f-9744-4546-9833-a25226c5359b · inbound
UniVocal: Unified Speech-Singing Code-Switching Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16dedabf-4f7f-4a58-8162-5d6a50496af0 · inbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9350765f-4c89-45dc-affd-64f9f4f2041b · inbound
UniVoice: A Unified Model for Speech and Singing Voice Generation VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 298ce71c-92cb-475a-911c-c4223f441028 · inbound
TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 054d7587-0a6b-4ae9-971a-4feb9c37c6a1 · inbound
MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca50811b-9982-4aba-a2e4-56b9bda8a131 · inbound
Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8b77bdb-d009-464a-840b-7cd5be7d4fbb · inbound
PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccf1f974-c7e1-4f05-bc22-628c2e7525fc · inbound
Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b86b6ea6-9515-4274-bc64-37edf17ea2f8 · inbound
CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c9e99ca-6521-446e-bb85-08545df4436f · inbound
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48ea64e8-2bb3-48f4-8f5a-65b0cecc7052 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 236
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a256841e-f3a6-48a0-8633-cfba7b2c8117 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 236
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17ef13f6-c1e3-4fa9-bdf7-dc79538eb0b0 · inbound
SSTMark: Robust Training-Free Semantic-Level Speech Watermarking VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41cc9bbd-15a6-4a9d-b01c-e1eafaa25407 · inbound
Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce2fba8b-b260-40ef-a172-decdda9b2d4c · inbound
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7878a799-da21-4452-a938-411c8de1e7f4 · inbound
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee3d8f0-ae27-4ffb-a62f-2f41c2b2fb4c · inbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.