Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2406.07855.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:51.952647Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T23:19:03.868044Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 50e7f70c-2504-4599-a672-a57d5f920719 · inbound
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 141fa9e8-c17f-4f10-b23b-8427fbae40d1 · inbound
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2057ab5a-29bd-459a-a087-f2b4b936c622 · inbound
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c099402e-4fee-4705-b271-6f0c51e6f718 · inbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 578d911b-13af-4ec7-899d-516c8e495232 · inbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae8e0e9-d909-44bc-b465-2e2e7fd95af7 · inbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24e8798-a39f-40c1-b133-d801cb3473be · inbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68b88b73-7536-4456-8910-c2cd7652884b · inbound
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d71bed5-e07f-464b-8b25-890a7a7847c6 · inbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da940102-74c5-4465-8666-46c1934a7c50 · inbound
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b58e0cd-015b-4b9c-a206-ecd55188d95f · inbound
Next Tokens Denoising for Speech Synthesis VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb6b0037-dbd3-4354-9c35-d3fcd46dd6fd · inbound
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e059189-d9c6-48b4-9101-5cd600e5de40 · inbound
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34601198-1263-48fe-a624-349a6fbd892b · inbound
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da6cff5b-5b4c-4595-b5bf-35424dc08763 · inbound
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1c7c0ad-3a2b-41df-98f2-1572fbc0aba3 · inbound
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 972ee4d0-1303-4922-bdb3-2f0576d74bae · inbound
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8db91ba4-9e23-47ce-b64e-c5d770b7165e · inbound
SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50d6a789-19ce-4787-9372-6fb4276a5cf6 · inbound
UniVocal: Unified Speech-Singing Code-Switching Synthesis VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1eb896f0-e2f4-47a9-896e-a36b0eeb781b · inbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8604b175-b2de-43f5-b579-ba4ca0e83a0b · inbound
MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.