Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:58:11.144143Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2412.16977.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:58:11.144143Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:50:27.874421Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T05:50:28.000843Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0f003225-09d2-40a8-bf73-1435baf886e8 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7023625b-d187-4e84-b25d-bd942764d633 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d196bf4-1044-4673-94e1-1349a915344d · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis SC- GlowTTS: An efficient zero-shot multi-speaker text-to-speech model,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 997b8d9d-7478-41f6-9d7d-903c54c9e970 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 53623a00-9ed0-4c1b-8d01-62471b23f976 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4e5832-a357-4e89-a686-a13a34685dbb · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis NoreSpeech: Knowledge distillation based conditional diffusion model for noise- robust expressive TTS,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 458f2004-8d48-41c8-b12b-c1c0decb78e1 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Noise-robust zero-shot text-to-speech synthesis condi- tioned on self-supervised speech-representation model with adapters,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0b9bc3ab-db5a-47c3-8a3f-df7d8a0260b2 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Acoustic matching by embedding impulse responses,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2addc2b7-91e2-43b2-b20d-29ef9d6733c5 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis DiffRENT: A diffusion model for recording envi- ronment transfer of speech,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 22d8bf33-250a-42cb-a233-8536a97f5afe · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Environment aware text-to-speech synthesis,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eea404e0-5fdd-4951-a37a-cf65c8c6819d · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Binary and ratio time-frequency masks for robust speech recognition,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9058edef-2e78-4fe1-9978-4b8896f6c2fc · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Ideal ratio mask estimation using deep neural networks for robust speech recognition,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 78e20f21-de4d-4e88-92f5-7c0e44746592 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Attention is all you need,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 309b264e-8afb-4e5b-b295-127615cc909f · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e20162a5-6a79-4010-b926-3c59b8ac7a72 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis MP-SENet: A speech enhancement model with parallel denoising of magnitude and phase spectra,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df3c7e4b-c566-4938-a541-bd268fa1696d · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ef85c216-da57-4ebf-9708-0be1bd5047f9 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fef8a35-d5cf-4738-b928-1aa3bfb702f1 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis V oxCeleb2: Deep speaker recognition,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7dd5b4b4-5be3-4798-a5db-1fcc9c96d43a · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis DDS: A new device-degraded speech dataset for speech enhancement,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e998ab52-587b-483c-a036-854e9e610d3d · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Decoupled Weight Decay Regularization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e267d25c-3388-49da-95fa-7df3125db38c · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis WavLM: Large-scale self-supervised pre- training for full stack speech processing,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bed7036-d487-41c8-aec8-c4d517eee552 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis FunASR: A fundamental end-to-end speech recognition toolkit,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8012d855-cba1-4012-b3c4-a21a5f9f5141 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Distance measures for speech processing,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76734146-ac1d-40dd-a9bc-c0fb50221643 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e6887e-5a6d-4924-8116-944771e1e6e6 · outbound
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis ViSQOL v3: An open source production ready objective speech and audio metric,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3836c19a-1375-4dcd-960d-f1e59a7a2acf · inbound
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.