Pith. sign in

Paper Citation Record · LEDGER

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis

As of 20 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2412.16977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16977 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:58:11.144143Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:50:27.874421Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:50:28.000843Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f003225-09d2-40a8-bf73-1435baf886e8 · outbound

This paper cites Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.456630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.051317Z digest=sha256:f20a279e2b963aafec9e7a42e4a34e5c9509ac2ca865c75cb1bb5cd032c07c9e

Observation 7023625b-d187-4e84-b25d-bd942764d633 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.055801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.055801Z digest=sha256:fc3f608787fdb63224d2bc90784fe58f7aa232e8faed2c89ebe14e7f02327ee0

Observation 0d196bf4-1044-4673-94e1-1349a915344d · outbound

This paper cites SC- GlowTTS: An efficient zero-shot multi-speaker text-to-speech model,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis SC- GlowTTS: An efficient zero-shot multi-speaker text-to-speech model,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.436653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.059944Z digest=sha256:652b7b4e90fdde598a0a1f33df93016ef4e06baa0814932b602bd0f495ecaaea

Observation 997b8d9d-7478-41f6-9d7d-903c54c9e970 · outbound

This paper cites YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.424010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.064642Z digest=sha256:ff7cc76ad015b757ef050f7b05ad59a70792222fe013f59bc27a1d016b93a05a

Observation 53623a00-9ed0-4c1b-8d01-62471b23f976 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.068685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.068685Z digest=sha256:bccdabc3a43ec93b0db19994a2b87d2192dfc6f80883170a3f292cdca65076a6

Observation db4e5832-a357-4e89-a686-a13a34685dbb · outbound

This paper cites NoreSpeech: Knowledge distillation based conditional diffusion model for noise- robust expressive TTS,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis NoreSpeech: Knowledge distillation based conditional diffusion model for noise- robust expressive TTS,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.409937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.073147Z digest=sha256:c1fe3a75d03ae7650bed0496bc04d34f494214372710269ef878003c251f4f69

Observation 458f2004-8d48-41c8-b12b-c1c0decb78e1 · outbound

This paper cites Noise-robust zero-shot text-to-speech synthesis condi- tioned on self-supervised speech-representation model with adapters,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Noise-robust zero-shot text-to-speech synthesis condi- tioned on self-supervised speech-representation model with adapters,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.395299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.077470Z digest=sha256:28f0752ab58d461ab994654c017f28cd18dbbe35d420fbd52a10f7ce0951cf6a

Observation 0b9bc3ab-db5a-47c3-8a3f-df7d8a0260b2 · outbound

This paper cites Acoustic matching by embedding impulse responses,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Acoustic matching by embedding impulse responses,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.382768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.081492Z digest=sha256:3df07f9823f270e879ba6affaf5a1c1f1cb8524c90a24bf65ce2af02ad85b668

Observation 2addc2b7-91e2-43b2-b20d-29ef9d6733c5 · outbound

This paper cites DiffRENT: A diffusion model for recording envi- ronment transfer of speech,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis DiffRENT: A diffusion model for recording envi- ronment transfer of speech,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.360285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.085125Z digest=sha256:471d521104d53bf65ec6e47b9a3f6d20897914f88dbcb06535da1b7c4a115172

Observation 22d8bf33-250a-42cb-a233-8536a97f5afe · outbound

This paper cites Environment aware text-to-speech synthesis,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Environment aware text-to-speech synthesis,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.346732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.088792Z digest=sha256:366291fb56c79df6659acae9b3eba0381528710464d51a317ee8cc28fff296cd

Observation eea404e0-5fdd-4951-a37a-cf65c8c6819d · outbound

This paper cites Binary and ratio time-frequency masks for robust speech recognition,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Binary and ratio time-frequency masks for robust speech recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.092470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.092470Z digest=sha256:e8fd479cd54d85d127cad3967fdd28a7519565f22de6f4028b9f8aefbb8e26ff

Observation 9058edef-2e78-4fe1-9978-4b8896f6c2fc · outbound

This paper cites Ideal ratio mask estimation using deep neural networks for robust speech recognition,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Ideal ratio mask estimation using deep neural networks for robust speech recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.323189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.096056Z digest=sha256:5e9c7fb8002710e5389fe0f323701065f98f8b7604a1941e76a1bdd2f650b6b0

Observation 78e20f21-de4d-4e88-92f5-7c0e44746592 · outbound

This paper cites Attention is all you need,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Attention is all you need,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.099891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.099891Z digest=sha256:cdc2016e2f99c510f2c649c7cfceff74c29441565bed2aca538fb7895fcea92e

Observation 309b264e-8afb-4e5b-b295-127615cc909f · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.299980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.103426Z digest=sha256:023a874f2c706db261aaa6b8d97567c1c762088ddd69d5f933e411f0b099381a

Observation e20162a5-6a79-4010-b926-3c59b8ac7a72 · outbound

This paper cites MP-SENet: A speech enhancement model with parallel denoising of magnitude and phase spectra,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis MP-SENet: A speech enhancement model with parallel denoising of magnitude and phase spectra,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.107212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.107212Z digest=sha256:537fca733c73812e31e921f6fd018dfb09245acf8363b52f58e2a8c8db7e4280

Observation df3c7e4b-c566-4938-a541-bd268fa1696d · outbound

This paper cites ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.280249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.110926Z digest=sha256:87b57ff5ce2c63bdb2d3917cc3a72da12beae0015a611f86f2f24f4eeaabfe63

Observation ef85c216-da57-4ebf-9708-0be1bd5047f9 · outbound

This paper cites Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.114835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.114835Z digest=sha256:77c8fae29b742cad6a7358d2f997d9bb532a83450bba08d84373ef5c499d09a4

Observation 1fef8a35-d5cf-4738-b928-1aa3bfb702f1 · outbound

This paper cites V oxCeleb2: Deep speaker recognition,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis V oxCeleb2: Deep speaker recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.267898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.118900Z digest=sha256:ba4385252e922c5bac36563de8efa9cad57f2000131013ebcc29df93636f1edf

Observation 7dd5b4b4-5be3-4798-a5db-1fcc9c96d43a · outbound

This paper cites DDS: A new device-degraded speech dataset for speech enhancement,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis DDS: A new device-degraded speech dataset for speech enhancement,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.254967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.122679Z digest=sha256:eb44722ba754772bcdec152e300f88a0972feffb1f408bcfde32503979b24b7a

Observation e998ab52-587b-483c-a036-854e9e610d3d · outbound

This paper cites Decoupled Weight Decay Regularization.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Decoupled Weight Decay Regularization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.126411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.126411Z digest=sha256:e3324eb6e956e3be8d868e55d9d3a442ec4d49c3595a736c2c4efbc8e3828b9e

Observation e267d25c-3388-49da-95fa-7df3125db38c · outbound

This paper cites WavLM: Large-scale self-supervised pre- training for full stack speech processing,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis WavLM: Large-scale self-supervised pre- training for full stack speech processing,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.130406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.130406Z digest=sha256:07cc38834a4d3b3b20933919e5c9fa653fbdd9101957b1226e51b0174c7805bd

Observation 1bed7036-d487-41c8-aec8-c4d517eee552 · outbound

This paper cites FunASR: A fundamental end-to-end speech recognition toolkit,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis FunASR: A fundamental end-to-end speech recognition toolkit,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.235232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.133763Z digest=sha256:4db4867c12ed4ae73009a9cdd2ad1da1ff1bb6ddf2d79c745064ce4778e0d4de

Observation 8012d855-cba1-4012-b3c4-a21a5f9f5141 · outbound

This paper cites Distance measures for speech processing,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Distance measures for speech processing,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.137152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.137152Z digest=sha256:2a8710daea44ead6ebe63c97f0e3a230b64c0271ee0d7e5778fad50ca346d1b1

Observation 76734146-ac1d-40dd-a9bc-c0fb50221643 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:11.140700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:11.140700Z digest=sha256:c8c81f63fe79f241500b08a05f0f5f62a8b450cadeedbff77a2bc2d4c4fbbacc

Observation a9e6887e-5a6d-4924-8116-944771e1e6e6 · outbound

This paper cites ViSQOL v3: An open source production ready objective speech and audio metric,.

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis ViSQOL v3: An open source production ready objective speech and audio metric,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:11.209542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:58:11.144143Z digest=sha256:7c1edac914fcf9fca24612a7199340b33bc9f2df8a38d20ee470120d5106d40f

Pith citing papers

Observation 3836c19a-1375-4dcd-960d-f1e59a7a2acf · inbound

In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion cites this paper.

In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:50:28.007478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:50:27.874421Z digest=sha256:363ad7e2be7ecac1435d0a8a8fae2f5c9a0abcbfa568bd7cda2eb4085cb9dacd