Pith. sign in

Paper Citation Record · LEDGER

Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2104.01409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.01409 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:37:49.997321Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:28:18.721905Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 69b0d3b4-99c5-4976-9e38-e90db2844ca6 · inbound

I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception cites this paper.

I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:37:25.334264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:37:25.334264Z digest=sha256:aaca4d58ef2195c321bc9f5de463ae935620fac95196ecd25da0121dcdd0cb68

Observation a71b7461-2d48-43c5-b9e0-0ccf15f548ad · inbound

BiDM: Pushing the Limit of Quantization for Diffusion Models cites this paper.

BiDM: Pushing the Limit of Quantization for Diffusion Models Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:10.455465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:10.455465Z digest=sha256:a5b7d74d0c93866b3a48ccf2de8747c4b827db224d36508d3232dd6832867cd6

Observation 25d5390d-fcd5-4d57-bfbf-717db62961ca · inbound

UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion Models cites this paper.

UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion Models Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:27.629372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:27.629372Z digest=sha256:b8a24914b03e4da6958569761881c34988013a4818ce570a865dfca62aab87f6

Observation caa7e61c-5546-4f50-bd15-c1fdc9ccb94e · inbound

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis cites this paper.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.742096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.742096Z digest=sha256:8af95ebe32cf7082c3c1bcf80bfe20357db1b22e22d0e7a4698783ba8cbfaaa0

Observation f2ba1678-4716-4563-8a31-4ac646a617a2 · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:22.491013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:22.491013Z digest=sha256:668bead9131aaa3571b1037bced985383d9504ccc4307d6b51d2bbd3c8f2b337

Observation 28eeff79-995a-4718-8680-a967211824a5 · inbound

Towards Flow-Matching-based TTS without Classifier-Free Guidance cites this paper.

Towards Flow-Matching-based TTS without Classifier-Free Guidance Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:37:49.997321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:37:49.997321Z digest=sha256:84c8d03d42058ea2ed13067ff8b03c00ad6a5b98c08f35f9817c502c8f7ecec9

Observation beb6ccda-a997-4dd8-86b0-3ae4f7f25889 · inbound

Language translation, and change of accent for speech-to-speech task using diffusion model cites this paper.

Language translation, and change of accent for speech-to-speech task using diffusion model Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-16T01:03:14.125761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:03:14.125761Z digest=sha256:3f0bb631bbc6e264ecc610212b63aa701c888cd2ba2ca679bcdd4d06bd66acbb

Observation 48f9d036-8c4d-4757-b766-4d72b26dd8fb · inbound

Constraint-Aware Diffusion Guidance for Robotics: Real-Time Obstacle Avoidance for Autonomous Racing cites this paper.

Constraint-Aware Diffusion Guidance for Robotics: Real-Time Obstacle Avoidance for Autonomous Racing Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:52.676400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:52.676400Z digest=sha256:e5537ab6250ef3490ed58415b9dce3827e9890c1a32acae9eaa6ae72da4caebe

Observation f6e01e88-bcb3-4361-aba5-afa975df48de · inbound

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning cites this paper.

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:01.278354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:01.278354Z digest=sha256:c03275d0b82f05366346cffed6e00c7dfacd5634eab64e280d0a3fb6c44719ff

Observation c636cfa4-0e30-4c23-86b8-2db483dbb3c3 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.528332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.528332Z digest=sha256:b05926fcd80808ec96c94e0f7abd57443281fdeadc175539f606acfff30f1567

Observation 6046f9ba-f728-4514-a924-9f35f1ed279c · inbound

Noise Aggregation Analysis Driven by Small-Noise Injection: Efficient Membership Inference for Diffusion Models cites this paper.

Noise Aggregation Analysis Driven by Small-Noise Injection: Efficient Membership Inference for Diffusion Models Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:05:57.947285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T06:02:41.596562Z digest=sha256:67221bdbf590d6600fd66b091246b9903362b1a736785fa47895139f94a0d6db

Observation eedb4db5-9011-4342-8c38-b03710722572 · inbound

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech cites this paper.

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:00.776378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:24:54.627215Z digest=sha256:e5e275092fb5603007be99385968c50fedc19a48a7623c352f9d7372c4f25e52

Observation 5070c683-7319-4e9e-9a25-75b95cf08c01 · inbound

Scaling Properties of Continuous Diffusion Spoken Language Models cites this paper.

Scaling Properties of Continuous Diffusion Spoken Language Models Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:28.056125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T03:43:10.132447Z digest=sha256:77f2f5f24e14d467fb448c2672489dfd7ac9d32c22f2918ae1b52c9cd3d2a47e

Observation 1548522f-cae0-4606-8915-4cccf0898484 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.812312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:85f45a854e0c48e3af4ccfdbd07093b4aa1929f8cd3d18262820878d86488590

Observation 61d39f3b-6218-466b-b649-3ce44a09eabf · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.723282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:3a5fb4681aa5ed0d41033f9a97d65aef5c4c6c3158209de38ea74a0d2f62ad3d

Observation 581e6858-b141-4dce-ad42-fa3ef2d7af1e · inbound

Autonomous Collaborative Learning Among an Ensemble of Tsetlin Machines with Consensus-Based Inference cites this paper.

Autonomous Collaborative Learning Among an Ensemble of Tsetlin Machines with Consensus-Based Inference Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T10:45:21.609565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:45:21.609565Z digest=sha256:fd4ea0729a3a64150b502485fb974602d766c7a57b89c9774fd78ea18e5376ca