Pith. sign in

Paper Citation Record · LEDGER

A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2303.13336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.13336 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:27:53.194410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.950402Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1aa292d8-1074-403e-871c-d66fb6638db9 · inbound

LatentSpeech: Latent Diffusion for Text-To-Speech Generation cites this paper.

LatentSpeech: Latent Diffusion for Text-To-Speech Generation A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:12.915500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:12.915500Z digest=sha256:318e7b97ae8054a67ac0dd5f310469f5a9dd6b1031a1aa6ffbaf59d03c2d3a25

Observation e051d1c3-7e5e-430b-9e02-eb6759013cb3 · inbound

Are audio DeepFake detection models polyglots? cites this paper.

Are audio DeepFake detection models polyglots? A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:11:20.708684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:11:20.708684Z digest=sha256:2e3113a56d0b690fe67e5737cd7bdf75c03ef9b4ab268cb6b9509a102c5cfcf5

Observation 785047e2-ff42-4ee9-a1b2-6dceb08f0623 · inbound

Boost-and-Skip: A Simple Guidance-Free Diffusion for Minority Generation cites this paper.

Boost-and-Skip: A Simple Guidance-Free Diffusion for Minority Generation A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:17:12.793801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:17:12.793801Z digest=sha256:c21d0453ff2b383e0cdea5c6915b885de85d37a07df5183ec97b6754a2a29734

Observation c57031df-7dd5-4e0b-8e73-f699b9d5cf01 · inbound

SyntheticPop: Attacking Speaker Verification Systems With Synthetic VoicePops cites this paper.

SyntheticPop: Attacking Speaker Verification Systems With Synthetic VoicePops A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T21:10:02.504068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:10:02.504068Z digest=sha256:f0f078aaf6c17d60df40ff9e9e26fcb9ea144753bc4df5b50f5d5e339d21f063

Observation f190b9e8-24dc-4db9-9fa0-3999412316aa · inbound

Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond cites this paper.

Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:53.194410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:27:53.194410Z digest=sha256:e53c9a9c3d4070c37a3f66d9c2b25b641ac1f32a101f52038ebaa22980004297

Observation decb9ae9-765b-439a-8be8-2b9f73f8a154 · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:58.845200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:58.845200Z digest=sha256:41542dd163cc6d6f55fb821dac286737c45cd298e43b27abd6320c31abc5e9c6

Observation ba40f52d-e387-48a7-9d1b-dcec9d791bdd · inbound

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching cites this paper.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.307090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.307090Z digest=sha256:8686d64b48e9811a30da53fffa086c4479f03bf3f71943970bfff840b27e39a7

Observation ed6b3a6c-29f5-478f-83a8-d6b4cdc25168 · inbound

The Effect of Stochasticity in Score-Based Diffusion Sampling: a KL Divergence Analysis cites this paper.

The Effect of Stochasticity in Score-Based Diffusion Sampling: a KL Divergence Analysis A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:40.247061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:40.247061Z digest=sha256:9253d834e94edf153c820425aac9cc4c96e5fb312484745ec8afbbb57c246ff2

Observation bdb366e9-50f0-4711-99f6-962014b3a2e2 · inbound

MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation cites this paper.

MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:18.263678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:57:18.263678Z digest=sha256:ca4916c3f9ee0eb8071ae37988970e413f9f3ea9597fc9770d4ebea47c1a3e27

Observation 605f7043-c24c-42d9-82e5-de471472d079 · inbound

Marco-Voice Technical Report cites this paper.

Marco-Voice Technical Report A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.655722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.655722Z digest=sha256:9085d3e743909efd71c778b38b3faea4273435adebfeae5e5d2c6310b4cc6114

Observation c4712f20-4b91-46b7-81db-1de49dcf08e2 · inbound

Permutation-Invariant Spectral Learning via Dyson Diffusion cites this paper.

Permutation-Invariant Spectral Learning via Dyson Diffusion A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:26.632802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:26.632802Z digest=sha256:1a4ee59009cdddaaab5071a4470cb41d3d6f94386d3d186a55e99428df0ebf50

Observation f121603c-f82c-4356-bc00-ee365d68637c · inbound

Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs cites this paper.

Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.720588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T16:23:50.751356Z digest=sha256:33c2a3a32c2b218b0f138d6a617fccc7f13c58e8477314f66afe080a08a0c090

Observation e561caa1-c2cb-4966-adf2-af054f964a90 · inbound

Grokking of Diffusion Models: Case Study on Modular Addition cites this paper.

Grokking of Diffusion Models: Case Study on Modular Addition A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.869487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:28:08.221886Z digest=sha256:39fb83b01a16a6757be72ee4bbad00dc9aa5e5967bfe3e7e42a72c8783e0a002

Observation b9d300d0-be12-4912-819b-708ed15ab9dc · inbound

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges cites this paper.

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:01.259983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T15:21:02.100545Z digest=sha256:4f02fa769a0ffc19316f7e291ba146b6252d4e48fdf055941ca1e031293673e5

Observation 94a39d1e-6cab-447d-8a3e-e1ea1d83cc1b · inbound

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges cites this paper.

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:12:28.683468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T07:08:43.088643Z digest=sha256:7e4f3bde8af8f47c72a77a0bdff23e45d82f0513c1c4d68e596255b8a049b0fb

Observation ea88926b-e53f-435c-8d21-638ff2aa0b9c · inbound

Inverse Design for Conditional Distribution Matching cites this paper.

Inverse Design for Conditional Distribution Matching A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:56:29.547249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:46:08.232616Z digest=sha256:78d38f18dc5251f709b2e1b09a88630958ade518cd13276057b45c6dccaf31ab

Observation 3af78a66-4dca-439b-a3ff-dd229d8d4773 · inbound

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining cites this paper.

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:52:35.108279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T18:53:10.978631Z digest=sha256:1da58b9df432887ed4c217ece6577f7d6c89529334613e5fcae78c239d2c50e3

Observation 62e40a7a-ec40-46d4-98a8-9300c21c4415 · inbound

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS cites this paper.

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:20:07.953940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-25T20:11:18.045777Z digest=sha256:5d3cc9d1f8fe49e837b7a4c63b66bd0c0b67ac921fed2a59a621eb5d013ca668

Observation 16c60f6a-ce96-486e-a433-54372c5c4485 · inbound

Simulation-free and finite-time diffusion model cites this paper.

Simulation-free and finite-time diffusion model A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:00:16.676024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:00:16.676024Z digest=sha256:75cec3a3de379a99d42b7557bbb8441696fde9b39f7a222407f15285225ffb8e