Pith. sign in

Paper Citation Record · LEDGER

A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2303.13336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.13336 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:10:02.504068Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.950402Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c57031df-7dd5-4e0b-8e73-f699b9d5cf01 · inbound

SyntheticPop: Attacking Speaker Verification Systems With Synthetic VoicePops cites this paper.

SyntheticPop: Attacking Speaker Verification Systems With Synthetic VoicePops A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T21:10:02.504068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:10:02.504068Z digest=sha256:ed92d50911ff1da206bdf9a8d34693e12e54ba5721c488eda1ba1e74677d5dd0

Observation decb9ae9-765b-439a-8be8-2b9f73f8a154 · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:58.845200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:58.845200Z digest=sha256:af694aca18eaaafb1765b27082777416326d4a307d8f19078a2a3968873362c0

Observation ba40f52d-e387-48a7-9d1b-dcec9d791bdd · inbound

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching cites this paper.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.307090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.307090Z digest=sha256:39ed86cdb57479cba38c5f6ddcc969fad8e35cfd7a3e2ab95da5495f942da446

Observation ed6b3a6c-29f5-478f-83a8-d6b4cdc25168 · inbound

The Effect of Stochasticity in Score-Based Diffusion Sampling: a KL Divergence Analysis cites this paper.

The Effect of Stochasticity in Score-Based Diffusion Sampling: a KL Divergence Analysis A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:40.247061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:40.247061Z digest=sha256:2ff55a3c78f5421edc09a9c3f3e34e69a05062dc50e4d9e1f33d5a33988cea43

Observation bdb366e9-50f0-4711-99f6-962014b3a2e2 · inbound

MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation cites this paper.

MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:18.263678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:57:18.263678Z digest=sha256:d987e3a5d6095625bebc7fd884b32ed153e0325d5fe73812b4bbae0a838e95e1

Observation 605f7043-c24c-42d9-82e5-de471472d079 · inbound

Marco-Voice Technical Report cites this paper.

Marco-Voice Technical Report A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.655722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.655722Z digest=sha256:e1aeb91f46c98db0537d8e29a53779a78af93b85c25769a9c78e4155c7379915

Observation c4712f20-4b91-46b7-81db-1de49dcf08e2 · inbound

Permutation-Invariant Spectral Learning via Dyson Diffusion cites this paper.

Permutation-Invariant Spectral Learning via Dyson Diffusion A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:26.632802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:26.632802Z digest=sha256:fd16a250491d4736369bb0ab99fe653532914ef39571fba7b63510803d974c1a

Observation f121603c-f82c-4356-bc00-ee365d68637c · inbound

Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs cites this paper.

Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.720588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:23:50.751356Z digest=sha256:a4e301e9e694ec6505505d7cc04f7c5e5e4116955ea9ea754c3e789563a27116

Observation e561caa1-c2cb-4966-adf2-af054f964a90 · inbound

Grokking of Diffusion Models: Case Study on Modular Addition cites this paper.

Grokking of Diffusion Models: Case Study on Modular Addition A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.869487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:28:08.221886Z digest=sha256:92fa2b9eef7bba0c05ef9c272c0eebea394d8a3070399af7791b1322da62de06

Observation b9d300d0-be12-4912-819b-708ed15ab9dc · inbound

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges cites this paper.

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:01.259983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T15:21:02.100545Z digest=sha256:cf31b06b44108b8286bbabb5dc77082ef20f622730e3528bc95bba33bb8ac66a

Observation 94a39d1e-6cab-447d-8a3e-e1ea1d83cc1b · inbound

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges cites this paper.

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:12:28.683468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T07:08:43.088643Z digest=sha256:644213c96a42462d02b5092422be9ac4256ca691d1917dfecf5a184e207310af

Observation ea88926b-e53f-435c-8d21-638ff2aa0b9c · inbound

Inverse Design for Conditional Distribution Matching cites this paper.

Inverse Design for Conditional Distribution Matching A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:56:29.547249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:46:08.232616Z digest=sha256:da27a5057395b82d337876ec22145d1f3f0de8083cfc2f1c7b50589b6d948e98

Observation 3af78a66-4dca-439b-a3ff-dd229d8d4773 · inbound

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining cites this paper.

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:52:35.108279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T18:53:10.978631Z digest=sha256:198c63c405df9ce268b147281f3b97f04f71385ac8e3bab76d5bb0cc8f1d5550

Observation 62e40a7a-ec40-46d4-98a8-9300c21c4415 · inbound

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS cites this paper.

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:20:07.953940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T20:11:18.045777Z digest=sha256:30506008c2889daa7ac125395903bfb8eb751a03421306b8b6cbf2420bbaff03