Pith. sign in

Paper Citation Record · LEDGER

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2301.12503.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.12503 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:58:34.577052Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:59:50.469153Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4244a785-d3d1-4482-8f44-0db941ed88c7 · inbound

DGSNA: Dynamic Generative Scene-based Noise Addition method cites this paper.

DGSNA: Dynamic Generative Scene-based Noise Addition method AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:45:46.270816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T17:43:47.086524Z digest=sha256:8161ef80284cd41af3b37262fa030647b8aafbd6843a11ed31d831d0dcc642cc

Observation 833355d6-6bc7-45e7-b768-abac05b45178 · inbound

How Far Are We from Generating Missing Modalities with Foundation Models? cites this paper.

How Far Are We from Generating Missing Modalities with Foundation Models? AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:15:33.620840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T08:15:12.947854Z digest=sha256:c5d45fd58b07e9c625c4339d81e67ffee55da135a96a3097597a83d63b065ba7

Observation 401ae98b-4024-41b3-ad2e-c5b3acd1d977 · inbound

Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models cites this paper.

Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:34.577052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:58:34.577052Z digest=sha256:f9b77a28fa7f016060e806bc31d34bdbdbaa17041dd180a8ed1c4b9b3766256d

Observation 3384e3c5-011a-4867-b15e-0451f2ba050a · inbound

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation cites this paper.

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:54.589254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:54.589254Z digest=sha256:7532f5ff07919958cd6bc61416451c7d729484f3510767f5cf5e2315b8892b8a

Observation 8afa710d-257f-449e-a116-60b3e028de26 · inbound

ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion cites this paper.

ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:42.399897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:42.399897Z digest=sha256:31274dd398f2637c117ed331da451cfba5890fccccecc3902a5677f91fe2c19d

Observation c1a90ca5-8730-466f-96ed-44324b7d35cd · inbound

Diffusion Models for Time Series Forecasting: A Survey cites this paper.

Diffusion Models for Time Series Forecasting: A Survey AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T15:58:50.494374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:58:50.494374Z digest=sha256:0a321fc264ac57fab14e8dff51d048a52294a3a8bcb91fcdc78651652821e81d

Observation 9139e409-ca8c-43d2-9ce7-0a57256d1a80 · inbound

CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers cites this paper.

CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:02.687551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:47:02.687551Z digest=sha256:592d2c1e1e3b1ba66241c7a50fc14fb5ad1d7056ddcb3dd7ca8f499a94559ec9

Observation 035f79e7-9d79-4d62-83f6-e690c2962950 · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:00.776212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:00.776212Z digest=sha256:c9139731dcb2faf6bccb66fc4ac5054062827f3d13592f06f54868a1d707adca

Observation 511d80a0-aa90-46af-93dc-954e0950af30 · inbound

Flow Matching Policy Gradients cites this paper.

Flow Matching Policy Gradients AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.495265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.495265Z digest=sha256:4ccc3729ed06cc85c64995db72cdddd60538a973dfe6c969e601dc9bab0da136

Observation d212bfcf-8c24-4765-b564-a318947e4b00 · inbound

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs cites this paper.

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:19:01.944976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:19:01.944976Z digest=sha256:d39e9244be59318be298f414456c01529fed4cf45e214b794a9504c0a5cf8da5

Observation eae0bf53-89bc-4d9b-adb5-4745a5dce62d · inbound

Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning cites this paper.

Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:14:46.151866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:14:46.151866Z digest=sha256:03e1f5f6f9940cf5119c5e8437c1c70c30dd359ab798b9b5c209cadeeb2260c8

Observation 1d75771a-4415-4798-b175-2d463d71af6d · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.913997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.913997Z digest=sha256:f99f9838561d34f4dea01b2525977224e0283509cc7a7b8c0380fd496955f089

Observation 64148a22-c14a-43ff-a71a-acb113314683 · inbound

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation cites this paper.

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:56.915651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:40:56.915651Z digest=sha256:bbc02b98923f8027ce1214edd2fc3afca2dc93833d49004c0a1fd3cb335f6ae9

Observation e0fee5c5-370f-447c-94d1-65406a061189 · inbound

Inference-time Scaling for Diffusion-based Audio Super-resolution cites this paper.

Inference-time Scaling for Diffusion-based Audio Super-resolution AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.134039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.134039Z digest=sha256:37aa0637dfb357878b3eefad7275b875dd715999c63834c45d70242a466cc992

Observation 986d2272-7240-4034-913a-4c095f3bd0ca · inbound

ASAudio: A Survey of Advanced Spatial Audio Research cites this paper.

ASAudio: A Survey of Advanced Spatial Audio Research AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:55.090118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:55.090118Z digest=sha256:e62e490c46737fc7320bca78b0267e29d0128cf62418ca7f40c66e768038d7cf

Observation b40285ca-fcc7-4c76-86fb-d4cc7f16590e · inbound

A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions cites this paper.

A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:10.648281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:10.648281Z digest=sha256:eb9460b882737b8c8ede02102bd07f65baea450108a25689bed091ff5f5ba4f7

Observation b0ef9092-7b6a-45a5-a957-900a8070585c · inbound

Audio-Guided Visual Editing with Complex Multi-Modal Prompts cites this paper.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.753877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.753877Z digest=sha256:a073ee3d3eec2f323e98ab89b07f5f8d356f0bd54dbc866f82f6fa309b569956

Observation 01204637-e03e-46bd-86aa-ef69bf55b79c · inbound

WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration cites this paper.

WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:36:37.617270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:36:37.617270Z digest=sha256:89f7187b0b256b66e5865814104ebead5077a5394b2c8087aaa7a4567505c809

Observation 29112a42-28a7-48ca-a089-aee4158f9105 · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.790250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:16fa6e487716a6a9897ae331563d75c727a94c2831ccbc448a7440d9027f9bf4

Observation cf197551-189e-4413-bfbb-75800a2a5d56 · inbound

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction cites this paper.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.198477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.198477Z digest=sha256:e0e6ea2c2c5de8889680fe9fc873da7becb90f24feabeb569adf26856215ae42

Observation 9ba08cfd-6e72-4426-9c96-74149d070b56 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.883026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.883026Z digest=sha256:ee22a0794d3709c77f274e7eea22911d6edf7ce275355891c9ef40335c820e0c

Observation 5cab15e9-428c-48d5-8f1c-1590be20ebf8 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.658946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T07:42:43.077644Z digest=sha256:cc577e5ba483c6c94d875395ffb204ba093c684f516ce0e9facc12a08433d996

Observation 7f33acef-db98-43a1-926c-d821ecfb8475 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.130192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T20:56:37.533183Z digest=sha256:b098a63b65e1c79c64a7e2b868a0399dd28337241ca7198d1e53042823cb0df5

Observation 34d78cc4-a91d-4143-a1b0-bab32f6be7fa · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:49:43.654091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:49:43.654091Z digest=sha256:267f38b55f218a5eee65e7d96713e8d157e3bd4874968b8b65beafa8d5d33761

Observation ad9ddeab-612d-4c4d-806d-6591e5de0714 · inbound

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion cites this paper.

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:57:42.900012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:55:32.044506Z digest=sha256:6d4015e55957294ec86b0e06f723a71e440ecba0a6377b0ae35f1f10b902d03a

Observation ad26e427-935b-4ab3-bf78-e31e4327eb8f · inbound

Dual-End Consistency Model cites this paper.

Dual-End Consistency Model AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:40:39.814660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:40:36.406150Z digest=sha256:d7d326d45393fcfc1dfbd45ee055901a48c8082c458c07eee4d4f6a2906c647f

Observation a5fe170f-ff47-4c70-9b9d-e78aba20a207 · inbound

Dual-End Consistency Model cites this paper.

Dual-End Consistency Model AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T01:04:20.684172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:04:20.684172Z digest=sha256:c9cf02877bbf759587c9d7ddbb94fdc2c1c6cfecfe23773846713b880175a6fc

Observation d8e4dd18-5751-40fe-bcf0-56855f426025 · inbound

Diffusion Models Memorize in Training -- and Generalize in Inference cites this paper.

Diffusion Models Memorize in Training -- and Generalize in Inference AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:54:07.907159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T10:52:31.849094Z digest=sha256:3d61f69d2fc59a2ab6dc4568a8754ece477155fd785bda9a46df2d5291409f88

Observation 6f6fbe12-1056-481f-a394-17d714e80274 · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:40.092528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:40.092528Z digest=sha256:26f54f046814e2fedc6f61e4d26f099992ee5fff0545e9aa51ad09bcd744c1ac

Observation 7d8d74b6-99e3-412b-a589-d61ee8b439db · inbound

Woosh: A Sound Effects Foundation Model cites this paper.

Woosh: A Sound Effects Foundation Model AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.805338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:51:08.144573Z digest=sha256:fac80af6621a878dc795435379ffaea928e8fde1bfbda81f0cc24dabd1f2b9d8

Observation 01351671-3543-4f01-a0f1-0ac538ec27e2 · inbound

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips cites this paper.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.725801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:bd43aa1974a4875613b9b0173141cd085eb239ae6b0d06f4e3b0310a6d402957

Observation 7a567944-8d5b-41e9-8911-b1b29291eb45 · inbound

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan cites this paper.

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.633260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:40:10.657590Z digest=sha256:7aa53e6e247149bfd20fdefd0f8f0fa689f697185b6fa41bc752154e4c4ae137

Observation 29ea7c1a-3e59-4755-a1a7-e76c6cf6c59b · inbound

Latent Fourier Transform cites this paper.

Latent Fourier Transform AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:21:07.061004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T03:45:07.892234Z digest=sha256:c492d555f27a421005d8cbc817f652abf2a8dcbc79ae40e62379a48ad5a13d85

Observation 01797019-b68d-4ae6-ac3f-1015173d9bd7 · inbound

ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis cites this paper.

ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:11:22.315786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:11:12.460695Z digest=sha256:c9f9ba95ce4eef46a542ae3b77406bfd13cca6a534f9bca9909ecd3c7963dbc4

Observation 39d88a97-f0b0-48b8-9ac7-7861846c0866 · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:44.540336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:7d3d51e48360430cfe74d0284c985351aa625dc030879c784e1ad92e2baca9a6

Observation 2d7908dd-aed0-4ad3-87ba-adc684434783 · inbound

Stage-adaptive audio diffusion modeling cites this paper.

Stage-adaptive audio diffusion modeling AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:06.288613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T17:21:34.140699Z digest=sha256:e231bb971dee133c6a8ae618da0d5efce6640f0a60b5647a341bd16cbd574d7d

Observation 9213e690-cbd2-494a-b509-1434c2bb22a9 · inbound

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems cites this paper.

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:28.002815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:58:26.634355Z digest=sha256:81a4e8db1a52610f42f96b2bf0ddc469fe4ffb58d351c3a57eaf77a896fed7b0

Observation a8c7fa61-6d59-4926-9f7a-96b99a4fd859 · inbound

DiffATS: Diffusion in Aligned Tensor Space cites this paper.

DiffATS: Diffusion in Aligned Tensor Space AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:26.589603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:38:12.633086Z digest=sha256:5907cf602240613e2880f8f7ab213d15e9b229654e0275766f345101a7dc45f3

Observation a5609b5b-8f1c-4de5-8d38-02ceb2028256 · inbound

HapticLDM: A Diffusion Model for Text-to-Vibrotactile Generation cites this paper.

HapticLDM: A Diffusion Model for Text-to-Vibrotactile Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:28.188916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:11:29.782569Z digest=sha256:81f457661b696ca62fdefcdceaa1c52c0187dcc1ff8f1239309591d53372b5cd

Observation 1e51d76d-72e3-48e3-9143-9a44fa42d66a · inbound

PoDAR: Power-Disentangled Audio Representation for Generative Modeling cites this paper.

PoDAR: Power-Disentangled Audio Representation for Generative Modeling AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:35.176549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:04:08.420600Z digest=sha256:2f7ef4f36a859874306ffd4dd9e5777ff574e13ee7e4f8c92480243b84a1a07a

Observation 875fd0db-094b-4cac-a577-9f4e223b7d81 · inbound

WavFlow: Audio Generation in Waveform Space cites this paper.

WavFlow: Audio Generation in Waveform Space AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.626883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T07:33:35.243337Z digest=sha256:e895eb61c66d3fe4f3607e698a07efdb179857c79dedce2f2df70d9f39527fc4

Observation 9e9be8ae-4fc9-4935-9355-e05862a46376 · inbound

EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction cites this paper.

EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.537734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T10:25:14.873276Z digest=sha256:3ba46c2312cd7ad9ca48d535b1e0e9d5079a96e023264de9f54a463b671b9d85

Observation 917ed80e-e9b2-4838-8169-7486d16d7403 · inbound

Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation cites this paper.

Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:15:01.184596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:57:35.126894Z digest=sha256:97ff81e597507839e9bb6c9632e0e448cd6e045c4fdcc2a350753d28f281880d

Observation d8b7e1ac-56f6-48d5-9137-35997613f8c2 · inbound

dMoE: dLLMs with Learnable Block Experts cites this paper.

dMoE: dLLMs with Learnable Block Experts AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.089912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T22:50:51.900169Z digest=sha256:cb1d1922546ce6ff32e60b603836073eaf5e10446a87ef55449ab0096ae02d2f

Observation b957d6e4-935a-412e-a3f5-2f0b3e30d1e4 · inbound

Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development cites this paper.

Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.253964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:06:45.886467Z digest=sha256:34cf2a90a1a36d5b34e2a4357ab38416000e1cc32eb0bd1fc2ae691454bf5d73

Observation 2987213b-e846-46ff-9bad-cb73d2f87e50 · inbound

Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics cites this paper.

Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:27:46.755613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:44:39.810293Z digest=sha256:548ca5e42e28cc467f0ce58d09b85e1d1057ddf95379dee22cb0fcfae77af34a

Observation b2be330d-b942-43c6-81c4-adebd15d638f · inbound

Net-Ev$^2$: A Generative Simulator for Network Event Evolution cites this paper.

Net-Ev$^2$: A Generative Simulator for Network Event Evolution AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.609498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:09:20.657051Z digest=sha256:803f70b6789eb1375e3edd6580caa9cc00c3d7c927f41a73139aff6abf943b97

Observation 6553f382-ad31-49dd-8943-a4627ad387a8 · inbound

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation cites this paper.

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.471182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T07:26:28.527338Z digest=sha256:979de74ce703403c1ff02356c4937c7c3efc6908d94b0baa12004861c7de3058

Observation 201afcc3-58ad-49f4-b226-f6f680c2fc2d · inbound

ARIA: Adaptive Region-Based Importance Allocation for Conditional Diffusion Distillation cites this paper.

ARIA: Adaptive Region-Based Importance Allocation for Conditional Diffusion Distillation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:44.954421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:41:11.951712Z digest=sha256:3b4bd0c6d8dddd3ae7c1fad64566d4e94bfc942a8137f13ca5189a063e65e972

Observation 1f58b6b3-7635-4d6a-b9b4-127dc290554b · inbound

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control cites this paper.

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:24:21.554371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:18:20.501369Z digest=sha256:743b2411f667ffefb394f34f88649db6ea9155ac619cd93667102dae721d17c7

Observation d870af0f-cf64-45bc-add2-1507338a977b · inbound

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control cites this paper.

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T17:03:01.432145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:03:01.432145Z digest=sha256:034f93c08765df087e9de93fbb87442246eb15c4c9e0c841e150f3887728615b

Observation 09db36e7-d0b1-4508-9387-bb85347cb181 · inbound

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models cites this paper.

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:05:37.221480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:47:53.308909Z digest=sha256:a4844d445305e206a12a839f37acf36bc347c627f11753b0cd395babd9c2f025

Observation eda6be9b-5c73-4779-a2c7-3b4ba7304da7 · inbound

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation cites this paper.

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.292439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T05:04:23.295592Z digest=sha256:1e575b9b33960ad6bbbc3a78ca16db28222f24759c9d8aa112f64f5e1bcba3e9

Observation 6b618a80-5245-450a-b915-ba7b6712dcc8 · inbound

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation cites this paper.

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T12:34:20.057072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:34:20.057072Z digest=sha256:4f57959ebf2efc14e04eb7c8c0169cf2ae60c52d4c3e5be51eef62aed3d223d0

Observation 46dabe77-610a-4901-8c3c-bfea8c1bbebd · inbound

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment cites this paper.

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T10:59:22.914974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:59:22.914974Z digest=sha256:1ca53d3e2eeb3684d439ca5a00d81f49014a2c73064ed80b8ab192577e088c94

Observation c0eb786b-3987-46b9-8017-92c2c3547882 · inbound

FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving cites this paper.

FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:40.274605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:40.274605Z digest=sha256:b16c033d826cfa016f11a59ffebdb597feaf8d1c9b06cb918a576eff7503e3d9

Observation 961127b7-5d90-47d2-a72e-207eddba16ef · inbound

Analytic Distribution of Classifier-Free Guidance for Schedule Design cites this paper.

Analytic Distribution of Classifier-Free Guidance for Schedule Design AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:59:45.840064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:59:45.840064Z digest=sha256:e807d9a53c4d235b60f2dd320d5ef56d7ae88e874dcb8cef8786a24f49baa65a

Observation 09755735-db25-461a-a381-ef03fee5233a · inbound

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling cites this paper.

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:48:09.922353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:48:09.922353Z digest=sha256:fec0cec8f17249eafe7b009f4583d1a211bd73ac745972501824d610b1225fcf

Observation 9bbb7541-3030-47ac-9cc4-64cb3ef81e9f · inbound

Amortized Moment Matching for Visual Generation cites this paper.

Amortized Moment Matching for Visual Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-30T18:58:28.076405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:58:28.076405Z digest=sha256:7d702124b8c73038920b2712be2fe8dc9c3f38b397e9d24b2d7a7f2474435b97

Observation 8e6509e6-dac5-48c8-9391-0c7876ebc211 · inbound

Exploring Efficient Waveform Diffusion Models for Foley Sound Generation cites this paper.

Exploring Efficient Waveform Diffusion Models for Foley Sound Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T12:46:58.947785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:46:58.947785Z digest=sha256:6e20778d80cc65517819e27fc2df3ee0d816c78f22276399dba7b060875319f9

Observation 5c16e57b-094a-4a42-a1dd-9b50a8f604f0 · inbound

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities cites this paper.

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:43:18.242065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:43:18.242065Z digest=sha256:65137ffd971a84661c8cba263530f3cb76d9a31c65791a13a89fed0ecd21c541