Pith. sign in

Paper Citation Record · LEDGER

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

As of 6 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 0 inbound Pith citation observations for arXiv:2509.06027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06027 v3

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T18:26:51.583145Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

81 of 81 outbound references displayed

  • verified exact18
  • verified fuzzy63
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7642b494-c035-44cc-ac4a-b05c2e9aefa6 · outbound

This paper cites A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.735534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:09f1a3767084e421c029f5a458ae3369973b1b0870d884d9ad60657d3020a223

Observation 0773b260-04c7-47f6-aa98-4d3df8433028 · outbound

This paper cites AudioLDM: Text-to-Audio generation with latent diffusion models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models AudioLDM: Text-to-Audio generation with latent diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.255407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:1eeb14042a6048b966f74b4c1c23bb3731b4096d48bdb4fb22e545cf22864b83

Observation 00ec06bb-1478-4b21-8bae-2f70238bbc95 · outbound

This paper cites Sound to visual scene generation by audio-to-visual latent alignment.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Sound to visual scene generation by audio-to-visual latent alignment

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.288142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:d48526575c915e9d237dfcb795d01143042620c5caea9ce29bc61c02f227a2c0

Observation 4375c37f-95b2-4e0d-b1f0-6abff235761e · outbound

This paper cites I hear your true colors: Image guided audio gen- eration.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models I hear your true colors: Image guided audio gen- eration

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.263743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:fc862f671f5d03272750a0a638205950b84db526adc769dc80737092eeef83d9

Observation f7977440-22bd-44ea-9fc1-4d2afd73ba7c · outbound

This paper cites FoleyGen: Visually-guided audio generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models FoleyGen: Visually-guided audio generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.277812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:71a3d2929137482d43458ed521c745fccf2f01854857362496fe4f94a4bce89d

Observation 645bc478-5778-45cd-b982-ce51f867ddd1 · outbound

This paper cites Taming visually guided sound generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Taming visually guided sound generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.259089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:9f581af2e717547013d7fdd1a5e83c52d9e312fadfe2ac3590789a17680e6dcd

Observation 269b90b1-4d25-4a80-9380-b3f3faa91e60 · outbound

This paper cites AudioGen: Textually guided audio generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models AudioGen: Textually guided audio generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.270694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:ae27f32aa3303d303284c013478d263a59a61056e87417ca1019b4f49b426603

Observation 773c6459-989e-4f7b-b31d-317b2f504356 · outbound

This paper cites Riffusion: Stable diffusion for real-time music generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Riffusion: Stable diffusion for real-time music generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.291649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:4ecde93b281c6b0b469fc2f48f0ce5f74e844d41d356b7257e744262607c6553

Observation 9440ba7a-6b32-4097-9fea-fe0e6e52c4a4 · outbound

This paper cites WavJourney: Compositional Audio Creation with Large Language Models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models WavJourney: Compositional Audio Creation with Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.730268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:2bef6695afcbf3d237c92d0d7bd2467c994443b03c1c85f9c0ec159783c8a86b

Observation 59d5ca74-22ca-4bfe-821a-d0f16ad18ccf · outbound

This paper cites Denoising diffusion implicit models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Denoising diffusion implicit models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.281587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:a69d841950fa7917280d54bb36f2fe9500d8da75dd267e0650650caa4fe0b444

Observation fec0267d-85c8-4d47-a872-d6c970337aa8 · outbound

This paper cites Denoising diffusion probabilistic models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Denoising diffusion probabilistic models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.267445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:5aaa8dab6dc374d93bb97d85db8ba2c2a310065ad674067baed4ada553875328

Observation 4e2dd7ff-7dd3-4af5-8423-79fc92a1e248 · outbound

This paper cites AudioSet: An ontology and human-labeled dataset for audio events.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models AudioSet: An ontology and human-labeled dataset for audio events

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.284844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:2fc4878eeefaed0c44e8025e3b7056927ffca218e7553d6659000d45cfd3fe83

Observation 485879b3-bd6e-4a4b-af0a-6cadf46c17b8 · outbound

This paper cites WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.251951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:4c48efcca78b185d0a0a7ea998bcd58def4643d26104aff584dc8a9cde542b4e

Observation 4771f2ac-97a3-40ef-b2a2-46101ba00a32 · outbound

This paper cites Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.659445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:bd0d376f32b50462d9ef65e095735c6d0bee8d02cd7ced99caf35dedae4c189f

Observation def8acb9-1deb-4742-84b3-da6f5c0bc2eb · outbound

This paper cites Sound-VECaps: Improving audio generation with visual enhanced captions.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Sound-VECaps: Improving audio generation with visual enhanced captions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.171759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:cb4c2d03cd9e1e3a993c6f5bd17dcec90c5c66c46125f3d9eacd487f0972b136

Observation c93ab04c-7266-4c66-9671-a3440115c573 · outbound

This paper cites AudioGen: textually guided audio generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models AudioGen: textually guided audio generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.077274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:4055700fcf5e81a3b4c4ef58ae4bbd3d5525883bc947c6723e7c12e370f8e240

Observation 64d810fb-0c2c-4331-918b-1500e8870872 · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Diffsound: Discrete diffusion model for text-to-sound generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.199051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:1491fa342ef74d779e9efa0243f108cfd244393701c31ebb5267d2cb362a8bc4

Observation 0648e3b5-858d-4cb1-a28f-fa27f9d9e31d · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.725561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:1b5211332062df4513f09a4a7431a3383a5be6c9be6e221b026f3225fd4e2e85

Observation dd7f6c52-3483-4957-b4a3-9fc7c3949ada · outbound

This paper cites Retrieval-augmented text-to-audio generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Retrieval-augmented text-to-audio generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.222572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:d1c929cedd5df52e04dca1b842321f13d5f8310d7f82e0529d65402ecb68c3c6

Observation b36eea5b-daff-4566-8a17-df09558ee8d2 · outbound

This paper cites Audioldm 2: Learning holistic audio generation with self-supervised pretraining.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Audioldm 2: Learning holistic audio generation with self-supervised pretraining

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.149148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:94b742113a26c7d6c6f82309d76ed496135978adabe7996120c1761c245537de

Observation 9f715069-664e-4936-8233-34602905603d · outbound

This paper cites DreamBooth: Fine tuning text-to-image diffusion models for subject- driven generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models DreamBooth: Fine tuning text-to-image diffusion models for subject- driven generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.064854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:5b7c7e4cff448295317446c06ff714b6868b145eab74355a72fe9eed796233e0

Observation 7cf666c1-444d-4e0e-b017-5323d65fc174 · outbound

This paper cites Zero-shot text-to-image generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Zero-shot text-to-image generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.213731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:bf59e23f2cbdd2d17edeb58b4da20014d480392a6c6022af54a94836aedeb715

Observation 5120f924-e567-4640-93cd-3b921fcf7ec3 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:31:44.720702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:d58cf5cdf6f924e354d2203358e051f2f3610067ca235845697deb370d78f43c

Observation f4a4da49-f86a-44ed-a980-95547428161b · outbound

This paper cites Improving image generation with better captions.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Improving image generation with better captions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.058062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:24489db42f9d883980636af6b70a0162cc3679753551d31c128f8f96c6ae728b

Observation 364242a6-1d18-4c1e-9b59-56f5b7d5c23d · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models High- resolution image synthesis with latent diffusion models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.139169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:d0cca80e17d4ad8ab3a31491985fc142f736eff59f6e1b2ec8ffad743a71085d

Observation 33e205bb-794a-4203-adae-f27bbdb5611d · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Scaling rectified flow transformers for high-resolution image synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.047641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:a3b91bb49684e10c621964635e8afce55e27a60d2ad7f833f6daba2272ae5501

Observation e97f410b-3b57-4ae2-99d8-bf9f93baad3b · outbound

This paper cites SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.716748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:dd0b8d23de21b17a4f5844cf7ab67329e5dad8f863762f9554afb1630b405d20

Observation 84c253a3-6161-48a6-9596-a06b6dd60bf5 · outbound

This paper cites Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.712246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:ecac0dd5a1998351fa6b48f8cac53d79f6a7c2ba6a43cffff2baa0c9b0a9dd87

Observation 9bfe36c9-d573-45d4-8b7f-1b632c67435f · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Adding conditional control to text-to-image diffusion models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.227898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:565721cea375eaf82e665ccc57684834c98967684807de40f5fce3061b434db4

Observation e0ff90d9-8a31-48ba-8bcf-767b4036f3d9 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models AudioCaps: Generating captions for audios in the wild

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.248530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:b83f189a44937167c96223b90f772128c301c139cb8f94057aaa2456d21c852c

Observation da4eef81-fa20-46f1-b060-f66cd40e8b9d · outbound

This paper cites Score-based generative modeling through stochastic differen- tial equations.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Score-based generative modeling through stochastic differen- tial equations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.239914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:417f8482b8567a48a5347319d03218000e2f1e2940e77201f1fa54942c681d3f

Observation 84db829c-4ee8-4434-8338-b9890dda6842 · outbound

This paper cites Diffusion models beat GANs on image synthesis.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Diffusion models beat GANs on image synthesis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.235929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:bc4aeb3bcc11d0006a007dc3799618ebf3452a9442cd37b577918f4331ab753d

Observation 2b328429-a9ad-4b12-8d30-6ef250b20ba0 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:31:44.708027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:8e8ee25837a6b900129464bbb37356da9f27ca3ec143a4bcfcee96c28d7997e5

Observation 191cfcc0-b6c5-46b6-886f-464dbf33bad8 · outbound

This paper cites Image super-resolution via iterative refinement.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Image super-resolution via iterative refinement

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.231941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:b28cfdc3e2ecbe0434b70e5b877b57a03f8076bb7c8010ffdd70e7fce5adfb20

Observation 14c63d36-f194-4579-9abb-1ddd971ec3c4 · outbound

This paper cites Wave- Grad: Estimating gradients for waveform generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Wave- Grad: Estimating gradients for waveform generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.223960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:044d92cf97adb35cec9ea0854b5a7b347bb0db52991e8a3b74da71425285a853

Observation 36ee6ca8-46eb-431f-8f47-4f546174e7e3 · outbound

This paper cites DiffWave: A versatile diffusion model for audio synthesis.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models DiffWave: A versatile diffusion model for audio synthesis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.220618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:91597e14bfb4c0f17a2d816479ea38fe6dbcdef2cadc3625f4ab3ad64ba3867c

Observation 09fdeafa-e13e-400f-a6fe-1852dfad22c6 · outbound

This paper cites Make-A-Video: Text-to-video generation without text-video data.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Make-A-Video: Text-to-video generation without text-video data

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.217378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:f4b7ff84fb11d9d635cc2fa8df5a26fc70aa6b7286e379777a90bf2c1a796b8b

Observation 2ca7df0f-7703-4f87-88ae-f0aaf1961e51 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Imagen Video: High Definition Video Generation with Diffusion Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:31:44.703532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:0f8dc4c43db6794b927d3a479d671c556f9f8f31194757df36526d1ba86a37ce

Observation 01ab0f5c-9026-496f-a7e5-1e2af8b3f46d · outbound

This paper cites Grad- TTS: A diffusion probabilistic model for text-to-speech.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Grad- TTS: A diffusion probabilistic model for text-to-speech

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.112359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:c8ca4edf2e8a045d167a95f19bb7ec0942a609a322adac367231e1ecb6f618e4

Observation 691bc4c5-fe5a-4381-94cf-5955b2957f4d · outbound

This paper cites ResGrad: Residual Denoising Diffusion Probabilistic Models for Text to Speech.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models ResGrad: Residual Denoising Diffusion Probabilistic Models for Text to Speech

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.699231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:b463221aa2ac256c09787c96b9fb41847546f48e2bdf390ca2adf2861bd3d963

Observation ddcd019c-fbd3-4b3f-be77-f134ddea8e19 · outbound

This paper cites Bilateral denoising diffusion models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Bilateral denoising diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.209647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:172c7754bb4b545ae4d7b7e8bf21dba5422566c7e9114c32701fcef7dcc3059a

Observation ff22e81a-af69-4650-a305-93bbca2d53b6 · outbound

This paper cites Priorgrad: Improving conditional denoising diffu- sion models with data-driven adaptive prior.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Priorgrad: Improving conditional denoising diffu- sion models with data-driven adaptive prior

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.206407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:6095e017df319ea6bda4a4b6dff6c17f057f10a6d0588845e1c46f93f4b1c1be

Observation 417d6878-168b-4a50-8ed0-6f717f805481 · outbound

This paper cites InferGrad: Improving diffusion models for vocoder by considering inference in training.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models InferGrad: Improving diffusion models for vocoder by considering inference in training

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.202371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:38aed17d42d807b35cb3d6e9eeb13d8bf007cfa4e926e37358f3b8295a3f38c1

Observation 543ca49d-1415-45cc-855a-9ad53852bbfb · outbound

This paper cites Acoustic scene generation with conditional SampleRNN.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Acoustic scene generation with conditional SampleRNN

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.102246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:bdd19e433f050fac7d31866d77135b5bd680c9f5ec3b43fcd54611a717fa5fdf

Observation 74c6d451-8002-4d7d-995e-6ab1e3746b46 · outbound

This paper cites Conditional sound generation using neural discrete time-frequency representation learning.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Conditional sound generation using neural discrete time-frequency representation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.195191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:b6b986a3db921158f3afa6736a78c1b235c0a214078929997a4dc7370384c280

Observation 537d1a66-23fe-4838-befe-b91f941ab9b8 · outbound

This paper cites Leveraging pre-trained AudioLDM for sound generation: A benchmark study.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Leveraging pre-trained AudioLDM for sound generation: A benchmark study

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.191439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:4255829dac91ddb7f1e2953db549c55aa8f928a07f10019351010bb8d05c6d38

Observation 857f599a-caa6-44ab-844a-a0b7ea530df5 · outbound

This paper cites HiFi-GAN: generative adversarial networks for efficient and high fidelity speech synthesis.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models HiFi-GAN: generative adversarial networks for efficient and high fidelity speech synthesis

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.187239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:4ede5b19f8517022986275c84fb213387c0f9838d2a882f2da48d4400ad8883e

Observation 78d9861a-4240-427d-9df8-433ea6bbe7da · outbound

This paper cites MelGAN: Generative adversarial networks for conditional waveform synthesis.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models MelGAN: Generative adversarial networks for conditional waveform synthesis

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.152572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:861483e256a1e9205a528d175d50df2398440b483c5fa681323f8a2464966eba

Observation 96a9e711-5b24-45b1-b03a-09daa8c6fbe0 · outbound

This paper cites BigVGAN: A universal neural vocoder with large-scale training.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models BigVGAN: A universal neural vocoder with large-scale training

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.183521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:406c08bab30d6b2c0351394b7855f1e2331ec076f3f8675690abaf5cd8f75041

Observation 7fcf4245-4b64-4aba-b077-49330d394572 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.179907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:f6e92ab7932fe2797673049eb0ec3875ac6f4b51133527126a9fe2b2079aac1f

Observation eb26b8d8-ce1c-4a41-a413-4f98fe9b0774 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.694499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:ece8913f858603421284afcaa75e981905aa1da80923539271620ecd3a16bc14

Observation 6f77b52d-fb04-4bf7-9477-8242d499816e · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.175751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:f9fc7209ba42f1c940d06bb7098df584932a2b620a6baa0804487c020ee777a8

Observation fbc9a665-1a62-406d-83f5-f360db4e653a · outbound

This paper cites Using pre-training can improve model robustness and uncertainty.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Using pre-training can improve model robustness and uncertainty

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.172322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:5cce5acb137ed004d9a5d0f687aa5670e13b7c5458d440d4ac26bcd445353f65

Observation 444931a6-241a-4c8d-bdd2-c112456962b4 · outbound

This paper cites Text-driven Foley sound generation with latent diffusion model.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Text-driven Foley sound generation with latent diffusion model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.168789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:4860aebc8c5981012c62a045b195373ca2e9652fbb6ea216b577063261bb4f3b

Observation ed599b40-1961-4146-9fb4-d896182b6c5f · outbound

This paper cites Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.165167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:2c166debe56e90dcda9682d74e2fb72da9c5a36ef40111396b90192c98ad5c3b

Observation 3f07dc39-3e6e-4fd0-982c-760fdc705d64 · outbound

This paper cites Anydoor: Zero-shot object-level image customization.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Anydoor: Zero-shot object-level image customization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.161780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:20e2f20de9a2dd81565f2700bae8d98f5be456463e197c21e4c0396c3ea922a9

Observation 14f12460-1806-4595-bed6-305bc8ae34ca · outbound

This paper cites Encoder-based domain tuning for fast personalization of text-to- image models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Encoder-based domain tuning for fast personalization of text-to- image models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.095618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:6d8e98330a1e8797cad4ed95d92df9f0aaa271650cfc51028dd8f6f44fb318c7

Observation e1a6eb64-33f7-4f90-94af-c29f4d27a22f · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Imagic: Text-based real image editing with diffusion models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.216582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:b988dced54dd0c3f6a599e090c43adf4c62d0da7dd0e7f5640716ee0fd91b9a0

Observation 2c003200-8fff-4523-9e5d-cb5e4ffe0f69 · outbound

This paper cites Key-locked rank one editing for text-to-image personalization.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Key-locked rank one editing for text-to-image personalization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.154507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:205d7b0fe8287215c516a8230160a6190a6674b2adeaa01b569943642e7ec859

Observation 2b10d3fb-e10f-48e3-adcc-de599506b002 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:31:44.689048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:a78230ad0d2042fd16715a645fed28a05fb9adc5a792d5dec7a11458bd2e49d9

Observation e7b12b0a-e3c5-4b2e-9019-f47054e01250 · outbound

This paper cites P+: Extended Textual Conditioning in Text-to-Image Generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models P+: Extended Textual Conditioning in Text-to-Image Generation

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.684811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:802e192f885420229f03910317085426c96fc3ac6716f5bbaa855f3e4d331bc1

Observation a2bd4efc-ed42-4966-8ea1-b39a580a0ea4 · outbound

This paper cites A neural space- time representation for text-to-image personalization.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models A neural space- time representation for text-to-image personalization

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.150864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:cede8c0624a0723f311c8024b2798395faafc94b5a30cbc4efcac329e2ae2d1c

Observation 77c16b7b-85b5-4c3c-a481-32fc4b828583 · outbound

This paper cites Improving expressivity of GNNs with subgraph- specific factor embedded normalization.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Improving expressivity of GNNs with subgraph- specific factor embedded normalization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.098795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:94f050417943efcf17478714cc6e73d0998f73bb5777492e97d0902fca35a458

Observation 9512552a-da7f-476a-a087-55485301f2cf · outbound

This paper cites BLIP-Diffusion: Pre-trained subject represen- tation for controllable text-to-image generation and editing.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models BLIP-Diffusion: Pre-trained subject represen- tation for controllable text-to-image generation and editing

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.175579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:5f3daf2916a8f8e35961cc22ea7138cbcad738570b1c5e1616f2294b82d804f3

Observation 1acfa72d-c962-432c-b7c6-241f855422f3 · outbound

This paper cites Cones: Concept Neurons in Diffusion Models for Customized Generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Cones: Concept Neurons in Diffusion Models for Customized Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.680433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:17887336eae140332fe8f89bcf862903fb224bdfbb1dc48d9be04dcb64431f70

Observation c23d6596-551d-4d67-98b4-26d5b4bbaafb · outbound

This paper cites HyperDreamBooth: Hypernetworks for fast personalization of text-to-image models.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models HyperDreamBooth: Hypernetworks for fast personalization of text-to-image models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.142967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:6b6f8cef372936bc1ea39656790364f401029a1a5ca9b4621ccad1aa0d7c1b76

Observation 42c3d6ff-c572-4bc6-8a63-59e6597a2bcc · outbound

This paper cites Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.139736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:eb68f5fedb99d1601b0c3feda51c8c239a4510228b5f15108ee37fb1892d616e

Observation 7609e581-e592-4a54-a4b9-7e49d379e565 · outbound

This paper cites FreeCustom: Tuning-free customized image generation for multi- concept composition.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models FreeCustom: Tuning-free customized image generation for multi- concept composition

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.102619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:91f6e8ede823f4222d7f007870baa5642a8ce43b4efdbfeeedbb757bbae3ad43

Observation 075329c6-7c43-4340-8b2c-a121b080ab41 · outbound

This paper cites KNN-Diffusion: Image generation via large-scale retrieval.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models KNN-Diffusion: Image generation via large-scale retrieval

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.135256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:c1e676dd8a30ef3378791b6a7056b39d4c6910228c3217540fd9ab6befc58a0d

Observation 459a3b5d-3726-49e3-8aae-289612e1e60d · outbound

This paper cites Re-Imagen: Retrieval- augmented text-to-image generator.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Re-Imagen: Retrieval- augmented text-to-image generator

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.131509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:09df8857cd825b418ed8ecdf12113e9f33357011eb22fd615179e58aa430edbd

Observation 0467721d-c8af-4480-8ed4-312651b5084a · outbound

This paper cites Learning transferable visual models from natural language supervision.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Learning transferable visual models from natural language supervision

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.128285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:7d6dba0facf14f1a1c83465bab04a8222498423bf4907d51bcadc8e377d0ecdd

Observation 0cfe089f-498b-465b-a01e-706c264240d9 · outbound

This paper cites T-CLAP: Temporal-enhanced contrastive language- audio pretraining.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models T-CLAP: Temporal-enhanced contrastive language- audio pretraining

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.124995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:1bb52eecddc1280bdab3705c8c6203630e985329b21bf3db07ec512dc411e0c5

Observation 49fb1234-c08f-4517-9e35-0a821f1f329f · outbound

This paper cites FlowSep: Language-Queried Sound Separation with Rectified Flow Matching.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models FlowSep: Language-Queried Sound Separation with Rectified Flow Matching

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.676201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:3b078955c96e9e18a8d13d494ccf8d87ef7fba50ef32757fcf2b77ef50822f0c

Observation 5e1be8c0-385a-42bf-8d7f-0378490c6428 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.672162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:6e3607891eb72e6bba48590d2fce11d283d7f577a6dbf64a1be9c498a70f8d37

Observation 5fb89305-277b-4cd3-bd86-c922a56c7183 · outbound

This paper cites Freesound technical demo.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Freesound technical demo

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.219526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:8a3e9bf6bdef4e96984893fdff22d50b2c26220f45e5c3eca618a4a8e1a4f90b

Observation 84c444cc-11ee-4fd7-bbce-5413f5578203 · outbound

This paper cites Deep convolutional neural networks and data augmentation for environmental sound classification.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Deep convolutional neural networks and data augmentation for environmental sound classification

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.118109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:958cc0308e83edd0960b14743746f99a3ce8604b952c4adb02a60fd21d335ccd

Observation d3f86821-871c-41b9-b778-5ede1e142f95 · outbound

This paper cites ESC: dataset for environmental sound classification.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models ESC: dataset for environmental sound classification

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.113602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:c46ad692308ec1f0f04bfb4152608acb1b10399822feb765a42cec9781c84739

Observation ce1c6aea-1512-4e81-a7d3-9dee95468717 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:31:44.667863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:35322c93ef5e63092154b9b7bde539e8de38eb7d6f501c00b1d63da34737e5ea

Observation 76c90938-b373-480b-a22d-50c109742089 · outbound

This paper cites PANNs: Large-scale pretrained audio neural networks for audio pattern recognition.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models PANNs: Large-scale pretrained audio neural networks for audio pattern recognition

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.106034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:168e5b0d756fdfed65f2e3ffb630df7df87bf739b9dd9ab83023bacc17bc0e26

Observation 16758743-0c6b-4c28-9a98-1d1d0c64bac3 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:31:44.663554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:fc3421eb39614150fc55390e594a47bceaa9f7ef31b0cdff06b080b94d2fc222

Observation f7f62ca0-6b6e-430a-a541-7281f3a6cea0 · outbound

This paper cites Decoupled weight decay regularization.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Decoupled weight decay regularization

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:32:48.110090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:2d1ad77ce53d420c77037c9713e38e202041b5867b50411de68fb1641e8d9c16

Pith citing papers

No inbound Pith citation observations are available.