Pith. sign in

REVIEW 8 cited by

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.12661 v1 pith:R7VJAQD7 submitted 2023-01-30 cs.SD cs.LGcs.MMeess.AS

classification cs.SDcs.LGcs.MMeess.AS
keywords audiomake-an-audioaudiosbehinddatadiffusiongenerationlarge-scale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio pairs, and the complexity of modeling long continuous audio data. In this work, we propose Make-An-Audio with a prompt-enhanced diffusion model that addresses these gaps by 1) introducing pseudo prompt enhancement with a distill-then-reprogram approach, it alleviates data scarcity with orders of magnitude concept compositions by using language-free audios; 2) leveraging spectrogram autoencoder to predict the self-supervised audio representation instead of waveforms. Together with robust contrastive language-audio pretraining (CLAP) representations, Make-An-Audio achieves state-of-the-art results in both objective and subjective benchmark evaluation. Moreover, we present its controllability and generalization for X-to-Audio with "No Modality Left Behind", for the first time unlocking the ability to generate high-definition, high-fidelity audios given a user-defined modality input. Audio samples are available at https://Text-to-Audio.github.io

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Activation Concentration: Characterizing Column-Level Output Sparsity Across Diffusion Model Architectures

    cs.AR 2026-05 unverdicted novelty 7.0 of 10

    First systematic column-level sparsity profiling across seven diffusion workloads reveals element-level sparsity overstates hardware savings by up to 78 points and identifies a three-way taxonomy of concentration vs. ...

  2. Generative Semantic Communication: Diffusion Models Beyond Bit Recovery

    cs.AI 2023-06 unverdicted novelty 7.0 of 10

    A generative semantic communication system that sends compressed semantic information and uses diffusion models with spatially-adaptive normalizations to reconstruct high-quality, semantically consistent images even u...

  3. FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

    cs.SD 2026-03 unverdicted novelty 6.0 of 10

    FoleyDirector introduces structured temporal scripts and a fusion module to enable precise timing control in DiT-based video-to-audio generation while preserving audio fidelity.

  4. Diffusion Models Memorize in Training -- and Generalize in Inference

    cs.LG 2026-03 unverdicted novelty 6.0 of 10

    Diffusion models overfit denoising loss at intermediate noise but generalize in inference as model error smooths the flow field and sampling paths avoid memorized noisy training data.

  5. SemanticAudio: Audio Generation and Editing in Semantic Space

    eess.AS 2026-01 conditional novelty 6.0 of 10

    SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...

  6. TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models

    cs.SD 2025-09 conditional novelty 6.0 of 10

    TTA-Bench offers a seven-dimension, 2,999-prompt evaluation of ten text-to-audio models with 118,000 human ratings, covering quality, robustness, fairness, bias, and toxicity.

  7. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

    cs.CV 2023-07 conditional novelty 6.0 of 10

    SDXL improves upon prior Stable Diffusion versions through a larger UNet backbone, dual text encoders, novel conditioning, and a refinement model, producing higher-fidelity images competitive with black-box state-of-t...

  8. SiPhy: Single-Image Physical Property Reasoning

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A single-image vision-language pipeline reports state-of-the-art mass, density, and stiffness predictions by combining CLIP features, a fine-tuned VLM, and depth-adaptive pseudo-voxel sampling.

Pith tools