REVIEW 8 cited by
Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio pairs, and the complexity of modeling long continuous audio data. In this work, we propose Make-An-Audio with a prompt-enhanced diffusion model that addresses these gaps by 1) introducing pseudo prompt enhancement with a distill-then-reprogram approach, it alleviates data scarcity with orders of magnitude concept compositions by using language-free audios; 2) leveraging spectrogram autoencoder to predict the self-supervised audio representation instead of waveforms. Together with robust contrastive language-audio pretraining (CLAP) representations, Make-An-Audio achieves state-of-the-art results in both objective and subjective benchmark evaluation. Moreover, we present its controllability and generalization for X-to-Audio with "No Modality Left Behind", for the first time unlocking the ability to generate high-definition, high-fidelity audios given a user-defined modality input. Audio samples are available at https://Text-to-Audio.github.io
Forward citations
Cited by 8 Pith papers
-
Activation Concentration: Characterizing Column-Level Output Sparsity Across Diffusion Model Architectures
First systematic column-level sparsity profiling across seven diffusion workloads reveals element-level sparsity overstates hardware savings by up to 78 points and identifies a three-way taxonomy of concentration vs. ...
-
Generative Semantic Communication: Diffusion Models Beyond Bit Recovery
A generative semantic communication system that sends compressed semantic information and uses diffusion models with spatially-adaptive normalizations to reconstruct high-quality, semantically consistent images even u...
-
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
FoleyDirector introduces structured temporal scripts and a fusion module to enable precise timing control in DiT-based video-to-audio generation while preserving audio fidelity.
-
Diffusion Models Memorize in Training -- and Generalize in Inference
Diffusion models overfit denoising loss at intermediate noise but generalize in inference as model error smooths the flow field and sampling paths avoid memorized noisy training data.
-
SemanticAudio: Audio Generation and Editing in Semantic Space
SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...
-
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
TTA-Bench offers a seven-dimension, 2,999-prompt evaluation of ten text-to-audio models with 118,000 human ratings, covering quality, robustness, fairness, bias, and toxicity.
-
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
SDXL improves upon prior Stable Diffusion versions through a larger UNet backbone, dual text encoders, novel conditioning, and a refinement model, producing higher-fidelity images competitive with black-box state-of-t...
-
SiPhy: Single-Image Physical Property Reasoning
A single-image vision-language pipeline reports state-of-the-art mass, density, and stiffness predictions by combining CLIP features, a fine-tuned VLM, and depth-adaptive pseudo-voxel sampling.
Discussion (0). Sign in to comment.