Pith. sign in

REVIEW 12 cited by

Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.01044 v1 pith:UNL67X6J submitted 2024-01-02 cs.SD cs.AIcs.CLeess.AS

classification cs.SDcs.AIcs.CLeess.AS
keywords auffusionalignmentmodelsdiffusionlanguagestudiesaigcaudio
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advancements in diffusion models and large language models (LLMs) have significantly propelled the field of AIGC. Text-to-Audio (TTA), a burgeoning AIGC application designed to generate audio from natural language prompts, is attracting increasing attention. However, existing TTA studies often struggle with generation quality and text-audio alignment, especially for complex textual inputs. Drawing inspiration from state-of-the-art Text-to-Image (T2I) diffusion models, we introduce Auffusion, a TTA system adapting T2I model frameworks to TTA task, by effectively leveraging their inherent generative strengths and precise cross-modal alignment. Our objective and subjective evaluations demonstrate that Auffusion surpasses previous TTA approaches using limited data and computational resource. Furthermore, previous studies in T2I recognizes the significant impact of encoder choice on cross-modal alignment, like fine-grained details and object bindings, while similar evaluation is lacking in prior TTA works. Through comprehensive ablation studies and innovative cross-attention map visualizations, we provide insightful assessments of text-audio alignment in TTA. Our findings reveal Auffusion's superior capability in generating audios that accurately match textual descriptions, which further demonstrated in several related tasks, such as audio style transfer, inpainting and other manipulations. Our implementation and demos are available at https://auffusion.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  2. EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Prompt-to-Prompt cross-attention control is adapted to autoregressive audio generation, enabling training-free music editing that outperforms a diffusion baseline.

  3. Sounding that Object: Interactive Object-Aware Image to Audio Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A latent diffusion audio model is trained to ground sound in image patches, then uses SAM segmentation masks at test time so users can generate audio for selected objects in a scene.

  4. TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A fast flow-matching text-to-audio model aligned via CLAP-ranked self-generated preference pairs reports state-of-the-art AudioCaps and human-evaluation scores.

  5. ETTA: Elucidating the Design Space of Text-to-Audio Models

    cs.SD 2024-12 conditional novelty 6.0 of 10

    ETTA, a text-to-audio model trained on a large synthetic caption dataset, outperforms public-data baselines on AudioCaps and MusicCaps and rivals proprietary-data systems.

  6. Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance

    cs.SD 2024-12 conditional novelty 6.0 of 10

    Smooth-Foley uses frame-level visual features and label-guided temporal conditions to generate continuous, synchronized audio for videos with moving or ambiguous sound sources.

  7. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single flow-matching transformer trained jointly on audio-video and audio-text data produces state-of-the-art public video-to-audio synthesis with a frame-level synchronization module.

  8. Video-Guided Foley Sound Generation with Multimodal Controls

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A video-guided diffusion model generates synchronized foley sound from text, audio, and video controls, using joint training on noisy internet videos and professional sound-effect libraries to reach 48kHz output.

  9. Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning

    eess.AS 2025-01 conditional novelty 4.0 of 10

    Combining global and local text embeddings in a diffusion UNet for music improves text adherence, while mean-pooling T5 local embeddings yields the best audio quality without extra parameters.

  10. Sound Scene Synthesis at the DCASE 2024 Challenge

    cs.AI 2025-01 conditional novelty 4.0 of 10

    Four text-to-audio systems were evaluated against a human reference in the DCASE 2024 Task 7 challenge, with a 36% quality gap and strong but small-sample FAD-to-human correlation.

  11. YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

    cs.SD 2024-12 reject novelty 4.0 of 10

    A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.

  12. Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation

    cs.SD 2024-12 reject novelty 4.0 of 10

    Mel-Refine boosts text-to-audio spectrogram sharpness by amplifying skip-connection high frequencies and attenuating backbone high frequencies at inference, but its reported 25% improvement is based on test-set-tuned ...

Pith tools