Pith. sign in

REVIEW 11 cited by

Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.13731 v2 pith:UFAP3XQO submitted 2023-04-24 eess.AS cs.AIcs.CLcs.SD

Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

classification eess.AS cs.AIcs.CLcs.SD
keywords encodermodelaudiodiffusiongenerationinstruction-tunedlanguagelatent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The immense scale of the recent large language models (LLM) allows many interesting properties, such as, instruction- and chain-of-thought-based fine-tuning, that has significantly improved zero- and few-shot performance in many natural language processing (NLP) tasks. Inspired by such successes, we adopt such an instruction-tuned LLM Flan-T5 as the text encoder for text-to-audio (TTA) generation -- a task where the goal is to generate an audio from its textual description. The prior works on TTA either pre-trained a joint text-audio encoder or used a non-instruction-tuned model, such as, T5. Consequently, our latent diffusion model (LDM)-based approach TANGO outperforms the state-of-the-art AudioLDM on most metrics and stays comparable on the rest on AudioCaps test set, despite training the LDM on a 63 times smaller dataset and keeping the text encoder frozen. This improvement might also be attributed to the adoption of audio pressure level-based sound mixing for training set augmentation, whereas the prior methods take a random mix.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

    cs.SD 2026-05 unverdicted novelty 7.0

    PlanAudio introduces a unified autoregressive LLM framework with semantic latent chain-of-thought for generating composite speech and sound audio from free-form text, plus a new benchmark.

  2. Omni2Sound: Towards Unified Video-Text-to-Audio Generation

    cs.SD 2026-01 unverdicted novelty 7.0

    A single DiT-based diffusion model unifies video-to-audio, text-to-audio, and joint video-text-to-audio generation, supported by a new 470k-pair dataset and three-stage progressive training that resolves task competition.

  3. AudioMoG: Guiding Audio Generation with Mixture-of-Guidance

    cs.SD 2025-09 unverdicted novelty 7.0

    AudioMoG is a mixture-of-guidance sampling technique that combines CFG and AG signals to outperform single-guidance baselines in text-to-audio generation at equivalent speed.

  4. Generative Semantic Communication: Diffusion Models Beyond Bit Recovery

    cs.AI 2023-06 unverdicted novelty 7.0

    A generative semantic communication system that sends compressed semantic information and uses diffusion models with spatially-adaptive normalizations to reconstruct high-quality, semantically consistent images even u...

  5. Auditing Training Data in Generative Music Models via Black-Box Membership Inference

    cs.LG 2026-05 unverdicted novelty 6.0

    Black-box membership inference on text-to-music models reaches up to 98.6% accuracy by training an auditor on semantic alignment patterns extracted from shadow-model generations.

  6. FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

    cs.SD 2026-03 unverdicted novelty 6.0

    FoleyDirector introduces structured temporal scripts and a fusion module to enable precise timing control in DiT-based video-to-audio generation while preserving audio fidelity.

  7. DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

    cs.SD 2025-09 unverdicted novelty 6.0

    DreamAudio generates audio clips that incorporate user-specified personalized audio events from reference samples while remaining aligned with text prompts.

  8. AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

    cs.SD 2026-06 unverdicted novelty 5.0

    A distilled multimodal diffusion model generates audio from text, video, or audio in four steps with claimed superior quality and ~25× fewer function evaluations.

  9. AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

    cs.SD 2026-06 unverdicted novelty 5.0

    AudioX-Turbo distills a Multimodal Diffusion Transformer into a 4-step student model for efficient multimodal anything-to-audio generation, trained on a new 9.2M-sample dataset IF-caps-Pro.

  10. Movie Gen: A Cast of Media Foundation Models

    cs.CV 2024-10 unverdicted novelty 5.0

    A 30B-parameter transformer and related models generate high-quality videos and audio, claiming state-of-the-art results on text-to-video, video editing, personalization, and audio generation tasks.

  11. Training-Free Multi-User Generative Semantic Communications via Null-Space Diffusion Sampling

    eess.SP 2024-05 unverdicted novelty 5.0

    Introduces a null-space diffusion sampling method for training-free multi-user generative semantic communications in OFDMA systems.