Pith. sign in

REVIEW 23 cited by

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.21037 v2 pith:VEX4F6XO submitted 2024-12-30 cs.SD cs.AIcs.CLeess.AS

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

classification cs.SD cs.AIcs.CLeess.AS
keywords preferenceaudiomodelstangofluxclap-rankedcrpoframeworkgeneration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce TangoFlux, an efficient Text-to-Audio (TTA) generative model with 515M parameters, capable of generating up to 30 seconds of 44.1kHz audio in just 3.7 seconds on a single A40 GPU. A key challenge in aligning TTA models lies in the difficulty of creating preference pairs, as TTA lacks structured mechanisms like verifiable rewards or gold-standard answers available for Large Language Models (LLMs). To address this, we propose CLAP-Ranked Preference Optimization (CRPO), a novel framework that iteratively generates and optimizes preference data to enhance TTA alignment. We demonstrate that the audio preference dataset generated using CRPO outperforms existing alternatives. With this framework, TangoFlux achieves state-of-the-art performance across both objective and subjective benchmarks. We open source all code and models to support further research in TTA generation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation

    eess.AS 2026-06 unverdicted novelty 7.0

    AudioCALM presents a continuous autoregressive framework with flow-matching prediction and A-MoME architecture that unifies speech, sound, and music generation while matching modality-specific state-of-the-art performance.

  2. AudioMoG: Guiding Audio Generation with Mixture-of-Guidance

    cs.SD 2025-09 unverdicted novelty 7.0

    AudioMoG is a mixture-of-guidance sampling technique that combines CFG and AG signals to outperform single-guidance baselines in text-to-audio generation at equivalent speed.

  3. Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping

    cs.CV 2025-05 unverdicted novelty 7.0

    A contrastive multimodal framework augments satellite-audio datasets with vision-language model sound descriptions to learn shared soundscape concepts for zero-shot retrieval and synthesis.

  4. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0

    Audex unifies audio understanding and generation on a strong text MoE backbone with multi-stage SFT plus text-only Cascade RL, matching open SOTA audio scores while mostly retaining text capability.

  5. SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

    cs.SD 2026-07 conditional novelty 6.0

    SynSFX provides a multi-generator sound-effect deepfake corpus showing speech detectors fail, joint training mitigates forgetting, but generalization to unseen generators remains poor due to artifact overfitting.

  6. RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

    cs.SD 2026-06 unverdicted novelty 6.0

    Hybrid two-stage diffusion transformer architecture for instruction-guided audio editing via rectified flow that performs joint attention at low resolution then alternates joint and cross-attention at high resolution ...

  7. Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

    cs.SD 2026-05 unverdicted novelty 6.0

    Dasheng AudioGen uses multi-view captions and a unified semantic-acoustic representation to enable end-to-end generation of mixed audio scenes from text descriptions.

  8. AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling

    cs.SD 2026-05 unverdicted novelty 6.0

    AuDirector is a self-reflective closed-loop multi-agent framework that generates immersive audio narratives with improved structural coherence, emotional expressiveness, and acoustic fidelity via identity-aware voice ...

  9. ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

    cs.MM 2026-04 unverdicted novelty 6.0

    ControlFoley introduces a unified framework for controllable video-to-audio generation using joint visual encoding, temporal-timbre decoupling, and robust multimodal training to handle cross-modal conflicts.

  10. RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing

    cs.SD 2025-09 unverdicted novelty 6.0

    RFM-Editing applies rectified flow matching to text-guided audio editing without masks or full captions and introduces a new overlapping multi-event audio dataset for training and evaluation.

  11. FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation

    eess.AS 2026-07 conditional novelty 5.5

    MeanFlow-anchored multi-representation FD post-training improves one-step text-to-audio quality without collapsing multi-step sampling.

  12. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 5.0

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  13. OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models

    cs.LG 2026-06 unverdicted novelty 5.0

    OTCache uses optimal transport to interpolate caching schedules between a graph-based reference and an Optuna-optimized anchor, delivering 3.66x-4.7x speedups on FLUX.1, Qwen-Image and HunyuanVideo with improved fidelity.

  14. RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

    cs.SD 2026-06 unverdicted novelty 5.0

    Hybrid two-stage diffusion transformer architecture for instruction-guided audio editing uses coarse joint attention at low resolution and refined alternating blocks at high resolution to improve performance and effic...

  15. RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

    cs.SD 2026-06 conditional novelty 5.0

    A coarse-to-fine hybrid MMDiT/DiT audio editor trained with rectified flow matching improves fidelity and cuts edit time versus prior instruction-guided baselines on synthetic overlapping-event tasks.

  16. AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

    cs.SD 2026-06 unverdicted novelty 5.0

    AudioX-Turbo distills a Multimodal Diffusion Transformer into a 4-step student model for efficient multimodal anything-to-audio generation, trained on a new 9.2M-sample dataset IF-caps-Pro.

  17. AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

    cs.SD 2026-06 unverdicted novelty 5.0

    A distilled multimodal diffusion model generates audio from text, video, or audio in four steps with claimed superior quality and ~25× fewer function evaluations.

  18. UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

    eess.AS 2026-05 unverdicted novelty 5.0

    UNISON introduces a unified latent diffusion framework with layer-wise LLM fusion and channel-mask task encoding for multiple speech and sound generation and editing tasks.

  19. ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

    eess.AS 2026-05 unverdicted novelty 5.0

    ImmersiveTTS proposes an environment-aware TTS system that integrates speech with environmental audio via multimodal diffusion transformer, joint attention, and domain-specific representation alignment, claiming super...

  20. AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling

    cs.SD 2026-05 unverdicted novelty 5.0

    AuDirector proposes a self-reflective closed-loop multi-agent framework with identity-aware pre-production, collaborative synthesis-correction, and human-guided refinement for coherent immersive audio storytelling.

  21. Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation

    cs.SD 2026-05 unverdicted novelty 5.0

    A one-step text-to-audio model using energy-distance training and contextual distillation outperforms prior fast baselines on AudioCaps and achieves up to 8.5x faster inference than the multi-step IMPACT system with c...

  22. Woosh: A Sound Effects Foundation Model

    cs.SD 2026-04 accept novelty 5.0

    Woosh is a new publicly released foundation model optimized for high-quality sound effect generation from text or video, showing competitive or better results than open alternatives like Stable Audio Open.

  23. STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

    eess.AS 2026-06 unverdicted novelty 4.0

    STAR-VAE introduces topology-aware regularization to reshape VAE latent geometry for audio, claiming to resolve the Rate-Distortion-Regularity Trilemma and achieve SOTA reconstruction.