Pith. sign in

REVIEW 39 cited by

Stable Audio Open

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14358 v2 pith:A7REK5XU submitted 2024-07-19 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords modelsmodelopentext-to-audioaccessibleacrossallowingarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 39 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books

    cs.HC 2026-08 conditional novelty 7.0 of 10

    Dramarrator introduces object-based audio editing, where characters and scenes are editable objects that propagate edits across all linked speech, sound effects, music, and ambience.

  2. Scaling Transformers for Low-Bitrate High-Quality Speech Coding

    eess.AS 2024-11 conditional novelty 7.0 of 10

    A scaled transformer codec with FSQ reaches state-of-the-art speech reconstruction at 400-700 bps, outperforming CNN/RVQ baselines.

  3. Doppelganger: Sound Effects and Their Synthetic Twins

    cs.SD 2026-07 accept novelty 6.5 of 10

    Instance-pair training matches synthetic sound-effect twins to their real sources on unseen events (~80% R@1), while class supervision degrades below the frozen baseline and the mapping stays generator-specific.

  4. Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Using matched neutral and target-swap prompts, ACE-Step 1.5 and Stable Audio 3 show real key control and partial beat control, while LeVo2 does not, and much four-beat agreement is just the models' default output.

  5. A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies

    eess.AS 2026-08 conditional novelty 6.0 of 10

    The deciding factors for modeling an audio latent are dependency horizon and conditional ambiguity, viewed jointly with the representation, not whether the latent is discrete or continuous.

  6. On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A single mean shift in the latent space of several neural codecs achieves competitive music bandwidth extension on some metrics, implying a largely linear structure.

  7. FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving

    cs.DC 2026-07 conditional novelty 6.0 of 10

    FlashDiff reduces diffusion serving latency by 30–97% and raises throughput 1.2–2.2× by adaptively skipping refinement of latent regions that no longer need it.

  8. Qwen-Music Technical Report

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Qwen-Music generates high-fidelity vocal songs via 25 Hz semantic tokens, Melody-CoT planning, and DiT rendering, claiming SOTA on 13/16 metrics and expert preference over proprietary systems.

  9. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  10. An Empirical Analysis of Task-Induced Encoder Bias in Fr\'echet Audio Distance

    eess.AS 2026-02 conditional novelty 6.0 of 10

    No single tested audio encoder catches all quality issues: reconstruction-trained encoders detect signal degradation, speech-trained encoders detect temporal order, and classification-trained encoders detect semantic ...

  11. SemanticAudio: Audio Generation and Editing in Semantic Space

    eess.AS 2026-01 conditional novelty 6.0 of 10

    SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...

  12. Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Amadeus generates symbolic music by autoregressively predicting note-level latents and decoding their attributes in parallel with a masked discrete diffusion model, yielding faster and more controllable generation tha...

  13. JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

    cs.SD 2025-07 conditional novelty 6.0 of 10

    JAM is a 530M-parameter flow-matching song generator that adds word- and phoneme-level timing control and duration control, achieving strong lyric fidelity and musicality scores when ground-truth timings are provided.

  14. WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling

    cs.SD 2025-07 conditional novelty 6.0 of 10

    WildFX generates multi-track audio datasets by rendering real DAW effect graphs with commercial plugins inside Docker, and demonstrates the pipeline on blind mixing-graph estimation.

  15. MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI

    cs.SD 2025-07 conditional novelty 6.0 of 10

    MusGO is a community-refined framework with 13 openness categories, applied to 16 music-generative models to produce a public openness leaderboard.

  16. AI-Generated Song Detection via Lyrics Transcripts

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Transcribing audio with Whisper and classifying the transcript with LLM2Vec detects AI-generated songs from audio alone, nearly matching clean-lyrics accuracy and beating audio-based detectors under perturbations and ...

  17. Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

    cs.SD 2025-06 conditional novelty 6.0 of 10

    OSSL is the first self-hosted, mood-annotated video-music dataset, and a video adapter on MusicGen-Medium improves film music generation over text-only baselines.

  18. In-the-wild Audio Spatialization with Flexible Text-guided Localization

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.

  19. EgoZero: Robot Learning from Smart Glasses

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Robot policies trained only on egocentric human videos from smart glasses transfer zero-shot to a Franka gripper, with 70% success across 7 manipulation tasks.

  20. Fast Text-to-Audio Generation with Adversarial Post-Training

    cs.SD 2025-05 conditional novelty 6.0 of 10

    ARC post-training speeds up text-to-audio generation to near-real-time speeds on GPUs and a few seconds on phones, without distillation or classifier-free guidance.

  21. Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Audio-SDS uses a pretrained text-to-audio diffusion model as a frozen critic to optimize parameters of FM synthesizers, impact simulators, and source separation latents, matching text prompts without task-specific training.

  22. SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation

    cs.SD 2025-05 conditional novelty 6.0 of 10

    SonicRAG uses an LLM to convert text, voice, or onomatopoeia into a script that retrieves and mixes existing audio assets into a new, high-fidelity sound effect.

  23. RenderBox: Expressive Performance Rendering with Text Control

    eess.AS 2025-02 conditional novelty 6.0 of 10

    RenderBox is a text-and-score conditioned diffusion model that renders expressive, controllable audio performances across piano, guitar, saxophone, violin, and orchestral instruments.

  24. TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A fast flow-matching text-to-audio model aligned via CLAP-ranked self-generated preference pairs reports state-of-the-art AudioCaps and human-evaluation scores.

  25. Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion Transformer generates up to 60 seconds of 44.1 kHz stereo audio from video, text, and audio prompts, with a learned loudness envelope for fine-grained control.

  26. ETTA: Elucidating the Design Space of Text-to-Audio Models

    cs.SD 2024-12 conditional novelty 6.0 of 10

    ETTA, a text-to-audio model trained on a large synthetic caption dataset, outperforms public-data baselines on AudioCaps and MusicCaps and rivals proprietary-data systems.

  27. FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

    cs.SD 2024-12 conditional novelty 6.0 of 10

    FolAI predicts an editable RMS envelope from silent video and uses it, with semantic embeddings, to condition a Stable Audio diffusion model for 44.1 kHz stereo foley generation.

  28. Learned Compression for Compressed Learning

    eess.IV 2024-12 conditional novelty 6.0 of 10

    WaLLoC couples an invertible wavelet packet transform with a linear-projection encoder and nonlinear decoder to produce low-dimensional, quantization-resilient codes for compressed-domain learning.

  29. SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers

    cs.LG 2024-11 conditional novelty 6.0 of 10

    SmoothCache uses calibration-measured layer error thresholds to skip redundant attention and feed-forward computations in Diffusion Transformers, achieving 8-71% speedup across image, video, and audio tasks.

  30. KVAE: Family of Tokenizers for Multimodal Generative Models

    cs.CV 2026-08 conditional novelty 5.0 of 10

    KVAE introduces image, video, and full-band audio tokenizers whose reconstruction and downstream generation quality is competitive with, and often better than, current open-source tokenizers in head-to-head tests.

  31. Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Boomerang sampling, applied to a pretrained music diffusion model, creates audio variations that improve beat tracking when training data is scarce and can change instruments via text prompts.

  32. SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A cascaded pipeline of audio compression, latent diffusion extraction, and generative correction achieves state-of-the-art target speech extraction quality and intelligibility on Libri2Mix and out-of-domain data.

  33. SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation

    cs.SD 2024-12 conditional novelty 5.0 of 10

    A caption-augmentation method that adds DSP-derived acoustic descriptors to text prompts gives a text-to-audio diffusion model controllable loudness, pitch, reverb, noise, brightness, fade, and duration.

  34. Video-Guided Foley Sound Generation with Multimodal Controls

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A video-guided diffusion model generates synchronized foley sound from text, audio, and video controls, using joint training on noisy internet videos and professional sound-effect libraries to reach 48kHz output.

  35. From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

    eess.AS 2025-04 conditional novelty 4.0 of 10

    Across five text-to-music models, aesthetic predictor scores, pairwise human preferences, and reference-based distribution metrics produce inconsistent rankings, so the choice of evaluation metric changes the winner.

  36. Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning

    eess.AS 2025-01 conditional novelty 4.0 of 10

    Combining global and local text embeddings in a diffusion UNet for music improves text adherence, while mean-pooling T5 local embeddings yields the best audio quality without extra parameters.

  37. Sound Scene Synthesis at the DCASE 2024 Challenge

    cs.AI 2025-01 conditional novelty 4.0 of 10

    Four text-to-audio systems were evaluated against a human reference in the DCASE 2024 Task 7 challenge, with a 36% quality gap and strong but small-sample FAD-to-human correlation.

  38. Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

    cs.AI 2024-12 conditional novelty 4.0 of 10

    The paper proposes learning from language feedback to synthesize multimodal preference pairs, but the evidence is weakened by an undefined improvement metric and small, unvalidated effect sizes.

  39. Video Diffusion Transformers are In-Context Learners

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.

Pith tools