REVIEW 39 cited by
Stable Audio Open
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.
Forward citations
Cited by 39 Pith papers
-
Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books
Dramarrator introduces object-based audio editing, where characters and scenes are editable objects that propagate edits across all linked speech, sound effects, music, and ambience.
-
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
A scaled transformer codec with FSQ reaches state-of-the-art speech reconstruction at 400-700 bps, outperforming CNN/RVQ baselines.
-
Doppelganger: Sound Effects and Their Synthetic Twins
Instance-pair training matches synthetic sound-effect twins to their real sources on unseen events (~80% R@1), while class supervision degrades below the frozen baseline and the mapping stays generator-specific.
-
Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping
Using matched neutral and target-swap prompts, ACE-Step 1.5 and Stable Audio 3 show real key control and partial beat control, while LeVo2 does not, and much four-beat agreement is just the models' default output.
-
A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies
The deciding factors for modeling an audio latent are dependency horizon and conditional ambiguity, viewed jointly with the representation, not whether the latent is discrete or continuous.
-
On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs
A single mean shift in the latent space of several neural codecs achieves competitive music bandwidth extension on some metrics, implying a largely linear structure.
-
FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving
FlashDiff reduces diffusion serving latency by 30–97% and raises throughput 1.2–2.2× by adaptively skipping refinement of latent regions that no longer need it.
-
Qwen-Music Technical Report
Qwen-Music generates high-fidelity vocal songs via 25 Hz semantic tokens, Melody-CoT planning, and DiT rendering, claiming SOTA on 13/16 metrics and expert preference over proprietary systems.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
An Empirical Analysis of Task-Induced Encoder Bias in Fr\'echet Audio Distance
No single tested audio encoder catches all quality issues: reconstruction-trained encoders detect signal degradation, speech-trained encoders detect temporal order, and classification-trained encoders detect semantic ...
-
SemanticAudio: Audio Generation and Editing in Semantic Space
SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...
-
Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music
Amadeus generates symbolic music by autoregressively predicting note-level latents and decoding their attributes in parallel with a masked discrete diffusion model, yielding faster and more controllable generation tha...
-
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
JAM is a 530M-parameter flow-matching song generator that adds word- and phoneme-level timing control and duration control, achieving strong lyric fidelity and musicality scores when ground-truth timings are provided.
-
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
WildFX generates multi-track audio datasets by rendering real DAW effect graphs with commercial plugins inside Docker, and demonstrates the pipeline on blind mixing-graph estimation.
-
MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI
MusGO is a community-refined framework with 13 openness categories, applied to 16 music-generative models to produce a public openness leaderboard.
-
AI-Generated Song Detection via Lyrics Transcripts
Transcribing audio with Whisper and classifying the transcript with LLM2Vec detects AI-generated songs from audio alone, nearly matching clean-lyrics accuracy and beating audio-based detectors under perturbations and ...
-
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
OSSL is the first self-hosted, mood-annotated video-music dataset, and a video adapter on MusicGen-Medium improves film music generation over text-only baselines.
-
In-the-wild Audio Spatialization with Flexible Text-guided Localization
A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.
-
EgoZero: Robot Learning from Smart Glasses
Robot policies trained only on egocentric human videos from smart glasses transfer zero-shot to a Franka gripper, with 70% success across 7 manipulation tasks.
-
Fast Text-to-Audio Generation with Adversarial Post-Training
ARC post-training speeds up text-to-audio generation to near-real-time speeds on GPUs and a few seconds on phones, without distillation or classifier-free guidance.
-
Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond
Audio-SDS uses a pretrained text-to-audio diffusion model as a frozen critic to optimize parameters of FM synthesizers, impact simulators, and source separation latents, matching text prompts without task-specific training.
-
SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation
SonicRAG uses an LLM to convert text, voice, or onomatopoeia into a script that retrieves and mixes existing audio assets into a new, high-fidelity sound effect.
-
RenderBox: Expressive Performance Rendering with Text Control
RenderBox is a text-and-score conditioned diffusion model that renders expressive, controllable audio performances across piano, guitar, saxophone, violin, and orchestral instruments.
-
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
A fast flow-matching text-to-audio model aligned via CLAP-ranked self-generated preference pairs reports state-of-the-art AudioCaps and human-evaluation scores.
-
Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control
A diffusion Transformer generates up to 60 seconds of 44.1 kHz stereo audio from video, text, and audio prompts, with a learned loudness envelope for fine-grained control.
-
ETTA: Elucidating the Design Space of Text-to-Audio Models
ETTA, a text-to-audio model trained on a large synthetic caption dataset, outperforms public-data baselines on AudioCaps and MusicCaps and rivals proprietary-data systems.
-
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
FolAI predicts an editable RMS envelope from silent video and uses it, with semantic embeddings, to condition a Stable Audio diffusion model for 44.1 kHz stereo foley generation.
-
Learned Compression for Compressed Learning
WaLLoC couples an invertible wavelet packet transform with a linear-projection encoder and nonlinear decoder to produce low-dimensional, quantization-resilient codes for compressed-domain learning.
-
SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers
SmoothCache uses calibration-measured layer error thresholds to skip redundant attention and feed-forward computations in Diffusion Transformers, achieving 8-71% speedup across image, video, and audio tasks.
-
KVAE: Family of Tokenizers for Multimodal Generative Models
KVAE introduces image, video, and full-band audio tokenizers whose reconstruction and downstream generation quality is competitive with, and often better than, current open-source tokenizers in head-to-head tests.
-
Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation
Boomerang sampling, applied to a pretrained music diffusion model, creates audio variations that improve beat tracking when training data is scarce and can change instruments via text prompts.
-
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
A cascaded pipeline of audio compression, latent diffusion extraction, and generative correction achieves state-of-the-art target speech extraction quality and intelligibility on Libri2Mix and out-of-domain data.
-
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
A caption-augmentation method that adds DSP-derived acoustic descriptors to text prompts gives a text-to-audio diffusion model controllable loudness, pitch, reverb, noise, brightness, fade, and duration.
-
Video-Guided Foley Sound Generation with Multimodal Controls
A video-guided diffusion model generates synchronized foley sound from text, audio, and video controls, using joint training on noisy internet videos and professional sound-effect libraries to reach 48kHz output.
-
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
Across five text-to-music models, aesthetic predictor scores, pairwise human preferences, and reference-based distribution metrics produce inconsistent rankings, so the choice of evaluation metric changes the winner.
-
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning
Combining global and local text embeddings in a diffusion UNet for music improves text adherence, while mean-pooling T5 local embeddings yields the best audio quality without extra parameters.
-
Sound Scene Synthesis at the DCASE 2024 Challenge
Four text-to-audio systems were evaluated against a human reference in the DCASE 2024 Task 7 challenge, with a 36% quality gap and strong but small-sample FAD-to-human correlation.
-
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
The paper proposes learning from language feedback to synthesize multimodal preference pairs, but the evidence is weakened by an undefined improvement metric and small, unvalidated effect sizes.
-
Video Diffusion Transformers are In-Context Learners
Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.
Discussion (0). Continue with ORCID to comment.