Pith. sign in

hub

Stable Audio Open

17 Pith papers cite this work. Polarity classification is still indexing.

17 Pith papers citing it
abstract

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.

hub tools

citation-role summary

background 1

citation-polarity summary

roles

background 1

polarities

background 1

representative citing papers

Moshi: a speech-text foundation model for real-time dialogue

eess.AS · 2024-09-17 · accept · novelty 7.0

Moshi is the first real-time full-duplex spoken large language model that casts dialogue as speech-to-speech generation using parallel audio streams and an inner monologue of time-aligned text tokens.

FSD50K-Solo: Automated Curation of Single-Source Sound Events

eess.AS · 2026-05-13 · unverdicted · novelty 6.0 · 2 refs

A curation pipeline combining diffusion-based synthetic mixtures with a discriminative classifier produces and releases FSD50K-Solo, a single-source subset of FSD50K that matches human expert labels on a test set.

GPC: Large-Scale Generative Pretraining for Transferable Motor Control

cs.CV · 2026-06-28 · unverdicted · novelty 5.0

GPC learns a motion vocabulary via Finite Scalar Quantization and end-to-end RL, then trains an autoregressive transformer for next-token control generation, achieving 99.98% motion reproduction success with emergent robustness.

UniVoice: A Unified Model for Speech and Singing Voice Generation

cs.SD · 2026-06-04 · unverdicted · novelty 5.0

UniVoice is a conditional flow matching model with a Diffusion Transformer backbone that unifies TTS and SVS via modality-specific encoders and a null melody token for speech, achieving 5.26% speech PER and 16.22% singing PER.

Instrumental Text-to-Music Generation with Auxiliary Conditioning Branches

cs.SD · 2026-05-20 · conditional · novelty 5.0

Auxiliary lyric and timbre branches improve instrumental text-to-music generation quality in a controlled DiT setting even with degenerate inputs, outperforming parameter-reallocated depth variants and external baselines in objective and MOS evaluations.

Woosh: A Sound Effects Foundation Model

cs.SD · 2026-04-02 · accept · novelty 5.0

Woosh is a new publicly released foundation model optimized for high-quality sound effect generation from text or video, showing competitive or better results than open alternatives like Stable Audio Open.

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention

cs.SD · 2025-02-06 · unverdicted · novelty 5.0

XAttnMark is a new neural audio watermarking method using partial parameter sharing, cross-attention for message retrieval, temporal conditioning, and a psychoacoustic TF masking loss that reports state-of-the-art detection and attribution robustness.

citing papers explorer

Showing 17 of 17 citing papers.