Proposes an attribution-aware compensation framework for generative music that derives closed-form payments from catalog-level attribution informativeness and quantifies welfare effects under competition.
ACE-Step 1.5: Pushing the boundaries of open-source music generation
12 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 12roles
background 1polarities
background 1representative citing papers
Audio-Oscar is a multi-agent system that coordinates specialist agents for generating audio from complex scene descriptions and introduces the ASG-Bench benchmark for evaluation.
HAIM is a new labeled dataset for granular tracking of AI interventions across music production stages, enabling evaluation beyond binary AI-or-human classification.
DEMON is a streaming diffusion engine that exposes denoising parameters as playable controls at up to 12.3 decoder completions per second via per-slot scheduling, shared state, source blending, and accelerated decoding.
A native quantized runtime runs Stable Audio 3 on commodity and Pi hardware with 8-bit quality within seed noise, 7× faster cold start, and bounded in-graph taste steering.
S2Accompanist is a 402M-parameter semantic-aware diffusion model that achieves SOTA on the ATTM Grand Challenge benchmark for music accompaniment generation via automated data processing and structure-guided VAE fine-tuning.
SongBench is a new fine-grained benchmark for song quality assessment with seven dimensions and an expert-annotated dataset of 11,717 samples showing high correlation with professional ratings.
LeVo 2 presents a hierarchical LLM-Diffusion model with progressive post-training stages to generate full-length songs that balance semantic planning, track-specific acoustics, and musicality.
Score-aware training uses alignment scores to route low-quality segments into high-noise regimes as implicit regularizers, enabling a 450M model to rank competitively in a text-to-music challenge with limited data.
SketchSong uses temporal sketch planning with high-level tokens and explicit modeling of four tracks (vocals, bass, drums, other) to generate more coherent songs than baselines.
SAME is a semantically regularized transformer autoencoder for music that delivers 4096x compression with open-weights release of large and small variants.
LaDA-Band generates complete vocal-aligned instrumental accompaniments with discrete masked diffusion, claiming simultaneous gains in audio fidelity, long-range coherence, and orchestration quality.
citing papers explorer
-
What's a Credit Worth? A Market Framework for Attribution-Aware Compensation in Generative Music
Proposes an attribution-aware compensation framework for generative music that derives closed-form payments from catalog-level attribution informativeness and quantifies welfare effects under competition.
-
Audio-Oscar: A Multi-Agent System for Complex Audio Scene Generation, Orchestration, and Refinement
Audio-Oscar is a multi-agent system that coordinates specialist agents for generating audio from complex scene descriptions and introduces the ASG-Bench benchmark for evaluation.
-
HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark
HAIM is a new labeled dataset for granular tracking of AI interventions across music production stages, enabling evaluation beyond binary AI-or-human classification.
-
DEMON: Diffusion Engine for Musical Orchestrated Noise
DEMON is a streaming diffusion engine that exposes denoising parameters as playable controls at up to 12.3 decoder completions per second via per-slot scheduling, shared state, source blending, and accelerated decoding.
-
A Quantized Native Runtime for On-Device Semantic Audio Generation
A native quantized runtime runs Stable Audio 3 on commodity and Pi hardware with 8-bit quality within seed noise, 7× faster cold start, and bounded in-graph taste steering.
-
S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation
S2Accompanist is a 402M-parameter semantic-aware diffusion model that achieves SOTA on the ATTM Grand Challenge benchmark for music accompaniment generation via automated data processing and structure-guided VAE fine-tuning.
-
SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
SongBench is a new fine-grained benchmark for song quality assessment with seven dimensions and an expert-annotated dataset of 11,717 samples showing high correlation with professional ratings.
-
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training
LeVo 2 presents a hierarchical LLM-Diffusion model with progressive post-training stages to generate full-length songs that balance semantic planning, track-specific acoustics, and musicality.
-
Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation
Score-aware training uses alignment scores to route low-quality segments into high-noise regimes as implicit regularizers, enabling a 450M model to rank competitively in a text-to-music challenge with limited data.
-
SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling
SketchSong uses temporal sketch planning with high-level tokens and explicit modeling of four tracks (vocals, bass, drums, other) to generate more coherent songs than baselines.
-
SAME: A Semantically-Aligned Music Autoencoder
SAME is a semantically regularized transformer autoencoder for music that delivers 4096x compression with open-weights release of large and small variants.
-
LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation
LaDA-Band generates complete vocal-aligned instrumental accompaniments with discrete masked diffusion, claiming simultaneous gains in audio fidelity, long-range coherence, and orchestration quality.