Pith. sign in

REVIEW 5 cited by

Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.11757 v3 pith:U6L7SYBO submitted 2023-01-27 cs.CL cs.LGcs.SDeess.AS

Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

classification cs.CL cs.LGcs.SDeess.AS
keywords musicmodeltextgenerationmodelsopen-sourcediffusionhttps
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent years have seen the rapid development of large generative models for text; however, much less research has explored the connection between text and another "language" of communication -- music. Music, much like text, can convey emotions, stories, and ideas, and has its own unique structure and syntax. In our work, we bridge text and music via a text-to-music generation model that is highly efficient, expressive, and can handle long-term structure. Specifically, we develop Mo\^usai, a cascading two-stage latent diffusion model that can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions. Moreover, our model features high efficiency, which enables real-time inference on a single consumer GPU with a reasonable speed. Through experiments and property analyses, we show our model's competence over a variety of criteria compared with existing music generation models. Lastly, to promote the open-source culture, we provide a collection of open-source libraries with the hope of facilitating future work in the field. We open-source the following: Codes: https://github.com/archinetai/audio-diffusion-pytorch; music samples for this paper: http://bit.ly/44ozWDH; all music samples for all models: https://bit.ly/audio-diffusion.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration

    cs.SD 2026-05 unverdicted novelty 7.0

    Polyphonia improves zero-shot stem-specific timbre transfer in polyphonic music by 15.5% target alignment via acoustic-informed attention calibration that uses probabilistic priors to set coarse boundaries.

  2. AudioMoG: Guiding Audio Generation with Mixture-of-Guidance

    cs.SD 2025-09 unverdicted novelty 7.0

    AudioMoG is a mixture-of-guidance sampling technique that combines CFG and AG signals to outperform single-guidance baselines in text-to-audio generation at equivalent speed.

  3. Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment

    cs.SD 2026-07 conditional novelty 6.0

    Beat-guided contrastive alignment of pretrained MotionBERT/MERT features plus ControlNet conditioning of AudioLDM improves dance–music alignment on AIST++ over a MusicGen textual-inversion baseline while remaining com...

  4. Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

    cs.SD 2026-06 unverdicted novelty 6.0

    A data-free streaming consistency distillation framework enables single-step autoregressive generation from text-to-music models for real-time interactive use while preserving timbre and rhythm via latent, spectral, a...

  5. NeuroSonic: Conditional Flow Matching for EEG-to-Speech Reconstruction

    cs.LG 2026-06 unverdicted novelty 5.0

    NeuroSonic introduces a conditional flow-matching framework that learns a deterministic transport from noise to speech conditioned on EEG, reporting up to 26.3% gains in perceptual quality over GAN, diffusion, and mea...