Pith. sign in

REVIEW 13 cited by

FLUX that Plays Music

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.00587 v2 pith:LDEPDHFS submitted 2024-09-01 cs.SD cs.CVeess.AS

classification cs.SDcs.CVeess.AS
keywords fluxmusicflowfluxmusicgithubhttpsinformationmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper explores a simple extension of diffusion-based rectified flow Transformers for text-to-music generation, termed as FluxMusic. Generally, along with design in advanced Flux\footnote{https://github.com/black-forest-labs/flux} model, we transfers it into a latent VAE space of mel-spectrum. It involves first applying a sequence of independent attention to the double text-music stream, followed by a stacked single music stream for denoised patch prediction. We employ multiple pre-trained text encoders to sufficiently capture caption semantic information as well as inference flexibility. In between, coarse textual information, in conjunction with time step embeddings, is utilized in a modulation mechanism, while fine-grained textual details are concatenated with the music patch sequence as inputs. Through an in-depth study, we demonstrate that rectified flow training with an optimized architecture significantly outperforms established diffusion methods for the text-to-music task, as evidenced by various automatic metrics and human preference evaluations. Our experimental data, code, and model weights are made publicly available at: \url{https://github.com/feizc/FluxMusic}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    Presents the ATTM grand challenge with efficiency and performance tracks for text-to-music generation using a public instrumental music dataset, evaluated via FAD, CLAP, a new CCS metric, and subjective tests.

  2. MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline

    cs.SD 2026-02 unverdicted novelty 7.0 of 10

    MIDI-SAG generates consistent long-form singing accompaniments by feeding symbolic MIDI timing, chords, and structure labels into a compositional pipeline built from pre-trained modules.

  3. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Audex unifies audio understanding and generation on a strong text MoE backbone with multi-stage SFT plus text-only Cascade RL, matching open SOTA audio scores while mostly retaining text capability.

  4. Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music

    cs.SD 2026-05 unverdicted novelty 6.0 of 10

    Introduces the first large-scale Persian music dataset and shows fine-tuned MusicGen produces compositions more aligned with Persian stylistic conventions via tag-based evaluation.

  5. SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering

    cs.SD 2025-08 unverdicted novelty 6.0 of 10

    SonicMaster is a text-conditioned flow-matching generative model for unified music restoration and mastering, trained on a dataset of simulated degradations across equalization, dynamics, reverb, amplitude, and stereo.

  6. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 5.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  7. ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

    eess.AS 2026-05 unverdicted novelty 5.0 of 10

    ImmersiveTTS proposes an environment-aware TTS system that integrates speech with environmental audio via multimodal diffusion transformer, joint attention, and domain-specific representation alignment, claiming super...

  8. Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods

    cs.SD 2026-05 accept novelty 5.0 of 10

    The paper introduces the ATTM Grand Challenge with a CC-licensed instrumental subset of MTG-Jamendo, two tracks, and evaluation via FAD, CLAP, and a new Concept Coverage Score to support academic text-to-music research.

  9. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

    eess.AS 2024-10 unverdicted novelty 5.0 of 10

    F5-TTS generates natural speech from text via flow matching on DiT with simple text padding, ConvNeXt refinement, and sway sampling, trained on 100K hours multilingual data.

  10. FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

    cs.SD 2026-07 reject novelty 4.0 of 10

    FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.

  11. TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization

    cs.SD 2025-08 reject novelty 4.0 of 10

    TinyMusician distills MusicGen and applies hand-picked mixed-precision quantization to make a 1.04 GB on-device music generator, but the headline '93% quality, 55% smaller' claims conflict with the paper's own tables.

  12. UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation

    cs.SD 2026-07 unverdicted novelty 3.0 of 10

    Text-embedding clustering for batch sampling outperforms audio-embedding clustering on objective metrics in low-data text-to-music generation, with moderate cluster counts best on metrics and larger counts better for ...

  13. Improving Text-to-Music Generation with Human Preference Rewards

    cs.SD 2026-06 unverdicted novelty 2.0 of 10

    A text-to-music model is improved by conditioning on and selecting with a human preference reward, where expert iteration on top outputs contributes the largest measured gains on 100 Song Describer prompts.

Pith tools