Pith. sign in

REVIEW 10 cited by

Mel-Band RoFormer for Music Source Separation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01809 v1 pith:I37FSKML submitted 2023-10-03 cs.SD eess.AS

classification cs.SDeess.AS
keywords band-splitbs-roformerbsrnnmodelschemeseparationmel-bandmel-roformer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, multi-band spectrogram-based approaches such as Band-Split RNN (BSRNN) have demonstrated promising results for music source separation. In our recent work, we introduce the BS-RoFormer model which inherits the idea of band-split scheme in BSRNN at the front-end, and then uses the hierarchical Transformer with Rotary Position Embedding (RoPE) to model the inner-band and inter-band sequences for multi-band mask estimation. This model has achieved state-of-the-art performance, but the band-split scheme is defined empirically, without analytic supports from the literature. In this paper, we propose Mel-RoFormer, which adopts the Mel-band scheme that maps the frequency bins into overlapped subbands according to the mel scale. In contract, the band-split mapping in BSRNN and BS-RoFormer is non-overlapping and designed based on heuristics. Using the MUSDB18HQ dataset for experiments, we demonstrate that Mel-RoFormer outperforms BS-RoFormer in the separation tasks of vocals, drums, and other stems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

    eess.AS 2026-06 unverdicted novelty 7.0 of 10

    HoliDubber introduces a patch-based autoregressive diffusion transformer for joint text-guided synthesis of speech and ambient audio in video dubbing, with a new benchmark showing outperformance over prior speech-only...

  2. Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    Presents the ATTM grand challenge with efficiency and performance tracks for text-to-music generation using a public instrumental music dataset, evaluated via FAD, CLAP, a new CCS metric, and subjective tests.

  3. Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    BandTok tokenizes Mel-spectrograms as independent time-frequency band tokens from a single codebook and pairs it with 2D RoPE in an autoregressive model to improve music generation over residual multi-codebook tokenizers.

  4. YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance

    eess.AS 2026-03 unverdicted novelty 7.0 of 10

    YingMusic-Singer-Plus is a diffusion model for singing voice synthesis that preserves melody from a reference clip while allowing flexible lyric changes without manual alignment, outperforming Vevo2 and introducing th...

  5. MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline

    cs.SD 2026-02 unverdicted novelty 7.0 of 10

    MIDI-SAG generates consistent long-form singing accompaniments by feeding symbolic MIDI timing, chords, and structure labels into a compositional pipeline built from pre-trained modules.

  6. StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems

    cs.SD 2026-07 conditional novelty 6.0 of 10

    StemFX predicts tokenized per-stem audio-effect chains with a jointly-trained Transformer encoder-decoder, beating contrastive and prior FX-encoding methods on effect-chain retrieval and real-mix style transfer.

  7. Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music

    cs.SD 2026-05 unverdicted novelty 6.0 of 10

    Introduces the first large-scale Persian music dataset and shows fine-tuned MusicGen produces compositions more aligned with Persian stylistic conventions via tag-based evaluation.

  8. YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance

    eess.AS 2026-03 conditional novelty 6.0 of 10

    A diffusion SVS model with curriculum training and GRPO edits lyrics while preserving melody without manual alignment, outperforming Vevo2 on LyricEditBench.

  9. Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods

    cs.SD 2026-05 accept novelty 5.0 of 10

    The paper introduces the ATTM Grand Challenge with a CC-licensed instrumental subset of MTG-Jamendo, two tracks, and evaluation via FAD, CLAP, and a new Concept Coverage Score to support academic text-to-music research.

  10. Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

    cs.SD 2026-07 conditional novelty 3.5 of 10

    MSST unifies training, validation, and inference for many music source-separation architectures and reports small quality gains from TTA, ensembling, and related engineering techniques.

Pith tools