Pith. sign in

REVIEW 3 cited by

RenderBox: Expressive Performance Rendering with Text Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.07711 v1 pith:T7GNN4HB submitted 2025-02-11 eess.AS cs.MM

RenderBox: Expressive Performance Rendering with Text Control

classification eess.AS cs.MM
keywords performanceexpressiverenderboxacrosscontrollablecontrolsintentmusic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Expressive music performance rendering involves interpreting symbolic scores with variations in timing, dynamics, articulation, and instrument-specific techniques, resulting in performances that capture musical can emotional intent. We introduce RenderBox, a unified framework for text-and-score controlled audio performance generation across multiple instruments, applying coarse-level controls through natural language descriptions and granular-level controls using music scores. Based on a diffusion transformer architecture and cross-attention joint conditioning, we propose a curriculum-based paradigm that trains from plain synthesis to expressive performance, gradually incorporating controllable factors such as speed, mistakes, and style diversity. RenderBox achieves high performance compared to baseline models across key metrics such as FAD and CLAP, and also tempo and pitch accuracy under different prompting tasks. Subjective evaluation further demonstrates that RenderBox is able to generate controllable expressive performances that sound natural and musically engaging, aligning well with prompts and intent.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PianoKontext: Expressive Performance Rendering from Deadpan Context

    cs.SD 2026-06 unverdicted novelty 6.0

    PianoKontext renders expressive piano performances from deadpan scores using flow matching in Music2Latent latent space with DTW alignment for paired training data.

  2. Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis

    cs.SD 2026-05 unverdicted novelty 6.0

    Break-the-Beat! renders drum MIDI audio that matches the timbre of a reference clip by fine-tuning a text-to-audio model with a content encoder and hybrid conditioning on a new paired dataset.

  3. RenCon 2025: Revival of the Expressive Performance Rendering Competition

    cs.MM 2026-05 unverdicted novelty 2.0

    RenCon 2025 revived the expressive performance rendering competition, attracting nine diverse systems whose online and live evaluations showed measurable advances alongside persistent gaps to human-level musical expression.