Pith. sign in

REVIEW 3 cited by

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.23869 v1 pith:E5PAMDLT submitted 2025-06-30 cs.SD cs.AIcs.LGeess.AS

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

classification cs.SD cs.AIcs.LGeess.AS
keywords symbolicmodelsclassificationcontrastivegenerationgenerativemodelmusic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We study the capabilities of generative autoregressive transformer models trained on large amounts of symbolic solo-piano transcriptions. After first pretraining on approximately 60,000 hours of music, we use a comparatively smaller, high-quality subset, to finetune models to produce musical continuations, perform symbolic classification tasks, and produce general-purpose contrastive MIDI embeddings by adapting the SimCLR framework to symbolic music. When evaluating piano continuation coherence, our generative model outperforms leading symbolic generation techniques and remains competitive with proprietary audio generation models. On MIR classification benchmarks, frozen representations from our contrastive model achieve state-of-the-art results in linear probe experiments, while direct finetuning demonstrates the generalizability of pretrained representations, often requiring only a few hundred labeled examples to specialize to downstream tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

    cs.SD 2026-04 unverdicted novelty 7.0

    BEAT tokenizes symbolic music by uniform beat steps with sparse per-beat pitch encodings, producing higher quality and more coherent music continuation and accompaniment than event-based tokenizations.

  2. Tipiano: Cascaded Piano Hand Motion Synthesis via Fingertip Priors

    cs.AI 2026-04 unverdicted novelty 6.0

    The Tipiano system synthesizes piano hand motions via cascaded fingertip priors, trajectory refinement, wrist estimation, and STGCN pose synthesis, achieving F1=0.910 and near motion-capture quality in user studies.

  3. BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

    cs.SD 2026-04 unverdicted novelty 5.0

    A uniform-temporal-step tokenization for symbolic music improves generation quality, efficiency, and long-range coherence over event-based alternatives.