Sequence tokens of a class-token-only contrastively trained ViT-1D carry usable beat, chord, and onset information that emerges during training without any frame-level supervision.
,zT k ] at layers k = 3, 6, 9, 12 (same as Section 5, denoted z3 to z12), along with tokens from a randomly initialized ViT-1D model, de- noted zr
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Emergent musical properties of a transformer under contrastive self-supervised learning
Sequence tokens of a class-token-only contrastively trained ViT-1D carry usable beat, chord, and onset information that emerges during training without any frame-level supervision.