A masked self-supervised model trained only on pitch, energy, and voice activity captures prosodic structure at multiple timescales, with random masking yielding the most generalizable representations.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
A masked self-supervised model trained only on pitch, energy, and voice activity captures prosodic structure at multiple timescales, with random masking yielding the most generalizable representations.