A masked self-supervised model trained only on pitch, energy, and voice activity captures prosodic structure at multiple timescales, with random masking yielding the most generalizable representations.
However, the extent to which this predictive capacity is contingent on the structure of lexical information or of its acoustic realisation is unclear
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
A masked self-supervised model trained only on pitch, energy, and voice activity captures prosodic structure at multiple timescales, with random masking yielding the most generalizable representations.