REVIEW 1 cited by
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Speech perception involves storing and integrating sequentially presented items. Recent work in cognitive neuroscience has identified temporal and contextual characteristics in humans' neural encoding of speech that may facilitate this temporal processing. In this study, we simulated similar analyses with representations extracted from a computational model that was trained on unlabelled speech with the learning objective of predicting upcoming acoustics. Our simulations revealed temporal dynamics similar to those in brain signals, implying that these properties can arise without linguistic knowledge. Another property shared between brains and the model is that the encoding patterns of phonemes support some degree of cross-context generalization. However, we found evidence that the effectiveness of these generalizations depends on the specific contexts, which suggests that this analysis alone is insufficient to support the presence of context-invariant encoding.
Forward citations
Cited by 1 Pith paper
-
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
Phonetic information becomes decodable in AV-HuBERT only about 20 ms before audio-only HuBERT, not the 100 to 300 ms visual lead in human speech, indicating AV-HuBERT's temporal dynamics are dominated by audio.
Discussion (0). Sign in to comment.