Pith. sign in

REVIEW 1 cited by

A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.08237 v1 pith:6O4KU3YB submitted 2024-05-13 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechtemporalencodingmodeldynamicsfoundlearningneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech perception involves storing and integrating sequentially presented items. Recent work in cognitive neuroscience has identified temporal and contextual characteristics in humans' neural encoding of speech that may facilitate this temporal processing. In this study, we simulated similar analyses with representations extracted from a computational model that was trained on unlabelled speech with the learning objective of predicting upcoming acoustics. Our simulations revealed temporal dynamics similar to those in brain signals, implying that these properties can arise without linguistic knowledge. Another property shared between brains and the model is that the encoding patterns of phonemes support some degree of cross-context generalization. However, we found evidence that the effectiveness of these generalizations depends on the specific contexts, which suggests that this analysis alone is insufficient to support the presence of context-invariant encoding.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

    eess.AS 2025-06 conditional novelty 6.0 of 10

    Phonetic information becomes decodable in AV-HuBERT only about 20 ms before audio-only HuBERT, not the 100 to 300 ms visual lead in human speech, indicating AV-HuBERT's temporal dynamics are dominated by audio.

Pith tools