Pith. sign in

Effective Context in Neural Speech Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Modern neural speech models benefit from having longer context, and many approaches have been proposed to increase the maximum context a model can use. However, few have attempted to measure how much context these models actually use, i.e., the effective context. Here, we propose two approaches to measuring the effective context, and use them to analyze different speech Transformers. For supervised models, we find that the effective context correlates well with the nature of the task, with fundamental frequency tracking, phone classification, and word classification requiring increasing amounts of effective context. For self-supervised models, we find that effective context increases mainly in the early layers, and remains relatively short -- similar to the supervised phone model. Given that these models do not use a long context during prediction, we show that HuBERT can be run in streaming mode without modification to the architecture and without further fine-tuning.

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Effective Context in Neural Speech Models

cs.SD · 2025-05-28 · conditional · novelty 6.0

Introduces truncation and Jacobian-based measures of effective context showing that self-supervised speech Transformers use a short, mostly local context and can be streamed with minor probing degradation.

citing papers explorer

Showing 1 of 1 citing paper.

  • Effective Context in Neural Speech Models cs.SD · 2025-05-28 · conditional · none · ref 2 · internal anchor

    Introduces truncation and Jacobian-based measures of effective context showing that self-supervised speech Transformers use a short, mostly local context and can be streamed with minor probing degradation.