Pith. sign in

Online Audio-Visual Autoregressive Speaker Extraction

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less explored. We first propose a lightweight visual frontend based on depth-wise separable convolution. Then, we propose a lightweight autoregressive acoustic encoder to serve as the second cue, to actively explore the information in the separated speech signal from past steps. Scenario-wise, for the first time, we study how the algorithm performs when there is a change in focus of attention, i.e., the target speaker. Experimental results on LRS3 datasets show that our visual frontend performs comparably to the previous state-of-the-art on both SkiM and ConvTasNet audio backbones with only 0.1 million network parameters and 2.1 MACs per second of processing. The autoregressive acoustic encoder provides an additional 0.9 dB gain in terms of SI-SNRi, and its momentum is robust against the change in attention.

citation-role summary

background 1

citation-polarity summary

fields

eess.AS 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Online Audio-Visual Autoregressive Speaker Extraction

eess.AS · 2025-06-02 · conditional · novelty 5.0

A 0.1M-parameter visual frontend plus an autoregressive acoustic encoder improves online audio-visual speaker extraction by up to 0.9 dB SI-SNRi on simulated LRS3 mixtures, with roughly 0.5 dB attributable to the acoustic encoder alone.

citing papers explorer

Showing 1 of 1 citing paper.

  • Online Audio-Visual Autoregressive Speaker Extraction eess.AS · 2025-06-02 · conditional · none · ref 1 · internal anchor

    A 0.1M-parameter visual frontend plus an autoregressive acoustic encoder improves online audio-visual speaker extraction by up to 0.9 dB SI-SNRi on simulated LRS3 mixtures, with roughly 0.5 dB attributable to the acoustic encoder alone.