Pith. sign in

REVIEW 2 cited by

Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.14778 v5 pith:CH72C5WP submitted 2023-10-23 cs.MM cs.SDeess.AS

classification cs.MMcs.SDeess.AS
keywords audio-visualtrackingspeakerdeeppastyearsaudioinformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide complementary information for localization and tracking. With audio and visual information, the Bayesian-based filter and deep learning-based methods can solve the problem of data association, audio-visual fusion and track management. In this paper, we conduct a comprehensive overview of audio-visual speaker tracking. To our knowledge, this is the first extensive survey over the past five years. We introduce the family of Bayesian filters and summarize the methods for obtaining audio-visual measurements. In addition, the existing trackers and their performance on the AV16.3 dataset are summarized. In the past few years, deep learning techniques have thrived, which also boost the development of audio-visual speaker tracking. The influence of deep learning techniques in terms of measurement extraction and state estimation is also discussed. Finally, we discuss the connections between audio-visual speaker tracking and other areas such as speech separation and distributed speaker tracking.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Introduces AVTrack dataset for audio-visual tracking in challenging human-centric scenes, demonstrating performance drops in existing methods.

  2. Region-Specific Audio Tagging for Spatial Sound

    eess.AS 2025-09 conditional novelty 6.0 of 10

    Region-specific audio tagging: a model trained to tag sound events within a specified angular or distance region, tested on a new simulated benchmark and STARSS23.

Pith tools