Pith. sign in

REVIEW 1 major objections 1 minor 55 references

MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

T0 review · 1 major / 1 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read MuVAP extends voice activity projection with face tracks and role-relative mapping to predict turn-taking from single-camera monaural recordings.

desk verdict MuVAP adds a single-camera multimodal grounding step and a new unedited corpus, but the Role-Relative Projection needs to show it preserves speaker identity for real next-speaker gains rather than just binary shift/hold. read the letter →

arxiv 2606.16731 v2 pith:GVJSP2AW submitted 2026-06-15 cs.SD cs.AIcs.HC

classification cs.SDcs.AIcs.HC
keywords multimodalturn-takingvoiceactivityprojectionmultipartyconversationnextspeakerpredictionaudiovisualcorpushuman-robotinteraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents MuVAP as a way to predict who will speak next in group conversations using only one microphone and one camera. It grounds the audio-based voice activity model in visible face tracks to make speaker-aware forecasts. A new mapping called Role-Relative Projection reduces the problem of any number of speakers to a simple current-versus-next decision. The authors created a 31-hour dataset of natural unedited conversations to train and test this. The model beats baselines on deciding when to hold or shift the floor and on naming the next speaker in two- and three-person talks.

What carries the argument

Role-Relative Projection, which maps any N-speaker interaction onto a fixed current versus next floor-holder state to address combinatorial complexity.

What would settle it

Collecting new recordings where single-camera face tracking frequently loses tracks due to movement or occlusion and checking whether prediction accuracy then falls below audio-only baselines.

Watch

Extended reading notes

Core claim

MuVAP is a causal multimodal framework that grounds acoustic predictions in face tracks from a single camera view, using Role-Relative Projection to map multiparty interactions onto a fixed current versus next floor-holder state, thereby enabling accurate shift-hold and next-speaker predictions from monaural audio in unedited wild settings, as demonstrated by outperformance of strong baselines on the Audio-Visual Conversation Corpus.

Load-bearing premise

Face tracks from a single camera can reliably ground the acoustic predictions and that the Role-Relative Projection preserves enough multiparty dynamics without speaker-specific modeling.

Editorial extensions

If this is right

  • MuVAP enables turn-taking prediction without complex microphone arrays or multi-camera setups.
  • It supports human-robot interaction scenarios with simple hardware.
  • Performance gains hold for both two- and three-speaker settings on shift-hold and next-speaker tasks.
  • The unedited nature of the new corpus avoids artifacts from editing cuts that break causal tracking.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If single-view face tracking works reliably, the same approach could apply to consumer video calls for automatic turn management.
  • The role mapping might allow scaling to four or more speakers without exponential growth in model complexity.
  • Combining this with existing speech recognition could lead to fully hands-free group conversation agents.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. The paper introduces MuVAP, a causal multimodal framework extending Voice Activity Projection by grounding acoustic predictions in face tracks from a single camera and monaural audio. It proposes Role-Relative Projection to map any N-speaker interaction onto a fixed current-versus-next floor-holder state to address combinatorial complexity, introduces the 31-hour Audio-Visual Conversation Corpus of unedited single-camera recordings, and claims that MuVAP outperforms strong baselines on Shift-Hold and next-speaker prediction tasks in two- and three-speaker settings.

Significance. If the empirical claims hold after addressing the output-space concern, the work would be significant for practical human-robot interaction by enabling speaker-aware turn-taking from minimal hardware. The new unedited dataset fills a gap for causal audiovisual modeling, and the multimodal grounding approach could generalize beyond current microphone-array or multi-camera setups.

major comments (1)
  1. [Abstract] Abstract: The central claim that Role-Relative Projection enables 'next-speaker prediction' while mapping N-speaker dynamics onto a fixed current-versus-next state requires explicit clarification on output representation. If the mapping assigns a single undifferentiated 'next' label without retaining speaker identity or role distinctions among candidate next-speakers, the task reduces to binary hold/shift detection; reported gains on next-speaker prediction would then be artifacts of the reduced output space rather than evidence that multimodal grounding identifies specific speakers.
minor comments (1)
  1. The abstract states that evaluations demonstrate outperformance but provides no quantitative metrics, baseline details, or ablation results; these must be presented with error analysis to allow verification of the performance claims.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the careful reading and for highlighting the need for explicit clarification on the output representation of Role-Relative Projection. We address the concern directly below and will revise the manuscript accordingly.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The central claim that Role-Relative Projection enables 'next-speaker prediction' while mapping N-speaker dynamics onto a fixed current-versus-next state requires explicit clarification on output representation. If the mapping assigns a single undifferentiated 'next' label without retaining speaker identity or role distinctions among candidate next-speakers, the task reduces to binary hold/shift detection; reported gains on next-speaker prediction would then be artifacts of the reduced output space rather than evidence that multimodal grounding identifies specific speakers.

    Authors: We agree that the abstract requires explicit clarification on this point. Role-Relative Projection reduces combinatorial complexity by mapping arbitrary speaker configurations to a fixed current-versus-next floor-holder state; however, the output is not an undifferentiated binary label. Because predictions are grounded in the face tracks extracted from the single-camera view, the model produces per-track probabilities that identify which specific visual track (and therefore which individual) corresponds to the current floor-holder and which corresponds to the next floor-holder. This retains speaker identity and is distinct from pure hold/shift classification. We will revise the abstract to state the output representation explicitly (i.e., that next-speaker prediction is performed over the set of detected face tracks). revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity; claims rest on new dataset and empirical baselines

full rationale

The paper introduces Role-Relative Projection as a modeling choice to reduce combinatorial complexity and evaluates MuVAP on Shift-Hold and next-speaker tasks using a newly collected unedited dataset. No equations or derivations reduce predictions to fitted inputs by construction, no load-bearing self-citations are invoked to justify uniqueness, and no ansatz is smuggled via prior work. The central results are presented as empirical outperformance against baselines rather than tautological redefinitions. This is the expected non-finding for an applied modeling paper whose validity can be checked externally via the reported metrics.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract-only view yields no explicit free parameters, axioms, or invented entities; the central claims rest on the unstated premise that single-camera face tracking is sufficiently reliable in the target domain.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild." pith.science (2026). https://pith.science/paper/GVJSP2AW

@misc{pith2026260616731,
  author       = {Pith},
  title        = {Pith review of: MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GVJSP2AW}},
  note         = {Machine review of arXiv:2606.16731}
}
read the original abstract

Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios. We introduce MuVAP, a causal multimodal framework that extends Voice Activity Projection by grounding acoustic predictions in face tracks, enabling speaker-aware turn-taking predictions from a monaural audio stream and a single camera view. To address the combinatorial complexity of modeling multiple speakers, we propose Role-Relative Projection, which maps any N-speaker interaction onto a fixed current versus next floor-holder state. Because existing audiovisual datasets contain disruptive editing cuts that break causal tracking, we introduce the Audio-Visual Conversation Corpus, a 31-hour dataset of unedited, single-camera multiparty conversations. Evaluations demonstrate that MuVAP outperforms strong baselines on Shift-Hold and next-speaker prediction tasks across two- and three-speaker settings.

Figures

Figures reproduced from arXiv: 2606.16731 by the authors.

Figure 1
Figure 1. Visualization demo1 of multiparty turn-taking predic￾tion. The top panel shows the video stream with tracking of speakers S0 (red), S1 (green), and S2 (yellow), with the sin￾gle channel audio stream below. The bottom panel shows the ground truth speech activity timelines, with the red vertical line marking the current time step t, leading to a turn shift from S1 to S0. The left panel shows the GlobalVAP (GVAP) shift… view at source ↗
Figure 2
Figure 2. Comparison of dataset continuity and structure. AVA￾ActiveSpeaker (Top) is limited to many isolated speaker seg￾ments rather than full scene, often relying on scripted cues that lack natural turn-taking dynamics. MSDWild (Middle) captures unconstrained scenes but suffers from editorial and jump cuts that disrupt the visual timeline needed for prediction. In con￾trast, AVCC (Bottom) preserves the unedited, continuous… view at source ↗
Figure 3
Figure 3. The difference between speaker based projection [15] and (Ours) Role-Relative Projection for 3-speaker conversa￾tions. pre-trained ASD model, as shown in [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Modular architecture of the MuVAP model. The embeddings from the VAP and ASD modules (ZVAP and ZASD) are used as input to the main module (right), which makes both GlobalVAP (GVAP) and SpeakerVAP (SVAP) predictions. capturing certain backchannel expressions within edit…
Figure 5
Figure 5. Figure 5: Visualization of the Active and Silent prediction sam￾pling methods. The yellow region indicates the required solo￾speaker duration (1000 ms) and offset (100 ms) used to prevent label leakage. The maximum gap duration is 3000 ms for both event types. Active events are …

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

55 extracted references · 3 canonical work pages

  1. [1]

    MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

    Introduction Turn-taking is a fundamental aspect of conversation. Since peo- ple cannot easily speak and listen simultaneously, they must coordinate their turns through a complex exchange of cues [1, 2, 3]. For example, a syntactically or semantically incom- plete phrase may signal a turn hold, whereas a complete phrase may be turn-yielding [4]. A filled ...

  2. [2]

    Active Speaker Detection vs

    Related Work 2.1. Active Speaker Detection vs. Turn-Taking We distinguish our work from multimodal Active Speaker De- tection (ASD), which aims to identify the active speaker among multiple faces in a video stream. Current ASD architectures, such as TalkNet [17] and LoCoNet [18], achieve impressive pre- cision on the A V A-ActiveSpeaker (A V A-AS) benchma...

  3. [3]

    Speakers use these visual sig- nals alongside prosodic contours to navigate turn-taking [25]

    to identify floor acquisition. Speakers use these visual sig- nals alongside prosodic contours to navigate turn-taking [25]. While sentence-level NSP [26] and event-based visual frame- Dataset Modality Duration Source Lang. Unedited Module Fisher Audio 1958h7m Telephone English✓V AP A V A-AS A V 38h5m Movie Multi✗ASD MSDWild A V 80h18m Vlogs/Wild Multi✗AS...

  4. [4]

    As illustrated in Figure 2, these sources contain editing artifacts like jump cuts that disrupt the temporal flow of conversation

    The A VCC Dataset Currently, publicly available ASD datasets [31, 16] suffer from domain gaps that limit their utility for turn-taking prediction in interactive settings (such as HRI). As illustrated in Figure 2, these sources contain editing artifacts like jump cuts that disrupt the temporal flow of conversation. This prevents models from learning the tr...

  5. [5]

    cur- rent floor holder

    Model We proposeMuV AP(Multimodal Multiparty V oice Activity Projection), a causal framework designed to forecast turn-taking dynamics for an arbitrary number of speakers. Given a single- channel audio waveformA ∈R T and a set of face tracks V={v 1, . . . , vN }forNdetected speakers, the model predicts two probability distributions at each time stept: • G...

  6. [6]

    The speaker with the highest activity is designated as the current/past floor holder

    Current Holder (S curr): We rank allNspeakers by their to- tal activity in the history bins. The speaker with the highest activity is designated as the current/past floor holder

  7. [7]

    social conductor

    Next Holder (S next): We rank the remainingN−1speak- ers by their activity in the future bins. The highest-ranked remaining speaker is designated as the primary next speaker. This reduction transforms the complex multiparty dynamic into a fixed pairwise state{S curr, Snext}. Following the standard V AP encoding method [11], we extract the binary activity ...

  8. [8]

    Each dataset is mapped to specific mod- ules to facilitate multi-stage training as shown in Table 1

    Implementation Our pipeline utilizes five distinct datasets: Fisher (Parts 1 & 2), MSDWild, W ASD, A V A-ActiveSpeaker, and our newly intro- duced A VCC dataset. Each dataset is mapped to specific mod- ules to facilitate multi-stage training as shown in Table 1. All modules were optimized using a Cosine Annealing learning rate scheduler with a linear warm...

Show all 55 references
  1. [9]

    To measure the usefulness of these learned embeddings, we de- fine a set of downstream tasks for evaluation

    Downstream tasks In the spirit of the original V AP model [11], MuV AP is trained to learn generic embeddings capturing turn-taking dynamics. To measure the usefulness of these learned embeddings, we de- fine a set of downstream tasks for evaluation. 6.1. Shift-Hold vs. Next S...

  2. [10]

    Results are reported as mean ± 95% confidence intervals over ten runs using random seeds 42–51

    Results We evaluate Shift-Hold Prediction using Macro-F1 (due to class imbalance), and Next Speaker Prediction (NSP) using accuracy. Results are reported as mean ± 95% confidence intervals over ten runs using random seeds 42–51. We compare MuV AP against two baselines:Majority...

  3. [11]

    Discussion The results reveal a clear division of labor between the modal- ities in our model. Acoustic, linguistic, and prosodic cues pro- vide the global rhythm of coordination, while the visual channel acts as an essential separation anchor during competitive over- lap and ...

  4. [12]

    To address the scaling bottlenecks of joint modeling, we proposed a Role- Relative Projection that compresses complex group dynamics into a scalable pairwise state

    Conclusion We introduced MuV AP, a multimodal framework for predict- ing turn-taking in unconstrained multiparty settings. To address the scaling bottlenecks of joint modeling, we proposed a Role- Relative Projection that compresses complex group dynamics into a scalable pairw...

  5. [13]

    The computations and data handling were enabled by the Berzelius resource provided by the Knut and Alice Wallenberg Foundation at the National Supercomputer Centre

    Acknowledgment This work was supported by the Wallenberg AI, Autonomous Systems and Software Program (W ASP) funded by the Knut and Alice Wallenberg Foundation (KAW), and the Swedish Re- search Council project 2020-03812. The computations and data handling were enabled by the ...

  6. [14]

    These tools were not used for writing any major parts of the paper

    Generative AI Use Disclosure Generative AI tools were used in this paper exclusively for edit- ing and polishing the text to remove grammatical errors. These tools were not used for writing any major parts of the paper

  7. [15]

    Some signals and rules for taking speaking turns in conversations,

    S. Duncan, “Some signals and rules for taking speaking turns in conversations,”Journal of Personality and Social Psychology, vol. 23, pp. 283–292, 08 1972

  8. [16]

    A simplest system- atics for the organization of turn-taking for conversation,

    H. Sacks, E. A. Schegloff, and G. Jefferson, “A simplest system- atics for the organization of turn-taking for conversation,”Lan- guage, vol. 50, no. 4, pp. 696–735, 1974

  9. [17]

    On the structure of speaker–auditor interaction dur- ing speaking turns,

    S. Duncan Jr, “On the structure of speaker–auditor interaction dur- ing speaking turns,”Language in Society, vol. 3, no. 2, pp. 161– 180, 1974

  10. [18]

    Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,

    C. E. Ford and S. A. Thompson, “Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,”Studies in interactional sociolinguistics, vol. 13, pp. 134–184, 1996

  11. [19]

    H. H. Clark,Using language. Cambridge University Press, 1996

  12. [20]

    Listeners’ responses to filled pauses in relation to floor apportionment,

    P. Ball, “Listeners’ responses to filled pauses in relation to floor apportionment,”British Journal of Social & Clinical Psychology, vol. 14, pp. 423–424, 11 1975

  13. [21]

    Duncan and D

    S. Duncan and D. Fiske,Face-to-face interaction: research, meth- ods and theory. Wiley, 1977

  14. [22]

    Universals and cultural variation in turn-taking in conversation,

    T. Stivers, N. J. Enfield, P. Brown, C. Englert, M. Hayashi, T. Heinemann, G. Hoymann, F. Rossano, J. P. De Ruiter, K.-E. Yoonet al., “Universals and cultural variation in turn-taking in conversation,”Proceedings of the National Academy of Sciences, vol. 106, no. 26, pp. 10 58...

  15. [23]

    Timing in turn-taking and its im- plications for processing models of language,

    S. C. Levinson and F. Torreira, “Timing in turn-taking and its im- plications for processing models of language,”Frontiers in psy- chology, vol. 6, p. 136034, 2015

  16. [24]

    Conversational interaction with social robots,

    G. Skantze, “Conversational interaction with social robots,” in Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction, 2021, p. 717

  17. [25]

    V oice Activity Projection: Self- supervised Learning of Turn-taking Events,

    E. Ekstedt and G. Skantze, “V oice Activity Projection: Self- supervised Learning of Turn-taking Events,” inProc. Interspeech 2022, 2022, pp. 5190–5194

  18. [26]

    Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,

    K. Inoue, D. Lala, G. Skantze, and T. Kawahara, “Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,” inProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the Association for Computational Lin...

  19. [27]

    Applying general turn-taking models to conversational human-robot interaction,

    G. Skantze and B. Irfan, “Applying general turn-taking models to conversational human-robot interaction,” inProceedings of the ACM/IEEE International Conference on Human-Robot Interac- tion (HRI), 2025

  20. [28]

    Multimodal turn analysis and prediction for multi-party conversations,

    M.-C. Lee, M. Trinh, and Z. Deng, “Multimodal turn analysis and prediction for multi-party conversations,” inProceedings of the 25th International Conference on Multimodal Interaction, 2023, pp. 436–444

  21. [29]

    Triadic multi- party voice activity projection for turn-taking in spoken dialogue systems,

    M. Elmers, K. Inoue, D. Lala, and T. Kawahara, “Triadic multi- party voice activity projection for turn-taking in spoken dialogue systems,” inProc. Interspeech 2025, 2025, pp. 3015–3019

  22. [30]

    MSDWild: Multi- modal Speaker Diarization Dataset in the Wild,

    T. Liu, S. Fan, X. Xiang, H. Song, S. Lin, J. Sun, T. Han, S. Chen, B. Yao, S. Liu, Y . Wu, Y . Qian, and K. Yu, “MSDWild: Multi- modal Speaker Diarization Dataset in the Wild,” inProc. Inter- speech 2022, 2022, pp. 1476–1480

  23. [31]

    Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection,

    R. Tao, Z. Pan, R. K. Das, X. Qian, M. Z. Shou, and H. Li, “Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection,” inProceedings of the 29th ACM International Conference on Multimedia, 2021, pp. 3927– 3935

  24. [32]

    Loconet: Long-short con- text network for active speaker detection,

    X. Wang, F. Cheng, and G. Bertasius, “Loconet: Long-short con- text network for active speaker detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2024, pp. 18 462–18 472

  25. [33]

    A V A Active Speaker: An audio-visual dataset for active speaker detection,

    J. Roth, S. Chaudhuri, O. Klejch, R. Marvin, A. Gallagher, L. Kaver, S. Ramaswamy, A. Stopczynski, C. Schmid, Z. Xiet al., “A V A Active Speaker: An audio-visual dataset for active speaker detection,” inICASSP 2020 IEEE International Conference on Acoustics, Speech and Signal ...

  26. [34]

    How much does prosody help turn- taking? Investigations using voice activity projection models,

    E. Ekstedt and G. Skantze, “How much does prosody help turn- taking? Investigations using voice activity projection models,” in Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. Edinburgh, UK: Association for Computational Linguist...

  27. [35]

    Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,

    R. Ishii, K. Otsuka, S. Kumano, M. Matsuda, and J. Yamato, “Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,” inProceedings of the 15th ACM on Inter- national conference on multimodal interaction, 2013, pp. 79–86

  28. [36]

    Gaze-enhanced multimodal turn-taking prediction in triadic conversations,

    S. Heo, C. Miller, C. Murdock, and M. Proulx, “Gaze-enhanced multimodal turn-taking prediction in triadic conversations,” in Proc. Interspeech 2025, 2025, pp. 1068–1072

  29. [37]

    Multimodal voice activ- ity prediction: Turn-taking events detection in expert-novice con- versation,

    K. Onishi, H. Tanaka, and S. Nakamura, “Multimodal voice activ- ity prediction: Turn-taking events detection in expert-novice con- versation,” inProceedings of the 11th International Conference on Human-Agent Interaction, 2023, pp. 13–21

  30. [38]

    Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,

    R. Ishii, K. Otsuka, S. Kumano, R. Higashinaka, and J. Tomita, “Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,”Multimodal Tech- nologies and Interaction, vol. 3, no. 4, p. 70, 2019

  31. [39]

    Who’s next? speaker-selection mech- anisms in multiparty dialogue,

    V . Petukhova and H. Bunt, “Who’s next? speaker-selection mech- anisms in multiparty dialogue,” inProceedings of the Workshop on the Semantics and Pragmatics of Dialogue, 2009, pp. 19–26

  32. [40]

    A computational study on sentence-based next speaker prediction in multiparty conversa- tions,

    M.-C. Lee, W. A. Li, and Z. Deng, “A computational study on sentence-based next speaker prediction in multiparty conversa- tions,” inProceedings of the 24th ACM International Conference on Intelligent Virtual Agents, 2024, pp. 1–4

  33. [41]

    Multimodal continuous turn-taking prediction using multiscale rnns,

    M. Roddy, G. Skantze, and N. Harte, “Multimodal continuous turn-taking prediction using multiscale rnns,” inProceedings of the 20th ACM International Conference on Multimodal Interac- tion, 2018, pp. 186–190

  34. [42]

    Visual cues enhance predictive turn- taking for two-party human interaction,

    S. O. Russell and N. Harte, “Visual cues enhance predictive turn- taking for two-party human interaction,” inFindings of the As- sociation for Computational Linguistics: ACL 2025, 2025, pp. 209–221

  35. [43]

    Multi-channel sequence-to-sequence neural diarization: Experimental results for the misp 2025 challenge,

    M. Cheng, F. Su, C. Li, J. Liu, and M. Li, “Multi-channel sequence-to-sequence neural diarization: Experimental results for the misp 2025 challenge,” inProc. Interspeech 2025, 2025, pp. 1898–1902

  36. [44]

    Enhancing gaze prediction in multi- party conversations via speaker-aware multimodal adaptation,

    M.-C. Lee and Z. Deng, “Enhancing gaze prediction in multi- party conversations via speaker-aware multimodal adaptation,” in Proceedings of the 27th International Conference on Multimodal Interaction, 2025, pp. 200–208

  37. [45]

    How to design a three- stage architecture for audio-visual active speaker detection in the wild,

    O. K ¨op¨ukl¨u, M. Taseska, and G. Rigoll, “How to design a three- stage architecture for audio-visual active speaker detection in the wild,”arXiv preprint arXiv:2106.03932, 2021

  38. [46]

    Ego4d: Around the world in 3,000 hours of egocentric video,

    K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liuet al., “Ego4d: Around the world in 3,000 hours of egocentric video,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18...

  39. [47]

    The ami meeting corpus,

    W. Kraaij, T. Hain, M. Lincoln, and W. Post, “The ami meeting corpus,” inProc. International Conference on Methods and Tech- niques in Behavioral Research, 2005, pp. 1–4

  40. [48]

    Masked face recognition challenge: The insightface track report,

    J. Deng, J. Guo, X. An, Z. Zhu, and S. Zafeiriou, “Masked face recognition challenge: The insightface track report,” inProceed- ings of the IEEE/CVF International Conference on Computer Vi- sion, 2021, pp. 1437–1444

  41. [49]

    Reti- naface: Single-shot multi-level face localisation in the wild,

    J. Deng, J. Guo, E. Ververas, I. Kotsia, and S. Zafeiriou, “Reti- naface: Single-shot multi-level face localisation in the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 5203–5212

  42. [50]

    Sample and com- putation redistribution for efficient face detection,

    J. Guo, J. Deng, A. Lattas, and S. Zafeiriou, “Sample and com- putation redistribution for efficient face detection,”arXiv preprint arXiv:2105.04714, 2021

  43. [51]

    Arcface: Additive an- gular margin loss for deep face recognition,

    J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive an- gular margin loss for deep face recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2019, pp. 4690–4699

  44. [52]

    The hungarian method for the assignment prob- lem,

    H. W. Kuhn, “The hungarian method for the assignment prob- lem,”Naval research logistics quarterly, vol. 2, no. 1-2, pp. 83– 97, 1955

  45. [53]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” inAdvances in Neural Information Processing Systems, I. Guyon, U. V . Luxburg, S. Bengio, H. Wal- lach, R. Fergus, S. Vishwanathan, and R. Garnett...

  46. [54]

    Talknce: Improving active speaker detection with talk- aware contrastive learning,

    C. Jung, S. Lee, K. Nam, K. Rho, Y . J. Kim, Y . Jang, and J. S. Chung, “Talknce: Improving active speaker detection with talk- aware contrastive learning,” inICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. ...

  47. [55]

    Turn-taking in conversational systems and human- robot interaction: A review,

    G. Skantze, “Turn-taking in conversational systems and human- robot interaction: A review,”Computer Speech & Language, vol. 67, p. 101178, 2021

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.