REVIEW 1 major objections 1 minor 55 references
MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild
T0 review · 1 major / 1 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read MuVAP extends voice activity projection with face tracks and role-relative mapping to predict turn-taking from single-camera monaural recordings.
desk verdict MuVAP adds a single-camera multimodal grounding step and a new unedited corpus, but the Role-Relative Projection needs to show it preserves speaker identity for real next-speaker gains rather than just binary shift/hold. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Role-Relative Projection, which maps any N-speaker interaction onto a fixed current versus next floor-holder state to address combinatorial complexity.
What would settle it
Collecting new recordings where single-camera face tracking frequently loses tracks due to movement or occlusion and checking whether prediction accuracy then falls below audio-only baselines.
Extended reading notes
Core claim
MuVAP is a causal multimodal framework that grounds acoustic predictions in face tracks from a single camera view, using Role-Relative Projection to map multiparty interactions onto a fixed current versus next floor-holder state, thereby enabling accurate shift-hold and next-speaker predictions from monaural audio in unedited wild settings, as demonstrated by outperformance of strong baselines on the Audio-Visual Conversation Corpus.
Load-bearing premise
Face tracks from a single camera can reliably ground the acoustic predictions and that the Role-Relative Projection preserves enough multiparty dynamics without speaker-specific modeling.
Editorial extensions
If this is right
- MuVAP enables turn-taking prediction without complex microphone arrays or multi-camera setups.
- It supports human-robot interaction scenarios with simple hardware.
- Performance gains hold for both two- and three-speaker settings on shift-hold and next-speaker tasks.
- The unedited nature of the new corpus avoids artifacts from editing cuts that break causal tracking.
Reading between the lines
- If single-view face tracking works reliably, the same approach could apply to consumer video calls for automatic turn management.
- The role mapping might allow scaling to four or more speakers without exponential growth in model complexity.
- Combining this with existing speech recognition could lead to fully hands-free group conversation agents.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MuVAP, a causal multimodal framework extending Voice Activity Projection by grounding acoustic predictions in face tracks from a single camera and monaural audio. It proposes Role-Relative Projection to map any N-speaker interaction onto a fixed current-versus-next floor-holder state to address combinatorial complexity, introduces the 31-hour Audio-Visual Conversation Corpus of unedited single-camera recordings, and claims that MuVAP outperforms strong baselines on Shift-Hold and next-speaker prediction tasks in two- and three-speaker settings.
Significance. If the empirical claims hold after addressing the output-space concern, the work would be significant for practical human-robot interaction by enabling speaker-aware turn-taking from minimal hardware. The new unedited dataset fills a gap for causal audiovisual modeling, and the multimodal grounding approach could generalize beyond current microphone-array or multi-camera setups.
major comments (1)
- [Abstract] Abstract: The central claim that Role-Relative Projection enables 'next-speaker prediction' while mapping N-speaker dynamics onto a fixed current-versus-next state requires explicit clarification on output representation. If the mapping assigns a single undifferentiated 'next' label without retaining speaker identity or role distinctions among candidate next-speakers, the task reduces to binary hold/shift detection; reported gains on next-speaker prediction would then be artifacts of the reduced output space rather than evidence that multimodal grounding identifies specific speakers.
minor comments (1)
- The abstract states that evaluations demonstrate outperformance but provides no quantitative metrics, baseline details, or ablation results; these must be presented with error analysis to allow verification of the performance claims.
Simulated Author's Rebuttal
We thank the referee for the careful reading and for highlighting the need for explicit clarification on the output representation of Role-Relative Projection. We address the concern directly below and will revise the manuscript accordingly.
read point-by-point responses
-
Referee: [Abstract] Abstract: The central claim that Role-Relative Projection enables 'next-speaker prediction' while mapping N-speaker dynamics onto a fixed current-versus-next state requires explicit clarification on output representation. If the mapping assigns a single undifferentiated 'next' label without retaining speaker identity or role distinctions among candidate next-speakers, the task reduces to binary hold/shift detection; reported gains on next-speaker prediction would then be artifacts of the reduced output space rather than evidence that multimodal grounding identifies specific speakers.
Authors: We agree that the abstract requires explicit clarification on this point. Role-Relative Projection reduces combinatorial complexity by mapping arbitrary speaker configurations to a fixed current-versus-next floor-holder state; however, the output is not an undifferentiated binary label. Because predictions are grounded in the face tracks extracted from the single-camera view, the model produces per-track probabilities that identify which specific visual track (and therefore which individual) corresponds to the current floor-holder and which corresponds to the next floor-holder. This retains speaker identity and is distinct from pure hold/shift classification. We will revise the abstract to state the output representation explicitly (i.e., that next-speaker prediction is performed over the set of detected face tracks). revision: yes
Circularity Check
No circularity; claims rest on new dataset and empirical baselines
full rationale
The paper introduces Role-Relative Projection as a modeling choice to reduce combinatorial complexity and evaluates MuVAP on Shift-Hold and next-speaker tasks using a newly collected unedited dataset. No equations or derivations reduce predictions to fitted inputs by construction, no load-bearing self-citations are invoked to justify uniqueness, and no ansatz is smuggled via prior work. The central results are presented as empirical outperformance against baselines rather than tautological redefinitions. This is the expected non-finding for an applied modeling paper whose validity can be checked externally via the reported metrics.
Assumptions & free parameters
Cite this review
Pith. "Pith review of MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild." pith.science (2026). https://pith.science/paper/GVJSP2AW
@misc{pith2026260616731,
author = {Pith},
title = {Pith review of: MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild},
year = {2026},
howpublished = {\url{https://pith.science/paper/GVJSP2AW}},
note = {Machine review of arXiv:2606.16731}
}
read the original abstract
Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios. We introduce MuVAP, a causal multimodal framework that extends Voice Activity Projection by grounding acoustic predictions in face tracks, enabling speaker-aware turn-taking predictions from a monaural audio stream and a single camera view. To address the combinatorial complexity of modeling multiple speakers, we propose Role-Relative Projection, which maps any N-speaker interaction onto a fixed current versus next floor-holder state. Because existing audiovisual datasets contain disruptive editing cuts that break causal tracking, we introduce the Audio-Visual Conversation Corpus, a 31-hour dataset of unedited, single-camera multiparty conversations. Evaluations demonstrate that MuVAP outperforms strong baselines on Shift-Hold and next-speaker prediction tasks across two- and three-speaker settings.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild
Introduction Turn-taking is a fundamental aspect of conversation. Since peo- ple cannot easily speak and listen simultaneously, they must coordinate their turns through a complex exchange of cues [1, 2, 3]. For example, a syntactically or semantically incom- plete phrase may signal a turn hold, whereas a complete phrase may be turn-yielding [4]. A filled ...
work page Pith review arXiv 2026
-
[2]
Active Speaker Detection vs
Related Work 2.1. Active Speaker Detection vs. Turn-Taking We distinguish our work from multimodal Active Speaker De- tection (ASD), which aims to identify the active speaker among multiple faces in a video stream. Current ASD architectures, such as TalkNet [17] and LoCoNet [18], achieve impressive pre- cision on the A V A-ActiveSpeaker (A V A-AS) benchma...
2000
-
[3]
Speakers use these visual sig- nals alongside prosodic contours to navigate turn-taking [25]
to identify floor acquisition. Speakers use these visual sig- nals alongside prosodic contours to navigate turn-taking [25]. While sentence-level NSP [26] and event-based visual frame- Dataset Modality Duration Source Lang. Unedited Module Fisher Audio 1958h7m Telephone English✓V AP A V A-AS A V 38h5m Movie Multi✗ASD MSDWild A V 80h18m Vlogs/Wild Multi✗AS...
-
[4]
As illustrated in Figure 2, these sources contain editing artifacts like jump cuts that disrupt the temporal flow of conversation
The A VCC Dataset Currently, publicly available ASD datasets [31, 16] suffer from domain gaps that limit their utility for turn-taking prediction in interactive settings (such as HRI). As illustrated in Figure 2, these sources contain editing artifacts like jump cuts that disrupt the temporal flow of conversation. This prevents models from learning the tr...
-
[5]
cur- rent floor holder
Model We proposeMuV AP(Multimodal Multiparty V oice Activity Projection), a causal framework designed to forecast turn-taking dynamics for an arbitrary number of speakers. Given a single- channel audio waveformA ∈R T and a set of face tracks V={v 1, . . . , vN }forNdetected speakers, the model predicts two probability distributions at each time stept: • G...
-
[6]
The speaker with the highest activity is designated as the current/past floor holder
Current Holder (S curr): We rank allNspeakers by their to- tal activity in the history bins. The speaker with the highest activity is designated as the current/past floor holder
-
[7]
social conductor
Next Holder (S next): We rank the remainingN−1speak- ers by their activity in the future bins. The highest-ranked remaining speaker is designated as the primary next speaker. This reduction transforms the complex multiparty dynamic into a fixed pairwise state{S curr, Snext}. Following the standard V AP encoding method [11], we extract the binary activity ...
-
[8]
Each dataset is mapped to specific mod- ules to facilitate multi-stage training as shown in Table 1
Implementation Our pipeline utilizes five distinct datasets: Fisher (Parts 1 & 2), MSDWild, W ASD, A V A-ActiveSpeaker, and our newly intro- duced A VCC dataset. Each dataset is mapped to specific mod- ules to facilitate multi-stage training as shown in Table 1. All modules were optimized using a Cosine Annealing learning rate scheduler with a linear warm...
Show all 55 references
-
[9]
To measure the usefulness of these learned embeddings, we de- fine a set of downstream tasks for evaluation
Downstream tasks In the spirit of the original V AP model [11], MuV AP is trained to learn generic embeddings capturing turn-taking dynamics. To measure the usefulness of these learned embeddings, we de- fine a set of downstream tasks for evaluation. 6.1. Shift-Hold vs. Next S...
-
[10]
Results are reported as mean ± 95% confidence intervals over ten runs using random seeds 42–51
Results We evaluate Shift-Hold Prediction using Macro-F1 (due to class imbalance), and Next Speaker Prediction (NSP) using accuracy. Results are reported as mean ± 95% confidence intervals over ten runs using random seeds 42–51. We compare MuV AP against two baselines:Majority...
-
[11]
Discussion The results reveal a clear division of labor between the modal- ities in our model. Acoustic, linguistic, and prosodic cues pro- vide the global rhythm of coordination, while the visual channel acts as an essential separation anchor during competitive over- lap and ...
-
[12]
To address the scaling bottlenecks of joint modeling, we proposed a Role- Relative Projection that compresses complex group dynamics into a scalable pairwise state
Conclusion We introduced MuV AP, a multimodal framework for predict- ing turn-taking in unconstrained multiparty settings. To address the scaling bottlenecks of joint modeling, we proposed a Role- Relative Projection that compresses complex group dynamics into a scalable pairw...
-
[13]
The computations and data handling were enabled by the Berzelius resource provided by the Knut and Alice Wallenberg Foundation at the National Supercomputer Centre
Acknowledgment This work was supported by the Wallenberg AI, Autonomous Systems and Software Program (W ASP) funded by the Knut and Alice Wallenberg Foundation (KAW), and the Swedish Re- search Council project 2020-03812. The computations and data handling were enabled by the ...
2020
-
[14]
These tools were not used for writing any major parts of the paper
Generative AI Use Disclosure Generative AI tools were used in this paper exclusively for edit- ing and polishing the text to remove grammatical errors. These tools were not used for writing any major parts of the paper
-
[15]
Some signals and rules for taking speaking turns in conversations,
S. Duncan, “Some signals and rules for taking speaking turns in conversations,”Journal of Personality and Social Psychology, vol. 23, pp. 283–292, 08 1972
1972
-
[16]
A simplest system- atics for the organization of turn-taking for conversation,
H. Sacks, E. A. Schegloff, and G. Jefferson, “A simplest system- atics for the organization of turn-taking for conversation,”Lan- guage, vol. 50, no. 4, pp. 696–735, 1974
1974
-
[17]
On the structure of speaker–auditor interaction dur- ing speaking turns,
S. Duncan Jr, “On the structure of speaker–auditor interaction dur- ing speaking turns,”Language in Society, vol. 3, no. 2, pp. 161– 180, 1974
1974
-
[18]
Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,
C. E. Ford and S. A. Thompson, “Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,”Studies in interactional sociolinguistics, vol. 13, pp. 134–184, 1996
1996
-
[19]
H. H. Clark,Using language. Cambridge University Press, 1996
1996
-
[20]
Listeners’ responses to filled pauses in relation to floor apportionment,
P. Ball, “Listeners’ responses to filled pauses in relation to floor apportionment,”British Journal of Social & Clinical Psychology, vol. 14, pp. 423–424, 11 1975
1975
-
[21]
Duncan and D
S. Duncan and D. Fiske,Face-to-face interaction: research, meth- ods and theory. Wiley, 1977
1977
-
[22]
Universals and cultural variation in turn-taking in conversation,
T. Stivers, N. J. Enfield, P. Brown, C. Englert, M. Hayashi, T. Heinemann, G. Hoymann, F. Rossano, J. P. De Ruiter, K.-E. Yoonet al., “Universals and cultural variation in turn-taking in conversation,”Proceedings of the National Academy of Sciences, vol. 106, no. 26, pp. 10 58...
2009
-
[23]
Timing in turn-taking and its im- plications for processing models of language,
S. C. Levinson and F. Torreira, “Timing in turn-taking and its im- plications for processing models of language,”Frontiers in psy- chology, vol. 6, p. 136034, 2015
2015
-
[24]
Conversational interaction with social robots,
G. Skantze, “Conversational interaction with social robots,” in Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction, 2021, p. 717
2021
-
[25]
V oice Activity Projection: Self- supervised Learning of Turn-taking Events,
E. Ekstedt and G. Skantze, “V oice Activity Projection: Self- supervised Learning of Turn-taking Events,” inProc. Interspeech 2022, 2022, pp. 5190–5194
2022
-
[26]
Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,
K. Inoue, D. Lala, G. Skantze, and T. Kawahara, “Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,” inProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the Association for Computational Lin...
2025
-
[27]
Applying general turn-taking models to conversational human-robot interaction,
G. Skantze and B. Irfan, “Applying general turn-taking models to conversational human-robot interaction,” inProceedings of the ACM/IEEE International Conference on Human-Robot Interac- tion (HRI), 2025
2025
-
[28]
Multimodal turn analysis and prediction for multi-party conversations,
M.-C. Lee, M. Trinh, and Z. Deng, “Multimodal turn analysis and prediction for multi-party conversations,” inProceedings of the 25th International Conference on Multimodal Interaction, 2023, pp. 436–444
2023
-
[29]
Triadic multi- party voice activity projection for turn-taking in spoken dialogue systems,
M. Elmers, K. Inoue, D. Lala, and T. Kawahara, “Triadic multi- party voice activity projection for turn-taking in spoken dialogue systems,” inProc. Interspeech 2025, 2025, pp. 3015–3019
2025
-
[30]
MSDWild: Multi- modal Speaker Diarization Dataset in the Wild,
T. Liu, S. Fan, X. Xiang, H. Song, S. Lin, J. Sun, T. Han, S. Chen, B. Yao, S. Liu, Y . Wu, Y . Qian, and K. Yu, “MSDWild: Multi- modal Speaker Diarization Dataset in the Wild,” inProc. Inter- speech 2022, 2022, pp. 1476–1480
2022
-
[31]
Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection,
R. Tao, Z. Pan, R. K. Das, X. Qian, M. Z. Shou, and H. Li, “Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection,” inProceedings of the 29th ACM International Conference on Multimedia, 2021, pp. 3927– 3935
2021
-
[32]
Loconet: Long-short con- text network for active speaker detection,
X. Wang, F. Cheng, and G. Bertasius, “Loconet: Long-short con- text network for active speaker detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2024, pp. 18 462–18 472
2024
-
[33]
A V A Active Speaker: An audio-visual dataset for active speaker detection,
J. Roth, S. Chaudhuri, O. Klejch, R. Marvin, A. Gallagher, L. Kaver, S. Ramaswamy, A. Stopczynski, C. Schmid, Z. Xiet al., “A V A Active Speaker: An audio-visual dataset for active speaker detection,” inICASSP 2020 IEEE International Conference on Acoustics, Speech and Signal ...
2020
-
[34]
How much does prosody help turn- taking? Investigations using voice activity projection models,
E. Ekstedt and G. Skantze, “How much does prosody help turn- taking? Investigations using voice activity projection models,” in Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. Edinburgh, UK: Association for Computational Linguist...
2022
-
[35]
Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,
R. Ishii, K. Otsuka, S. Kumano, M. Matsuda, and J. Yamato, “Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,” inProceedings of the 15th ACM on Inter- national conference on multimodal interaction, 2013, pp. 79–86
2013
-
[36]
Gaze-enhanced multimodal turn-taking prediction in triadic conversations,
S. Heo, C. Miller, C. Murdock, and M. Proulx, “Gaze-enhanced multimodal turn-taking prediction in triadic conversations,” in Proc. Interspeech 2025, 2025, pp. 1068–1072
2025
-
[37]
Multimodal voice activ- ity prediction: Turn-taking events detection in expert-novice con- versation,
K. Onishi, H. Tanaka, and S. Nakamura, “Multimodal voice activ- ity prediction: Turn-taking events detection in expert-novice con- versation,” inProceedings of the 11th International Conference on Human-Agent Interaction, 2023, pp. 13–21
2023
-
[38]
Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,
R. Ishii, K. Otsuka, S. Kumano, R. Higashinaka, and J. Tomita, “Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,”Multimodal Tech- nologies and Interaction, vol. 3, no. 4, p. 70, 2019
2019
-
[39]
Who’s next? speaker-selection mech- anisms in multiparty dialogue,
V . Petukhova and H. Bunt, “Who’s next? speaker-selection mech- anisms in multiparty dialogue,” inProceedings of the Workshop on the Semantics and Pragmatics of Dialogue, 2009, pp. 19–26
2009
-
[40]
A computational study on sentence-based next speaker prediction in multiparty conversa- tions,
M.-C. Lee, W. A. Li, and Z. Deng, “A computational study on sentence-based next speaker prediction in multiparty conversa- tions,” inProceedings of the 24th ACM International Conference on Intelligent Virtual Agents, 2024, pp. 1–4
2024
-
[41]
Multimodal continuous turn-taking prediction using multiscale rnns,
M. Roddy, G. Skantze, and N. Harte, “Multimodal continuous turn-taking prediction using multiscale rnns,” inProceedings of the 20th ACM International Conference on Multimodal Interac- tion, 2018, pp. 186–190
2018
-
[42]
Visual cues enhance predictive turn- taking for two-party human interaction,
S. O. Russell and N. Harte, “Visual cues enhance predictive turn- taking for two-party human interaction,” inFindings of the As- sociation for Computational Linguistics: ACL 2025, 2025, pp. 209–221
2025
-
[43]
Multi-channel sequence-to-sequence neural diarization: Experimental results for the misp 2025 challenge,
M. Cheng, F. Su, C. Li, J. Liu, and M. Li, “Multi-channel sequence-to-sequence neural diarization: Experimental results for the misp 2025 challenge,” inProc. Interspeech 2025, 2025, pp. 1898–1902
2025
-
[44]
Enhancing gaze prediction in multi- party conversations via speaker-aware multimodal adaptation,
M.-C. Lee and Z. Deng, “Enhancing gaze prediction in multi- party conversations via speaker-aware multimodal adaptation,” in Proceedings of the 27th International Conference on Multimodal Interaction, 2025, pp. 200–208
2025
-
[45]
How to design a three- stage architecture for audio-visual active speaker detection in the wild,
O. K ¨op¨ukl¨u, M. Taseska, and G. Rigoll, “How to design a three- stage architecture for audio-visual active speaker detection in the wild,”arXiv preprint arXiv:2106.03932, 2021
2021
-
[46]
Ego4d: Around the world in 3,000 hours of egocentric video,
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liuet al., “Ego4d: Around the world in 3,000 hours of egocentric video,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18...
2022
-
[47]
The ami meeting corpus,
W. Kraaij, T. Hain, M. Lincoln, and W. Post, “The ami meeting corpus,” inProc. International Conference on Methods and Tech- niques in Behavioral Research, 2005, pp. 1–4
2005
-
[48]
Masked face recognition challenge: The insightface track report,
J. Deng, J. Guo, X. An, Z. Zhu, and S. Zafeiriou, “Masked face recognition challenge: The insightface track report,” inProceed- ings of the IEEE/CVF International Conference on Computer Vi- sion, 2021, pp. 1437–1444
2021
-
[49]
Reti- naface: Single-shot multi-level face localisation in the wild,
J. Deng, J. Guo, E. Ververas, I. Kotsia, and S. Zafeiriou, “Reti- naface: Single-shot multi-level face localisation in the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 5203–5212
2020
-
[50]
Sample and com- putation redistribution for efficient face detection,
J. Guo, J. Deng, A. Lattas, and S. Zafeiriou, “Sample and com- putation redistribution for efficient face detection,”arXiv preprint arXiv:2105.04714, 2021
2021
-
[51]
Arcface: Additive an- gular margin loss for deep face recognition,
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive an- gular margin loss for deep face recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2019, pp. 4690–4699
2019
-
[52]
The hungarian method for the assignment prob- lem,
H. W. Kuhn, “The hungarian method for the assignment prob- lem,”Naval research logistics quarterly, vol. 2, no. 1-2, pp. 83– 97, 1955
1955
-
[53]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” inAdvances in Neural Information Processing Systems, I. Guyon, U. V . Luxburg, S. Bengio, H. Wal- lach, R. Fergus, S. Vishwanathan, and R. Garnett...
2017
-
[54]
Talknce: Improving active speaker detection with talk- aware contrastive learning,
C. Jung, S. Lee, K. Nam, K. Rho, Y . J. Kim, Y . Jang, and J. S. Chung, “Talknce: Improving active speaker detection with talk- aware contrastive learning,” inICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. ...
2024
-
[55]
Turn-taking in conversational systems and human- robot interaction: A review,
G. Skantze, “Turn-taking in conversational systems and human- robot interaction: A review,”Computer Speech & Language, vol. 67, p. 101178, 2021
2021
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.