REVIEW 3 major objections 5 minor 48 references
Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper argues that the near-absence of incremental dialogue management is the bottleneck for natural human-robot conversation, and that large language models cannot fill the gap because their processing is monotonic.
desk verdict A genuinely useful review whose central claim is more brittle than it looks: the 'very little incremental DM' finding depends on where you draw the DM boundary, and the paper draws it on both sides. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing distinction is restart-incremental versus update-incremental processing. A restart-incremental model re-reads the growing prefix on each new input and recomputes; an update-incremental model maintains a state that is revised with each word. The paper also leans on the incremental unit framework, in which each piece of information produced by a module is a node in a network that can be added, revoked, or committed. Revocation is the mechanism that makes non-monotonic dialogue possible, and it is exactly what the dialogue manager must absorb: when a recognized hypothesis such as gray is replaced by green, the manager's state and action plans need to change. This is the lens through which the authors judge LLMs, whose autoregressive output is incremental but whose committed tokens cannot be taken back, and through which they frame their requirements.
What would settle it
A systematic bibliographic search that counts systems with explicit, timing-aware decision-making modules operating at word-level increments, regardless of whether the authors label them dialogue managers, would settle it: if such systems are numerous, the scarcity claim is false. A smaller behavioral test is to build two robots, one whose timing decisions are made by the dialogue manager and one where timing is fixed by ASR endpointing, and measure perceived naturalness and task completion; if no advantage appears, the paper's central motivation fails.
Extended reading notes
Core claim
At the center is a scarcity result. Surveying the literature on interactive systems, the authors find ample incremental work in automatic speech recognition, natural language understanding, natural language generation, and turn-taking, but state plainly that there is very little research on incremental dialogue management. They define dialogue management as dialogue state tracking plus action and response selection, and their desiderata explicitly assign it the additional responsibility of timing: deciding when to act, not just what to do. The papers that do exist, including an incremental information-state manager, the DIUM manager that produces self-corrections, a time-board manager, and policy-learning studies in fast-paced games, are described as promising but largely abandoned. The paper's positive claim is that this neglected component becomes the bottleneck when robots interact with people, because a robot must be able to act before the user finishes speaking, revise a plan when a recognized word is revoked, and signal understanding through early action.
Load-bearing premise
The scarcity finding rests on a definitional premise: dialogue management is defined as state tracking plus action selection, and timing decisions are assigned to the dialogue manager; if timing and revision are instead handled by adjacent components such as ASR endpointing, NLG output buffers, or robot-level replanners, the claimed gap could be an artifact of that division of labor.
Editorial extensions
If this is right
- Incremental dialogue management would let a robot begin moving toward an object while the user is still finishing the request, using the motion itself as a backchannel that signals understanding.
- A proper incremental dialogue manager should replace reliance on ASR endpointing: pauses would become one signal among many, with the manager deciding when silence or content is enough to act.
- Evaluating dialogue managers incrementally requires new data: datasets annotated at the word level for when a decision should be made, not only for what the final decision is.
- LLM-based dialogue managers must be made update-incremental, or paired with an external mechanism that can retract and repair output, rather than restarting on every growing prefix.
Reading between the lines
- The reported scarcity may partly reflect where the authors draw the line: if timing and revision are distributed across ASR endpointing, NLG output buffers, and robot-level replanners, then the literature contains more incremental decision-making than a narrow dialogue-manager search will find.
- The desiderata imply an evaluation metric the community does not yet have: a cost-sensitive score that rewards acting early and correcting course, rather than only final task success or final response accuracy.
- The robot replanning literature, which already handles revoking and replacing action plans under new observations, is a natural source of mechanisms for the revision half of incremental dialogue management.
- If full-duplex LLM dialogue agents continue to improve, the paper's requirements suggest a concrete test: does the model decide when to act from a maintained, revisable state, or does it simply wait for a complete utterance?
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper surveys incremental spoken dialogue systems with a focus on dialogue management (DM) and argues that, unlike ASR, NLU, and NLG, incremental DM has received very little research attention. The authors motivate incrementality by the needs of human-robot interaction, explain frameworks such as Incremental Units and the restart vs. update distinction, review incremental modules and full systems, and propose desiderata for incremental DM on robotic platforms, including responsibility for timing, acting on incomplete information, and making fast, small decisions concurrently.
Significance. If the central scarcity claim is robust, the paper identifies a genuinely under-explored area with practical importance: incremental decision-making for dialogue and robot action management in the age of LLMs. The paper is useful as a synthesis: it collects a broad set of references, clearly explains the IU framework and the restart/update distinction, connects incremental processing to non-monotonicity and replanning, and offers concrete desiderata that could guide future work. The authors are appropriately careful in places, phrasing the finding as 'very little research' rather than 'none'. However, the main contribution is the gap claim itself, and that claim currently rests on an implicit and internally inconsistent boundary between DM and adjacent modules, so the paper's significance depends on the strength of that boundary.
major comments (3)
- [§3.1.3, §3.2, §4.1] The paper's scope for DM is internally inconsistent. Section 3.1.3 explicitly states that DM is about not only what to do but also when to do it, and Section 4.1 lists 'Incremental DM is responsible for timing' as a desideratum. Under this definition, several systems reviewed in Section 3.1.3—Raux and Eskenazi (2009) deciding whether to grab, release, wait for, or keep the floor; end-of-turn prediction models; Roddy and Harte (2020) generating response timings; and voice activity projection methods—are incremental decision-making contributions. Yet Section 3.2 excludes these works from its count of incremental DM research and concludes that very little such research exists. The scarcity claim is therefore an artifact of the classification unless the paper either includes timing-focused work as incremental DM or explicitly justifies a narrower definition that excludes timing.
- [§1, §3.2] The central claim is a quantitative statement about the literature, but the review does not report a systematic selection method, search protocol, inclusion criteria, or corpus size. The abstract states 'We find that there is very little research on incremental dialogue management,' which reads as an empirical finding, yet no procedure is described for identifying the set of works surveyed. The claim is not independently verifiable and could reflect the authors' selection rather than the actual state of the field. The authors should either soften the claim to the set of works they examined or add an explicit methodology and inclusion criteria.
- [§2.2.1, §3.1.2, §5] The title and motivation emphasize robot action management, but the review does not survey the robotics replanning literature that it identifies as relevant. Section 2.2.1 and Section 3.1.2 mention replanning (Cashmore et al., 2019; Garrett et al., 2020; Zhou et al., 2023) as an analogue to revising DM decisions, and the conclusion argues that incremental decision-making is needed for robots, yet no incremental action management or replanning work is actually reviewed. Without such coverage, the paper cannot support its implicit claim that incremental decision-making in embodied settings is under-researched; the gap may be a result of excluding this literature rather than a property of the field.
minor comments (5)
- [§1] There is a typo in 'people antrhopomorphize robots'—should be 'anthropomorphize'.
- [§2.1] The phrase 'as the first word is uttered is uttered' contains a duplicated 'is uttered' and should be corrected.
- [§2.3.4] The sentence 'Ghigi et al. (2014b) showed that that incremental dialogue strategy' has a doubled 'that'.
- [§4.1] The text 'V oice Activity Detection' has unusual spacing in 'Voice Activity Detection'; this appears to be a formatting artifact and should be fixed.
- [Figure 2] The figure caption includes a raw URL instead of a formal citation to the Localized Narrative dataset paper; a proper reference would be more appropriate.
Circularity Check
No significant circularity: the gap claim is a survey assessment, not a derived or self-citation-forced result.
full rationale
This paper is a literature review rather than a derivation, so the standard circularity patterns (fitted inputs renamed as predictions, self-citation chains forcing a conclusion, uniqueness theorems imported from the authors) do not apply. The central claim that there is very little research on incremental dialogue management is a survey-based assessment supported by a broad set of external references (e.g., Buß et al. 2010, Selfridge et al. 2012, Yaghoubzadeh et al. 2015, Zilka and Jurčíček 2015), and the definition of dialogue management as covering state tracking, action/response selection, and timing is grounded in the externally cited Information State approach of Traum and Larsson (2003), not only in the authors' own work. Self-citations such as Schlangen and Skantze's IU framework and Lison and Kennington's prior DM papers are background or toolkit citations and are not used to infer the gap. A possible objection that turn-taking and end-of-turn prediction work reviewed in Section 3.1.3 should count as incremental DM is a classification or validity concern about the survey, not a circular reduction of the paper's claim to its own definition by construction.
Assumptions & free parameters
assumptions (3)
- domain assumption Incremental processing (word-level or below, with non-monotonic revision) is the appropriate standard for natural human-robot dialogue.
- domain assumption The reviewed literature is representative of the field of incremental dialogue management.
- domain assumption Causal language models are monotonic and trained on complete sentences, making them incapable of incremental revision.
Cite this review
Pith. "Pith review of Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models." pith.science (2026). https://pith.science/paper/XIIL4KX3
@misc{pith2026250100953,
author = {Pith},
title = {Pith review of: Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/XIIL4KX3}},
note = {Machine review of arXiv:2501.00953}
}
read the original abstract
Efforts towards endowing robots with the ability to speak have benefited from recent advancements in natural language processing, in particular large language models. However, current language models are not fully incremental, as their processing is inherently monotonic and thus lack the ability to revise their interpretations or output in light of newer observations. This monotonicity has important implications for the development of dialogue systems for human--robot interaction. In this paper, we review the literature on interactive systems that operate incrementally (i.e., at the word level or below it). We motivate the need for incremental systems, survey incremental modeling of important aspects of dialogue like speech recognition and language generation. Primary focus is on the part of the system that makes decisions, known as the dialogue manager. We find that there is very little research on incremental dialogue management, offer some requirements for practical incremental dialogue management, and the implications of incremental dialogue for embodied, robotic platforms in the age of large language models.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[8]
Multimodal end-of-turn prediction in multi-party meetings
Iwan De Kok and Dirk Heylen. Multimodal end-of-turn prediction in multi-party meetings. In Proceedings of the 2009 international conference on Multimodal interfaces, pages 91–98,
work page 2009
-
[9]
Association for Computational Linguistics. David DeVault and David Traum. A method for the approximation of incremental understanding of explicit utterance meaning using predictive models in finite domains. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pag...
work page 2013
-
[11]
Towards LLM-driven dialogue state tracking
Yujie Feng, Zexin Lu, Bo Liu, Liming Zhan, and Xiao-Ming Wu. Towards LLM-driven dialogue state tracking. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 739–755, Stroudsburg, PA, USA,
work page 2023
-
[12]
Initiating human-robot interactions using incremental speech adaptation
Kerstin Fischer, Lakshadeep Naik, Rosalyn M Langedijk, Timo Baumann, Matou ˇs Jel´ınek, and Oskar Palinko. Initiating human-robot interactions using incremental speech adaptation. InCom- panion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction , HRI ’21 Companion, pages 421–425, New York, NY , USA, March
work page 2021
-
[14]
A syntactic language model based on incremental ccg parsing
Hany Hassan, Khalil Simaan, and Andy Way. A syntactic language model based on incremental ccg parsing. In 2008 IEEE Workshop on Spoken Language Technology, SLT 2008 - Proceedings, pages 205–208. Institute of Electrical and Electronics Engineers,
work page 2008
-
[15]
Kyuyeon Hwang and Wonyong Sung
Association for Computational Linguis- tics. Kyuyeon Hwang and Wonyong Sung. Character-level incremental speech recognition with recurrent neural networks. In ICASSP , IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, volume 2016-May, pages 5335–5339, January
work page 2016
-
[18]
Patrick Kahardipraja, Brielen Madureira, and David Schlangen
Association for Computational Linguis- tics. Patrick Kahardipraja, Brielen Madureira, and David Schlangen. TAPIR: Learning adaptive revision for incremental natural language understanding with a two-pass model. In Findings of the As- sociation for Computational Linguistics: ACL 2023 , pages 4173–4197, Stroudsburg, PA, USA, July
work page 2023
-
[19]
Casey Kennington, Kotaro Funakoshi, Yuki Takahashi, and Mikio Nakano
Associa- tion for Computational Linguistics. Casey Kennington, Kotaro Funakoshi, Yuki Takahashi, and Mikio Nakano. Probabilistic multiparty dialogue management for a game master robot. In Proceedings of the 2014 ACM/IEEE interna- tional conference on Human-robot interaction - HRI ’14 , pages 200–201, Bielefeld, Germany, 2014a. ACM. Casey Kennington, Spyro...
work page 2014
Show all 48 references
-
[20]
Spyros Kousidis, Casey Kennington, Timo Baumann, Hendrik Buschmeier, Stefan Kopp, and Stefan Schlangen
Association for Computational Linguistics. Spyros Kousidis, Casey Kennington, Timo Baumann, Hendrik Buschmeier, Stefan Kopp, and Stefan Schlangen. Situationally aware in-car information presentation using incremental speech gener- ation: Safer, and more effective. In Proceedin...
2014
-
[22]
Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems
26 Bing Liu, G ¨okhan T¨ur, Dilek Hakkani-Tur, Pararth Shah, and Larry Heck. Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems. InProceed- ings of the 2018 Conference of the North American Chapter of the Association for C...
2018
-
[26]
It’s not the way you look, it’s how you move: Validating a general scheme for robot affective behaviour
Jekaterina Novikova, Gang Ren, and Leon Watts. It’s not the way you look, it’s how you move: Validating a general scheme for robot affective behaviour. In Julio Abascal, Simone Barbosa, Mirko Fetter, Tom Gross, Philippe Palanque, and Marco Winckler, editors, Human-Computer Int...
2015
-
[29]
Connect- ing vision and language with localized narratives
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, and Vittorio Ferrari. Connect- ing vision and language with localized narratives. In Computer Vision – ECCV 2020 , volume 12350 LNCS, pages 647–664. Springer International Publishing,
2020
-
[30]
Konferenz Elektronische Sprachsignalverarbeitung 2019, ESSV , pages 134–140, Dresden, March
2019
-
[31]
Entity-consistent end-to-end task-oriented dialogue system with kb retriever
28 Libo Qin, Yijia Liu, Wanxiang Che, Haoyang Wen, Yangming Li, and Ting Liu. Entity-consistent end-to-end task-oriented dialogue system with kb retriever. In Proceedings of the 2019 Con- ference on Empirical Methods in Natural Language Processing and the 9th International Joi...
2019
-
[32]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever
doi: 10.1109/TASLP.2023.3244503. Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision,
2023
-
[33]
Towards universal dialogue state tracking
Liliang Ren, Kaige Xie, Lu Chen, and Kai Yu. Towards universal dialogue state tracking. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages 2780–2786,
2018
-
[34]
Investigating speech features for continuous turn-taking prediction using lstms
M Roddy, Gabriel Skantze, and N Harte. Investigating speech features for continuous turn-taking prediction using lstms. In 19th Annual Conference of the International Speech Communication, INTERSPEECH 2018; Hyderabad International Convention Centre (HICC) Hyderabad; India; 2 S...
2018
-
[35]
Stephen Roller, Y-Lan Boureau, Jason Weston, Antoine Bordes, Emily Dinan, Angela Fan, David Gunning, Da Ju, Margaret Li, Spencer Poff, et al
Association for Computational Linguistics. Stephen Roller, Y-Lan Boureau, Jason Weston, Antoine Bordes, Emily Dinan, Angela Fan, David Gunning, Da Ju, Margaret Li, Spencer Poff, et al. Open-domain conversational agents: Current progress, open problems, and future directions. a...
2006 arXiv
-
[36]
Ethan Selfridge, Iker Arizmendi, Peter Heeman, and Jason Williams
Association for Computational Linguistics. Ethan Selfridge, Iker Arizmendi, Peter Heeman, and Jason Williams. Continuously predicting and processing barge-in during a live spoken dialogue task. In Proceedings of the SIGDIAL 2013 Conference, pages 384–393, Metz, France, August
2013
-
[37]
Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning
Pararth Shah, Dilek Hakkani-Tur, Bing Liu, and G¨okhan T¨ur. Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computa...
2018
-
[38]
Challenging neural dialogue models with nat- ural data: Memory networks fail on incremental phenomena
Igor Shalyminov, Arash Eshghi, and Oliver Lemon. Challenging neural dialogue models with nat- ural data: Memory networks fail on incremental phenomena. In SEMDIAL 2017 (SaarDial) Workshop on the Semantics and Pragmatics of Dialogue, ISCA, August
2017
-
[39]
ProgPrompt: Generating situated robot task plans us- ing large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. ProgPrompt: Generating situated robot task plans us- ing large language models. In 2023 IEEE International Conference on Robotics and Automat...
2023
-
[40]
Guided dialog policy learning: Reward estima- tion for multi-domain task-oriented dialog
Ryuichi Takanobu, Hanlin Zhu, and Minlie Huang. Guided dialog policy learning: Reward estima- tion for multi-domain task-oriented dialog. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Nat...
2019
-
[42]
doi: 10.18653/v1/W18-5032
Association for Com- putational Linguistics. doi: 10.18653/v1/W18-5032. URL https://aclanthology.org/ W18-5032. Herwin Van Welbergen, Dennis Reidsma, and Stefan Kopp. An incremental multimodal realizer for behavior co-articulation and coordination. In Lecture Notes in Computer...
-
[43]
Nicholas Thomas Walker, Stefan Ultes, and Pierre Lison
Association for Computational Lin- guistics. Nicholas Thomas Walker, Stefan Ultes, and Pierre Lison. Graphwoz: Dialogue management with conversational knowledge graphs. arXiv preprint arXiv:2211.12852,
-
[44]
flexdiam – flexible dialogue management for problem- aware, incremental spoken interaction for all user groups (demo paper)
Ramin Yaghoubzadeh and Stefan Kopp. flexdiam – flexible dialogue management for problem- aware, incremental spoken interaction for all user groups (demo paper). Proceedings of the 7th Workshop on Speech and Language Processing for Assistive Technologies (SLPAT 2016),
2016
-
[45]
A robotic agent in a virtual environment that performs situated incremental understanding of navigational utterances
Takashi Yamauchi, Mikio Nakano, and Kotaro Funakoshi. A robotic agent in a virtual environment that performs situated incremental understanding of navigational utterances. In SIGdial 2013, pages 369–371, Metz, France, August
2013
-
[46]
32 Shiquan Yang, Rui Zhang, and Sarah Erfani
Association for Computational Linguistics. 32 Shiquan Yang, Rui Zhang, and Sarah Erfani. Graphdialog: Integrating graph knowledge into end- to-end task-oriented dialogue systems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP),...
2020
-
[47]
Easy things first: Installments improve referring expression gen- eration for objects in photographs
Sina Zarrieß and David Schlangen. Easy things first: Installments improve referring expression gen- eration for objects in photographs. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL 2016),
2016
-
[48]
A probabilistic end-to-end task-oriented dialog model with latent belief states towards semi-supervised learning
Yichi Zhang, Zhijian Ou, Min Hu, and Junlan Feng. A probabilistic end-to-end task-oriented dialog model with latent belief states towards semi-supervised learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9207–921...
2020
-
[2002]
A hybrid approach to dialogue management based on probabilistic rules
Pierre Lison. A hybrid approach to dialogue management based on probabilistic rules. Comput. Speech Lang., 34(1):232–255, 2015a. Pierre Lison. A hybrid approach to dialogue management based on probabilistic rules. Computer Speech & Language, 34(1):232–255, 2015b. Pierre Lison ...
2016
-
[2003]
Pydial: A multi- domain statistical dialogue system toolkit
Stefan Ultes, Lina M Rojas Barahona, Pei-Hao Su, David Vandyke, Dongho Kim, Inigo Casanueva, Paweł Budzianowski, Nikola Mrk ˇsi´c, Tsung-Hsien Wen, Milica Gasic, et al. Pydial: A multi- domain statistical dialogue system toolkit. In Proceedings of ACL 2017, System Demonstratio...
2017
-
[2008]
TurnGPT: A transformer-based language model for predict- ing turn-taking in spoken dialog
Erik Ekstedt and Gabriel Skantze. TurnGPT: A transformer-based language model for predict- ing turn-taking in spoken dialog. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2981–2990, Stroudsburg, PA, USA, November
2020
-
[2009]
Recognising conversa- tional speech: What an incremental ASR should do for a dialogue system and how to get there
Timo Baumann, Casey Kennington, Julian Hough, and David Schlangen. Recognising conversa- tional speech: What an incremental ASR should do for a dialogue system and how to get there. 20 In Proceedings of the International Workshop Series on Spoken Dialogue Systems Technology (I...
2016
-
[2011]
An incremental response policy in an automatic word-game
Eli Pincus and David Traum. An incremental response policy in an automatic word-game. InCEUR Workshop Proceedings, volume 1943, pages 1–8,
1943
-
[2012]
Okko Buß and David Schlangen
Association for Com- putational Linguistics. Okko Buß and David Schlangen. {DIUM} – an incremental dialogue manager that can produce self-corrections. In Proceedings of semdial 2011 (Los Angelogue), Proceedings of semdial 2011 (Los Angelogue),
2011
-
[2013]
Multiwoz–a large-scale multi-domain wizard-of-oz dataset for task- oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Inigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gaˇsi´c. Multiwoz–a large-scale multi-domain wizard-of-oz dataset for task- oriented dialogue modelling. arXiv preprint arXiv:1810.00278,
-
[2014]
The InproTK 2012 release
Timo Baumann and David Schlangen. The InproTK 2012 release. In NAACL-HLT Workshop on Future directions and needs in the Spoken Dialog Community: Tools and Data (SDCTD
2012
-
[2015]
Masaya Ohagi, Tomoya Mizumoto, and Katsumasa Yoshikawa
Springer International Publishing. Masaya Ohagi, Tomoya Mizumoto, and Katsumasa Yoshikawa. Investigation of look-ahead tech- niques to improve response time in spoken dialogue system. In Interspeech 2024, pages 3580– 3584,
2024
-
[2017]
ISBN 21 9781450355438
Association for Computing Machinery. ISBN 21 9781450355438. doi: 10.1145/3136755.3154480. URL https://doi.org/10.1145/ 3136755.3154480. Philip R Cohen and Lucian Galescu. A planning-based explainable collaborative dialogue system. arXiv [cs.AI], February
-
[2018]
Incremental processing in the age of non-incremental encoders: An empirical assessment of bidirectional models for incremental NLU
Brielen Madureira and David Schlangen. Incremental processing in the age of non-incremental encoders: An empirical assessment of bidirectional models for incremental NLU. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pages 357–374,
2020
-
[2019]
User sim- ulation in dialogue systems using inverse reinforcement learning
Senthilkumar Chandramohan, Matthieu Geist, Fabrice Lefevre, and Olivier Pietquin. User sim- ulation in dialogue systems using inverse reinforcement learning. In Interspeech 2011, pages 1025–1028,
2011
-
[2020]
Towards incremental transformers: An empirical analysis of transformer models for incremental NLU
Patrick Kahardipraja, Brielen Madureira, and David Schlangen. Towards incremental transformers: An empirical analysis of transformer models for incremental NLU. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages 1178–1189, Online ...
2021
-
[2021]
Caelan Reed Garrett, Chris Paxton, Tomas Lozano-Perez, Leslie Pack Kaelbling, and Dieter Fox
Association for Computing Ma- chinery. Caelan Reed Garrett, Chris Paxton, Tomas Lozano-Perez, Leslie Pack Kaelbling, and Dieter Fox. Online replanning in belief space for partially observable task and motion problems. In 2020 IEEE International Conference on Robotics and Autom...
2020
-
[2022]
A finite-state turn-taking model for spoken dialog systems
Antoine Raux and Maxine Eskenazi. A finite-state turn-taking model for spoken dialog systems. In Proceedings of human language technologies: The 2009 annual conference of the North Ameri- can chapter of the association for computational linguistics, pages 629–637,
2009
-
[2023]
doi: 10.1162/tacl a 00534
ISSN 2307-387X. doi: 10.1162/tacl a 00534. URL https://doi.org/10.1162/tacl_a_00534. Yuya Chiba and Ryuichiro Higashinaka. Investigating the impact of incremental processing and voice activity projection on spoken dialogue systems. In Proceedings of the 31st International Conf...
-
[2024]
Towards deep end-of-turn prediction for situated spoken dialogue systems
Angelika Maier, Julian Hough, and David Schlangen. Towards deep end-of-turn prediction for situated spoken dialogue systems. In Proc. Interspeech 2017, pages 1676–1680, 2017a. Angelika Maier, Julian Hough, and David Schlangen. Towards deep end-of-turn prediction for situated s...
2017
-
[2025]
Real-time and continuous turn-taking prediction using voice activity projection
Koji Inoue, Bing’er Jiang, Erik Ekstedt, Tatsuya Kawahara, and Gabriel Skantze. Real-time and continuous turn-taking prediction using voice activity projection. arXiv [cs.CL], January 2024a. 24 Koji Inoue, Divesh Lala, Gabriel Skantze, and Tatsuya Kawahara. Yeah, un, oh: Conti...
2019
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.