REVIEW 4 major objections 5 minor 91 references
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper argues that symbolic annotations over multi-floor human-robot dialogue can make the dimensions of meaning needed for common ground accessible to autonomous systems, enabling flexible two-way dialogue for robot teams.
desk verdict A solid, transparent corpus-and-schema paper whose abstract oversells the autonomous-dialogue payoff; the missing Experiment 4 results are the obvious fix. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the three-layer annotation schema over multi-floor dialogue. The first layer, Dialogue-AMR, augments Abstract Meaning Representation with a speech-act root (e.g., command-SA, offer-SA), tense and aspect features, and a normalized lexicon in which varied surface forms like 'turn,' 'rotate,' and 'pivot' map to a single robot concept such as Rotation, so that unconstrained language can be distilled into executable behavior parameters. The second layer, Dialogue Structure, segments the dialogue into Transaction Units, which are clusters of utterances from multiple speakers and floors that together realize one intent, and labels each utterance with a relation type and an antecedent, including translation relations that capture how the dialogue manager passes instructions to the navigation module and relays acknowledgements back. The third layer, Visual Context, annotates photo-request strategies and LIDAR exploration maps, linking language to the shared visual environment. Together these layers are intended to give a dialogue system the information it needs to establish common ground and initiate repairs when the commander's and robot's views of the environment diverge.
What would settle it
Re-run the original SCOUT Commander utterances through an automated pipeline that replaces the two human wizards with an automated dialogue manager using the dialogue-structure annotations and an autonomous navigation module using the Dialogue-AMR-derived behavior specifications, then measure how often the robot executes the commander's intent successfully compared with the wizard-mediated original. If the automated execution rate falls far below the wizard-mediated rate on the same utterances, the corpus would not transfer to real autonomous components; conversely, matching or exceeding wizard performance would support the paper's central claim.
Extended reading notes
Core claim
The paper's central claim is that a robot can only act as a cooperative partner in a physically situated task if it can access several distinct dimensions of meaning at once: propositional content and illocutionary force within a single utterance, the discourse relations between utterances across multiple conversational floors, and the grounding of language in images and LIDAR. The authors demonstrate this through a multi-layer annotation of the SCOUT corpus, in which Commander utterances are parsed into Dialogue-AMR graphs with a speech-act root, grouped into Transaction Units with typed antecedent relations (including translations between the Commander-DM floor and the DM-RN floor), and aligned with photo-request strategies and annotated LIDAR exploration maps. They further report that these annotations have been used to train three successive dialogue systems, ScoutBot, MultiBot, and JUDI, which accept natural-language navigation instructions and respond with feedback or clarification, with grounding added through an AMR-based pipeline. The claim is therefore not just that the annotations describe human-robot dialogue, but that they make the dimensions of meaning accessible to autonomous systems in a form that supports implementation.
Load-bearing premise
The load-bearing premise is that the human wizards who stood in for the dialogue manager and the robot navigator during data collection interpret and execute commander instructions about as well as the automated components they stand in for, so that the annotated patterns will transfer to real deployed systems.
Editorial extensions
If this is right
- A system built from these annotations can take a commander's natural-language instruction, distill it to a robot behavior and its parameters through Dialogue-AMR, and pass it to a separate navigation module, mirroring the multi-floor architecture of a deployed system.
- The fully annotated dialogue-structure data, covering all 89,056 utterances in SCOUT, lets a dialogue manager learn instruction-response pairs, so it can provide feedback or ask for clarification without hand-written rules.
- The visual-context annotations give a robot the ability to notice when a commander's last photo no longer matches the robot's current position, making repair dialogue possible.
- The progression from ScoutBot to MultiBot to JUDI is evidence that the annotation approach generalizes from one robot to heterogeneous robot teams and from lab conditions to situations without a cloud connection.
Reading between the lines
- Editorial inference: the multi-floor 'translation' relations are a general mechanism for any relayed instruction chain, so the schema could find use in human-only teams where a coordinator passes orders to field operators, a setting the paper does not discuss.
- Editorial inference: if the Dialogue-AMR normalization is as consistent as reported, it could be used not only to parse commands but to generate clarification questions, for instance composing 'Do you mean the door on the left?' from the speech-act and action parameters.
- Editorial inference: the paper's finding that commanders who used the cardinal or degrees-of-rotation photo strategies were more successful suggests a robot could learn to propose such strategies proactively, but the paper only reports the correlation, not a causal test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes a multi-layer symbolic annotation framework applied to SCOUT, a Wizard-of-Oz (WoZ) human-robot dialogue corpus in a remote search-and-navigation domain. Four layers are presented: Standard-AMR and Dialogue-AMR for utterance-level propositional semantics and illocutionary force; Dialogue Structure annotations (transaction units and relations) for multi-floor dialogue; and two visual-context annotations (photo-requesting strategies and LIDAR Exploration Maps). The paper reports inter-annotator agreement for Dialogue Structure (antecedents α=0.79–0.94, relation types 0.83–0.93, TUs 0.65–0.85 with the modified schema) and Smatch 86.6% for Dialogue-AMR on 290 utterances, and it discusses the ScoutBot, MultiBot, and JUDI dialogue systems as implementations that use the annotations. The central claim is that these annotations make the dimensions of meaning needed for common ground accessible to autonomous systems, enabling flexible two-way dialogue for human-robot teams.
Significance. If the annotation schemas prove usable for developing autonomous dialogue systems, this is a substantial resource: a large (89,056 utterances, 278 dialogues) multi-floor, multi-modal corpus with layered annotations, released on GitHub, plus concrete worked examples and explicit discussion of architectural challenges. Strong points include the reported IAA on dialogue structure, the public corpus release, the reproducibility of the annotation guidelines, and the implemented prototypes (ScoutBot, MultiBot, JUDI) that at least partially exploit the dialogue-structure annotations. However, the visual-context annotation is still small-scale (30 dialogues) with no agreement metrics, and the transfer from WoZ data collection to fully autonomous components is asserted rather than demonstrated, with the systems described as lacking grounding and world models and the AMR-based grounding still ongoing. The paper is best read as a corpus and schema description with an aspirational end-to-end goal, not as a demonstrated autonomous dialogue system.
major comments (4)
- [Section 2, Section 6] The abstract claims the annotations 'enable physical robots to autonomously engage with humans in bi-directional dialogue and navigation,' and Section 2 asserts that the multi-WoZ setup 'mimics how the physical motion component of an autonomous system would not speak directly to a human user.' However, the fourth experiment that replaced the DM-Wizard with an automated dialogue manager is mentioned but no results are reported. The systems described in Section 6.1 (ScoutBot, MultiBot) are retrieval-based and explicitly lack grounding and world models; Section 6.2 states that AMR-based grounding is 'ongoing' with evaluation strategy still to be determined. The evidence therefore does not yet support the strong claim. Please either report results from Experiment 4 (even descriptive comparisons of dialogue-structure distributions or commander behavior) or temper the claims throughout to 'support the development of' and 'progress toward' such systems.
- [Section 5.1, Section 5.2] The Visual Context annotations are presented as a component of the multi-modal common-ground contribution, but no inter-annotator agreement is reported for either annotation task. Photo-requesting strategies were analyzed by one annotator, and Exploration Maps were annotated by one annotator and verified by a second, with no quantitative reliability measure such as Cohen's kappa or Krippendorff's alpha. Given the emphasis on IAA for the other annotation layers, the lack of agreement data makes the reliability of the visual annotations difficult to assess. Please report agreement metrics if available, or explicitly label these as preliminary exploratory analyses whose reliability remains to be established.
- [Section 3.2, Section 3.1.1, Section 4.1] The IAA figures for Dialogue-AMR and Dialogue Structure are computed on the same SCOUT corpus on which the annotation guidelines were iteratively refined, as described in Sections 3.1.1 and 4.1. This measures internal consistency of the final guidelines on the development corpus, not external validity or generalizability to other domains or to fully autonomous behavior. The paper does not currently acknowledge this limitation. Please add an explicit statement of this point, and, if available, report any held-out or cross-domain annotation results (for example, the Minecraft Dialogue Corpus annotation mentioned in Section 3.2) that would speak to the schema's transferability.
- [Section 3.3, Table 3, Table A1, Figure 11] The annotation notation is inconsistent in the core examples. Section 3.1 states that new speech-act relations are denoted with a '-SA' suffix, and Figure 11 uses 'command-SA' with a numbered roleset 'go-02'; Table A2 likewise lists Command-SA and go-02. However, Table 3 and Table A1 show 'command-00' and 'go-01' in the Dialogue-AMR annotations for the same type of utterances. The reader cannot tell whether this reflects an intentional version difference or a typographical error. Please harmonize the notation across the paper, or explicitly state the mapping between the '-SA' relations and the numbered relation forms.
minor comments (5)
- [Section 2 vs Section 5.2] The corpus size is given as 278 dialogues in Section 2 but as 287 dialogues in Section 5.2 ('until all 287 dialogues are completed'); these numbers should be reconciled.
- [Section 3.2] Typo: 'comparisong' should be 'comparing'; Section 3.1.3 and Table 2 contain 'repetoire' instead of 'repertoire'; Section 1 has 'interlocuter' instead of 'interlocutor'; Section 3.1.1 has 'Damsl guidelines' instead of 'DAMSL guidelines'.
- [Section 5.1.1] The first bullet point in the photo-requesting strategies begins with 'F ront' due to a formatting issue; this should read 'Front'.
- [Section 4.2.2] The phrase 'a translation-right' appears without the noun 'relation' in one sentence; this is grammatically awkward and should be rephrased.
- [Section 6.1] The description of how Dialogue Structure annotations are used as training data is at the level of 'instruction-response pairs' but does not specify which annotation layers (relations, TUs, antecedents) are leveraged in the retrieval model and the dialogue policy; a sentence clarifying the mapping would strengthen the reproducibility of the system description.
Circularity Check
No circularity: corpus and annotation-schema paper with no derived predictive claim; iterative schema development and WoZ transfer are validity limitations, not circular reductions.
full rationale
This is a resource/corpus description paper, not a derivation with fitted parameters or predicted quantities. The annotation layers (Dialogue-AMR, Dialogue Structure, visual context) are descriptive schemas, and the paper's central contribution is the released SCOUT annotations plus a unified description of the layers. No equation or quantitative claim is derived from an input that already contains the output. The iteratively developed speech-act and dialogue-relation categories are evaluated by inter-annotator agreement on the same corpus; IAA is an internal consistency measure, and the paper does not present it as an external predictive validation. That is a generalization caveat, not a circular step. Self-citations (Traum et al. 2018; Bonial et al. 2019, 2021, 2023) document earlier schema versions, but the present paper fully specifies the categories (Figures 7 and 9, Tables 2, 7, A1, A2), so the argument does not reduce to an unverified self-citation chain. The ScoutBot/MultiBot/JUDI systems are offered as implemented evidence that the annotations can support dialogue management; the paper explicitly acknowledges remaining gaps ('these systems lack the grounding and world model components', 'we currently have not achieved' cross-component communication), which is an omitted evaluation rather than a circularity. The Wizard-of-Oz multi-floor setup is asserted to mimic a future automated architecture, and the absence of reported results from the fourth experiment (automated DM) is a real transfer-validity risk, but this is an empirical assumption, not a definitional equivalence. Under the hard rules, no specific reduction of a claim to its own input by construction can be exhibited, so the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The Wizard-of-Oz multi-floor setup faithfully mimics a fully automated dialogue manager and robot navigator.
- domain assumption Commander speech under WOz deception is representative of natural language to a real autonomous robot.
- domain assumption Existing semantic and dialogue act taxonomies (AMR, DAMSL, ISO 24617-2, Searle) provide a suitable foundation for the new schemas.
- domain assumption Smatch and Krippendorff's alpha are valid reliability metrics for these annotation types.
invented entities (5)
-
Dialogue-AMR speech act relations (Command-SA, Assert-SA, Offer-SA, Request-SA, etc.)
-
Transaction Unit spanning multiple conversational floors
-
Translation relations (translation-left, translation-right and subtypes)
-
:completable aspectual feature
-
Exploration Map
Cite this review
Pith. "Pith review of Human-Robot Dialogue Annotation for Multi-Modal Common Ground." pith.science (2026). https://pith.science/paper/VQENZHYE
@misc{pith2026241112829,
author = {Pith},
title = {Pith review of: Human-Robot Dialogue Annotation for Multi-Modal Common Ground},
year = {2026},
howpublished = {\url{https://pith.science/paper/VQENZHYE}},
note = {Machine review of arXiv:2411.12829}
}
read the original abstract
In this paper, we describe the development of symbolic representations annotated on human-robot dialogue data to make dimensions of meaning accessible to autonomous systems participating in collaborative, natural language dialogue, and to enable common ground with human partners. A particular challenge for establishing common ground arises in remote dialogue (occurring in disaster relief or search-and-rescue tasks), where a human and robot are engaged in a joint navigation and exploration task of an unfamiliar environment, but where the robot cannot immediately share high quality visual information due to limited communication constraints. Engaging in a dialogue provides an effective way to communicate, while on-demand or lower-quality visual information can be supplemented for establishing common ground. Within this paradigm, we capture propositional semantics and the illocutionary force of a single utterance within the dialogue through our Dialogue-AMR annotation, an augmentation of Abstract Meaning Representation. We then capture patterns in how different utterances within and across speaker floors relate to one another in our development of a multi-floor Dialogue Structure annotation schema. Finally, we begin to annotate and analyze the ways in which the visual modalities provide contextual information to the dialogue for overcoming disparities in the collaborators' understanding of the environment. We conclude by discussing the use-cases, architectures, and systems we have implemented from our annotations that enable physical robots to autonomously engage with humans in bi-directional dialogue and navigation.
Figures
Reference graph
Works this paper leans on
-
[1]
Allen, J. and M. Core Draft, 1997. Draft of DAMSL: dialog act markup in several layers. available at: http://www.cs.rochester.edu/research/trains/annotation
1997
-
[2]
Nivre, and E
Allwood, J., J. Nivre, and E. Ahlsen 1992. On the semantics and pragmatics of linguistic feedback. In Journal of Semantics , Volume 9
1992
-
[3]
Bell III, P
Arvidson, R.E., J.F. Bell III, P. Bellutta, N.A. Cabrol, J. Catalano, J. Cohen, L.S. Crumpler, D. Des Marais, T. Estlin, W. Farrand, et al. 2010. Spirit mars rover mission: Overview and selected results from the northern home plate winter haven to the side of scamander crater. Journal of Geophysical Research: Planets\/ 115\/ (E7)
2010
-
[4]
Bonial, S
Banarescu, L., C. Bonial, S. Cai, M. Georgescu, K. Griffitt, U. Hermjakob, K. Knight, P. Koehn, M. Palmer, and N. Schneider 2013. A bstract M eaning R epresentation for sembanking. In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse , pp.\ 178--186
2013
-
[5]
Abrams, A.L
Bonial, C., M. Abrams, A.L. Baker, T. Hudson, S.M. Lukin, D. Traum, and C.R. Voss 2021. Context is key: Annotating situated dialogue relations in multi-floor dialogue. In Proceedings of the 25th Workshop on the Semantics and Pragmatics of Dialogue
2021
-
[6]
Abrams, D
Bonial, C., M. Abrams, D. Traum, and C. Voss 2021, June. Builder, we have done it: Evaluating & extending dialogue- AMR NLU pipeline for two collaborative domains. In Proceedings of the 14th International Conference on Computational Semantics (IWCS) , Groningen, The Netherlands (online), pp.\ 173--183. Association for Computational Linguistics
2021
-
[7]
Badarau, K
Bonial, C., B. Badarau, K. Griffitt, U. Hermjakob, K. Knight, T. O’Gorman, M. Palmer, and N. Schneider 2018. Abstract meaning representation of constructions: The more we include, the better the representation. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)
2018
-
[8]
Donatelli, S.M
Bonial, C., L. Donatelli, S.M. Lukin, S. Tratz, R. Artstein, D. Traum, and C.R. Voss 2019, August. Augmenting A bstract M eaning R epresentation for human-robot dialogue. In Proceedings of the First International Workshop on Designing Meaning Representations (DMR) , Florence, Italy, pp.\ 199--210
2019
Show all 91 references
-
[9]
Foresta, N
Bonial, C., J. Foresta, N. Fung, C. Hayes, P. Osteen, J. Arkin, B. Hedegaard, and T. Howard 2023. Amr for grounded human-robot communication. In Proceedings of the Designing Meaning Representation 2023 Workshop at IWCS 2023
2023
-
[10]
Hudson, L
Bonial, C., T. Hudson, L. Donatelli, D. Traum, C. Voss, M. Abrams, and A. Blodgett 2023. Dialogue-amr annotation guidelines. Technical Report ARL-TR-9648, DEVCOM Army Research Laboratory, Adelphi, Maryland, United States
2023
-
[11]
Marge, A
Bonial, C., M. Marge, A. Foots, F. Gervits, C.J. Hayes, C. Henry, S.G. Hill, A. Leuski, S.M. Lukin, P. Moolchandani, K.A. Pollard, D. Traum, and C.R. Voss 2017. Laying down the yellow brick road: Development of a wizard-of-oz interface for collecting human-robot dialogue. In S...
2017
-
[12]
Traum, C
Bonial, C., D. Traum, C. Henry, S.M. Lukin, M. Marge, R. Artstein, K.A. Pollard, A. Foots, A.L. Baker, and C.R. Voss 2019. Dialogue structure annotation guidelines for Army Research Laboratory (ARL) human-robot dialogue corpus. Technical Report ARL-TR-8833, DEVCOM Army Researc...
2019
-
[13]
Donatelli, J
Bonial, C.N., L. Donatelli, J. Ervin, and C.R. Voss 2019. Abstract M eaning R epresentation for human-robot dialogue. Volume 2, pp.\ 236--246
2019
-
[14]
Palmer, J
Bonn, J., M. Palmer, J. Cai, and K. Wright-Bettner 2020. Spatial AMR : Expanded spatial annotation in the context of a grounded minecraft corpus. In Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020)
2020
-
[15]
Duvallet, J
Boularias, A., F. Duvallet, J. Oh, and A. Stentz 2015. Grounding spatial relations for outdoor robot navigation. In 2015 IEEE International Conference on Robotics and Automation (ICRA) , pp.\ 1976--1982. IEEE
2015
-
[16]
Chebotar, C
Brohan, A., Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, et al. 2023. Do as i can, not as i say: Grounding language in robotic affordances. In Conference on robot learning , pp.\ 287--318. PMLR
2023
-
[17]
Donatelli, K
Brutti, R., L. Donatelli, K. Lai, and J. Pustejovsky 2022. Abstract meaning representation for gesture. In Proceedings of the Thirteenth Language Resources and Evaluation Conference , pp.\ 1576--1583
2022
-
[18]
Alexandersson, J.W
Bunt, H., J. Alexandersson, J.W. Choe, A.C. Fang, K. Hasida, V. Petukhova, A. Popescu-Belis, and D. Traum 2012. ISO 24617-2: A semantically-based standard for dialogue annotation. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'...
2012
-
[19]
Petukhova, E
Bunt, H., V. Petukhova, E. Gilmartin, C. Pelachaud, A. Fang, S. Keizer, and L. Prévot 2020, May. The ISO standard for dialogue act annotation, second edition. In Proceedings of The 12th Language Resources and Evaluation Conference , Marseille, France, pp.\ 549--558. European L...
2020
-
[20]
Cai, S. and K. Knight 2013. Smatch: An evaluation metric for semantic feature structures. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , Volume 2, pp.\ 748--752
2013
-
[21]
Reddy, D.R
Camilli, R., C.M. Reddy, D.R. Yoerger, B.A. Van Mooy, M.V. Jakuba, J.C. Kinsey, C.P. McIntyre, S.P. Sylva, and J.V. Maloney 2010. Tracking hydrocarbon plume transport and biodegradation at deepwater horizon. In Science , Volume 330, pp.\ 201--204. American Association for the ...
2010
-
[22]
Isard, S
Carletta, J., A. Isard, S. Isard, J. Kowtko, G. Doherty-Sneddon, and A. Anderson 1996. HCRC dialogue structure coding manual. Technical Report 82, Human Communication Research Centre, University of Edinburgh
1996
-
[23]
Arkin, C
Chen, Y., J. Arkin, C. Dawson, Y. Zhang, N. Roy, and C. Fan 2024. Autotamp: Autoregressive task and motion planning with llms as translators and checkers. In 2024 IEEE International Conference on Robotics and Automation (ICRA) , pp.\ 6695--6702. IEEE
2024
-
[24]
Epsimos, G
Chiou, M., G.T. Epsimos, G. Nikolaou, P. Pappas, G. Petousakis, S. M \"u hl, and R. Stolkin 2022. Robot-assisted nuclear disaster response: Report and insights from a field exercise. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp.\ 4545...
2022
-
[25]
Sanfilippo, and S
Chitikena, H., F. Sanfilippo, and S. Ma 2023. Robotics in search and rescue (SAR) operations: An ethical and design perspective framework for response phase. In Applied Sciences , Volume 13
2023
-
[26]
Choi, W.S., Y.J. Heo, D. Punithan, and B.T. Zhang 2022. Scene graph parsing via abstract meaning representation in pre-trained language models. In Proceedings of the 2nd Workshop on Deep Learning on Graphs for Natural Language Processing (DLG4NLP 2022) , pp.\ 30--35
2022
-
[27]
Heo, and B.T
Choi, W.S., Y.J. Heo, and B.T. Zhang 2022. Sgram: Improving scene graph parsing via abstract meaning representation. In arXiv preprint arXiv:2210.08675
2022 arXiv
-
[28]
Clark, H.H. and C.R. Marshall 1981. Definite reference and mutual knowledge. In A. K. Joshi, B. L. Webber, and I. A. Sag (Eds.), Elements of Discourse Understanding . Cambridge University Press
1981
-
[29]
Clark, H.H. and E.F. Schaefer 1989. Contributing to discourse. In Cognitive Science , Volume 13, pp.\ 259--294
1989
-
[30]
Clark, H.H. and D. Wilkes-Gibbs 1986. Referring as a collaborative process. In Cognition , Volume 22, pp.\ 1--39
1986
-
[31]
2023, Mar
Clearpath Robotics . 2023, Mar. Clearpath Husky UGV
2023
-
[32]
Condon, S. and C. Cech 1992. Manual for coding decision-making interactions. unpublished manuscript, updated May 1995, available at: \ ftp://sls-ftp.lcs.mit.edu/pub/multiparty/coding\ schemes/condon\
1992
-
[33]
Core, M.G. and J. Allen 1997. Coding dialogs with the damsl annotation scheme. In AAAI fall symposium on communicative action in humans and machines , Volume 56, pp.\ 28--35. Boston, MA
1997
-
[34]
Artstein, G
DeVault, D., R. Artstein, G. Benn, T. Dey, E. Fast, A. Gainer, K. Georgila, J. Gratch, A. Hartholt, M. Lhommet, G. Lucas, S.C. Marsella, M. Fabrizio, A. Nazarian, S. Scherer, G. Stratou, A. Suri, D. Traum, R. Wood, Y. Xu, A. Rizzo, and L.P. Morency 2014. S im S ensei K iosk: A...
2014
-
[35]
Rezaee, and F
Dipta, S.R., M. Rezaee, and F. Ferraro. 2022. Semantically-informed hierarchical event modeling
2022
-
[36]
Lai, and J
Donatelli, L., K. Lai, and J. Pustejovsky 2020, December. A two-level interpretation of modality in human-robot dialogue. In Proceedings of the 28th International Conference on Computational Linguistics , Barcelona, Spain (Online), pp.\ 4222--4238. International Committee on C...
2020
-
[37]
Regan, W
Donatelli, L., M. Regan, W. Croft, and N. Schneider 2018. Annotation of tense and aspect semantics for sentential AMR . In Proceedings of the Joint Workshop on Linguistic Annotation Multiword Expressions and Constructions (LAW-MWE-CxG-2018) , pp.\ 96--108
2018
-
[38]
Schneider, W
Donatelli, L., N. Schneider, W. Croft, and M. Regan 2019. Tense and aspect semantics for sentential AMR . In Proceedings of the Society for Computation in Linguistics , Volume 2, pp.\ 346--348
2019
-
[39]
Drew, D.S. 2021. Multi-agent systems for search and rescue applications. In Current Robotics Reports , Volume 2, pp.\ 189--200. Springer
2021
-
[40]
Xia, M.S.M
Driess, D., F. Xia, M.S.M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence. 2023. Palm-e: An ...
2023
-
[41]
Edelsky, C. 1981. Who's got the floor? In Language in Society , Volume 10, pp.\ 383--421. Cambridge University Press
1981
-
[42]
Anschober, R
Edlinger, R., M. Anschober, R. Froschauer, and A. N \"u chter 2022. Intuitive hri approach with reliable and resilient wireless communication for rescue robots and first responders. In 2022 31st IEEE International Conference on Robot and Human Interactive Communication (RO-MAN...
2022
-
[43]
Grosz, B.J. and C.L. Sidner. 1986. Attention, intention, and the structure of discourse. Computational Linguistics\/ 12\/ (3): 175--204
1986
-
[44]
Dadvar, B
Habibian, S., M. Dadvar, B. Peykari, A. Hosseini, M.H. Salehzadeh, A.H. Hosseini, and F. Najafi 2021. Design and implementation of a maxi-sized mobile robot ( Karo ) for rescue missions. In Robomech Journal , Volume 8, pp.\ 1. Springer
2021
-
[45]
Harnad, S. 1990. The symbol grounding problem. In Physica D: Nonlinear Phenomena , Volume 42, pp.\ 335--346
1990
-
[46]
Traum, S.C
Hartholt, A., D. Traum, S.C. Marsella, A. Shapiro, G. Stratou, A. Leuski, L.P. Morency, and J. Gratch. 2013, August. All together now: Introducing the virtual human toolkit, In Intelligent Virtual Agents: 13th International Conference, IVA 2013, E dinburgh, UK , August 29--31,...
2013
-
[47]
Halme, and A
Heikkil \"a , S.S., A. Halme, and A. Schiele 2012. Affordance-based indirect task communication for astronaut-robot cooperation. In Journal of field robotics , Volume 29, pp.\ 576--600. Wiley Online Library
2012
-
[48]
Holder, E. 2017. Defining ``soldier intent'' in a human–robot natural language interaction context. Technical Report ARL-TR-8195, Army Research Laboratory
2017
-
[49]
Howard, T.M., N. Roy, J. Fink, J. Arkin, R. Paul, D. Park, S. Roy, D. Barber, R. Bendell, K. Schmeckpeper, J. Tian, J. Oh, M. Wigness, L. Quang, B. Rothrock, J. Nash, M.R. Walter, F. Jentsch, and E. Stump 2021. An intelligence architecture for grounded language communication w...
2021
-
[50]
Kuwajerwala, Q
Jatavallabhula, K.M., A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, S. Li, G. Iyer, S. Saryazdi, N. Keetha, A. Tewari, et al. 2023. Conceptfusion: Open-set multimodal 3d mapping. In arXiv preprint arXiv:2302.07241
2023 arXiv
-
[51]
Sato, and Y
Kanazawa, K., N. Sato, and Y. Morita 2023. Considerations on interaction with manipulator in virtual reality teleoperation interface for rescue robots. In 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) , pp.\ 386--391
2023
-
[52]
Kang, S., C. Cho, J. Lee, D. Ryu, C. Park, K.C. Shin, and M. Kim 2003. Robhaz-dt2: Design and integration of passive double tracked mobile manipulator system for explosive ordnance disposal. In Proceedings 2003 IEEE/RSJ International Conference on Intelligent Robots and System...
2003
-
[53]
Higgins, P
Kebe, G.Y., P. Higgins, P. Jenkins, K. Darvish, R. Sachdeva, R. Barron, J. Winder, D. Engel, E. Raff, F. Ferraro, et al. 2021. A spoken language dataset of descriptions for speech-based grounded language learning. In Advances in neural information processing systems
2021
-
[54]
Krippendorff, K. 1980. Content Analysis: An Introduction to Its Methodology , Chapter 12, pp.\ 129--154. Beverly Hills, CA: Sage
1980
-
[55]
Leuski, A. and D. Traum 2011. NPCEditor : Creating virtual human dialogue using information retrieval techniques. In AI Magazine , Volume 32, pp.\ 42--56
2011
-
[56]
Bonial, M
Lukin, S.M., C.N. Bonial, M. Marge, T. Hudson, C.J. Hayes, K.A. Pollard, A. Baker, A. Foots, R. Artstein, F. Gervits, M. Abrams, C. Henry, L. Donatelli, A. Leuski, S.G. Hill, D. Traum, and C.R. Voss 2024. SCOUT : A situated and multi-modal human-robot dialogue corpus. In The J...
2024
-
[57]
Gervits, C.J
Lukin, S.M., F. Gervits, C.J. Hayes, A. Leuski, P. Moolchandani, J.G. Rogers III, C.S. Amaro, M. Marge, C.R. Voss, and D. Traum 2018, July. ScoutBot : A dialogue system for collaborative navigation. In Proceedings of the 56th Annual Meeting of the Association for Computational...
2018
-
[58]
Pollard, C
Lukin, S.M., K.A. Pollard, C. Bonial, T. Hudson, R. Artstein, C. Voss, and D. Traum 2023. Navigating to success in multi-modal human-robot collaboration: Corpus and analysis. In IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN)
2023
-
[59]
Bonial, B
Marge, M., C. Bonial, B. Byrne, T. Cassidy, A.W. Evans, S.G. Hill, and C. Voss 2016. Applying the wizard-of-oz technique to multimodal human-robot dialogue. In Proceedings of IEEE RO-MAN
2016
-
[60]
Bonial, A
Marge, M., C. Bonial, A. Foots, C. Hayes, C. Henry, K. Pollard, R. Artstein, C. Voss, and D. Traum 2017. Exploring variation of natural human commands to a robot in a collaborative navigation task. In Proceedings of the First Workshop on Language Grounding for Robotics , pp.\ 58--66
2017
-
[61]
Bonial, S
Marge, M., C. Bonial, S. Lukin, and C. Voss 2023. Bot language. Summary Technical Report, Oct 2016--Sep 2021 ARL-TR-9656, DEVCOM Army Research Laboratory
2023
-
[62]
Nogar, C.J
Marge, M., S. Nogar, C.J. Hayes, S.M. Lukin, J. Bloecker, E. Holder, and C. Voss 2019, June. A research platform for multi-robot dialogue with humans. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics (Demonst...
2019
-
[63]
Mavridis, N. 2015. A review of verbal and non-verbal human--robot interactive communication. In Robotics and Autonomous Systems , Volume 63, pp.\ 22--35. Elsevier
2015
-
[64]
Murphy, R.R. 2014. Disaster robotics. In MIT press
2014
-
[65]
Kiribayashi, Y
Nagatani, K., S. Kiribayashi, Y. Okada, K. Otake, K. Yoshida, S. Tadokoro, T. Nishimura, T. Yoshida, E. Koyanagi, M. Fukushima, and S. Kawatsuma 2013. Emergency response to the nuclear accident at the F ukushima D aiichi nuclear power plants using mobile rescue robots. In Jour...
2013
-
[66]
Jayannavar, and J
Narayan-Chen, A., P. Jayannavar, and J. Hockenmaier 2019, July. Collaborative dialogue in M inecraft. In A. Korhonen, D. Traum, and L. M \`a rquez (Eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Florence, Italy, pp.\ 5405--5415...
2019
-
[67]
Passonneau, R. 2006. Measuring agreement on set-valued items ( MASI ) for semantic and pragmatic annotation. In Proc. of LREC
2006
-
[68]
Ghoshal, G
Povey, D., A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely 2011. The kaldi speech recognition toolkit. In IEEE workshop on automatic speech recognition and understanding . IEEE Sig...
2011
-
[69]
Prasad, R. and H. Bunt 2015. Semantic relations in discourse: The current state of ISO 24617-8. In Proceedings 11th Joint ACL-ISO Workshop on Interoperable Semantic Annotation (ISA-11) , pp.\ 80--92
2015
-
[70]
Taipalmaa, B.C
Queralta, J.P., J. Taipalmaa, B.C. Pullinen, V.K. Sarker, T.N. Gia, H. Tenhunen, M. Gabbouj, J. Raitoharju, and T. Westerlund 2020. Collaborative multi-robot search and rescue: Planning, coordination, perception, and active vision. In Ieee Access , Volume 8, pp.\ 191617--191643. IEEE
2020
-
[71]
Haviland, S
Rana, K., J. Haviland, S. Garg, J. Abou-Chakra, I. Reid, and N. Suenderhauf 2023. Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning. In 7th Annual Conference on Robot Learning
2023
-
[72]
Turner, and C.H
Reinsch, N.L., J.W. Turner, and C.H. Tinsley 2008. Multicommunicating: A practice whose time has come? In Academy of Management Review , Volume 33, pp.\ 391--403. Academy of Management
2008
-
[73]
Dixit, A
Ren, A.Z., A. Dixit, A. Bodrova, S. Singh, S. Tu, N. Brown, P. Xu, L. Takayama, F. Xia, J. Varley, et al. 2023. Robots that ask for help: Uncertainty alignment for large language model planners. In Proceedings of Machine Learning Research , Volume 229
2023
-
[74]
Ryu, D., S. Kang, M. Kim, and J.B. Song 2004. Multi-modal user interface for teleoperation of robhaz-dt2 field robot system. In 2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)(IEEE Cat. No. 04CH37566) , Volume 1, pp.\ 168--173. IEEE
2004
-
[75]
Searle, J.R. 1969. Speech acts: An essay in the philosophy of language . Cambridge University Press
1969
-
[76]
Thomason, D
Shridhar, M., J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox 2020. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pp.\ 10740--10749
2020
-
[77]
Silver, T., S. Dan, K. Srinivas, J.B. Tenenbaum, L. Kaelbling, and M. Katz 2024. Generalized planning in pddl domains with pretrained large language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Volume 38, pp.\ 20256--20264
2024
-
[78]
Sinclair, J.M. and R.M. Coulthard. 1975. Towards an analysis of Discourse: The English used by teachers and pupils. Oxford University Press
1975
-
[79]
Kil, T.Y
Song, C.H., J. Kil, T.Y. Pan, B.M. Sadler, W.L. Chao, and Y. Su 2022. One step at a time: Long-horizon vision-and-language navigation with milestones. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp.\ 15482--15491
2022
-
[80]
Gopalan, H
Tellex, S., N. Gopalan, H. Kress-Gazit, and C. Matuszek. 2020. Robots that use language. Annual Review of Control, Robotics, and Autonomous Systems\/ 3: 25--55
2020
-
[81]
Henry, S
Traum, D., C. Henry, S. Lukin, R. Artstein, F. Gervits, K. Pollard, C. Bonial, S. Lei, C. Voss, M. Marge, C. Hayes, and S. Hill 2018, May 7-12, 2018. Dialogue structure annotation for multi-floor interaction. In N. Calzolari, K. Choukri, C. Cieri, T. Declerck, S. Goggi, K. Has...
2018
-
[82]
Traum, D. and S. Larsson. 2003. The information state approach to dialogue management, In Current and New Directions in Discourse and Dialogue , eds. van Kuppevelt, J. and R. Smith, 325--353
2003
-
[83]
Traum, D.R. 1994. A Computational Theory of Grounding in Natural Language Conversation . Ph.\ D. thesis, Department of Computer Science, University of Rochester. Also available as TR 545, Department of Computer Science, University of Rochester
1994
-
[84]
Traum, D.R. and C.H. Nakatani 1999. A two-level approach to coding dialogue for discourse structure: Activities of the 1998 working group on higher-level structures. In Proceedings of ACL 1999 Workshop: Towards Standards and Tools for Discourse Tagging , pp.\ 101--108
1999
-
[85]
Marquez, S
Valmeekam, K., M. Marquez, S. Sreedharan, and S. Kambhampati 2023. On the planning abilities of large language models: A critical investigation. In Advances in Neural Information Processing Systems , Volume 36
2023
-
[86]
Antone, E
Walter, M.R., M. Antone, E. Chuangsuwanich, A. Correa, R. Davis, L. Fletcher, E. Frazzoli, Y. Friedman, J. Glass, J.P. How, et al. 2015. A situationally aware voice-commandable robotic forklift working alongside people in unstructured outdoor environments. In Journal of Field ...
2015
-
[87]
Pizarro, M.V
Williams, S.B., O.R. Pizarro, M.V. Jakuba, C.R. Johnson, N.S. Barrett, R.C. Babcock, G.A. Kendrick, P.D. Steinberg, A.J. Heyward, P.J. Doherty, et al. 2012. Monitoring of benthic reference sites: using an autonomous underwater vehicle. In IEEE Robotics & Automation Magazine , ...
2012
-
[88]
Yamauchi, B.M. 2004. Packbot: a versatile platform for military robotics. In Unmanned ground vehicle technology VI , Volume 5422, pp.\ 228--237. SPIE
2004
-
[89]
Younger, D.H. 1967. Recognition and parsing of context-free languages in time n3. In Information and control , Volume 10, pp.\ 189--208. Elsevier
1967
-
[90]
Zhang, Y., J. Yang, J. Pan, S. Storks, N. Devraj, Z. Ma, K.P. Yu, Y. Bao, and J. Chai 2022. DANLI : Deliberative agent for following natural language instructions. In arXiv preprint arXiv:2210.12485
2022 arXiv
-
[91]
write newline
" write newline "" before.all 'output.state := FUNCTION output.doi doi empty skip "doi:" doi * "" * output if FUNCTION format.archive archivePrefix empty "" archivePrefix ":" * if FUNCTION format.primaryClass primaryClass empty "" " [" primaryClass * "] " * if FUNCTION format....
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.