REVIEW 4 major objections 5 minor 13 references
Analysing Explanation-Related Interactions in Collaborative Perception-Cognition-Communication-Action
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read In the TeamCollab emergency-response simulation, most explanation-related messages are clarification requests, not 'why' questions, and communication volume correlates with task accuracy.
desk verdict The headline result is largely coded into the codebook, so the main claim doesn't stand; still, it's a clear, honest pilot that would benefit from reanalysis and significant revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The coding scheme, a three-stage qualitative labelling procedure with four annotators and a codebook, maps each message to one of nine labels grouped into four categories from an interactive-XAI taxonomy (select/clarify, mutate/simulate, dialogue/progress, dialogue/answer). The select/clarify category carries the main argument: because auto, doing, and features—the three most frequent labels—all fall under it, the paper concludes that clarification is the dominant explanation need. Inter-annotator agreement (61% of messages having at least three of four annotators agree) provides the reliability check for this machinery.
What would settle it
Re-annotate the same 1,000 messages with an independent codebook that distinguishes clarification-seeking messages (e.g., explicit questions like 'what are you doing?') from statements that merely share information; if select/clarify is no longer the majority category, the paper's ranking of explanation needs would collapse. Alternatively, run a controlled experiment with a mandatory one-message-per-action condition to test whether increased communication volume actually raises the percentage of dangerous objects collected.
Extended reading notes
Core claim
The paper's central discovery is that in human-human collaboration on a time-pressured physical task, the explanation-related communication that actually occurs is dominated by the select/clarify mode: autocompleted sharing of sensed object attributes, 'what are you doing' updates, and discussion of object features. Explicit justification questions ('why is X dangerous?') essentially did not occur, and annotators found no majority-agreement cases for the 'why' label. The paper also establishes a positive Pearson correlation (0.30) between the number of messages a team sends and the fraction of collected objects that are actually dangerous, which it interprets as communication enabling agents to cross-check sensor readings and avoid wasting effort on benign objects.
Load-bearing premise
The finding that most explanation-related messages seek clarification largely follows from the coding decision to treat routine coordination messages (autocompleted sensor-sharing, 'doing' updates, and feature talk) as select/clarify 'explanation-related'; under a narrower definition of explanation, the majority might vanish.
Editorial extensions
If this is right
- Robot design implication: explanation capabilities for collaborative physical tasks should prioritise clarification and confirmation dialogues, such as answering 'what are you doing?' and sharing sensor data, over long-form causal explanations.
- Performance implication: communication volume is not a net cost; teams that communicate more are more accurate in identifying dangerous objects, so encouraging message exchange can improve outcomes measured by accuracy.
- The near-absence of explicit 'why' questions suggests that implicit explanation dialogues—messages that function as clarifications in context—are what matter in emergency-response-style teamwork.
- The positive team-level correlation justifies using the TeamCollab simulation to test explanation-equipped robot teammates in the paper's stated next step of human-robot interaction experiments.
Reading between the lines
- This suggests that a robot that simply broadcasts sensor data and periodically states its current action would satisfy most of the observed 'explanation' demand, before any causal reasoning is added—an inference the paper does not draw.
- A testable extension is to give one robot a clarification-only dialogue policy and another a causal-explanation policy in the same simulation, and compare team accuracy, to directly test the paper's design recommendation.
- The correlation between message volume and accuracy is modest (PCC 0.30), so a sensitivity analysis separating coordination messages from clarification-seeking ones would clarify whether clarification specifically drives accuracy—an analysis the paper does not report.
- Because the task is physical and time-pressured, the low frequency of 'why' questions may not transfer to planning-heavy or diagnosis-heavy collaborative tasks, where causal explanation needs could be stronger.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript analyses 2,607 chat messages from 20 TeamCollab sessions to identify which explanation-related interaction types humans use when collaborating in a simulated emergency-response task. A subset of 1,000 messages was coded by four annotators using a codebook derived from a two-level XAI taxonomy (select/clarify, mutate/simulate, dialogue/progress, dialogue/answer). The main claims are that most explanation-related messages seek clarification in decisions or actions (Section IV.B) and that the volume of communication correlates with team performance, specifically a negative correlation with total objects collected and a positive correlation with the percentage of collected objects that are dangerous (Section IV.A).
Significance. If the findings were valid, they would provide actionable guidance for robot explainability in collaborative settings, suggesting that clarification and confirmation capabilities are more important than deep causal explanation. The paper addresses a relevant question and its dataset of human-human communication in a simulated task is a useful resource. The iterative codebook development with multiple annotators is a strength, and the authors are transparent about the annotators' observations that explicit 'why' questions were rare. However, the central qualitative claim is largely a consequence of the coding scheme, and the quantitative performance claim is not backed by inferential statistics. The significance for XAI design is therefore currently not established.
major comments (4)
- [§III.C, Table I; §IV.B, Fig. 3] The main qualitative result that most explanation-related messages seek clarification is a direct consequence of the codebook definitions. Table I assigns the labels 'doing', 'features', and 'auto' to the category 'select/clarify', and Fig. 3 shows that these three labels are the most frequent in the majority-label distribution. Therefore the statement that 'most explanation-related messages seek clarification in the decisions or actions taken' follows by construction rather than from an independent categorization of the dialogue. This circularity is acknowledged implicitly in the discussion ('This is likely why messages labelled with auto and features are quite frequent'), but the manuscript still presents the finding as an empirical discovery. A re-analysis that distinguishes information transmission from explanation-seeking behavior, or a substantially qualified claim, would be needed.
- [§IV.A, Fig. 2] The three Pearson correlations (PCC = -0.28, 0.05, 0.30) are reported without p-values, confidence intervals, or any inferential test. With n = 20 sessions, none of these correlations is likely to reach conventional significance levels, and the text overstates the evidence by saying the data 'confirm that messages have an impact on the performance of our simulated task'. At minimum, exact p-values (or bootstrap intervals) and a clear statement about statistical power are required before drawing conclusions about the relationship between communication volume and performance.
- [§III.C.1, §IV.B] Inter-annotator agreement is reported as '61% of samples were labelled with at least three annotators agreeing on the label', with 5% unclassified. This is not a chance-corrected measure such as Cohen's kappa or Fleiss' kappa, so it does not substantiate the claim of a 'good level of inter-annotator agreement'. Without such a measure, the reliability of the majority labels that drive the central distribution in Fig. 3 is not established. The authors should report kappa statistics and, if agreement is low, discuss how disagreements were resolved.
- [§III.C, Table I] The construct validity of mapping 'auto' messages to the 'select/clarify' category is weak. Autocompleted messages are one-click sensor broadcasts (e.g., 'Object 18 ... Status Danger'), not user-driven selections or clarifications. Similarly, 'doing' messages are action updates and coordination statements, and 'features' messages report object attributes. Labeling these as explanation-related 'select/clarify' conflates information transmission with asking for or giving explanations. This directly affects the headline claim about what types of explanations robots need, because the dominant categories may not correspond to explanation behavior at all.
minor comments (5)
- [§IV.B] The paper says 'qualitative linguist analysis'; this should be 'linguistic analysis'.
- [§III.C.1] The procedure mentions 'inter-code agreement' but the paper later uses 'inter-annotator agreement'; please use consistent terminology throughout.
- [Fig. 2] The figures show scatter plots without any indication of uncertainty or regression lines; adding regression lines and confidence bands would improve readability, though the statistical issues in the major comments would still need to be addressed.
- [Fig. 3] The caption 'Majority Label Distribution ( ≥ 3 agreement level)' could clarify how the 'unclassified or no similarities' 5% is handled; currently it is unclear whether 'unclassified' is a separate bar or omitted from the denominator.
- [References] Reference [2] is a CHI 2023 paper and should include the page range or article number for completeness, per the reference style used elsewhere.
Circularity Check
Headline claim that most explanation-related messages seek clarification is an artifact of the codebook mapping: auto, doing, and features are pre-classified as select/clarify and dominate the label distribution.
-
self definitional
[Section III.C (Table I) and Section IV.B (Fig. 3 and surrounding text)]
"auto select/clarify Autocompleted messages used to share all sensed details for a given object via a button on the UI. ... The most frequent label is autocompleted messages, which participants can easily send by pressing a UI button to share all sensed attributes for an object. ... It is worth noting that most of the labelled messages fall into the select/clarify category (auto, doing, features), followed by dialogue/answer (confirm)."
Table I pre-assigns the labels 'doing', 'features' and 'auto' to the category 'select/clarify'. The majority-label distribution in Fig. 3 is dominated by exactly those labels (plus 'confirm'). The paper's headline finding that 'most explanation-related messages seek clarification in the decisions or actions taken' is therefore a direct consequence of the codebook mapping: the largest category is largest because the high-frequency labels were defined into it. The reduction is visible in the paper's own sentence 'most of the labelled messages fall into the select/clarify category (auto, doing, features)'. In particular, 'auto' messages are one-click sensor broadcasts (e.g., 'Object 18 ... Status Danger: dangerous, Prob.
full rationale
The central qualitative claim is circular in the specific sense that the category 'select/clarify' was defined to include the labels that dominate the data; the paper's own summary sentence makes this explicit. The performance analysis (Section IV.A) is not circular: it is an independent correlational analysis of message volume versus collected-object metrics. However, the three Pearson coefficients are reported without p-values or confidence intervals for 20 sessions, so the statement that messages 'confirm that messages have an impact' is statistically under-supported; that is a correctness/evidence concern, not circularity. I found no load-bearing self-citation: the TeamCollab paper [1] is cited only to describe the simulation environment, and the XAI taxonomy [2] is an external source. The circularity score is driven by the headline abstract claim reducing to the codebook assignment.
Assumptions & free parameters
assumptions (3)
- domain assumption The coding scheme's assignment of messages to explanation-related categories is a valid measure of explanation-seeking behavior.
- domain assumption The 1,000 analyzed messages are representative of the 2,607-message corpus.
- domain assumption Pearson correlations computed over session-level aggregates are meaningful evidence of a relationship.
Cite this review
Pith. "Pith review of Analysing Explanation-Related Interactions in Collaborative Perception-Cognition-Communication-Action." pith.science (2026). https://pith.science/paper/PZV527SS
@misc{pith2026241112483,
author = {Pith},
title = {Pith review of: Analysing Explanation-Related Interactions in Collaborative Perception-Cognition-Communication-Action},
year = {2026},
howpublished = {\url{https://pith.science/paper/PZV527SS}},
note = {Machine review of arXiv:2411.12483}
}
read the original abstract
Effective communication is essential in collaborative tasks, so AI-equipped robots working alongside humans need to be able to explain their behaviour in order to cooperate effectively and earn trust. We analyse and classify communications among human participants collaborating to complete a simulated emergency response task. The analysis identifies messages that relate to various kinds of interactive explanations identified in the explainable AI literature. This allows us to understand what type of explanations humans expect from their teammates in such settings, and thus where AI-equipped robots most need explanation capabilities. We find that most explanation-related messages seek clarification in the decisions or actions taken. We also confirm that messages have an impact on the performance of our simulated task.
Figures
Reference graph
Works this paper leans on
-
[1]
TeamCollab: A framework for col- laborative Perception-Cognition-Communication-Action,
J. de Gortari Briseno, R. Para ´c, L. Ardon, M. Roig Vilamala, D. Furelos-Blanco, L. Kaplan, V . K. Mishra, F. Cerutti, A. Preece, A. Russo, and M. Srivastava, “TeamCollab: A framework for col- laborative Perception-Cognition-Communication-Action,” inProc 27th International Conference on Information Fusion , in press 2024
work page 2024
-
[2]
A. Bertrand, T. Viard, R. Belloum, J. R. Eagan, and W. Maxwell, “On selective, mutable and dialogic XAI: a review of what users say about different types of interactive explanations,” in Proc 2023 CHI Conference on Human Factors in Computing Systems , April 2023, pp. 1–21
work page 2023
-
[3]
Building Cooperative Embodied Agents Modularly with Large Language Models,
H. Zhang, W. Du, J. Shan, Q. Zhou, Y . Du, J. B. Tenenbaum, T. Shu, and C. Gan, “Building Cooperative Embodied Agents Modularly with Large Language Models,” arXiv preprint, vol. arXiv:2307.02485, 2023
arXiv 2023
-
[4]
Supporting human-AI teams: Transparency, explain- ability, and situation awareness,
M. R. Endsley, “Supporting human-AI teams: Transparency, explain- ability, and situation awareness,” Computers in Human Behavior , vol. 140, p. 107574, 2023
work page 2023
-
[5]
Questioning the AI: informing design practices for explainable AI user experiences,
Q. V . Liao, D. Gruen, and S. Miller, “Questioning the AI: informing design practices for explainable AI user experiences,” in Proc 2020 CHI conference on human factors in computing systems , 2020, pp. 1–15
work page 2020
-
[6]
Principles of Explanation in Human-AI Systems
S. T. Mueller, E. S. Veinott, R. R. Hoffman, G. Klein, L. Alam, T. Mamun, and W. J. Clancey, “Principles of explanation in human-AI systems,” arXiv preprint arXiv:2102.04972 , 2021
work page Pith review arXiv 2021
-
[7]
One explanation does not fit all: The promise of interactive explanations for machine learning transparency,
K. Sokol and P. Flach, “One explanation does not fit all: The promise of interactive explanations for machine learning transparency,” KI- K¨unstliche Intelligenz, vol. 34, no. 2, pp. 235–250, 2020
2020
-
[8]
Explanation in artificial intelligence: Insights from the social sciences,
T. Miller, “Explanation in artificial intelligence: Insights from the social sciences,” Artificial intelligence, vol. 267, pp. 1–38, 2019
2019
Show all 13 references
-
[9]
Implicit commu- nication of actionable information in human-AI teams,
C. Liang, J. Proft, E. Andersen, and R. A. Knepper, “Implicit commu- nication of actionable information in human-AI teams,” in Proc 2019 CHI conference on human factors in computing systems , 2019, pp. 1–13
2019
-
[10]
ThreeDWorld: A platform for interactive multi-modal physical simulation,
C. Gan et al, “ThreeDWorld: A platform for interactive multi-modal physical simulation,” in Proc Advances in Neural Information Pro- cessing Systems (NeurIPS) Track on Datasets and Benchmarks , 2021
2021
-
[11]
Roboclean: Contextual language grounding for human-robot interactions in specialised low- resource environments,
C. Fuentes, M. Porcheron, and J. E. Fischer, “Roboclean: Contextual language grounding for human-robot interactions in specialised low- resource environments,” in Proc 5th International Conference on Conversational User Interfaces , 2023, pp. 1–11
2023
-
[12]
Developing a natural language dialogue system: Wizard of Oz studies,
M. Kullasaar, E. Vutt, and M. Koit, “Developing a natural language dialogue system: Wizard of Oz studies,” in Proc First International IEEE Symposium Intelligent Systems , vol. 1. IEEE, 2002, pp. 202– 207
2002
-
[13]
De- veloping and using a codebook for the analysis of interview data: An example from a professional development research project,
J. T. DeCuir-Gunby, P. L. Marshall, and A. W. McCulloch, “De- veloping and using a codebook for the analysis of interview data: An example from a professional development research project,” Field methods, vol. 23, no. 2, pp. 136–155, 2011
2011
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.