{"id":"5726b117-3d83-41d2-95f2-ea3bcd372d86","arxiv_id":"2505.03674","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In a human-AI teamwork game, showing the AI's inferred goals improved players' perceived understanding of the teammate but not objective performance or overall satisfaction.","lead":"This study ran an online experiment where people played a tool-fetching game with an AI teammate that either showed, showed on demand, or hid its guesses about the player's goal. Sharing the AI's inferred goals did not improve task performance or satisfaction, but participants said it helped them understand and steer the teammate.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Subjective collaboration benefit rests entirely on an unreported LLM thematic analysis with no codebook or reliability check; with the satisfaction ANOVA null, the perceived-collaboration claim lacks quantitative support.","rationale":"The paper's strongest claim has two parts: objective performance does not improve, and perceived collaboration does. The first part is credible: the reported ANOVAs for performance and duration are clearly non-significant and the distributions look similar. The second part is the load-bearing concern. The only quantitative subjective measure, satisfaction, is non-significant, so the paper must rely on the thematic analysis to support 'enhanced perceived collaboration.' That analysis is described only as ChatGPT-4o with a structured prompt, reviewed by the first author. There is no codebook, no prompt text, no response count, and no reliability metric. This is exactly the reader's weakest assumption, and I agree with it. If the themes are not reproducible or are biased by the prompt, the abstract overclaims and the contribution reduces to a null effect with exploratory observations. The proposed check—independent human coding with a pre-registered codebook and inter-rater reliability—would settle whether the subjective claim is real. Because this concern is already captured by the reader's conditional verdict, I recommend no change: the paper should be accepted only if the authors either supply the missing qualitative evidence or substantially temper the perceived-collaboration claim.","tokens_in":9554,"tokens_out":3314,"duration_ms":37422,"concrete_test":"Release the exact ChatGPT prompt and the de-identified open-ended responses; have two independent human coders, blind to condition, apply a pre-registered codebook (themes: early signaling, fetcher-step minimization, trial-and-error, confusion, etc.) to all responses; report Cohen's kappa or Krippendorff's alpha and compare theme frequencies across conditions. If the themes and condition differences do not replicate with acceptable reliability, the subjective/perceived-collaboration conclusion is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that goal-sharing 'boosts collaboration perception' is not supported by the paper's quantitative results and rests entirely on an unvalidated qualitative pipeline. In §V.B, the satisfaction ANOVA is non-significant (F=2.24, p=0.11), and no direct measure of perceived collaboration or trust is reported. The only evidence for subjective benefit is the thematic analysis in §V.B.b, which was performed by ChatGPT-4o with an unspecified structured prompt and 'reviewed and validated' by a single author. No codebook, prompt, response counts, or inter-rater reliability are provided, and the percentages (e.g., 40%, 45%, 50%) cannot be audited. If those themes are prompt-sensitive or not reproducible, the abstract's 'fosters trust and enhances perceived collaboration' collapses; the remaining result is a clean null on performance and satisfaction. The Discussion §VI.B asserts participants 'reported feeling more in control and satisfied' without citing any measured construct, further indicating an overreach from exploratory qualitative comments.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a between-subjects user study (N = 279 after exclusions) in a worker–fetcher game, comparing three conditions: no goal recognition (NR), continuously visible viable goals (VG), and viable goals on demand (VGod). The authors find no statistically significant differences across conditions in objective performance (steps), task duration, cognitive load (NASA-TLX), or explanation satisfaction. They also find no differences in need for cognition or AI attitude across groups. On this null basis, the paper argues that sharing an agent's inferred goals does not improve performance but does provide subjective value, based on a thematic analysis of open-ended responses performed with ChatGPT-4o and reviewed by the first author. The abstract and conclusion further state that goal sharing 'fosters trust and enhances perceived collaboration.'","tokens_in":9860,"tokens_out":2837,"duration_ms":29388,"significance":"If the subjective-benefit claim were well supported, this paper would be a valuable empirical demonstration of a perception–performance disconnect in ad-hoc human-agent teamwork, complementing earlier results by Pérez-D'Arpino et al. and Le Guillou et al. The objective null results, the absence of differences in cognitive load, and the pre-registration-like reporting of covariates (N4C, AIAS-4) are useful contributions, and the paper's candid discussion of evaluation failure modes in Section VI.E is a strength. However, the central positive claim about perceived collaboration currently rests on an unvalidated, non-reproducible qualitative analysis, and no direct measure of trust or perceived collaboration is reported. Consequently, the significance is at present limited to the null performance result and the identification of open design questions.","major_comments":[{"comment":"The central claim that goal sharing 'boosts collaboration perception' rests entirely on the thematic analysis in §V.B.b. This analysis provides no codebook, no full prompt text, no response counts, and no inter-rater or inter-method reliability metrics; the reported percentages (40%, 45%, 50%, etc.) cannot be audited. Because the quantitative satisfaction ANOVA in §V.B is non-significant (F = 2.24, p = 0.11), and no direct measure of perceived collaboration or trust is reported, the qualitative pipeline is load-bearing for the paper's main conclusion. The authors should either provide a fully auditable qualitative analysis (including the exact prompt, coding rules, response corpus, and reliability assessment) or explicitly reframe the perceived-collaboration claim as an exploratory hypothesis requiring further confirmation.","section":"V.B.b"},{"comment":"The Discussion states that 'Participants reported feeling more in control and satisfied with the agent's actions when they had access to the agent's perceived goals,' but the manuscript reports no measured construct of 'control' and no significant difference in satisfaction scores. This sentence overstates what the data show. The authors should either support this claim with a measured variable or remove/qualify it, since it directly feeds the abstract's assertion that goal sharing 'fosters trust and enhances perceived collaboration,' for which no trust metric is reported anywhere in the study.","section":"VI.B"},{"comment":"Effect sizes and a power or sensitivity analysis are absent throughout the results. Given the ambitious claim that goal sharing provides 'no performance improvement,' the null ANOVAs in §V.A (F = 0.036, p = 0.965) and §V.C (F = 1.455, p = 0.235) should be accompanied by effect sizes (e.g., partial η²) and an indication of the minimum detectable effect at the obtained sample size. Without this, the reader cannot distinguish a precise null from an underpowered one, which is load-bearing for the paper's interpretation of the performance result.","section":"V.A"},{"comment":"The participant exclusion procedure is not specified: the paper reports that after filtering for 'incomplete data, attention check failures, and outliers' 279 of 313 participants remained, but it does not define the outlier criterion, the attention-check threshold, or the incomplete-data rule. This affects the reproducibility of the central null findings and should be reported in the methods section.","section":"IV.A"}],"minor_comments":[{"comment":"The paper repeatedly types 'ANOV A' with an internal space (e.g., 'ANOV A tests confirmed', 'We conducted a one-way ANOV A'); this should be corrected to 'ANOVA' consistently.","section":"Throughout Sections V and VI"},{"comment":"The caption says 'adjusted s.t. higher scores indicate better outcomes,' but the NASA-TLX is conventionally scored with higher scores indicating worse load; the adjustment procedure should be described in the text or figure caption, otherwise readers may misinterpret the direction of the scale.","section":"Figure 3"},{"comment":"The threshold is defined as 'h = 1 −η' with η = 0.05, so h = 0.95; stating this explicitly would avoid ambiguity about the probability threshold used for goal commitment.","section":"III.B"},{"comment":"The manuscript states that 'The full survey is available in the supplementary material,' but no supplementary file is included in the submission; the authors should either provide the survey or delete this statement.","section":"IV.C"},{"comment":"The analysis section mentions SHAP values in XGBoost models and latent profile analysis, but no results from these analyses appear in the paper; please clarify whether these were exploratory and, if so, report them briefly or state that they are omitted for space.","section":"IV.D"},{"comment":"Reference [17] lists the venue as 'Adaptive and Learning Agents Workshop at AAMAS 202' with no year; the publication year is missing.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid empirical null-result study with an under-supported positive claim. The qualitative analysis in §V.B.b is the only evidence for 'perceived collaboration,' and it is not reported in a way that allows verification; without it the paper reduces to a clean null result on performance, satisfaction, and cognitive load. I would advise asking the authors to either make the qualitative analysis fully auditable (prompt, codebook, response counts, reliability) or soften the central claim to an exploratory finding, and to add effect sizes for the null results. With those changes the paper could be acceptable for the venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a solid, large-N user study showing that showing a human teammate the AI's inferred goals does not improve task performance or satisfaction. The on-demand condition is a nice addition to prior work. But the abstract overclaims a boost in perceived collaboration: the only support is a ChatGPT-generated thematic analysis with no codebook, no prompt, no inter-rater reliability, and the satisfaction ANOVA is null (F=2.24, p=0.11). That overreach is the main problem.\n\nWhat's new: the VGod condition (a button to see the agent's viable goals on demand) is a reasonable extension of earlier intention-sharing studies. The sample is large (279 after filtering), attention checks are used, and the nulls on performance, duration, and cognitive load are clean and credible. The failure-modes section in VI.E is genuinely useful for the community. I believe the null result.\n\nWhere it's soft: first, the perception claim. The abstract says goal-sharing 'fosters trust and enhances perceived collaboration,' but no trust measure was collected, and the satisfaction ANOVA is non-significant. The thematic analysis is not auditable: ChatGPT-4o with an unspecified structured prompt, 'validated' by the first author, with percentages that can't be traced. The paper even admits in V.B that satisfaction wasn't significant, yet Discussion VI.B says participants 'reported feeling more in control and satisfied' with no measured construct. That is a real overstatement.\n\nSecond, missing effect sizes and power analysis make it hard to know how confident the nulls are. Third, the central result already appears in the cited literature (Perez-D'Arpino et al., Le Guillou et al.), so novelty is incremental—the on-demand condition plus the domain.\n\nIf the authors either make the thematic analysis reproducible (codebook, prompts, reliability) or explicitly label it exploratory and remove 'fosters trust' from the abstract, this becomes a useful data point. As written, the abstract oversells.\n\nWho's this for? People working on XAI transparency and human-agent teamwork who want another case where perceived help and objective performance diverge. Not a breakthrough, but a legitimate empirical study.\n\nMy recommendation: send to peer review—it deserves a serious referee—but expect major revision on the subjective-benefit claim. I wouldn't cite it for 'perceived collaboration' until the qualitative pipeline is transparent.","headline":"Clean null result on goal-sharing transparency, but the perceived-collaboration claim rests on an unauditable LLM thematic analysis and overstates what the data support.","tokens_in":10202,"tokens_out":2489,"would_cite":false,"duration_ms":23179,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Seeing an AI teammate's inferred goals changes how collaboration feels, while measured performance stays the same.","keywords":["human-agent teamwork","ad hoc teamwork","goal recognition","explainable AI","information sharing","cognitive load","user study","thematic analysis"],"falsifier":"Pre-register a replication with two independent human coders blind to condition who code the open-ended responses from a fixed codebook; if inter-rater agreement is weak or the condition differences in themes disappear, the perceived-collaboration claim is unsupported. A larger preregistered study with satisfaction as the primary endpoint would also settle whether the null comparison (F=2.24, p=0.11) persists.","tokens_in":9371,"feed_emoji":"🤖","tokens_out":9446,"duration_ms":84824,"temperature":0.7,"pith_summary":"This paper reports a 279-participant experiment in which a person moved a worker toward one of four goal stations while an AI fetcher inferred the goal and fetched a tool; some participants could see the agent's displayed viable-goal beliefs at all times, on demand, or not at all. The paper claims that giving people access to the agent's inferred goals did not significantly improve task performance, completion time, satisfaction, or cognitive load, but qualitative thematic analysis indicates participants used the information to adapt strategically and felt the collaboration was more understandable. The intended contribution is to show a gap between perceived collaboration and objective performance: transparency about an AI teammate's mind can foster trust and perceived teamwork without delivering measurable efficiency gains.","feed_headline":"Seeing AI goals boosts perceived collaboration, not performance","feed_subtitle":"Seeing the agent's inferred goals changed strategies and trust, but task steps and satisfaction stayed flat","key_machinery":"The central object is the viable-goals display: the agent maintains a probability distribution over the four numbered goal stations, multiplies a station's probability by a learning rate close to zero when the worker moves away from it, renormalizes, and treats a station as the inferred goal once its probability exceeds a threshold. In the viable-goals condition this set of still-possible stations was always visible; in the on-demand condition it appeared when the participant pressed a button; in the no-recognition condition it was never shown. This display is the treatment that changes what information the human teammate can act on, and it is the mechanism the authors argue supports strategic adaptation and perceived transparency. The rest of the machinery is the between-subjects user study with standardized questionnaires and a language-model-assisted thematic analysis of open-ended responses, reviewed by the first author, used to compare the three conditions.","core_discovery":"The paper's central claim is that in ad-hoc human-agent teams, sharing the agent's inferred beliefs about the human's goal changes the subjective experience of collaboration more than it changes objective outcomes. Across three conditions, the objective metrics of steps and duration showed no statistically significant differences, and satisfaction scores did not differ significantly either. Yet thematic analysis of open-ended responses found that participants with access to the agent's beliefs described more strategic behavior, such as minimizing the fetcher's path or using trial-and-error, and the authors interpret this as evidence that goal sharing fosters trust and enhances perceived collaboration. The discovery is the dissociation itself: perceived collaboration can improve even when task performance and reported satisfaction do not.","pith_inferences":["A testable extension would decouple perceived transparency from real inference: showing a plausible but arbitrary set of viable goals would reveal whether the subjective benefit comes from the content of the belief or from the mere impression that the agent has a readable mind.","The finding that on-demand participants reported more trial-and-error suggests that optional access may encourage exploration; behavioral metrics such as path entropy or the timing of information requests could test this without relying on self-report.","If this dissociation generalizes, goal-sharing should be deployed where trust and perceived support are the product, such as assistive or educational AI, and avoided where raw speed and minimal steps dominate the objective."],"forward_implications":["In real-time collaborative tasks, showing users an AI teammate's inferred beliefs should not be expected to reduce completion time or the number of steps needed.","Goal-sharing can be treated as a trust and user-experience feature: it can make collaboration feel more transparent and controllable without measurably burdening the user.","Satisfaction scores and perceived collaboration can diverge, so future evaluations should measure the two separately rather than treating them as one construct.","Because average cognitive load did not differ across conditions, the lack of performance gain is not explained by overall overload; localized cognitive spikes and misinterpretation remain plausible explanations."],"supporting_citations":[{"why":"Provides the closest prior result: intention-sharing increased trust and perceived collaboration but not task performance, the pattern this study replicates.","marker":"[6]"},{"why":"Shows AI intention communication can raise trust without improving objective performance, extending the perception-performance distinction.","marker":"[7]"},{"why":"Documents a perception/performance gap in human-robot interaction where information requests are perceived as interruptions without objective interference.","marker":"[8]"},{"why":"Supplies cognitive load theory, used to explain why extra information might fail to improve performance.","marker":"[10]"},{"why":"Shows AI explanations can raise perceived usefulness and trust without improving learning, an analogue of the subjective/objective split.","marker":"[13]"},{"why":"Origin of the tool-fetching ad-hoc teamwork domain and communication-value framing that the experiment modifies.","marker":"[15]"},{"why":"The NASA-TLX instrument used to measure cognitive load across conditions.","marker":"[18]"}],"fun_headline_variants":["Goal sharing boosts perceived collaboration, not task performance","AI goals change perception, not outcomes","Sharing inferred goals boosts trust, not efficiency","Inferred goals: better vibes, same step count","Goal transparency changes perceived teamwork, not scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that goal sharing improves perceived collaboration rests on themes generated by one language model and vetted by one author, with no independent coding or reliability check, so the themes could reflect the tool's assumptions rather than what participants actually experienced.","fun_headline_variants_meta":{"raw":{"variants":["Goal sharing boosts perceived collaboration, not task performance","AI goals change perception, not outcomes","Sharing inferred goals boosts trust, not efficiency","Inferred goals: better vibes, same step count","Goal transparency changes perceived teamwork, not scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001068,"raw_usage":{"total_tokens":4422,"prompt_tokens":841,"completion_tokens":3581,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":3511}},"tokens_in":457,"tokens_out":3581,"duration_ms":23013,"temperature":1.0,"reasoning_tokens":3511,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:44:15.357913+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Pre-register a replication with two independent human coders blind to condition who code the open-ended responses from a fixed codebook; if inter-rater agreement is weak or the condition differences in themes disappear, the perceived-collaboration claim is unsupported. A larger preregistered study with satisfaction as the primary endpoint would also settle whether the null comparison (F=2.24, p=0.11) persists.","supporting_citations":[{"cited_title":"Experimental assessment of human-robot teaming for multi-step remote manipulation with expert operators,","cited_arxiv_id":null,"evidence_quote":"Provides the closest prior result: intention-sharing increased trust and perceived collaboration but not task performance, the pattern this study replicates."},{"cited_title":"Trusting artificial agents: Communication trumps performance,","cited_arxiv_id":null,"evidence_quote":"Shows AI intention communication can raise trust without improving objective performance, extending the perception-performance distinction."},{"cited_title":"Exploring the cost of interruptions in human-robot teaming,","cited_arxiv_id":null,"evidence_quote":"Documents a perception/performance gap in human-robot interaction where information requests are perceived as interruptions without objective interference."},{"cited_title":"Toward personalized xai: A case study in intelligent tutoring systems,","cited_arxiv_id":null,"evidence_quote":"Shows AI explanations can raise perceived usefulness and trust without improving learning, an analogue of the subjective/objective split."},{"cited_title":"Expected value of communication for planning in ad hoc teamwork,","cited_arxiv_id":null,"evidence_quote":"Origin of the tool-fetching ad-hoc teamwork domain and communication-value framing that the experiment modifies."},{"cited_title":"Nasa task load index (tlx). volume 1.0; paper and pencil package,","cited_arxiv_id":null,"evidence_quote":"The NASA-TLX instrument used to measure cognitive load across conditions."}],"review_version":1}