{"id":"4c41a022-9385-4b1e-914c-9d31a952e5d2","arxiv_id":"2505.06428","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Passengers understood AV driving scenarios better when given a static text explanation, with or without a follow-up question feature, than when they only watched driving videos or only asked questions.","lead":"This paper reports two lab studies on what passengers ask an autonomous vehicle and whether answering those questions helps them understand its driving decisions. Participants who received a written explanation plus a question-and-answer feature scored higher on a driving-scenario quiz than those who only watched videos.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Q&A adds no measurable benefit over static explanations; the significant effect is driven by static text, undercutting the central interactivity claim.","rationale":"The reader's weakest assumption was that SHRP2 human driving footage stands in for AV behavior, which is a genuine external-validity threat. However, the more immediate problem is internal to Study 2: the experimental design already contains the evidence needed to test whether Q&A matters, and that evidence points against the paper's central claim. The pairwise comparisons show Q&A-only is indistinguishable from baseline and static+Q&A is indistinguishable from static, so the significant omnibus effect is attributable to the static text rather than to answering end-user questions. A 2x2 factorial analysis would make this explicit and would directly test the paper's central assertion. The quiz's derivation from the same metadata as the static explanations (Appendix B) compounds the issue, since the measured 'understanding' may be recall of answer content; Section 4.5.4 acknowledges this possibility. The taxonomy from Study 1 and the existence of a positive static-explanation effect are reasonable contributions, and the data are not fabricated or internally inconsistent, so rejection is not warranted. The paper should be accepted only after reframing claims away from the 'importance of answering end-user questions' and toward the demonstrated value of static explanations, or after new evidence shows an incremental Q&A effect. This supports the reader's CONDITIONAL verdict, hence UNCHANGED, but for a different primary reason: not the human-driver simulation validity assumption, but the lack of any measurable incremental effect of the interactive Q&A component.","tokens_in":33322,"tokens_out":5763,"duration_ms":60953,"concrete_test":"Run a 2x2 factorial ANOVA on Study 2's driving-scenario comprehension scores, with factors Static (present/absent) and Q&A (present/absent), including the interaction, using the same ART-C and Holm-Bonferroni framework. If the Q&A main effect and the Static x Q&A interaction are not significant while the Static main effect is, the manuscript's conclusion that answering follow-up questions improves comprehension should be revised to credit static explanations, and the title/abstract should be softened accordingly. As a robustness check, re-score the quiz on held-out scenario facts not present in the static narratives to test whether the static advantage reflects memorization rather than transferable understanding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Study 2's design is effectively a 2x2 factorial (static explanation present/absent, Q&A present/absent), but the analysis is reported only as a one-way ANOVA with pairwise contrasts. The pairwise results show Q&A-only is not significantly better than baseline (p=0.467) and static+Q&A is not significantly better than static (p=0.467, estimate -6.3 points). Therefore the reliable driver of the significant omnibus effect F(3,70)=18.31 is the presence of static text: both static and static+Q&A beat baseline and Q&A-only, but adding Q&A to static adds nothing measurable, and Q&A without static adds nothing. This directly undercuts the abstract's and introduction's claim that answering end-users' follow-up questions improves comprehension of AV decisions: the data support 'static natural-language explanations improve quiz scores,' not 'interactivity matters.' Moreover, the comprehension quiz is constructed from the same expert metadata embedded in the static narratives (Appendix B), so the static advantage may partly reflect answer memorization rather than understanding of AI decisions. The paper's own Section 4.5.4 concedes the assessment may reward having the information readily available. The AI-literacy result is only marginal (p=0.0518), so it cannot carry the interactive claim. A conditional acceptance requiring a factorial reanalysis and reframing is warranted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports two user studies on end-user explainability for autonomous vehicles (AVs). Study 1 (N=17) uses a Wizard-of-Oz conversational interface with SHRP2 naturalistic driving videos, presented as if recorded from a fully autonomous vehicle, to elicit and taxonomize passenger questions about AI-driven AV behavior. Five question categories are identified, including situational awareness, hypothesis testing, contesting decisions, trust repair, and general AI curiosity. Study 2 (N=83) compares four between-subjects conditions—baseline, static text explanation, Q&A explanation, and static+Q&A—on driving-scenario comprehension, AI literacy, and workload. The authors report a significant overall effect of condition on scenario comprehension (F(3,70)=18.31, p<0.001), a marginal effect on AI literacy (p=0.0518), and no workload differences. They conclude that interactive text-based explanations improve end-users' understanding of AV decisions and inform the design of human-centered XAI.","tokens_in":33525,"tokens_out":3785,"duration_ms":40608,"significance":"If the central claim were supported, the paper would make a useful contribution to human-centered XAI for autonomous driving: it provides a taxonomy of passenger questions, a prototype conversational XAI system, and an attempt at objective comprehension assessment. The authors should also be credited for transparent reporting of effect sizes, Holm-Bonferroni-corrected ART-c contrasts, open prototype code, and an explicit discussion of LLM answer accuracy. However, the headline claim that answering end-user questions improves comprehension is not supported by the data as analyzed. The reliable result is that static narrative explanations improve quiz performance; adding Q&A to static explanations yields no measurable gain, and Q&A alone yields no significant gain over baseline. The paper's value therefore depends on substantially reframing the contribution from 'interactivity matters' to 'static natural-language explanations help, while the additional benefit of follow-up Q&A remains to be demonstrated.'","major_comments":[{"comment":"The abstract and introduction state that interactive text-based explanations improved comprehension of AV decisions, but the pairwise contrasts in Table 2 do not support this. Q&A alone is not significantly different from baseline (p=0.467), and static+Q&A is not significantly different from static alone (p=0.467, estimate -6.3 points). The significant omnibus effect F(3,70)=18.31 is carried by the presence of static text: both static and static+Q&A beat baseline and Q&A-only. The manuscript should be rewritten so that the central claim is 'static natural-language explanations improved comprehension scores,' with interactivity reported as an unconfirmed or limited result.","section":"Abstract and §4.5.1, Table 2"},{"comment":"Study 2's design is effectively a 2x2 factorial (static explanation present/absent, Q&A present/absent), but the analysis is reported only as a one-way ANOVA with pairwise contrasts. Because the central question is whether Q&A adds value beyond static text, the authors should also report a factorial ANOVA with the interaction term. The current analysis leaves the key interaction untested and the interpretation underdetermined, especially given the non-significant static vs. static+Q&A contrast.","section":"§4.1 and §4.5.1"},{"comment":"The driving-scenario comprehension quiz was constructed directly from the expert-annotated metadata, and the same metadata were used to generate the static narratives and the LLM prompts. Participants in the static and static+Q&A conditions therefore received the answers to many quiz questions in advance. The paper itself concedes in §4.5.4 that the assessment may reward having the information readily available. As a result, the measured comprehension gains may reflect verbatim information availability rather than understanding of AI decision-making. The authors should either add a transfer task with novel scenarios, or explicitly reframe the measure as an 'information availability' check and temper the claims about understanding.","section":"§4.1 and Appendix B"},{"comment":"The AI literacy result is marginal (p=0.0518), and none of the corrected pairwise comparisons reaches significance. This result cannot carry the paper's interactive-explanation claim. If the authors wish to claim an AI literacy benefit for static+Q&A, they need a preregistered and adequately powered test; the current study was powered for a large effect (f=0.4) on the primary outcome, and the AI literacy data should be treated as exploratory.","section":"§4.5.2 and Table 3"},{"comment":"Participants were told the videos were recorded from a fully autonomous vehicle, but the SHRP2 NDS footage shows human drivers, including human errors such as distracted driving and unsafe maneuvers. The explanations and quiz questions consequently attribute human errors to the AI. The paper should explicitly discuss whether human naturalistic driving incidents can stand in for AI-driven AV decisions; if actual AV behavior differs, the passenger question taxonomy and the measured comprehension effects may not transfer. At minimum, this is a boundary condition on all conclusions and should be flagged in the limitations.","section":"§3.4 and §4.3"}],"minor_comments":[{"comment":"The reported question-category percentages do not sum to 100% (15.5 + 18.5 + 18.5 + 23.9 + 23.3 = 99.7). Please clarify whether the remaining 0.3% is rounding or a small uncategorized residue.","section":"§3.6"},{"comment":"The x-axis label 'Not Releted' is misspelled; it should be 'Not Related'.","section":"Figure 6"},{"comment":"The terms 'driving scenario comprehension' and 'task expertise' are used interchangeably. Please define the relationship between the two at first use to avoid confusion.","section":"§4.1 and §4.5.1"},{"comment":"Question 3 in the AI literacy quiz is phrased as 'Select all that apply,' but the scoring method for this question is not described; please specify whether full or partial credit was awarded.","section":"Appendix C, Q3"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-executed in many respects, but its central claim is over-stated relative to the data. The Study 1 taxonomy is a solid empirical contribution, and the Study 2 data do show a robust effect of static text explanations. However, the 'interactivity improves understanding' narrative is not supported by the pairwise contrasts, and the outcome measure shares content with the intervention. A major revision focused on reframing the claims and adding a factorial analysis would make the contribution honest and publishable, but the current framing should not be accepted as is."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, usable empirical paper, but the central claim about interactive Q&A is not what the data show. The reusable asset is the passenger question taxonomy from Study 1; the reliable result from Study 2 is that static natural-language explanations raise driving-scenario comprehension. Q&A alone does not beat baseline, and adding Q&A to static does not beat static.\n\nWhat the paper does well: the Wizard-of-Oz elicitation of free-form questions is well executed and yields five categories (situational awareness, hypothesis testing, contesting decisions, trust repair, general AV/AI understanding) that should be useful to anyone designing in-vehicle explanation systems. The second study actually measures comprehension with an objective quiz rather than relying only on trust ratings, annotates LLM answer accuracy, and ships the prototype code. That is real, citable work.\n\nSoft spots, in order. First, the study is effectively a 2x2 (static present/absent; Q&A present/absent) but it is analyzed and framed as a one-way comparison of four interfaces. The pairwise contrasts undermine the abstract's “interactive text-based explanations effectively improved comprehension”: Q&A vs baseline is p=0.47, static+Q&A vs static is p=0.47. The significant omnibus effect is driven by the static text. The paper should report the factorial analysis and say plainly that interactivity, on this evidence, adds no measurable comprehension benefit. Second, the comprehension quiz items are built from the same expert metadata embedded in the static narratives, so part of the static gain may be answer availability. Section 4.5.4 concedes this. It does not break the static-explanation result, but it does mean “understanding” is partly “had the information at hand.” Third, the videos are SHRP2 naturalistic footage of human drivers, including human errors, while participants were told they were riding in a fully autonomous vehicle. The question taxonomy may transfer, but the assumption that AV errors will resemble human near-crashes is unexamined. Minor: the LLM is a fixed 2023 model, so the exact interaction is not reproducible, though the authors do report accuracy and release code.\n\nWho this is for: HCI and human-centered XAI researchers working on AV explanation interfaces. It deserves a serious referee, but the revision should re-center the contribution on the taxonomy and the static-explanation effect, with the Q&A result presented as exploratory. I'd send it out.","headline":"Useful passenger question taxonomy and a solid demonstration that static explanatory text improves comprehension, but the paper's interactivity claim is not supported by its own pairwise contrasts.","tokens_in":34056,"tokens_out":2751,"would_cite":true,"duration_ms":26922,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Passengers understand autonomous-driving decisions best when a static explanation is followed by a Q&A dialogue.","keywords":["Explainable AI","Autonomous vehicles","Conversational XAI","User-led explanations","AI literacy","Human-centered XAI","Interactive explanations","Passenger questions"],"falsifier":"Re-run the Study 2 experiment using footage and explanation content drawn from actual autonomous vehicles, such as disengagement or incident logs from operating self-driving fleets, rather than human naturalistic driving data; the central claim would be refuted if the static+Q&A condition no longer outperformed the video-only baseline on the scenario-comprehension quiz.","tokens_in":33109,"feed_emoji":"🚗","tokens_out":12301,"duration_ms":100954,"temperature":0.7,"pith_summary":"This paper asks what passengers actually want to know about the AI driving their self-driving car, and whether answering those questions improves their understanding. In two user studies the authors first collected over 300 spontaneous passenger questions, which fall into five needs: tracking the vehicle's context and status, diagnosing events, contesting decisions, repairing trust after incidents, and learning what the AI can and cannot do. They then tested four explanation formats and found that a static text explanation followed by an interactive question-and-answer dialogue produced the highest comprehension of driving scenarios, significantly above both watching videos with no explanation and a question-answer interface without the static seed. The same combined format also raised AI literacy scores marginally relative to baseline. The authors argue that next-generation explainable AI for autonomous vehicles should be dialogue-based and grounded in the passenger's own questions rather than a fixed set of engineer-authored justifications.","feed_headline":"Answering riders' follow-up questions boosts AV comprehension","feed_subtitle":"A two-study experiment finds that written explanations plus interactive Q&A help passengers understand driving decisions","key_machinery":"The load-bearing machinery is the two-stage conversational explainable AI setup: a static natural-language explanation, created by rewriting expert narratives from the vehicle's perspective, that seeds the dialogue, followed by a chat interface that answers follow-up questions using a large language model grounded in expert-annotated scenario metadata. The four experimental conditions isolate the contribution of each stage. The measurement instrument is a multiple-choice scenario-comprehension quiz built from the expert metadata, plus an AI-literacy quiz based on established AI literacy competencies. The question taxonomy from the initial Wizard-of-Oz study supplies the categories of needs that such a system must be able to handle.","core_discovery":"The central claim is that end-user understanding of AI-driven autonomous vehicle decisions is best served by layering live question answering on top of a static, task-focused explanation rather than by any single explanation format. On a driving-scenario comprehension quiz, the static and static+Q&A conditions both outperformed the video-only baseline, and static+Q&A also outperformed Q&A-only, which did not differ significantly from baseline; the overall effect of condition was significant, $F(3,70)=18.31$, $p<0.001$. The accompanying qualitative finding is that passenger questions cluster into five categories—AV context and status, hypothesizing and diagnosing events, contesting decisions, repairing trust, and learning system capabilities—most of which existing engineer-oriented explainable AI tools do not answer. The paper concludes that explanation systems should support a dialogue, seed the conversation to compensate for passengers not knowing what to ask, and cover the socio-technical context rather than only model internals.","pith_inferences":["The seed-then-answer pattern suggests a general design rule for explainable AI beyond driving: give the user a narrative of what happened first, because users often do not know what to ask, and then let follow-up questions target their specific gaps.","The five-category taxonomy may be reusable as an auditing instrument for explanation systems in other high-stakes AI contexts where users need to diagnose, contest, and rebuild trust, such as medical or financial decision aids.","A testable extension is to vary the order and length of the static seed (full narrative versus a one-sentence summary) to find the minimum context needed before the question-and-answer stage becomes effective.","Because participants still failed many technical AI-literacy items even in the best condition, an explanation system that aims to raise literacy may need to proactively introduce technical topics rather than wait for users to ask about them."],"forward_implications":["A conversational explanation interface that only answers questions, without first giving a static account of what happened, does not measurably improve scenario comprehension, so designers should treat a seed explanation as necessary context for useful follow-up questions.","Passengers will ask to diagnose, contest, and repair trust after critical events, so AV explanation systems should include post-incident functions such as damage assessment, fault attribution, and next-action planning, not just 'why' justifications.","LLM-generated answers that stay within the scenario metadata are highly accurate for scenario-related questions, while answers to general AI questions are much less reliable, so the knowledge base must be expanded before conversational explainable AI can teach broader AI literacy.","The five-category taxonomy gives designers a concrete checklist for evaluating whether an explanation interface supports situation awareness, diagnosis, contestation, trust repair, and learning about capabilities.","Because workload did not differ across conditions, richer interactive explanations can be added without imposing additional cognitive burden on passengers."],"supporting_citations":[{"why":"Supplies the naturalistic driving videos plus expert-annotated metadata and narratives used as stimuli and answer content in both studies.","marker":"[79]"},{"why":"Provides the AI literacy competencies that the AI-literacy assessment and its category analysis are based on.","marker":"[56]"},{"why":"Is the large language model used in Study 2 to generate conversational answers to participants' questions.","marker":"[67]"},{"why":"Supports the Wizard-of-Oz method used in Study 1 to elicit spontaneous passenger questions.","marker":"[54]"},{"why":"Grounds the analysis of contrastive and contesting question types, such as 'why not' and 'what-if' questions.","marker":"[60]"},{"why":"Provides the Align Rank Transform used to run ANOVA on the non-normal comprehension and literacy scores.","marker":"[95]"},{"why":"Provides the ART-c post-hoc contrasts with Holm-Bonferroni corrections that establish which conditions differ.","marker":"[25]"},{"why":"Supplies the three-level situation awareness model used to interpret passengers' questions about perception, comprehension, and projection.","marker":"[26]"}],"fun_headline_variants":["Adding rider Q&A to AV explanations boosts comprehension","Riders' questions, when answered, clarify AV decisions","Answering passengers' AV questions boosts comprehension","Interactive Q&A helps riders grasp autonomous vehicle choices","Q&A with riders unlocks better AV decision understanding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Participants were told the videos were recorded from a fully autonomous vehicle, but the footage actually shows human drivers, including human errors such as distraction and unsafe maneuvers, so the study assumes that human naturalistic driving incidents and expert narratives about them are a valid stand-in for genuinely AI-made driving decisions and their explanations.","fun_headline_variants_meta":{"raw":{"variants":["Adding rider Q&A to AV explanations boosts comprehension","Riders' questions, when answered, clarify AV decisions","Answering passengers' AV questions boosts comprehension","Interactive Q&A helps riders grasp autonomous vehicle choices","Q&A with riders unlocks better AV decision understanding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001525,"raw_usage":{"total_tokens":6086,"prompt_tokens":904,"completion_tokens":5182,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":5110}},"tokens_in":520,"tokens_out":5182,"duration_ms":33983,"temperature":1.0,"reasoning_tokens":5110,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:42:09.814320+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Study 2 experiment using footage and explanation content drawn from actual autonomous vehicles, such as disengagement or incident logs from operating self-driving fleets, rather than human naturalistic driving data; the central claim would be refuted if the static+Q&A condition no longer outperformed the video-only baseline on the scenario-comprehension quiz.","supporting_citations":[{"cited_title":"Perez, Koji Dan, Tomohiro Shimamiya, Toshihiro Hashimoto, Masahiro Kimura, Sachiko Yamada, and Toshiaki Seo","cited_arxiv_id":null,"evidence_quote":"Supplies the naturalistic driving videos plus expert-annotated metadata and narratives used as stimuli and answer content in both studies."},{"cited_title":"Large, Gary Burnett, and Leigh Clark","cited_arxiv_id":null,"evidence_quote":"Supports the Wizard-of-Oz method used in Study 1 to elicit spontaneous passenger questions."}],"review_version":1}