{"id":"8c8f98e1-6519-4e9e-9bd7-c466655ac843","arxiv_id":"2411.11449","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A qualitative study finds modular explanations support shared understanding and argument-building in groups, while individuals engage more deeply and perform better on tasks, trading depth for exchange.","lead":"Does explainable AI work differently for people deciding alone versus in groups? This interview study with 43 AI novices shows group discussion built shared understanding and arguments, while individuals dug deeper and solved more tasks, pointing to a real trade-off for public AI deliberation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The positive-understanding conclusion rests on an untested post-hoc 'calibrating understanding' construct; flat self-reports and the lack of a no-explanation control leave the explanation's learning effect unestablished.","rationale":"The reader's concern is the right one. The paper's comparative process claims are supported by transcribed interactions and by the reported task-performance difference, so those should not be dismissed. But the interpretive leap to 'explanations improved understanding' is the load-bearing part of RQ1 and of the abstract's claim that explanations support novices. Table 3 directly undercuts a naive before/after reading; Section 4.1.3 invokes calibrating understanding to reconcile this, and Section 5.2 acknowledges that the construct is not independently measured. The paper cites prior work (Keil 2006; Cheng et al. 2019), but the construct has no probe in this study. The concrete test I propose would turn the post-hoc reconciliation into a testable prediction: pre/post task performance and a perceived-information-scope item should move in the predicted pattern. If they do not, the correct conclusion is not that the explanation design failed entirely, but that the empirical support is limited to interaction processes and argumentation, not measured learning. Because the reader already conditioned acceptance on this concern, I keep the verdict unchanged.","tokens_in":46826,"tokens_out":4602,"duration_ms":51567,"concrete_test":"Run a follow-up or reanalysis in which participants complete the four task questions immediately after the newspaper article (pre-test), then again after the explanation phase (post-test), along with a new item asking 'How much information do you currently have about the system?' If post-test task scores improve while understanding self-ratings stay flat and perceived information scope increases, the calibration interpretation is supported. If verbal improvement claims are made without task improvement, or if perceived information scope does not track verbal claims, the positive-understanding conclusion should be downgraded to a process-only claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that individual and group settings provide different grounds for understanding AI systems depends on the claim that the explanations actually supported understanding. The study's direct quantitative measures do not show this: Table 3 shows most participants reported unchanged understanding after the explanation phase, and there is no no-explanation control condition. The paper bridges this gap with a construct, 'calibrating understanding' (Sections 4.1.3 and 5.2), under which participants rate understanding relative to the currently visible information scope rather than their prior state. The construct is plausible and cites prior work (Keil 2006; Cheng et al. 2019), but it is post hoc and not independently measured. The verbal reports of improvement that the construct is invoked to rescue are exactly the reports most vulnerable to social desirability and ambiguity about what 'understanding' means. If calibrating understanding is not the correct explanation, the flat self-reports cannot be dismissed, and the remaining evidence for explanation-driven understanding is task performance, which was not measured before the explanation phase and is itself confounded by the education imbalance between single-interview and focus-group samples. The comparative process findings (groups share, individuals focus) remain, but the causal 'support understanding' component of the central claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a task-based interview study (8 focus groups and 12 single interviews; 43 AI-novice participants) examining how a modular, question-driven explanation design—36 question–answer pairs about the Austrian AMS employment-prediction algorithm, organized into data, system details, usage, and context—supports understanding and deliberation in individual versus collective settings. The authors combine before/after self-reports, four factual task questions, and thematic analysis of transcripts. They report that groups used explanations to build shared understanding, source arguments, and sometimes experienced process loss, while individuals engaged more deeply and performed better on the tasks; self-reported understanding mostly stayed flat, which the paper reconciles by introducing a 'calibrating understanding' construct. The paper closes with design recommendations for XAI for public deliberation.","tokens_in":46985,"tokens_out":5328,"duration_ms":56286,"significance":"If the central claims hold, this is a useful contribution to human-centered XAI for public-sector AI: it provides a rare empirical comparison of one-to-one versus many-to-one explanation use, includes decision-subject focus groups, and ships a transparent codebook, full explanation set, and task materials. The qualitative process findings—groups share, outsource, and argue; individuals focus and calculate—are well illustrated and likely to inform future design. However, the causal component ('explanations improved understanding') is not established by the quantitative measures, and the task-performance comparison is confounded. The value of the paper is therefore mainly in the descriptive process account and design implications, which are credible after the claims are appropriately narrowed.","major_comments":[{"comment":"The conclusion that the explanations improved understanding rests on the post-hoc, unmeasured construct 'calibrating understanding.' The quantitative self-reports are mostly flat (e.g., 9 of 12 single-interview participants report no change, and many focus-group participants do too), and there is no no-explanation control. The verbal reports used to instantiate calibration are the measures most vulnerable to social desirability and to ambiguity about the meaning of 'understanding.' Because the calibration process is not independently assessed, the flat self-reports cannot simply be set aside. Please either operationalize calibration (e.g., the information-scope rating suggested in §5.2) and collect it in a follow-up, or revise the abstract and conclusion to state that explanations were perceived as helpful and supported deliberation processes, not that they demonstrably improved understanding.","section":"§4.1.3, §5.2, Table 3"},{"comment":"The claim that 'participants in single interviews performed better in the study tasks' is confounded by education and recruitment. The single-interview sample is predominantly university-educated (11 of 12), whereas the focus-group sample includes more vocational and secondary-school participants. The manuscript acknowledges the imbalance in §6 but does not report the promised comparison restricted to university-educated participants, and task performance was not measured before the explanation phase. The observed difference may reflect pre-existing knowledge or education rather than the social setting.","section":"§4.1.3, Tables 1–3, §6"},{"comment":"The procedural asymmetry between settings undermines the comparative task-performance result. Groups had 15 minutes of orientation, 15 minutes of tasks, and a separate 10-minute group decision phase, while individuals had 20 minutes of orientation and 20 minutes of tasks with no decision phase. Focus-group participants may have spent task-phase time discussing rather than answering, and the collective decision phase could have changed their later engagement. The comparative interpretation should either treat task scores as descriptive only or analyze performance on a comparable time/phase basis.","section":"§3.3.1–3.3.2"},{"comment":"The headline comparative finding—individual and group settings support different understanding facets—is well supported as a qualitative account. However, the conclusion then asserts that 'the explanations had a positive effect on understanding' (also echoed in the abstract). This stronger causal statement is not load-bearing for the facet-difference finding and should be separated from it: the process data support claims about how explanations were used, not that the explanation phase caused a measurable increase in understanding.","section":"§4.1.5, §5.2"}],"minor_comments":[{"comment":"The level labels are inconsistent: §3.2.1 says 'base level, level 2, level 3,' while §3.2.2 says 'base level, level 1, level 2'; Figure 1 uses 'Base/Level 2/Level 3.' Please standardize the nomenclature throughout.","section":"§3.2.1, §3.2.2, Figure 1"},{"comment":"The reference to 'P3' appears to be an erroneous participant label; all other participant labels in the paper are of the form S1–S12 or focus-group IDs (e.g., S3), so please correct this reference.","section":"§5.3"},{"comment":"Table 4 is hard to read because the three decision columns are not clearly separated; consider restructuring it so Decision I, Group Decision, and Decision II are visually distinct and the color coding is described in a print-accessible way.","section":"Table 4"},{"comment":"Table 1 lists F2's education as 'n/a'; if these data are missing, state that explicitly in the table note rather than leaving a bare value.","section":"Table 1"},{"comment":"There is a typo/capitalization error in the sentence beginning 'Thus, The article served as...'—'The' should be lowercase.","section":"§3.4.1"}],"recommendation":"major_revision","confidential_remarks":"This is a solid qualitative HCI contribution with unusually transparent supplementary materials. The main issue is that the causal 'understanding improvement' claim is overstated relative to the design, and the comparative task-performance claim is confounded by demographic and procedural differences. I believe the paper can be made publishable by reframing the central claims as descriptive process findings and by explicitly de-emphasizing the causal interpretation of the flat self-reports. The 'calibrating understanding' construct is interesting but needs either independent operationalization or a clearly labeled speculative status; otherwise it will continue to draw justified skepticism."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, transparent qualitative study of a genuinely under-explored question—whether modular explanations for AI novices do different work in individual versus group deliberation. The descriptive core holds up. The causal 'explanations improved understanding' component is weaker than the framing suggests, and the authors mostly know it.\n\nWhat is actually new: it is the first in-person comparison of focus groups and solo interviews for deliberation about a public-sector AI deployment, using a modular, question-driven explanation set (36 Q&A pairs across data, system details, usage, and context, at three depth levels). The materials are in the appendix—real reproducibility for qualitative work. The analysis is careful: codebook, verbatim excerpts, triangulation of self-reports, task performance, and verbalizations. The central finding, that groups build shared understanding and source arguments while individuals engage more deeply and score better on factual tasks, is well supported by the reported interactions and the task data. Citations are honest: they build on their own prior information-needs work and position the study carefully against Chiang et al.'s null result on groups versus individuals.\n\nSoft spots, in proportion. The 'calibrating understanding' construct (Section 5.2) reconciles flat self-reported understanding with verbal claims of improvement. It is plausible and cites prior work, but it is post hoc and not independently measured. The stress-test concern is fair but narrow: the comparative finding—different settings support different facets of understanding—does not rest on that construct. What is genuinely weaker is the claim that the explanations themselves caused improved understanding. There is no no-explanation control, task performance was measured only after the explanation phase, and the single-interview sample skews university-educated relative to the focus groups. The authors acknowledge both. Group H's strong performance shows education is not the whole story, but the group-versus-individual comparison stays suggestive rather than conclusive. The Group A groupthink analysis is a model of restraint: they identify partial features without over-labeling, which is the right level of caution.\n\nWho it is for: XAI researchers and people working on participatory formats for public-sector AI—mini-publics, citizen forums, and similar deliberation spaces. The design suggestions in Figure 8 are concrete and actionable. It deserves a serious referee; with revision pressure on the understanding measurement and the remaining confound, this becomes a solid contribution. Send it to review.","headline":"A transparent, well-run qualitative study of a genuinely open question in XAI—the group/individual depth-versus-exchange trade-off holds up, but the causal 'explanations improved understanding' claim is weaker than the framing admits.","tokens_in":47547,"tokens_out":5047,"would_cite":true,"duration_ms":50081,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Same AI explanations help groups and solo users differently","keywords":["explainable AI","AI novices","group deliberation","shared understanding","modular explanations","focus groups","public sector AI","decision confidence"],"falsifier":"Run the same explanation phase with a pre/post factual understanding test plus an information-scope question, comparing individuals and groups against a no-explanation control; if verbal claims of improved understanding appear without measured information gain, or appear equally with unrelated material, the calibrating-understanding account fails.","tokens_in":1578,"feed_emoji":"💬","tokens_out":1517,"duration_ms":88353,"temperature":0.7,"pith_summary":"This paper asks whether explanations of an AI system help people without technical backgrounds understand and deliberate about it differently when they work alone rather than in a group. Based on eight focus groups and twelve individual interviews using a modular set of 36 question-answer explanations about an employment-scoring algorithm, it claims that individual and group settings support different facets of understanding. Groups used the explanations to build shared understanding and to find arguments for and against deployment, while individuals engaged more deeply, performed better on factual study tasks, and said they missed exchanging views with others. The paper concludes that explainable AI design should treat the two settings as complementary rather than assuming one explanation format fits both.","feed_headline":"Same AI explanations help groups and solo users differently","feed_subtitle":"With 43 AI novices, groups built shared understanding and debate material; solo readers solved more factual tasks.","key_machinery":"The central object is a modular, question-driven explanation collection: 36 question-answer pairs grouped into four information categories (data, system details, usage, and context), each subdivided into topics and three levels of detail, printed as physical A5 sheets that participants can sort, exchange, point to, and read aloud. The design lets users select information according to their interests, supports different levels of completeness and soundness, and is intended to work for both solo reading and collaborative interaction. The analysis maps participants' interactions onto known mechanisms of collaborative success and failure and onto facets of understanding, which lets the paper argue that each setting activates different facets and that explanations and social dynamics jointly determine whether groups reach a working understanding or abandon it.","core_discovery":"Individual and group settings provide different grounds for understanding AI systems: groups realize cognitive and social mechanisms of collaborative success that produce shared understanding, while individuals develop focused, self-directed engagement that supports applying information to tasks. With the same question-driven modular explanation design, participants in groups located information together, shared it, debated interpretations, and delegated difficult material to more competent members, whereas solo participants read more intensively, requested comparable or more explanations, and calculated precise answers that no focus group completed. The paper also finds that explanations feed deliberation: groups used them to source reasoned arguments and to surface productive disagreement, while solo participants used them for internal deliberation and several changed their deployment decisions. At the same time, a concurrence-seeking dynamic resembling aspects of groupthink led one group to follow a minority position, showing that social dynamics can override explanation content. To reconcile mostly unchanged self-reported understanding with participants' verbal claims of improvement, the paper introduces a post-hoc process called calibrating understanding, in which people judge their understanding relative to the information they now know exists.","pith_inferences":["If calibrating understanding is real, self-report-only evaluation of explainable AI will systematically understate explanation benefits; future studies should measure perceived information scope alongside self-reported understanding.","The finding that individuals solved tasks better while groups deliberated better suggests a two-phase format of individual preparation followed by group deliberation, which the paper suggests as an ideal combination but does not itself test.","A testable extension would give the same modular explanations to groups with a structured opposing role, since the paper attributes the concurrence-seeking outcome partly to the absence of a devil's advocate voice.","The physical A5 format may itself matter, because shared understanding relied on sorting, exchanging, and pointing at sheets; whether these benefits survive a digital version is an open question."],"forward_implications":["Explanation designs for AI novices should not assume one format fits both solo and group deliberation; group settings need supports for shared understanding and argumentation, while solo settings need ways to compensate for the missing exchange of perspectives.","Because individuals outperformed groups on factual study tasks, deployment decisions that hinge on technical details may be better prepared individually before being discussed collectively.","The same modular explanation collection can support deliberation without a group: solo participants used it for internal deliberation, and several changed their deployment decisions after reading the materials.","Group outcomes depend on the social dynamic as much as on the explanations: familiar, trusting groups bridged individual understanding gaps, while groups with low trust or discouragement abandoned understanding.","Explanations that supply all four information categories give groups material for reasoned arguments and disagreement, but they do not by themselves prevent concurrence-seeking behavior such as the groupthink-like pattern observed in one focus group."],"supporting_citations":[{"why":"Supplies the cognitive and social mechanisms of collaborative success and failure used to interpret group interactions with the explanations.","marker":"[97]"},{"why":"Supplies the six facets of understanding framework that lets the paper argue each setting supports different facets.","marker":"[129]"},{"why":"Provides the account of outsourcing and of locating and filling understanding gaps that underlies shared understanding and the calibrating-understanding concept.","marker":"[64]"},{"why":"The prior comparison of group and individual AI-assisted decision making that this study extends to deployment deliberation.","marker":"[22]"},{"why":"Grounds the four information categories (data, system details, usage, and context) in prior work on AI novices' information needs.","marker":"[107]"},{"why":"Documents white-box explanations raising objective understanding while lowering self-reported understanding, motivating the calibration explanation.","marker":"[21]"},{"why":"Supplies the elements of deliberation used to identify sourced arguments, disagreement, and deliberation quality in the groups.","marker":"[116]"},{"why":"Provides the groupthink criteria used to interpret the concurrence-seeking dynamic in Group A.","marker":"[55]"},{"why":"Documents discrepancies between self-reported trust and understanding and observed behavior that the calibrating-understanding account is invoked to explain.","marker":"[100]"}],"fun_headline_variants":["Explanations aid groups and solo users in different ways","For AI novices, groups gain shared understanding, solos ace tasks","Groupthink risk: AI explanations can be overridden by social dynamics","Individual vs group AI deliberation: depth vs shared insight","Same AI explanations, different outcomes for individuals and groups"],"cache_read_input_tokens":49664,"weakest_assumption_plain":"The conclusion that explanations improved participants' understanding rests on the assumption that their verbal claims of better understanding reflect genuine learning rather than politeness or confusion, since the paper's calibrating-understanding mechanism is introduced after the fact and is not independently measured.","fun_headline_variants_meta":{"raw":{"variants":["Explanations aid groups and solo users in different ways","For AI novices, groups gain shared understanding, solos ace tasks","Groupthink risk: AI explanations can be overridden by social dynamics","Individual vs group AI deliberation: depth vs shared insight","Same AI explanations, different outcomes for individuals and groups"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000491,"raw_usage":{"total_tokens":2423,"prompt_tokens":960,"completion_tokens":1463,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":1379}},"tokens_in":576,"tokens_out":1463,"duration_ms":10204,"temperature":1.0,"reasoning_tokens":1379,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:30:16.899140+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same explanation phase with a pre/post factual understanding test plus an information-scope question, comparing individuals and groups against a no-explanation control; if verbal claims of improved understanding appear without measured information gain, or appear equally with unrelated material, the calibrating-understanding account fails.","supporting_citations":[{"cited_title":"Nokes-Malach, J","cited_arxiv_id":null,"evidence_quote":"Supplies the cognitive and social mechanisms of collaborative success and failure used to interpret group interactions with the explanations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the account of outsourcing and of locating and filling understanding gaps that underlies shared understanding and the calibrating-understanding concept."},{"cited_title":"Information That Matters: Exploring Information Needs of People Affected by Algorithmic Decisions","cited_arxiv_id":"2401.13324","evidence_quote":"Grounds the four information categories (data, system details, usage, and context) in prior work on AI novices' information needs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the elements of deliberation used to identify sourced arguments, disagreement, and deliberation quality in the groups."}],"review_version":1}