{"id":"ac458a3f-5975-47c3-a64a-2eb582ea37c7","arxiv_id":"2501.10909","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A multi-step transparent AI workflow improves appropriate reliance only for users who already decide intermediate sub-facts accurately; otherwise it causes under-reliance.","lead":"This paper reports a 233-person experiment on whether showing an AI's sub-steps helps people decide when to trust it in composite fact-checking. The multi-step transparent workflow helped mostly when users were already accurate on the sub-steps, and it added cognitive load.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'misleading-advice advantage' rests on unadjusted task-level accuracy and is contradicted by Task 1; no inferential test supports the claim.","rationale":"The reader's H3 confound concern is valid: the median split on AR-Intermediate could reflect general fact-checking ability rather than explicit consideration of intermediate steps. However, the first part of the strongest claim is even more directly at risk. Table 4's task-level percentages are between-subjects, unadjusted, and not accompanied by any inferential test, and Task 1 contradicts the proposed mechanism. A formal interaction analysis would settle whether Tasks 6 and 7 reflect a real contextual advantage or chance variation. Since the concern is addressable through reanalysis, the conditional verdict remains appropriate, but the paper as written does not demonstrate the misleading-advice claim.","tokens_in":30294,"tokens_out":8741,"duration_ms":101186,"concrete_test":"Fit a mixed-effects logistic regression on per-task final decision accuracy with fixed effects for condition, a misleading-advice indicator, and their interaction, plus random intercepts for tasks and participants. Estimate the MST-vs-Control contrast on the three misleading-advice tasks. If the contrast is non-significant or becomes negative when Task 1 is included, the abstract claim that MST outperforms one-step collaboration when AI advice is misleading is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim that an MST workflow can outperform one-step collaboration when AI advice is misleading is not supported by the reported analyses. Table 4 shows the three misleading-advice tasks (IDs 1, 6, 7), but only Tasks 6 and 7 favor MSTworkflow; in Task 1, Control participants scored 33.3% versus 14.5% for MSTworkflow and 28.1% for MSTworkflow+, even though Task 1's intermediate AI advice is also misleading. No mixed-effects model, interaction test, or task-level significance test is reported, so the claim rests on selected descriptive percentages. Table 5 further shows that Control significantly outperforms all MST conditions on overall Team Performance and Team Performance-wid, the latter being the metric most relevant to disagreement with AI advice. The paper's own limitations acknowledge transferability and cognitive load but do not flag this selective task-level evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a preregistered between-subjects experiment (N = 233) on AI-assisted composite fact-checking, comparing four conditions: a one-step workflow with final AI advice (Control), one-step plus global transparency (MST-GT), a multi-step transparent workflow where users verify each decomposed sub-fact before seeing AI advice (MSTworkflow), and the same workflow plus mandatory annotation of document usefulness (MSTworkflow+). The authors test three hypotheses: H1 (global transparency increases reliance), H2 (the MST workflow increases appropriate reliance), and H3 (accurate intermediate user decisions increase appropriate reliance on the final AI advice). They report that H1 is supported, H2 is not, and H3 is supported through a median split on intermediate accuracy. The broader claim is that the MST workflow can outperform one-step collaboration specifically when AI advice is misleading, and that fine-grained appropriate reliance is a prerequisite for that benefit. The measurement framework and the use of non-parametric tests are strengths, but several reported conclusions are not backed by the analyses as presented, and the central contextual claim rests on selected descriptive percentages rather than an inferential test.","tokens_in":30492,"tokens_out":7343,"duration_ms":71906,"significance":"If the results were robust, the paper would make a useful contribution to human-AI decision making. It introduces fine-grained measures of appropriate reliance (AR-Intermediate and AR-Evidence), uses a realistic LLM/RAG-based fact-checking task, preregisters hypotheses, and provides an a priori power analysis with enough participants (230 required, 233 analyzed). Public data and code on OSF, attention-check filtering, and an honest discussion of cognitive-load trade-offs are additional strengths. The paper also moves beyond the one-step decision paradigm that dominates the literature. However, the reported analyses contain inconsistencies — most notably for H1 and for the misleading-advice claim — and the H3 analysis has a confound that the current reporting does not address. The measurement framework is valuable, but the headline conclusions need to be re-derived or substantially moderated.","major_comments":[{"comment":"The text states that 'compared to Control, MST-GT showed significantly higher Agreement Fraction' and uses this to support H1, but the reported post-hoc comparison in Table 5 is 'Control, MST-GT > MSTworkflow, MSTworkflow+.' That notation establishes only that Control and MST-GT both exceed the two MST conditions; it does not establish a significant MST-GT > Control difference. The paper therefore does not report the pairwise test that its H1 conclusion depends on, and the claim of support for H1 is unsupported as written.","section":"Section 5.2.1, Table 5"},{"comment":"The headline claim that the MST workflow outperforms one-step collaboration when AI advice is misleading is not supported by the reported analyses. Of the three tasks with misleading AI advice in Table 4, only Task 6 (25.8% vs 9.3%) and, weakly, Task 7 (62.9% vs 59.3%) favor MSTworkflow over Control; in Task 1, Control achieves 33.3% vs 14.5% for MSTworkflow and 28.1% for MSTworkflow+. No task-level significance test, mixed-effects model, or condition-by-task interaction is reported. Moreover, Table 5 shows that Control significantly outperforms all MST conditions on Team Performance-wid, the metric most directly tied to decisions where users disagree with AI advice. The abstract and Section 6.1 should either be supported by such an analysis or substantially moderated.","section":"Abstract; Section 6.1; Tables 4 and 5"},{"comment":"The support for H3 rests on a post-hoc median split of AR-Intermediate and suffers from a circularity/confound problem. AR-Intermediate is the accuracy of users' sub-fact verifications; users who are accurate on sub-facts are likely to be accurate on final fact-check decisions, and RAIR and Team Performance-wid are constructed from final-decision correctness relative to AI advice. Table 9 shows that the high-AR group simultaneously outperforms the low-AR group on Team Performance, Agreement Fraction, Team Performance-wid, and RAIR, which is exactly the pattern expected from a general-ability confound. A preregistered test of H3 or a sensitivity analysis that controls for overall task accuracy is needed before the causal interpretation in Section 6.1 ('explicit considerations ... indicate that an MST decision workflow can be effective') is warranted.","section":"Section 5.2.2; Table 9"},{"comment":"The exclusion rule for participants with 'three or more such indications' of frequent switching is post-hoc in the sense that the paper does not state it was preregistered (the preregistration link is hidden) and no robustness check is reported with these participants included. This rule differentially removes participants who exhibit strong reliance changes, which is precisely the behavior analyzed for H2 and the reliance results. The authors should either show that the exclusion was preregistered or rerun the key analyses with and without these participants.","section":"Section 4.4"}],"minor_comments":[{"comment":"The key word 'Mutli-step' is misspelled and should be 'Multi-step'.","section":"Key Words"},{"comment":"Footnote 1 states 'URL hidden to preserve anonymity'; because the paper's preregistration claims are load-bearing, an anonymized link or a registration number should be provided.","section":"Footnote 1"},{"comment":"The post-hoc notation in Table 5 (e.g., 'Control> MST-GT, MSTworkflow> MSTworkflow+') is ambiguous; report adjusted p-values and clearer pairwise notation.","section":"Table 5"},{"comment":"The median split is described as 'evenly re-split' but the exact group sizes and the split rule (median vs. top/bottom half) should be stated.","section":"Section 5.2.2"},{"comment":"Figure 4 uses '**' without a caption definition; the caption should state that '**' denotes p < 0.017.","section":"Figure 4"},{"comment":"The header 'Switch Faction' in Table 8 is a typo for 'Switch Fraction'.","section":"Table 8"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from a more disciplined separation of confirmatory and exploratory analyses. The current draft presents exploratory findings (the median split, the task-level misleading-advice patterns) in the same register as the preregistered tests. I would ask the editor to return the manuscript for a reanalysis of the misleading-advice claim and the H3 confound; if the reanalysis does not support the current conclusions, the central claims should be explicitly downgraded rather than presented as established findings."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is about a real gap: no one has measured appropriate reliance at intermediate steps of a multi-step transparent workflow for fact-checking. The authors do this with a preregistered four-condition study (N=233), open code/data, and a reasonable power analysis. That is a legitimately new empirical contribution.\n\nThe design has strengths. The MST conditions separate global transparency (decomposed steps) from local transparency (retrieved evidence), and the AR-Intermediate and AR-Evidence measures are a step forward for studying reliance in composite tasks. The finding that MST reduces overall reliance, and that this can cut both ways, is believable. The cognitive load data on the annotation intervention is also useful.\n\nThe soft spots are real but fixable. First, H1 is not supported by the table. The text says MST-GT increased reliance relative to Control, but the post-hoc groups Control and MST-GT together. That is an internal inconsistency. Second, the H3 median split is exploratory, and the AR-Intermediate/RAIR correlation is partly definitional: users who are accurate on sub-facts will tend to be right on the final call, so the dependence of team performance on AR-Intermediate does not establish that explicit consideration of intermediate steps is the mechanism. Third, the misleading-advice claim is shaky. The stress-test is right: Tasks 6 and 7 favor MST, but Task 1 (also misleading advice) goes the other way, and no inferential test is reported. The abstract's 'can outperform one-step collaboration when advice is misleading' is over-claimed. The exclusion of frequent switchers is also worth disclosing as a sensitivity analysis, not just a filter.\n\nNone of this sinks the paper. The empirical core is solid enough to be interesting, and the limitations section is honest about transferability. But the conclusions need to be realigned with the data, and the authors should add task-level modeling or at least a clear statement that the misleading-advice result is mixed.\n\nI would send this to peer review. It deserves referee time. The authors have the material to fix it, and the contribution—fine-grained reliance in multi-step workflows—is worth having in the literature.","headline":"A genuinely new multi-step reliance study with a solid empirical core, but the misleading-advice headline and two hypothesis claims need tightening before the results are reliable.","tokens_in":30980,"tokens_out":2466,"would_cite":true,"duration_ms":27503,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Following the AI's own decomposition steps helps humans catch misleading AI advice, but only when users get the intermediate sub-fact checks right.","keywords":["human-AI collaboration","appropriate reliance","multi-step decision workflow","transparency","composite fact-checking","task decomposition","large language models","cognitive load"],"falsifier":"Re-run the study with a separate pre-test of each participant's unaided fact-checking accuracy on comparable composite claims, and check whether the high-AR-Intermediate advantage on team performance, RAIR, and agreement fraction survives once that general-ability score is controlled for; if the advantage vanishes, the mediator is general competence rather than consideration of the intermediate steps.","tokens_in":30133,"feed_emoji":"🤝","tokens_out":7296,"duration_ms":64826,"temperature":0.7,"pith_summary":"This paper tries to establish when and why showing a human the AI's internal multi-step reasoning improves joint decision making. Using composite fact-checking, where a claim must be verified through several sub-facts, it tests a Multi-Step Transparent (MST) workflow in which users follow the same decomposed steps as an LLM-based fact-checker, inspect the evidence retrieved at each step, and only then receive the AI's final verdict. In a between-subjects study with 233 crowd workers, the workflow did not beat one-step collaboration on average, but it did on the tasks where the AI's advice was misleading. The authors' central claim is that fine-grained appropriate reliance, meaning correct decisions at the intermediate steps, is a prerequisite for the workflow to help: users who got the sub-facts right benefited, while those who did not under-relied on correct AI advice. They also find that a cognitive-forcing nudge, annotating each document's usefulness, backfired by raising cognitive load.","feed_headline":"Multi-step AI advice wins when the AI is wrong","feed_subtitle":"Giving users the AI's sub-steps beats one-shot advice only when users verify each step and the advice is misleading.","key_machinery":"The central construct is AR-Intermediate, the accuracy of a user's verdicts on the three decomposed sub-facts, scored against expert-annotated ground truth at each step. It does double duty: it is the fine-grained measure of appropriate reliance the paper introduces, and it is the variable that separates users for whom the MST workflow succeeds from those for whom it produces under-reliance. The AI system is ProgramFC, an LLM pipeline that decomposes a composite claim into three sub-facts, verifies each against BM25-retrieved Wikipedia evidence, and aggregates the sub-verdicts into a final prediction; the MST workflow makes that pipeline visible and executable by the user, step by step.","core_discovery":"The paper's claim, stated on its own terms, is that fine-grained appropriate reliance at the level of intermediate AI steps is what makes a multi-step transparent workflow effective, and that this manifests most clearly exactly where one-step collaboration fails: when the AI's advice is wrong. Three preregistered hypotheses were tested. H1, that showing users the AI's decomposed steps and intermediate answers (global transparency) increases reliance, was supported. H2, that the MST workflow increases appropriate reliance relative to one-step collaboration, was not supported on average. H3, that more accurate intermediate user decisions produce more appropriate reliance on the final advice, was supported: in a median split on AR-Intermediate, high-consideration users showed higher team performance, higher agreement with the AI, and higher RAIR. Across the three tasks where the AI's final verdict was misleading, MST conditions matched or exceeded the one-step conditions' accuracy, while on the easy tasks the one-step conditions did better. The authors conclude that there is no one-size-fits-all decision workflow for optimal human-AI collaboration.","pith_inferences":["If AR-Intermediate mostly tracks general fact-checking competence rather than step-wise consideration, the practical lever for improving human-AI teams would be training or selection rather than workflow design; the paper's own data, where high-AR users do better on every measure at once, is consistent with that reading.","A testable extension the authors gesture at but do not run: an adaptive workflow that engages the multi-step form only when the AI's own confidence is low could capture the misleading-advice benefit without the under-reliance cost seen on easy tasks.","AR-Evidence, agreement with experts on document usefulness, correlated with team performance but not with the appropriate-reliance measures, hinting that evidence-level transparency supports accuracy without calibrating reliance; separating those two functions could sharpen interface design."],"forward_implications":["Providing decomposed steps and intermediate answers as explanations increases reliance on AI advice, extending the known over-reliance effect of explainable AI to multi-step transparency (H1).","On tasks where the AI's final advice is misleading, users in the MST workflow match or beat one-step users, so the workflow's benefit is context-specific rather than general (H2).","Users with low AR-Intermediate under-rely on correct AI advice, and their team performance, agreement fraction, and RAIR all drop, turning the workflow into a liability rather than an aid.","The document-usefulness annotation intervention raises mental demand and frustration, lowering team performance and appropriate reliance, a cognitive-load cost that offsets its intended forcing effect.","The MST workflow lowers user confidence after seeing AI advice and intermediate answers, consistent with the critical mindset that mitigates over-reliance in the misleading-advice setting."],"supporting_citations":[{"why":"Supplies ProgramFC, the LLM fact-checker whose task decomposition, intermediate answers, and BM25-retrieved evidence define the MST workflow.","marker":"[80]"},{"why":"Provides the FEVEROUS-S composite fact-checking dataset from which the ten experimental tasks were selected.","marker":"[5]"},{"why":"Defines RAIR and RSR, the appropriate-reliance measures used to test H2 and H3.","marker":"[93]"},{"why":"Supplies the agreement-fraction and switch-fraction reliance measures and the earlier illusion-of-competence finding this study builds on.","marker":"[41]"},{"why":"Establishes cognitive forcing functions, the design precedent for the document-usefulness annotation in MSTworkflow+.","marker":"[13]"},{"why":"Shows that reasoning-process transparency increases user trust, the empirical basis for H1.","marker":"[105]"},{"why":"Maps the one-step human-AI decision-making design space that this work extends to multi-step workflows.","marker":"[57]"},{"why":"Provides the trust-in-automation framework linking trust to reliance, and the basis for the trust questionnaire and performance-bonus incentive design.","marker":"[61]"}],"fun_headline_variants":["When AI is wrong, showing sub-steps helps humans","Multi-step AI wins only when its advice is wrong","Step-wise AI boosts teams when advice misleads","AI sub-steps help only if users check each step","Fine-grained reliance: AI transparency aids wrong advice"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that a user's accuracy on the sub-fact verdicts measures how carefully they considered the AI's intermediate steps, rather than merely measuring their general fact-checking ability.","fun_headline_variants_meta":{"raw":{"variants":["When AI is wrong, showing sub-steps helps humans","Multi-step AI wins only when its advice is wrong","Step-wise AI boosts teams when advice misleads","AI sub-steps help only if users check each step","Fine-grained reliance: AI transparency aids wrong advice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000232,"raw_usage":{"total_tokens":1545,"prompt_tokens":1059,"completion_tokens":486,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":675,"completion_tokens_details":{"reasoning_tokens":409}},"tokens_in":675,"tokens_out":486,"duration_ms":6145,"temperature":1.0,"reasoning_tokens":409,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:50:07.241094+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the study with a separate pre-test of each participant's unaided fact-checking accuracy on comparable composite claims, and check whether the high-AR-Intermediate advantage on team performance, RAIR, and agreement fraction survives once that general-ability score is controlled for; if the advantage vanishes, the mediator is general competence rather than consideration of the intermediate steps.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies ProgramFC, the LLM fact-checker whose task decomposition, intermediate answers, and BM25-retrieved evidence define the MST workflow."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines RAIR and RSR, the appropriate-reliance measures used to test H2 and H3."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the agreement-fraction and switch-fraction reliance measures and the earlier illusion-of-competence finding this study builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that reasoning-process transparency increases user trust, the empirical basis for H1."}],"review_version":1}