{"id":"5f986e91-55c4-42ba-a08f-7a1bfc82ad7a","arxiv_id":"2508.08158","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In an online study, participants frequently spent little time on AI explanations, so the assumption that human oversight catches AI errors by reading explanations is weak.","lead":"This human-computer interaction study ran an online experiment on trust in AI decision support and found that many participants spent little time reading the AI's explanations. The authors then explore which factors lead people to consider explanations carefully, and whether that changes their openness to the AI's suggestion.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Construct validity of the 'consideration' measure is the load-bearing concern; time spent may not reflect scrutiny.","rationale":"The reader identified the same load-bearing premise: the time-to-consideration link. My review agrees that this is the weakest point. The paper's novelty is modest and the finding is plausible, but the central claim that users often fail to scrutinize explanations rests on how 'scrutiny' was measured. Without validation of the proxy, the descriptive finding ('little time') does not automatically support the interpretive claim ('did not consider in detail'). The warning in the abstract that participants 'did not always consider it in detail' is the bridge from time to oversight failure, and that bridge is exactly where the construct validity problem sits. This is not a fatal objection from the abstract alone; it is a call for verification in the full text. Since the reader already marked UNVERDICTED with low confidence, my read does not move the verdict. If anything, it sharpens why the verdict cannot be higher without access to the methods. I agree with the reader's weakest_assumption and see no need to adjust the verdict.","tokens_in":752,"tokens_out":2243,"duration_ms":30819,"concrete_test":"In the full paper, identify the exact operationalization of 'considering the explanation in detail' (expected: time spent viewing the explanation). Check whether the authors provide any validation that this measure reflects genuine scrutiny, e.g., a correlation between time spent and performance on a comprehension question about the explanation, or between time spent and the likelihood of overriding an incorrect AI recommendation. If no such validation exists, reanalyze the data using an alternative measure of engagement (e.g., whether the participant clicked to reveal additional details, or answered a follow-up question correctly) and see whether the reported relationships hold. If the alternative measure yields different conclusions, the central claim is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central empirical claim is that participants 'spent little time on the explanation and did not always consider it in detail.' The second clause is likely inferred from the first, i.e., time-on-page is the operationalization of 'considering in detail.' This is a construct-validity assumption: low time is treated as evidence of low cognitive engagement, and factors predicting time are treated as factors predicting careful consideration. But low time could result from high user expertise, simple or well-designed explanations, or interface layout that makes the explanation easy to process quickly. It could also reflect study disengagement (e.g., clicking through to finish) rather than a genuine lack of scrutiny. If the measure is not validated against an independent indicator of comprehension or error detection, the conclusion that users fail to oversee AI suggestions in detail is not supported. The exploratory analysis of what impacts 'careful consideration' then rests on the same shaky proxy. This is the weakest link because the entire oversight argument depends on the link between time and scrutiny.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract reports an online study on trust in explainable AI decision support systems. The authors expected that users would consider AI explanations in enough detail to catch AI errors, but they were 'surprised' to find that many participants spent little time on the explanation and did not always consider it in detail. The paper frames an exploratory analysis of which factors affect how carefully participants consider explanations and whether that consideration affects their willingness to change their mind based on the AI suggestion. The core claim is that the oversight benefit of explanations is frequently unrealized in practice because users do not engage with them.","tokens_in":933,"tokens_out":2179,"duration_ms":27546,"significance":"If the empirical finding holds, the paper challenges a foundational assumption in XAI research and practice: that providing explanations leads to attentive human oversight. This is a valuable and potentially influential result, as it would imply that current explainability evaluation metrics and interface designs may overstate the practical value of explanations. The strength of the paper is its clear articulation of the assumption and a falsifiable behavioral prediction. However, because only the abstract is available for review, the methodological grounding—sample, measures, statistical analysis, and construct validation—is entirely absent. The significance of the finding cannot be assessed without these details. The exploratory framing also signals post-hoc analysis, which limits inferential strength unless corrected in the full paper.","major_comments":[{"comment":"The central claim that participants 'did not always consider [the explanation] in detail' appears to be operationalized by time spent on the explanation, but the abstract provides no evidence that time is a valid measure of cognitive engagement. Short viewing time could reflect user expertise, explanation simplicity, or interface layout rather than insufficient scrutiny, or it could indicate general disengagement from the study. Without validation against an independent indicator of comprehension or error detection (e.g., follow-up questions, accuracy on adversarial cases), the conclusion that users fail to oversee AI suggestions in detail is not supported. This is the load-bearing measurement-validity issue.","section":"Abstract"},{"comment":"No sample size, participant population, task description, or experimental design is reported. It is impossible to assess whether the finding is robust, how large an effect is claimed, or to whom the result generalizes. The abstract alone cannot support a journal-level empirical claim; a full methods section with these details is essential.","section":"Abstract"},{"comment":"The exploratory analysis is described only in qualitative terms ('what factors impact...'), with no reported factors, effect sizes, confidence intervals, or statistical tests. The phrase 'exploratory' in the abstract itself signals that the analysis is post-hoc, which raises the risk of overfitting to the observed data. The full paper must report whether these analyses were pre-specified or, if exploratory, provide appropriate corrections and interpretational caution.","section":"Abstract"},{"comment":"Generalization: the abstract implies a broad conclusion about human oversight of AI, but the study is a low-stakes online experiment. The boundary conditions are not discussed. For instance, participants facing real high-stakes AI oversight (e.g., medical or judicial decisions) may exhibit far more careful explanation engagement. The paper should either temper the generalization or provide evidence that the online setting is representative.","section":"Abstract"}],"minor_comments":[{"comment":"The title is a question that the abstract does not directly answer. The abstract states the surprising disengagement finding but does not state whether participants were 'open to changing their mind' or whether the exploratory analysis found such an effect. Clarify the headline result.","section":"Abstract"},{"comment":"The abstract uses 'we were surprised'—this is colloquial for a formal report. Recommend a neutral phrasing such as 'contrary to expectations.'","section":"Abstract"},{"comment":"The term 'explainable DSS' is introduced without definition; spell out 'decision support system' at first use.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"This appears to be an abstract-only submission; I assume the full manuscript exists. If not, the paper is not yet in a reviewable state. The central claim is plausible and important, but the complete absence of methodological detail in the available text makes a rigorous evaluation impossible. The revision must supply the full study report and, in particular, validate the engagement measure. If the full paper already contains these details, my recommendation should be reevaluated in light of them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick note on 2508.08158. The thing to know: it's an abstract-only exploratory study that claims people often spend little time on AI explanations in an online decision-support task, and the authors are transparent that the analysis is exploratory. The genuine contribution is that they put the 'users will read explanations carefully' assumption on the table and test it directly. That premise underlies a lot of human-oversight rhetoric, so the paper earns credit for examining it empirically, even if the finding is in line with known automation-bias literature.\n\nWhat's done well: the authors state the assumption explicitly, report their surprise honestly, and label the predictor analysis as exploratory rather than confirmatory. That's the right epistemic framing for this kind of study.\n\nThe soft spots, in order of importance. First, the central construct: 'spent little time on the explanation' is being used as a proxy for 'did not consider it in detail.' That link isn't obvious. Low time could mean the explanation was easy to process, the participant already knew the domain, or the interface made it skimmable. It could also mean overall disengagement with the study rather than with the explanation. Without an independent measure of comprehension or error detection, the headline claim is only as strong as that proxy. Second, the online study is low-stakes; generalizing to high-stakes AI oversight is a stretch, and the authors don't address that. Third, the exploratory results about predictors are hypothesis-generating by their own admission.\n\nMy review is abstract-only, so no sample sizes, tests, or effect sizes. That keeps my confidence low, not the paper's fault. The stress-test note has it right: the measure of 'consideration' is the load-bearing assumption. If the full paper doesn't validate the time-to-scrutiny link, the conclusion about oversight failure is undercut.\n\nWho's it for: XAI researchers, HCI folks, and people designing AI-governance mechanisms that assume users engage with explanations. It's not a methods breakthrough, and it's not a null result that kills a research program. But it deserves a serious referee, mainly to push for measurement validation and to check how the exploratory findings are framed. I'd send it out, not desk-reject; just don't let the abstract stand alone.","headline":"A transparent exploratory study that tests an important oversight assumption, but its 'little time = little consideration' proxy needs validation before the headline claim lands.","tokens_in":1394,"tokens_out":3116,"would_cite":false,"duration_ms":33497,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that AI explanations only enable human oversight if users actually pause to consider them, and an online study found that many participants did not.","keywords":["AI explanations","decision support","human oversight","user engagement","trust","online study","explainable AI","opinion change"],"falsifier":"A concrete check: run the same decision-support task but ask participants unexpected recall or comprehension questions about the explanation after each trial. If people who spent very little time still answer those questions as accurately as those who spent longer, the inference that short time equals shallow consideration collapses. Alternatively, if the pattern reverses for high-stakes scenarios with real consequences, the finding is an artifact of the study context.","tokens_in":645,"feed_emoji":"🧠","tokens_out":3078,"duration_ms":36926,"temperature":0.7,"pith_summary":"The paper tries to establish that AI explanations only enable human oversight if users actually engage with them, and that in an online decision-support experiment this condition often failed: many participants spent little time on the explanation and did not consider it in detail. The authors frame this as a surprising finding that questions a core assumption of explainable AI. They then present an exploratory analysis of what factors predict careful consideration of explanations and whether that carefulness makes participants more willing to change their minds based on the AI's suggestion.","feed_headline":"Many users barely read AI explanations, study finds","feed_subtitle":"An online decision-support experiment suggests explanation-based human oversight fails when engagement is short.","key_machinery":"The central machinery is the study's measurement of 'consideration': behavioral indicators of how long and how carefully participants engaged with the explanation. This operationalization carries the argument, because the observed short engagement times are the evidence that the oversight assumption fails, and the subsequent correlational analysis uses that measure to separate engaged from disengaged users.","core_discovery":"On the paper's own terms, the discovery is that the link between providing an explanation and achieving human oversight is broken by inattention: in the studied online setting, participants frequently did not spend enough time on the AI's explanation to meaningfully evaluate it. The reported analysis explores which factors, such as characteristics of the explanation or the decision context, drive how thoroughly participants engage, and whether deeper engagement corresponds to a greater openness to revising their initial judgment. The implied claim is that explanation quality alone is insufficient; the user's willingness or ability to attend to the explanation is the bottleneck.","pith_inferences":["The reported low engagement may partly stem from the low stakes of the online task; in high-stakes settings such as medical or financial decisions, users might read explanations far more carefully, making the deficiency context-dependent rather than universal.","Short viewing time might not always mean low scrutiny; a user with a clear interface could quickly grasp the explanation, so recall or comprehension checks would be a stronger test than time alone.","A testable extension: forcing users to spend a minimum time on the explanation or requiring them to summarize it before accepting the AI's suggestion should increase error-catching, if inattention is the true cause.","Because the paper describes its analysis as exploratory, its factor associations are hypothesis-generating; a pre-registered replication with direct manipulation of stakes and obstacles would clarify which patterns are robust."],"forward_implications":["If the finding holds, explanation-based oversight fails whenever users skim or ignore the explanation, so the accuracy of AI oversight depends on interface design and incentives as much as on explanation content.","Systems cannot assume that a user who saw an explanation actually processed it; explanations may need to be gated, summarized, or tested to ensure consideration.","Studies of explainable AI that measure only the presence of explanations likely overstate their effect, so engagement should be measured directly.","Designers may need to adapt explanations to the limited time and attention users are willing to give them, for example by making key points visible at a glance."],"supporting_citations":[],"fun_headline_variants":["AI explanations often ignored, study shows","When users skip AI explanations, oversight fails","Users rarely read AI explanations, study finds","Explanations don't guarantee human oversight in AI","The problem with AI explanations: users don't read them"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The study assumes that the measured time spent (or similar behavioral trace) genuinely reflects how carefully participants considered the explanation, and that the behavior observed in a low-stakes online task represents how users would oversee AI in real-world settings.","fun_headline_variants_meta":{"raw":{"variants":["AI explanations often ignored, study shows","When users skip AI explanations, oversight fails","Users rarely read AI explanations, study finds","Explanations don't guarantee human oversight in AI","The problem with AI explanations: users don't read them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000252,"raw_usage":{"total_tokens":1329,"prompt_tokens":608,"completion_tokens":721,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":352,"completion_tokens_details":{"reasoning_tokens":665}},"tokens_in":352,"tokens_out":721,"duration_ms":8812,"temperature":1.0,"reasoning_tokens":665,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:36:25.766007+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: run the same decision-support task but ask participants unexpected recall or comprehension questions about the explanation after each trial. If people who spent very little time still answer those questions as accurately as those who spent longer, the inference that short time equals shallow consideration collapses. Alternatively, if the pattern reverses for high-stakes scenarios with real consequences, the finding is an artifact of the study context.","supporting_citations":[],"review_version":1}