Pith. sign in

REVIEW 4 major objections 4 minor 79 references

In AI-assisted fact-checking, the rhetorical pattern of the assistant's advice is associated with significant differences in accuracy gains, confidence, time, and preference — with step-by-step scaffolded explanations producing the largest

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 17:26 UTC pith:QEPO4NHW

load-bearing objection Solid exploratory study; the accuracy headline doesn't survive the fixed 65% false base rate, but the confidence/time/preference results are real. the 4 major comments →

arxiv 2607.17627 v1 pith:QEPO4NHW submitted 2026-07-20 cs.HC cs.CL

It Matters How You Say It: Exploring Rhetorical Patterns for AI-Assisted Information Evaluation

classification cs.HC cs.CL
keywords rhetorical patternsAI-assisted information evaluationfact verificationmisinformation detectionhuman-AI interactionreflective thinkingconfidence calibrationuser preferences
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether the way an AI assistant delivers advice — not just what it says — changes how well people evaluate claims. In a study of 98 participants verifying 40 claims with on-demand help, it compares eight rhetorical styles, from plain 'oracle' verdicts to Socratic questioning, deliberate misdirection, and step-by-step scaffolds. The central finding is that rhetorical style is associated with significant differences in accuracy gains, confidence, time on task, and user preference. Step-by-step scaffolded explanations produced the largest accuracy gains, and even adversarial styles that distorted or mocked the truth improved accuracy modestly. At the same time, users preferred the easy, concise styles, not the ones that helped most — a preference-performance split with direct implications for AI assistant design.

Core claim

The paper's central claim is that rhetorical presentation is not a neutral wrapper on AI advice: it changes measurable outcomes in information evaluation. Across eight patterns, a pattern-by-time interaction showed that accuracy gains differed significantly (χ²(7)=83.2, p<0.001), final confidence differed (χ²(7)=54.7, p<0.001), time on task differed (χ²(7)=29.3, p<0.001), and preferences differed (Friedman χ²(7)=160.4, p<0.001). Scaffold Explanation had the largest point-estimate accuracy gain (Δ=0.2156), and Scaffold Explanation, Oracle, and Triggering Distrust showed significant before-to-after accuracy improvements. The paper interprets the adversarial gains as possible motivated scrutiny

What carries the argument

The central object is a set of eight 'rhetorical patterns' — templates for how an AI advisor phrases its response — operationalized for fact-checking: Oracle (direct gold-standard verdict), Scaffold Explanation (step-by-step logical deduction), Socratic Questioning (probing questions about missing evidence), Alternative Framing (tangential context that distracts), Interpretive Alternative (competing inferences that challenge assumptions), Information Distortion (omitting key facts), Triggering Distrust (deliberate absurdities), and Intentional Misleading (plausible but incorrect framing). Each pattern is instantiated on five manually crafted claims; participants could request advice and then

Load-bearing premise

The claim set is unbalanced (26 of 40 claims are false), and truth value is fixed per rhetorical pattern, so a pattern that nudges users toward skepticism can appear to improve accuracy without actually improving their ability to tell true from false — the authors acknowledge this response-bias account as a plausible alternative explanation for the adversarial patterns' gains.

What would settle it

A balanced replication of the same study with equal numbers of true and false claims per pattern (or reporting accuracy separately by truth value, e.g., sensitivity and specificity) would settle the strongest version of the claim: if adversarial patterns no longer beat the Oracle baseline once response bias is controlled, the 'motivated scrutiny' reading fails, and the accuracy advantage reduces to a bias toward 'false'.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If rhetorical style is a real lever, then AI assistants can be designed to improve user evaluation without changing the underlying information — for example, by defaulting to scaffolded explanations.
  • Confidence inflation under ineffective patterns means an assistant can make users feel more certain while not improving their judgments; calibration, not just accuracy, must be a design target.
  • The accuracy-preference divergence implies that optimizing AI assistants for user satisfaction alone could systematically undermine evaluation quality, since the most-preferred style (Alternative Framing) showed no significant accuracy gain.
  • The tentative finding that adversarial styles did not hurt accuracy suggests that detectable manipulation can trigger scrutiny that extends to the claim, opening a design space of legible friction — though the paper is explicit this needs confirmation.
  • No single pattern dominates: direct Oracle answers were fast and accurate, but oracle-quality information is rarely available in real settings, making contemplative styles the practically relevant design space.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The response-bias explanation the authors raise is testable: if the claim set were balanced 50/50 true/false, skeptical biases would cancel out; a pattern's true value would appear as discrimination, not overall accuracy. This is the natural next experiment.
  • The paper's 'legible friction' principle generalizes beyond fact-checking to any AI advisory context where premature closure is dangerous: the design principle is to make the reason for pause visible to the user.
  • The preference data suggest that revealed preference in AI tool choice may be a poor proxy for epistemic welfare; users may need affordances to choose 'deep' versus 'fast' modes depending on stakes.
  • The manual construction of advice per pattern, with only five claims per pattern, means the effect sizes are likely noisy; a replication with LLM-generated advice across many more claims per pattern could separate the rhetoric effect from idiosyncratic phrasing of individual examples.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper reports a within-subject online experiment (n=98) in which participants evaluated 40 true/false claims from Fool-Me-Twice with optional AI advice cast in one of eight rhetorical patterns (Oracle, Scaffold Explanation, Socratic Questioning, Triggering Distrust, etc.). The authors use GEE models with cluster-robust inference to compare accuracy gains, final confidence, time on task, and preference rankings across patterns. They report that the pattern×time interaction is significant (χ²(7)=83.2, p<0.001), that Scaffold Explanation has the largest accuracy gain (Δ=0.2156), and that preferences diverge from performance (Alternative Framing most preferred; Oracle mid-ranked). The paper frames these as preliminary associations rather than causal effects and lists several limitations, including fixed truth values per pattern and the absence of manipulation checks.

Significance. The question is timely and the design has real strengths: within-subject assignment to eight conditions, explicit hypotheses (H1–H6), cluster-robust GEE inference with Holm corrections, and a unusually candid limitations section. The accuracy-preference divergence (H4 not supported) and the qualitative reflection data are useful contributions for HCI and human-AI interaction. If the accuracy result were secured against the response-bias confound, the paper would be a meaningful step toward treating rhetorical style as a design variable. As it stands, the central accuracy claim is not yet established, because the design cannot separate improved true/false discrimination from a simple skepticism bias or from changes in information content.

major comments (4)
  1. [§V.A.1, Fig. 4; §III.A.2; §VII] The primary accuracy result is a before-to-after gain in raw percent correct. With 14 true and 26 false claims, 'false' is the correct response 65% of the time, and truth value is fixed per pattern with only five claims per cell. A pattern that nudges participants toward skepticism can therefore raise measured accuracy without improving true/false discrimination, and this bias is most plausible precisely for Triggering Distrust (Δ=0.1772) and Information Distortion (Δ=0.1317). The authors acknowledge this in §VII, but the acknowledgment does not resolve it: the pattern×time interaction χ²(7)=83.2 and the per-pattern gains are the main evidence for the 'how you say it matters' claim. The paper never reports the per-pattern true/false composition. I request a reanalysis reporting accuracy separately by claim truth value, a signal-detection or sensitivity measure, and either a balanced/rota
  2. [§III.C] Advice requests were voluntary (2,987/3,875 trials, 77.1%), and the models are fit on the advice-requested subset. Because participants decide whether to request advice after seeing the claim, pattern comparisons on this subset may be biased by claim difficulty and by participants' initial confidence or skepticism. The sentence 'we retain all trials in the primary analysis and fit every model on the advice-requested subset' is internally contradictory: if the accuracy models are fit only on the subset, no analysis actually uses all trials. Please report advice-request rates by pattern, compare characteristics of requested vs. non-requested trials, and, if possible, use selection weights or a model on the full trial set. Without this, even the descriptive accuracy comparisons are difficult to interpret.
  3. [§IV, §VII] Several conditions—Intentional Misleading, Information Distortion, Triggering Distrust, and to some degree Alternative Framing—alter information content or completeness, not only wording. The accuracy gains from these conditions can therefore be attributed to the information supplied/withheld rather than to any rhetorical pattern. The paper acknowledges this confound in §VII, and the discussion wisely limits some causal language, but the abstract and conclusion still state that 'how AI systems communicate mattered as much as what they communicate.' That overstates what the design can establish. The revised version should either hold information completeness constant (e.g., a wording-only condition per pattern) or explicitly restrict the rhetorical claim to the patterns that primarily vary phrasing.
  4. [§III.C, §III.A.2] The GEE models cluster on participant but not on claim. With only five claims per pattern and no claim rotation, pattern effects are partially claims effects; the reported χ² statistics and p-values treat the claim sample as a fixed feature rather than as a source of sampling variability. At minimum, report cluster-robust inference by claim (or a mixed model with random claims), and reinterpret the pattern-level comparisons as exploratory and specific to the 40 selected claims. This is not a demand for a new experiment, but the inferential language should match the design's degrees of freedom.
minor comments (4)
  1. [Figures 3, 4, 6] The figures use 'Selective Distortion' while the text, abstract, and index terms use 'Information Distortion'. Please make the labels consistent.
  2. [§III.A.3] There is a typo: 'an pattern' should be 'a pattern.' Also, 'five claims for eight patterns, totaling 40 claim advice pairs' should be '40 claims,' not '40 claim advice pairs.'
  3. [§V.A.1, Fig. 4] The star annotations in Figure 4 are not all self-explanatory; add a legend or table mapping each pattern to its Holm-corrected p-value and Δ estimate so the reader can verify the claims without reconstructing the tests.
  4. [§III.A.3, §IV] The Oracle condition is called a 'rhetorical pattern,' but it is defined as the absence of rhetorical modification. This is fine as a baseline, but the wording should be consistent—e.g., 'directive baseline'—so readers do not conflate the baseline with the seven rhetorical manipulations.

Circularity Check

0 steps flagged

No significant circularity: the study is an empirical experiment with hypotheses drawn from external literature, not a derivation that reduces to its inputs.

full rationale

This paper is an empirical user study, not a formal derivation. The eight rhetorical patterns are operationalized from external linguistics and psychology literature (Tversky & Kahneman, Druckman, Cialdini, Schwenk, etc.), and each hypothesis is stated before the results are reported. Notably, H2 and H6 are explicitly not supported by the data, which indicates the outcomes were not constructed to match the hypotheses. The accuracy, confidence, and time-on-task comparisons are measured outcomes analyzed with GEE models; the reported chi-square statistics are inferential tests, not definitions of the outcomes. The closest thing to a self-referential concern is that one co-author co-created the Fool-Me-Twice dataset, but that dataset is an external, published resource with fixed ground-truth labels and was not tailored to this study, so it does not make the results circular. The paper's own §VII limitation—that the 65% false base rate and truth-value-fixed-per-pattern design allow skepticism bias to raise measured accuracy without improving discrimination—is an acknowledged threat to causal interpretation, but it is an internal validity caveat, not a circular derivation. No prediction is fitted from a subset of the data and then renamed as an independent finding. Therefore no circular step is present.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No free parameters or invented entities; the central claims rest on design assumptions about pattern operationalization, the contemplation proxy, claim sampling, and sample representativeness, most of which the authors flag in Limitations.

axioms (4)
  • domain assumption Time on task is a valid behavioral proxy for contemplation and critical thinking.
    RQ2 operationalizes reflection as evaluation time; the paper cites prior uses but acknowledges it is a proxy, not a direct measure (§V-B1).
  • domain assumption Manually constructed advice accurately instantiates the intended rhetorical patterns.
    Two authors generated each pattern's advice via negotiated agreement with no manipulation check to confirm participants perceived the intended style (§III-A3, §VII).
  • domain assumption The claim sample's difficulty and truth distribution are balanced across pattern cells.
    Claims were sampled by difficulty from Fool-Me-Twice, but the 14:26 true/false split is fixed per pattern, so skepticism can masquerade as accuracy improvement (§III-A2, §VII).
  • domain assumption The Prolific sample represents the target user population.
    The sample skews educated and digitally literate; the authors caution that results may not generalize to populations more vulnerable to misinformation (§VII).

pith-pipeline@v1.3.0-alltime-deepseek · 18954 in / 12328 out tokens · 118782 ms · 2026-08-01T17:26:46.395670+00:00 · methodology

0 comments
read the original abstract

Prior work on AI-assisted information evaluation has largely focused on what AI systems communicate, comparing explanation types and formats, with responses predominantly cast in directive rhetoric where the system delivers a verdict and the user passively accepts it. While debate-style interactions have recently shown promise in prompting critical evaluation over deference, the rhetorical patterns that structure AI responses and how they might induce reflection, uncertainty, or independent reasoning remain largely unexamined. To address this, we investigated eight rhetorical patterns known to induce contemplation: Intentional Misleading, Interpretive Alternative, Scaffold Explanation, Triggering Distrust, Information Distortion, Alternative Framing, Socratic Questioning, and an Oracle baseline. Through a within-subject study with n=98 participants on a hint-on-demand fact verification task, we observed preliminary evidence that Scaffold Explanation were associated with the highest accuracy gains, and encouraging deeper reflection. Surprisingly, the adversarial conditions also improved accuracy modestly. Participants preferred Alternative Framing most and Interpretive Alternative least, largely due to the latter's perceived time cost. We discuss the implications of designing conversational agents with varied rhetorical styles and the trade-offs among user performance, satisfaction, and contemplation.

Figures

Figures reproduced from arXiv: 2607.17627 by Jonathan K. Kummerfeld, Jonathan May, Jordan Lee Boyd-Graber, Sadra Sabouri, Souti Chattopadhyay, Zeinabsadat Saghi.

Figure 1
Figure 1. Figure 1: An LLM assistant asked to verify a social media claim [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: This screenshot shows a claim (top), an pattern and [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Answer revisions (a) and confidence shifts (b) across patterns. Dashed lines mark the mean of each series. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Per-participant change in accuracy (∆ = after − before) for each pattern; violins show the full distribution of ∆ and the dashed line shows no change. Stars indicate significance levels: *** p < 0.001, ** p < 0.01, * p < 0.05, ns = not significant. the 40 claims before receiving any advice. After they received advice, the average percentage increased to 72.34%. Participants’ accuracy improved the most afte… view at source ↗
Figure 5
Figure 5. Figure 5: Rhetorical patterns ranked by preference (1 = most [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Task time (a) and usability ratings (b) across patterns. [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

79 extracted references · 8 linked inside Pith

  1. [1]

    PEP 744 – JIT compilation,

    B. Bucher, “PEP 744 – JIT compilation,” https://peps.python.org/ pep-0744/, 2023, accessed: 2025

  2. [2]

    Prevalence of health misinfor- mation on social media: systematic review,

    V . Suarez-Lledo and J. Alvarez-Galvez, “Prevalence of health misinfor- mation on social media: systematic review,”Journal of medical Internet research, vol. 23, no. 1, p. e17187, 2021

  3. [3]

    Trends in the diffusion of misinformation on social media,

    H. Allcott, M. Gentzkow, and C. Yu, “Trends in the diffusion of misinformation on social media,”Research & Politics, vol. 6, no. 2, p. 2053168019848554, 2019

  4. [4]

    Combating misinformation in the age of llms: Opportunities and challenges,

    C. Chen and K. Shu, “Combating misinformation in the age of llms: Opportunities and challenges,”AI Magazine, 2023

  5. [5]

    Designing effective ai explanations for misinformation detection: A comparative study of content, social, and combined explanations,

    Y . Gong, Y . Liu, L. Shang, N. Wei, and D. Wang, “Designing effective ai explanations for misinformation detection: A comparative study of content, social, and combined explanations,”Proceedings of the ACM on Human-Computer Interaction, vol. 9, no. 7, pp. 1–37, 2025

  6. [6]

    Reliability matters: Exploring the effect of ai explanations on misinformation detection with a warning,

    H. Seo, S. Lee, D. Lee, and A. Xiong, “Reliability matters: Exploring the effect of ai explanations on misinformation detection with a warning,” inProceedings of the international AAAI conference on web and social media, vol. 18, 2024, pp. 1395–1407

  7. [7]

    Effect of expla- nation conceptualisations on trust in ai-assisted credibility assessment,

    S. Pareek, N. Van Berkel, E. Velloso, and J. Goncalves, “Effect of expla- nation conceptualisations on trust in ai-assisted credibility assessment,” Proceedings of the ACM on Human-Computer Interaction, vol. 8, no. CSCW2, pp. 1–31, 2024

  8. [8]

    Context, credibility, and control: User reflections on ai assisted misinformation tools,

    V . Sangwan and H. Makitalo, “Context, credibility, and control: User reflections on ai assisted misinformation tools,”arXiv preprint arXiv:2506.22940, 2025

  9. [9]

    Debating with more persuasive llms leads to more truthful answers,

    A. Khan, J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rockt ¨aschel, and E. Perez, “Debating with more persuasive llms leads to more truthful answers,”arXiv preprint arXiv:2402.06782, 2024

  10. [11]

    Leveraging llms for detecting and modeling the propagation of misinformation in social networks,

    P. Santra, “Leveraging llms for detecting and modeling the propagation of misinformation in social networks,” inProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024, pp. 3073–3073

  11. [12]

    Using grok to avoid personal attacks while cor- recting misinformation on x,

    K. M. Caramancion, “Using grok to avoid personal attacks while cor- recting misinformation on x,”arXiv preprint arXiv:2601.04251, 2026

  12. [13]

    Synthetic lies: Understanding ai-generated misinformation and evaluating algorithmic and human solutions,

    J. Zhou, Y . Zhang, Q. Luo, A. G. Parker, and M. De Choudhury, “Synthetic lies: Understanding ai-generated misinformation and evaluating algorithmic and human solutions,” inProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, ser. CHI ’23. New York, NY , USA: Association for Computing Machinery,

  13. [14]

    A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,

    L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qinet al., “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,” arXiv preprint arXiv:2311.05232, 2023

  14. [15]

    Toward mitigating misin- formation and social media manipulation in llm era,

    Y . Zhang, K. Sharma, L. Du, and Y . Liu, “Toward mitigating misin- formation and social media manipulation in llm era,” inCompanion Proceedings of the ACM on Web Conference 2024, 2024, pp. 1302– 1305

  15. [16]

    Designing social media to foster user engagement in challenging misinformation: a cross-cultural comparison between the uk and arab countries,

    M. Noman, S. Gurgun, K. Phalp, and R. Ali, “Designing social media to foster user engagement in challenging misinformation: a cross-cultural comparison between the uk and arab countries,”Humanities and Social Sciences Communications, vol. 11, no. 1, 2024

  16. [17]

    Credibility of misinformation source moderates the effectiveness of corrective messages on social media,

    H.-K. Zeng, S.-Y . Lo, and S.-C. S. Li, “Credibility of misinformation source moderates the effectiveness of corrective messages on social media,”Public Understanding of Science, vol. 33, no. 5, pp. 587–603, 2024

  17. [18]

    Borchers and H

    T. Borchers and H. Hundley,Rhetorical theory: An introduction. Wave- land Press, 2018

  18. [19]

    The framing of decisions and the psychology of choice,

    A. Tversky and D. Kahneman, “The framing of decisions and the psychology of choice,”science, vol. 211, no. 4481, pp. 453–458, 1981

  19. [20]

    The implications of framing effects for citizen competence,

    J. N. Druckman, “The implications of framing effects for citizen competence,”Political behavior, vol. 23, pp. 225–256, 2001

  20. [21]

    The science of persuasion,

    R. B. Cialdini, “The science of persuasion,”Scientific American, vol. 284, no. 2, pp. 76–81, 2001

  21. [22]

    Persuasion,

    D. J. O’keefe, “Persuasion,” inThe handbook of communication skills. Routledge, 2006, pp. 333–352

  22. [23]

    Contemplation and conversation: Subtle influences on moral decision making,

    B. C. Gunia, L. Wang, L. Huang, J. Wang, and J. K. Murnighan, “Contemplation and conversation: Subtle influences on moral decision making,”Academy of Management Journal, vol. 55, no. 1, pp. 13–33, 2012

  23. [24]

    Conversational argu- mentation in decision making: Chinese and us participants in face-to- face and instant-messaging interactions,

    C. O. Stewart, L. D. Setlock, and S. R. Fussell, “Conversational argu- mentation in decision making: Chinese and us participants in face-to- face and instant-messaging interactions,”Discourse processes, vol. 44, no. 2, pp. 113–139, 2007

  24. [25]

    Impact of guidance and interaction strategies for llm use on learner performance and perception,

    H. Kumar, I. Musabirov, M. Reza, J. Shi, A. Kuzminykh, J. J. Williams, and M. Liut, “Impact of guidance and interaction strategies for llm use on learner performance and perception,”arXiv preprint arXiv:2310.13712, 2023

  25. [26]

    Management consulting in the artificial intelligence–llm era,

    S. K. Mohan, “Management consulting in the artificial intelligence–llm era,”Management Consulting Journal, vol. 7, no. 1, pp. 9–24, 2024

  26. [27]

    Healthcare copilot: Eliciting the power of general llms for medical consultation,

    Z. Ren, Y . Zhan, B. Yu, L. Ding, and D. Tao, “Healthcare copilot: Eliciting the power of general llms for medical consultation,”arXiv preprint arXiv:2402.13408, 2024

  27. [28]

    Measuring and benchmarking large language models’ capabilities to generate persuasive language,

    A. B. Pauli, I. Augenstein, and I. Assent, “Measuring and benchmarking large language models’ capabilities to generate persuasive language,” arXiv preprint arXiv:2406.17753, 2024

  28. [29]

    Debate chatbots to facilitate critical thinking on youtube: Social identity and conversational style make a difference,

    T. Tanprasert, S. S. Fels, L. Sinnamon, and D. Yoon, “Debate chatbots to facilitate critical thinking on youtube: Social identity and conversational style make a difference,” inProceedings of the CHI Conference on Human Factors in Computing Systems, 2024, pp. 1–24

  29. [30]

    Reflection in theory and reflection in practice: An exploration of the gaps in reflection support among personal informatics apps,

    J. Cho, T. Xu, A. Zimmermann-Niefield, and S. V oida, “Reflection in theory and reflection in practice: An exploration of the gaps in reflection support among personal informatics apps,” inProceedings of the 2022 CHI Conference on Human Factors in Computing Systems, ser. CHI ’22. New York, NY , USA: Association for Computing Machinery,

  30. [31]

    Supporting self-reflection at scale with large language models: Insights from randomized field experiments in classrooms,

    H. Kumar, R. Xiao, B. Lawson, I. Musabirov, J. Shi, X. Wang, H. Luo, J. J. Williams, A. N. Rafferty, J. Stamper, and M. Liut, “Supporting self-reflection at scale with large language models: Insights from randomized field experiments in classrooms,” inProceedings of the Eleventh ACM Conference on Learning @ Scale, ser. L@S ’24. New York, NY , USA: Associa...

  31. [32]

    Self-reflection in llm agents: Effects on problem-solving performance,

    M. Renze and E. Guven, “Self-reflection in llm agents: Effects on problem-solving performance,”arXiv preprint arXiv:2405.06682, 2024

  32. [33]

    Fact checking: Task definition and dataset construction,

    A. Vlachos and S. Riedel, “Fact checking: Task definition and dataset construction,” inProceedings of the ACL 2014 workshop on language technologies and computational social science, 2014, pp. 18–22

  33. [34]

    A survey on automated fact-checking,

    Z. Guo, M. Schlichtkrull, and A. Vlachos, “A survey on automated fact-checking,”Transactions of the Association for Computational Lin- guistics, vol. 10, pp. 178–206, 2022

  34. [35]

    A browser extension for in-place signaling and assessment of misinformation,

    F. Jahanbakhsh and D. R. Karger, “A browser extension for in-place signaling and assessment of misinformation,” inProceedings of the CHI Conference on Human Factors in Computing Systems, 2024, pp. 1–21

  35. [36]

    A mystery for you: A fact-checking game enhanced by large language models (llms) and a tangible interface,

    H. Tang and M. Singha, “A mystery for you: A fact-checking game enhanced by large language models (llms) and a tangible interface,” in Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems, ser. CHI EA ’24. New York, NY , USA: Association for Computing Machinery, 2024. [Online]. Available: https://doi.org/10.1145/3613905.3648110

  36. [37]

    Fool me twice: Entailment from Wikipedia gamification,

    J. Eisenschlos, B. Dhingra, J. Bulian, B. B ¨orschinger, and J. Boyd- Graber, “Fool me twice: Entailment from Wikipedia gamification,” inProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. B...

  37. [38]

    A motivational interviewing chatbot with generative reflections for increasing readiness to quit smoking: Iterative development study,

    A. Brown, A. T. Kumar, O. Melamed, I. Ahmed, Y . H. Wang, A. Deza, M. Morcos, L. Zhu, M. Maslej, N. Minianet al., “A motivational interviewing chatbot with generative reflections for increasing readiness to quit smoking: Iterative development study,”JMIR Mental Health, vol. 10, p. e49132, 2023

  38. [39]

    Enhancing motivation for change in substance use disorder treatment,

    S. Abuseet al., “Enhancing motivation for change in substance use disorder treatment,” inTreatment improvement protocol (TIP) series no

  39. [40]

    Prolific. ac—a subject pool for online experiments,

    S. Palan and C. Schitter, “Prolific. ac—a subject pool for online experiments,”Journal of behavioral and experimental finance, vol. 17, pp. 22–27, 2018

  40. [41]

    Longitudinal data analysis using gener- alized linear models,

    K.-Y . Liang and S. L. Zeger, “Longitudinal data analysis using gener- alized linear models,”Biometrika, vol. 73, no. 1, pp. 13–22, 1986

  41. [42]

    Substance Abuse and Mental Health Services Administration (US), 2019

  42. [43]

    Tannen,The argument culture: Stopping America’s war of words

    D. Tannen,The argument culture: Stopping America’s war of words. Ballantine Books, 1999

  43. [44]

    Effects of devil’s advocacy and dialectical inquiry on decision making: A meta-analysis,

    C. R. Schwenk, “Effects of devil’s advocacy and dialectical inquiry on decision making: A meta-analysis,”Organizational behavior and human decision processes, vol. 47, no. 1, pp. 161–176, 1990

  44. [45]

    Interpersonal deception theory,

    D. B. Buller and J. K. Burgoon, “Interpersonal deception theory,” Communication theory, vol. 6, no. 3, pp. 203–242, 1996

  45. [46]

    A review of cognitive dissonance theory in management research: Opportunities for further development,

    A. S. Hinojosa, W. L. Gardner, H. J. Walker, C. Cogliser, and D. Gullifor, “A review of cognitive dissonance theory in management research: Opportunities for further development,”Journal of Management, vol. 43, no. 1, pp. 170–199, 2017

  46. [47]

    I. Copi, C. Cohen, and D. Flage,Essentials of logic. Routledge, 2016

  47. [48]

    Culture, dialectics, and reasoning about contradiction

    K. Peng and R. E. Nisbett, “Culture, dialectics, and reasoning about contradiction.”American psychologist, vol. 54, no. 9, p. 741, 1999

  48. [49]

    Narration as a human communication paradigm: The case of public moral argument,

    W. R. Fisher, “Narration as a human communication paradigm: The case of public moral argument,”Communications Monographs, vol. 51, no. 1, pp. 1–22, 1984

  49. [50]

    Chain-of-thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhouet al., “Chain-of-thought prompting elicits reasoning in large language models,”Advances in neural information processing systems, vol. 35, pp. 24 824–24 837, 2022

  50. [51]

    Deductive reasoning,

    P. N. Johnson-Laird, “Deductive reasoning,”Annual review of psychol- ogy, vol. 50, no. 1, pp. 109–135, 1999

  51. [52]

    Use and effectiveness of billboards: Perspectives from selective-perception theory and retail- gravity models,

    C. R. Taylor, G. R. Franke, and H.-K. Bang, “Use and effectiveness of billboards: Perspectives from selective-perception theory and retail- gravity models,”Journal of advertising, vol. 35, no. 4, pp. 21–34, 2006

  52. [53]

    Framing as a theory of media effects,

    D. A. Scheufele, “Framing as a theory of media effects,”Journal of communication, vol. 49, no. 1, pp. 103–122, 1999

  53. [54]

    Reductio ad absurdum from a dialogical perspective,

    C. Dutilh Novaes, “Reductio ad absurdum from a dialogical perspective,” Philosophical Studies, vol. 173, pp. 2605–2628, 2016

  54. [55]

    Critical thinking: The art of socratic questioning,

    R. Paul and L. Elder, “Critical thinking: The art of socratic questioning,” Journal of developmental education, vol. 31, no. 1, p. 36, 2007

  55. [56]

    The self-consistency model of subjective confidence,

    A. Koriat, “The self-consistency model of subjective confidence,”Psy- chological Review, vol. 119, no. 1, pp. 80–113, 2012

  56. [57]

    Argumentation,

    F. H. Van Eemeren, F. H. van Eemeren, S. Jackson, and S. Jacobs, “Argumentation,”Reasonableness and effectiveness in argumentative discourse: Fifty contributions to the development of Pragma-dialectics, pp. 3–25, 2015

  57. [58]

    Self-confidence and performance on tests of cognitive abilities,

    L. Stankov and J. D. Crawford, “Self-confidence and performance on tests of cognitive abilities,”Intelligence, vol. 25, no. 2, pp. 93–109, 1997

  58. [59]

    The trouble with overconfidence,

    D. A. Moore and P. J. Healy, “The trouble with overconfidence,” Psychological Review, vol. 115, no. 2, pp. 502–517, 2008

  59. [60]

    How to measure metacognition,

    S. M. Fleming and H. C. Lau, “How to measure metacognition,” Frontiers in Human Neuroscience, vol. 8, p. 443, 2014

  60. [61]

    Does the whole exceed its parts? the effect of ai explanations on complementary team performance,

    G. Bansal, T. Wu, J. Zhou, R. Fok, B. Nushi, E. Kamar, M. T. Ribeiro, and D. Weld, “Does the whole exceed its parts? the effect of ai explanations on complementary team performance,” inProceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 2021, pp. 1–16

  61. [62]

    Understanding the effect of accuracy on trust in machine learning models,

    M. Yin, J. Wortman Vaughan, and H. Wallach, “Understanding the effect of accuracy on trust in machine learning models,” inProceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019, pp. 1–12

  62. [63]

    Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision mak- ing,

    Y . Zhang, Q. V . Liao, and R. K. Bellamy, “Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision mak- ing,” inProceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 2020, pp. 295–305

  63. [64]

    Cognitive reflec- tion, decision biases, and response times,

    C. Al ´os-Ferrer, M. Garagnani, and S. H ¨ugelsch¨afer, “Cognitive reflec- tion, decision biases, and response times,”Frontiers in Psychology, vol. 7, 2016

  64. [65]

    Manipulating response times in the cognitive reflection test: Time delay boosts deliberation, time pressure hinders it,

    E. Bilancini, L. Boncinelli, and T. Celadin, “Manipulating response times in the cognitive reflection test: Time delay boosts deliberation, time pressure hinders it,”Journal of Behavioral and Experimental Economics, 2024

  65. [66]

    The time course of conflict on the cognitive reflection test

    E. Travers, J. J. Rolison, and A. Feeney, “The time course of conflict on the cognitive reflection test.”Cognition, vol. 150, pp. 109–18, 2016

  66. [67]

    The interplay between reflective thinking, critical think- ing, self-monitoring, and academic achievement in higher education,

    A. Ghanizadeh, “The interplay between reflective thinking, critical think- ing, self-monitoring, and academic achievement in higher education,” Higher Education, vol. 74, no. 1, pp. 101–114, 2017

  67. [68]

    Exploring internal structure of a performance-based critical thinking assessment for new students in higher education,

    K. Kleemola, H. Hyytinen, and A. Toom, “Exploring internal structure of a performance-based critical thinking assessment for new students in higher education,”Assessment & Evaluation in Higher Education, vol. 47, no. 4, pp. 556–569, 2021

  68. [69]

    How do self-regulation and effort in test-taking contribute to undergraduate students’ critical thinking performance?

    H. Hyytinen, K. Nissinen, K. Kleemola, J. Ursin, and A. Toom, “How do self-regulation and effort in test-taking contribute to undergraduate students’ critical thinking performance?”Studies in Higher Education, vol. 49, no. 2, pp. 192–205, 2023

  69. [70]

    The time on task effect in reading and problem solving is moderated by task difficulty and skill: Insights from a computer-based large-scale assessment

    F. Goldhammer, J. Naumann, A. Stelter, K. T ´oth, H. R ¨olke, and E. Klieme, “The time on task effect in reading and problem solving is moderated by task difficulty and skill: Insights from a computer-based large-scale assessment.”Journal of Educational Psychology, vol. 106, pp. 608–626, 2014

  70. [71]

    Automation bias in intelligent time critical decision support systems,

    M. L. Cummings, “Automation bias in intelligent time critical decision support systems,” inDecision making in aviation. Routledge, 2017, pp. 289–294

  71. [72]

    Testing repli- cability and generalizability of the time on task effect,

    R. J. Kr ¨amer, M. Koch, J. Levacher, and F. Schmitz, “Testing repli- cability and generalizability of the time on task effect,”Journal of Intelligence, vol. 11, 2023

  72. [73]

    Exploring a behavioral model of “positive friction

    Z. Chen and R. Schmidt, “Exploring a behavioral model of “positive friction” in human-AI interaction,”arXiv preprint, 2024

  73. [74]

    From friction to synergy: The complex interplay of human creativity and AI,

    M. DeSchryver, D. Henriksen, and S. Leahy, “From friction to synergy: The complex interplay of human creativity and AI,”Possibility Studies & Society, 2025

  74. [75]

    Design frictions for mindful interactions: The case for microboundaries,

    A. Cox, S. Gould, M. Cecchinato, I. Iacovides, and I. Renfree, “Design frictions for mindful interactions: The case for microboundaries,” in Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems, 2016

  75. [76]

    Towards a reflection in creative experience questionnaire,

    C. Ford and N. Bryan-Kinns, “Towards a reflection in creative experience questionnaire,” inProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, ser. CHI ’23. New York, NY , USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3544548.3581077

  76. [77]

    Social sycophancy: A broader understanding of llm sycophancy,

    M. Cheng, S. Yu, C. Lee, P. Khadpe, L. Ibrahim, and D. Jurafsky, “Social sycophancy: A broader understanding of llm sycophancy,”arXiv preprint arXiv:2505.13995, 2025

  77. [78]

    Better slow than sorry: In- troducing positive friction for reliable dialogue systems,

    M. Inan, A. Sicilia, S. Dey, V . Dongre, T. Srinivasan, J. Thomason, G. Tur, D. Hakkani-Tur, and M. Alikhani, “Better slow than sorry: In- troducing positive friction for reliable dialogue systems,”arXiv preprint, 2025

  78. [2022]

    Available: https://doi.org/10.1145/3491102.3501991

    [Online]. Available: https://doi.org/10.1145/3491102.3501991

  79. [2023]

    Available: https://doi.org/10.1145/3544548.3581318

    [Online]. Available: https://doi.org/10.1145/3544548.3581318