Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Fine-Grained Appropriate Reliance: Human-AI Collaboration with a Multi-Step Transparent Decision Workflow for Complex Task Decomposition

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Following the AI's own decomposition steps helps humans catch misleading AI advice, but only when users get the intermediate sub-fact checks right.

desk verdict A genuinely new multi-step reliance study with a solid empirical core, but the misleading-advice headline and two hypothesis claims need tightening before the results are reliable. read the letter →

arxiv 2501.10909 v1 pith:5CK2NT7S submitted 2025-01-19 cs.AI cs.HC

classification cs.AIcs.HC
keywords human-AIcollaborationappropriatereliancemulti-stepdecisionworkflowtransparencycompositefact-checkingtaskdecompositionlargelanguagemodelscognitiveload
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish when and why showing a human the AI's internal multi-step reasoning improves joint decision making. Using composite fact-checking, where a claim must be verified through several sub-facts, it tests a Multi-Step Transparent (MST) workflow in which users follow the same decomposed steps as an LLM-based fact-checker, inspect the evidence retrieved at each step, and only then receive the AI's final verdict. In a between-subjects study with 233 crowd workers, the workflow did not beat one-step collaboration on average, but it did on the tasks where the AI's advice was misleading. The authors' central claim is that fine-grained appropriate reliance, meaning correct decisions at the intermediate steps, is a prerequisite for the workflow to help: users who got the sub-facts right benefited, while those who did not under-relied on correct AI advice. They also find that a cognitive-forcing nudge, annotating each document's usefulness, backfired by raising cognitive load.

What carries the argument

The central construct is AR-Intermediate, the accuracy of a user's verdicts on the three decomposed sub-facts, scored against expert-annotated ground truth at each step. It does double duty: it is the fine-grained measure of appropriate reliance the paper introduces, and it is the variable that separates users for whom the MST workflow succeeds from those for whom it produces under-reliance. The AI system is ProgramFC, an LLM pipeline that decomposes a composite claim into three sub-facts, verifies each against BM25-retrieved Wikipedia evidence, and aggregates the sub-verdicts into a final prediction; the MST workflow makes that pipeline visible and executable by the user, step by step.

What would settle it

Re-run the study with a separate pre-test of each participant's unaided fact-checking accuracy on comparable composite claims, and check whether the high-AR-Intermediate advantage on team performance, RAIR, and agreement fraction survives once that general-ability score is controlled for; if the advantage vanishes, the mediator is general competence rather than consideration of the intermediate steps.

Watch

Extended reading notes

Core claim

The paper's claim, stated on its own terms, is that fine-grained appropriate reliance at the level of intermediate AI steps is what makes a multi-step transparent workflow effective, and that this manifests most clearly exactly where one-step collaboration fails: when the AI's advice is wrong. Three preregistered hypotheses were tested. H1, that showing users the AI's decomposed steps and intermediate answers (global transparency) increases reliance, was supported. H2, that the MST workflow increases appropriate reliance relative to one-step collaboration, was not supported on average. H3, that more accurate intermediate user decisions produce more appropriate reliance on the final advice, was supported: in a median split on AR-Intermediate, high-consideration users showed higher team performance, higher agreement with the AI, and higher RAIR. Across the three tasks where the AI's final verdict was misleading, MST conditions matched or exceeded the one-step conditions' accuracy, while on the easy tasks the one-step conditions did better. The authors conclude that there is no one-size-fits-all decision workflow for optimal human-AI collaboration.

Load-bearing premise

The paper assumes that a user's accuracy on the sub-fact verdicts measures how carefully they considered the AI's intermediate steps, rather than merely measuring their general fact-checking ability.

Editorial extensions

If this is right

  • Providing decomposed steps and intermediate answers as explanations increases reliance on AI advice, extending the known over-reliance effect of explainable AI to multi-step transparency (H1).
  • On tasks where the AI's final advice is misleading, users in the MST workflow match or beat one-step users, so the workflow's benefit is context-specific rather than general (H2).
  • Users with low AR-Intermediate under-rely on correct AI advice, and their team performance, agreement fraction, and RAIR all drop, turning the workflow into a liability rather than an aid.
  • The document-usefulness annotation intervention raises mental demand and frustration, lowering team performance and appropriate reliance, a cognitive-load cost that offsets its intended forcing effect.
  • The MST workflow lowers user confidence after seeing AI advice and intermediate answers, consistent with the critical mindset that mitigates over-reliance in the misleading-advice setting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If AR-Intermediate mostly tracks general fact-checking competence rather than step-wise consideration, the practical lever for improving human-AI teams would be training or selection rather than workflow design; the paper's own data, where high-AR users do better on every measure at once, is consistent with that reading.
  • A testable extension the authors gesture at but do not run: an adaptive workflow that engages the multi-step form only when the AI's own confidence is low could capture the misleading-advice benefit without the under-reliance cost seen on easy tasks.
  • AR-Evidence, agreement with experts on document usefulness, correlated with team performance but not with the appropriate-reliance measures, hinting that evidence-level transparency supports accuracy without calibrating reliance; separating those two functions could sharpen interface design.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper reports a preregistered between-subjects experiment (N = 233) on AI-assisted composite fact-checking, comparing four conditions: a one-step workflow with final AI advice (Control), one-step plus global transparency (MST-GT), a multi-step transparent workflow where users verify each decomposed sub-fact before seeing AI advice (MSTworkflow), and the same workflow plus mandatory annotation of document usefulness (MSTworkflow+). The authors test three hypotheses: H1 (global transparency increases reliance), H2 (the MST workflow increases appropriate reliance), and H3 (accurate intermediate user decisions increase appropriate reliance on the final AI advice). They report that H1 is supported, H2 is not, and H3 is supported through a median split on intermediate accuracy. The broader claim is that the MST workflow can outperform one-step collaboration specifically when AI advice is misleading, and that fine-grained appropriate reliance is a prerequisite for that benefit. The measurement framework and the use of non-parametric tests are strengths, but several reported conclusions are not backed by the analyses as presented, and the central contextual claim rests on selected descriptive percentages rather than an inferential test.

Significance. If the results were robust, the paper would make a useful contribution to human-AI decision making. It introduces fine-grained measures of appropriate reliance (AR-Intermediate and AR-Evidence), uses a realistic LLM/RAG-based fact-checking task, preregisters hypotheses, and provides an a priori power analysis with enough participants (230 required, 233 analyzed). Public data and code on OSF, attention-check filtering, and an honest discussion of cognitive-load trade-offs are additional strengths. The paper also moves beyond the one-step decision paradigm that dominates the literature. However, the reported analyses contain inconsistencies — most notably for H1 and for the misleading-advice claim — and the H3 analysis has a confound that the current reporting does not address. The measurement framework is valuable, but the headline conclusions need to be re-derived or substantially moderated.

major comments (4)
  1. [Section 5.2.1, Table 5] The text states that 'compared to Control, MST-GT showed significantly higher Agreement Fraction' and uses this to support H1, but the reported post-hoc comparison in Table 5 is 'Control, MST-GT > MSTworkflow, MSTworkflow+.' That notation establishes only that Control and MST-GT both exceed the two MST conditions; it does not establish a significant MST-GT > Control difference. The paper therefore does not report the pairwise test that its H1 conclusion depends on, and the claim of support for H1 is unsupported as written.
  2. [Abstract; Section 6.1; Tables 4 and 5] The headline claim that the MST workflow outperforms one-step collaboration when AI advice is misleading is not supported by the reported analyses. Of the three tasks with misleading AI advice in Table 4, only Task 6 (25.8% vs 9.3%) and, weakly, Task 7 (62.9% vs 59.3%) favor MSTworkflow over Control; in Task 1, Control achieves 33.3% vs 14.5% for MSTworkflow and 28.1% for MSTworkflow+. No task-level significance test, mixed-effects model, or condition-by-task interaction is reported. Moreover, Table 5 shows that Control significantly outperforms all MST conditions on Team Performance-wid, the metric most directly tied to decisions where users disagree with AI advice. The abstract and Section 6.1 should either be supported by such an analysis or substantially moderated.
  3. [Section 5.2.2; Table 9] The support for H3 rests on a post-hoc median split of AR-Intermediate and suffers from a circularity/confound problem. AR-Intermediate is the accuracy of users' sub-fact verifications; users who are accurate on sub-facts are likely to be accurate on final fact-check decisions, and RAIR and Team Performance-wid are constructed from final-decision correctness relative to AI advice. Table 9 shows that the high-AR group simultaneously outperforms the low-AR group on Team Performance, Agreement Fraction, Team Performance-wid, and RAIR, which is exactly the pattern expected from a general-ability confound. A preregistered test of H3 or a sensitivity analysis that controls for overall task accuracy is needed before the causal interpretation in Section 6.1 ('explicit considerations ... indicate that an MST decision workflow can be effective') is warranted.
  4. [Section 4.4] The exclusion rule for participants with 'three or more such indications' of frequent switching is post-hoc in the sense that the paper does not state it was preregistered (the preregistration link is hidden) and no robustness check is reported with these participants included. This rule differentially removes participants who exhibit strong reliance changes, which is precisely the behavior analyzed for H2 and the reliance results. The authors should either show that the exclusion was preregistered or rerun the key analyses with and without these participants.
minor comments (6)
  1. [Key Words] The key word 'Mutli-step' is misspelled and should be 'Multi-step'.
  2. [Footnote 1] Footnote 1 states 'URL hidden to preserve anonymity'; because the paper's preregistration claims are load-bearing, an anonymized link or a registration number should be provided.
  3. [Table 5] The post-hoc notation in Table 5 (e.g., 'Control> MST-GT, MSTworkflow> MSTworkflow+') is ambiguous; report adjusted p-values and clearer pairwise notation.
  4. [Section 5.2.2] The median split is described as 'evenly re-split' but the exact group sizes and the split rule (median vs. top/bottom half) should be stated.
  5. [Figure 4] Figure 4 uses '**' without a caption definition; the caption should state that '**' denotes p < 0.017.
  6. [Table 8] The header 'Switch Faction' in Table 8 is a typo for 'Switch Fraction'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are empirical comparisons using externally anchored measures, not derivations that reduce to their own inputs.

full rationale

This is an empirical human-subjects study, not a formal derivation chain. The central claims (MST workflow helps in misleading-advice contexts; fine-grained accuracy correlates with reliance measures) are supported by task-level accuracy tables and non-parametric tests against tasks sampled from the public FEVEROUS-S benchmark. The reliance measures RAIR/RSR are adopted from prior work, including co-authored papers [41,93], but they are measurement definitions rather than load-bearing empirical premises, and they are applied to newly collected behavioral data; the citations do not supply a forbidden uniqueness theorem or ansatz. The H3 analysis (Table 6) splits participants on AR-Intermediate and shows higher RAIR/Team Performance in the high group. This could reflect a general-ability confound, and the paper's causal language ('consideration') outruns the measure, which is a construct-validity and statistical-inference concern, not a circularity by construction: AR-Intermediate (accuracy on sub-fact judgments) and RAIR (final-decision reliance relative to AI advice) are not definitionally linked by any equation in the paper. The misleading-advice claim also rests on selected descriptive rows in Table 4 and is contradicted by Task 1, but that is an evidentiary weakness, not a reduction of a prediction to a fitted input. No parameter is fitted and then renamed as a prediction, and no central premise is justified solely by a self-citation. Therefore the derivation chain, such as it is, is self-contained with respect to circularity, and the appropriate score is 0.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claims rest on the validity of expert annotation, the correctness of the AI's decompositions, and the assumption that AR-Intermediate captures consideration rather than general ability. The free parameters are analysis thresholds and task-selection choices that shape the results; the H3 conclusion is most sensitive to the median split and the exclusion rules.

free parameters (4)
  • AR-Intermediate median split = 0.73 (sample median, top/bottom 50%)
    Used to define 'high consideration' vs 'low consideration' groups in H3 (Section 5.2.2). The threshold is chosen post hoc and the conclusion depends on this split.
  • Switch exclusion threshold = 3 or more switches
    Participants switching from initial agreement to opposite AI advice 3+ times were excluded (Section 4.4). This removes frequent switchers and can bias reliance measures.
  • Minimum task completion time = 15 minutes
    Participants finishing in under 15 minutes were excluded as low-effort (Section 4.4). This threshold is arbitrary and could select for slower, more careful users.
  • Target AI accuracy for task selection = 70%
    Ten tasks were selected such that ProgramFC accuracy is 70%, close to its 67.7% overall (Section 4.2). This shapes the mix of correct/misleading AI advice.
assumptions (4)
  • domain assumption Expert annotations of sub-fact correctness and document usefulness are accurate ground truth.
    AR-Intermediate and AR-Evidence are computed against expert labels (Section 4.3.1). If expert labels are wrong, fine-grained measures and all H3 results are invalid.
  • domain assumption The AI's decomposed steps on the 10 selected tasks are all correct.
    Only tasks whose decomposed steps an author judged correct were retained (Section 4.2). Residual decomposition errors would mean users verify wrong sub-facts.
  • domain assumption RAIR/RSR are valid operationalizations of appropriate reliance.
    Adopted from prior work [41,93] without local validation; these measures assume binary final decisions and known ground truth.
  • domain assumption AR-Intermediate measures explicit consideration rather than general competence.
    H3 is interpreted causally ('prerequisite', 'cause under-reliance') from a correlational median split (Sections 5.2.2, 6.1). High-AR users also outperform on every outcome, consistent with a general-ability confound.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fine-Grained Appropriate Reliance: Human-AI Collaboration with a Multi-Step Transparent Decision Workflow for Complex Task Decomposition." pith.science (2026). https://pith.science/paper/5CK2NT7S

@misc{pith2026250110909,
  author       = {Pith},
  title        = {Pith review of: Fine-Grained Appropriate Reliance: Human-AI Collaboration with a Multi-Step Transparent Decision Workflow for Complex Task Decomposition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5CK2NT7S}},
  note         = {Machine review of arXiv:2501.10909}
}
read the original abstract

In recent years, the rapid development of AI systems has brought about the benefits of intelligent services but also concerns about security and reliability. By fostering appropriate user reliance on an AI system, both complementary team performance and reduced human workload can be achieved. Previous empirical studies have extensively analyzed the impact of factors ranging from task, system, and human behavior on user trust and appropriate reliance in the context of one-step decision making. However, user reliance on AI systems in tasks with complex semantics that require multi-step workflows remains under-explored. Inspired by recent work on task decomposition with large language models, we propose to investigate the impact of a novel Multi-Step Transparent (MST) decision workflow on user reliance behaviors. We conducted an empirical study (N = 233) of AI-assisted decision making in composite fact-checking tasks (i.e., fact-checking tasks that entail multiple sub-fact verification steps). Our findings demonstrate that human-AI collaboration with an MST decision workflow can outperform one-step collaboration in specific contexts (e.g., when advice from an AI system is misleading). Further analysis of the appropriate reliance at fine-grained levels indicates that an MST decision workflow can be effective when users demonstrate a relatively high consideration of the intermediate steps. Our work highlights that there is no one-size-fits-all decision workflow that can help obtain optimal human-AI collaboration. Our insights help deepen the understanding of the role of decision workflows in facilitating appropriate reliance. We synthesize important implications for designing effective means to facilitate appropriate reliance on AI systems in composite tasks, positioning opportunities for the human-centered AI and broader HCI communities.

Figures

Figures reproduced from arXiv: 2501.10909 by the authors.

Figure 1
Figure 1. Illustration of the multi-step workflow on the composite fact-checking tasks using the ProgramFC [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Screenshots of the composite fact-checking task interface with the MST workflow. (A) The starting [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. An illustration of the procedure that participants followed in our study. [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Bar plot illustrating the distribution of the cognitive load across different experimental conditions in [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Line plot illustrating the confidence dynamics among users after receiving the AI advice (and explana [PITH_FULL_IMAGE:figures/full_fig_p020_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.

Reference graph

Works this paper leans on

116 extracted references · 60 canonical work pages · cited by 1 Pith paper

  1. [1]

    Amina Adadi and Mohammed Berrada. 2018. Peeking inside the black-box: a survey on explainable artificial intelligence (XAI). IEEE access 6 (2018), 52138–52160

  2. [2]

    NIST AI. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0)

  3. [3]

    Kumar Akash, Griffon McMahon, Tahira Reid, and Neera Jain. 2020. Human trust-based feedback control: Dynamically varying automation transparency to optimize human-machine interactions. IEEE Control Systems Magazine 40, 6 (2020), 98–116

  4. [4]

    Jennifer Allen, Antonio A Arechar, Gordon Pennycook, and David G Rand. 2021. Scaling up fact-checking using the wisdom of crowds. Science advances 7, 36 (2021), eabf4393

  5. [5]

    Rami Aly, Zhijiang Guo, Michael Sejr Schlichtkrull, James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Oana Cocarascu, and Arpit Mittal. 2021. FEVEROUS: Fact Extraction and VERification Over Unstructured and Structured information. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)

  6. [6]

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. In Proceedings of the 2019 chi conference on human factors in computing systems . 1–13. , Vol. 1, No. 1, Article . Publication date: January 2025. Human-AI Colla...

  7. [7]

    Ahmer Arif, John J Robinson, Stephanie A Stanek, Elodie S Fichet, Paul Townsend, Zena Worku, and Kate Starbird

  8. [8]

    Isabelle Augenstein, Timothy Baldwin, Meeyoung Cha, Tanmoy Chakraborty, Giovanni Luca Ciampaglia, David Corney, Renee DiResta, Emilio Ferrara, Scott Hale, Alon Halevy, et al. 2024. Factuality challenges in the era of large language models and opportunities for fact-checking. Nature Machine Intelligence (2024), 1–12

Show all 116 references
  1. [9]

    Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI Conference on Human Fact...

  2. [10]

    Michael S Bernstein, Greg Little, Robert C Miller, Björn Hartmann, Mark S Ackerman, David R Karger, David Crowell, and Katrina Panovich. 2010. Soylent: a word processor with a crowd inside. In Proceedings of the 23nd annual ACM symposium on User interface software and technolo...

  3. [11]

    Astrid Bertrand, Rafik Belloum, James R Eagan, and Winston Maxwell. 2022. How cognitive biases affect XAI-assisted decision-making: A systematic review. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society . 78–91

  4. [12]

    Zana Buçinca, Phoebe Lin, Krzysztof Z Gajos, and Elena L Glassman. 2020. Proxy tasks and subjective measures can be misleading in evaluating explainable ai systems. In Proceedings of the 25th international conference on intelligent user interfaces. 454–464

  5. [13]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1 (2021), 1–21

  6. [14]

    Wouter Bulten, Maschenka Balkenhol, Jean-Joël Awoumou Belinga, Américo Brilhante, Aslı Çakır, Lars Egevad, Martin Eklund, Xavier Farré, Katerina Geronatsiou, Vincent Molinié, et al. 2021. Artificial intelligence assistance significantly improves Gleason grading of prostate bio...

  7. [15]

    Canyu Chen and Kai Shu. 2024. Combating misinformation in the age of llms: Opportunities and challenges. AI Magazine 45, 3 (2024), 354–368

  8. [16]

    Jessie YC Chen, Shan G Lakhmani, Kimberly Stowers, Anthony R Selkowitz, Julia L Wright, and Michael Barnes. 2018. Situation awareness-based agent transparency and human-autonomy teaming effectiveness. Theoretical issues in ergonomics science 19, 3 (2018), 259–282

  9. [17]

    Chun-Wei Chiang and Ming Yin. 2022. Exploring the Effects of Machine Learning Literacy Interventions on Laypeople’s Reliance on Machine Learning Models. In IUI 2022: 27th International Conference on Intelligent User Interfaces, Helsinki, Finland, March 22 - 25, 2022 , Giulio J...

  10. [18]

    Chun-Wei Chiang and Ming Yin. 2021. You’d Better Stop! Understanding Human Reliance on Machine Learning Models under Covariate Shift. In 13th ACM Web Science Conference 2021 . 120–129

  11. [19]

    Yu-Liang Chou, Catarina Moreira, Peter Bruza, Chun Ouyang, and Joaquim Jorge. 2022. Counterfactuals and causability in explainable artificial intelligence: Theory, algorithms, and applications. Information Fusion 81 (2022), 59–83

  12. [20]

    Michael Chromik, Malin Eiband, Felicitas Buchner, Adrian Krüger, and Andreas Butz. 2021. I think i get your point, AI! the illusion of explanatory depth in explainable AI. In 26th International Conference on Intelligent User Interfaces . 307–317

  13. [22]

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al . 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25, 70 (2024), 1–53

  14. [23]

    Lacey Colligan, Henry WW Potts, Chelsea T Finn, and Robert A Sinkin. 2015. Cognitive workload changes for nurses transitioning from a legacy system with paper documentation to a commercial electronic health record. International journal of medical informatics 84, 7 (2015), 469–476

  15. [24]

    António Correia, Andrea Grover, Daniel Schneider, Ana Paula Pimentel, Ramon Chaves, Marcos Antonio De Almeida, and Benjamim Fonseca. 2023. Designing for Hybrid Intelligence: A Taxonomy and Survey of Crowd-Machine Interaction. Applied Sciences 13, 4 (2023), 2198

  16. [25]

    Ewart J De Visser, Marieke MM Peeters, Malte F Jung, Spencer Kohn, Tyler H Shaw, Richard Pak, and Mark A Neerincx

  17. [26]

    Tim Draws, Alisa Rieger, Oana Inel, Ujwal Gadiraju, and Nava Tintarev. 2021. A checklist to combat cognitive biases in crowdsourcing. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol. 9. 48–59. , Vol. 1, No. 1, Article . Publication date: Janu...

  18. [27]

    Upol Ehsan, Q Vera Liao, Michael Muller, Mark O Riedl, and Justin D Weisz. 2021. Expanding explainability: Towards social transparency in ai systems. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–19

  19. [28]

    Upol Ehsan and Mark O Riedl. 2020. Human-centered explainable ai: towards a reflective sociotechnical approach. In International Conference on Human-Computer Interaction . Springer, 449–466

  20. [29]

    Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining ....

  21. [30]

    Franz Faul, Edgar Erdfelder, Axel Buchner, and Albert-Georg Lang. 2009. Statistical power analyses using G* Power 3.1: Tests for correlation and regression analyses. Behavior research methods 41, 4 (2009), 1149–1160

  22. [31]

    Heike Felzmann, Eduard Fosch-Villaronga, Christoph Lutz, and Aurelia Tamò-Larrieux. 2020. Towards transparency by design for artificial intelligence. Science and engineering ethics 26, 6 (2020), 3333–3361

  23. [32]

    Raymond Fok and Daniel S Weld. 2023. In search of verifiability: Explanations rarely enable complementary performance in AI-advised decision making. AI Magazine (2023)

  24. [33]

    Ujwal Gadiraju, Ricardo Kawase, and Stefan Dietze. 2014. A taxonomy of microtasks on the web. In Proceedings of the 25th ACM conference on Hypertext and social media . 218–223

  25. [34]

    Ben Green and Yiling Chen. 2021. Algorithmic risk assessments can alter human decision-making processes in high-stakes government contexts. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–33

  26. [35]

    Madeleine Grunde-McLaughlin, Michelle S Lam, Ranjay Krishna, Daniel S Weld, and Jeffrey Heer. 2023. Designing LLM Chains by Adapting Techniques from Crowdsourcing Workflows. arXiv preprint arXiv:2312.11681 (2023)

  27. [36]

    Andrew Guess, Brendan Nyhan, and Jason Reifler. 2018. Selective exposure to misinformation: Evidence from the consumption of fake news during the 2016 US presidential campaign. European Research Council 9, 3 (2018), 4

  28. [37]

    Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics 10 (2022), 178–206

  29. [38]

    Gaole He, Nilay Aishwarya, and Ujwal Gadiraju. 2025. Is Conversational XAI All You Need? Human-AI Decision Making With a Conversational XAI Assistant. In Proceedings of the 30th International Conference on Intelligent User Interfaces

  30. [39]

    Gaole He, Stefan Buijsman, and Ujwal Gadiraju. 2023. How Stated Accuracy of an AI System and Analogies to Explain Accuracy Affect Human Reliance on the System. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (2023)

  31. [40]

    Gaole He, Gianluca Demartini, and Ujwal Gadiraju. 2025. Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems

  32. [41]

    Gaole He, Lucie Kuiper, and Ujwal Gadiraju. 2023. Knowing About Knowing: An Illusion of Human Competence Can Hinder Appropriate Reliance on AI Systems. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–18

  33. [42]

    Patrick Hemmer, Max Schemmer, Michael Vössing, and Niklas Kühl. 2021. Human-AI Complementarity in Hybrid Intelligence Systems: A Structured Literature Review. PACIS (2021), 78

  34. [43]

    Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. 2021. Towards accountability for machine learning datasets: Practices from software engineering and infrastructure. In Proceedings of the 2021 ACM ...

  35. [44]

    Md Saroar Jahan and Mourad Oussalah. 2023. A systematic review of hate speech automatic detection using natural language processing. Neurocomputing 546 (2023), 126232

  36. [45]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. Comput. Surveys 55, 12 (2023), 1–38

  37. [46]

    Shan Jiang and Christo Wilson. 2018. Linguistic signals under misinformation and fact-checking: Evidence from user comments on social media. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–23

  38. [47]

    Patricia K Kahr, Gerrit Rooks, Martijn C Willemsen, and Chris CP Snijders. 2024. Understanding Trust and Re- liance Development in AI Advice: Assessing Model Accuracy, Model Explanations, and Experiences from Previous Interactions. ACM Transactions on Interactive Intelligent S...

  39. [48]

    Ece Kamar and Lydia Manikonda. 2017. Complementing the Execution of AI Systems with Human Computation.. In AAAI Workshops

  40. [49]

    Alexandra D Kaplan, Theresa T Kessler, J Christopher Brill, and Peter A Hancock. 2023. Trust in artificial intelligence: Meta-analytic findings. Human factors 65, 2 (2023), 337–359

  41. [50]

    Davinder Kaur, Suleyman Uslu, Kaley J Rittichier, and Arjan Durresi. 2022. Trustworthy artificial intelligence: a review. ACM computing surveys (CSUR) 55, 2 (2022), 1–38. , Vol. 1, No. 1, Article . Publication date: January 2025. Human-AI Collaboration with A Multi-step Transp...

  42. [51]

    Jooyeon Kim, Behzad Tabibian, Alice Oh, Bernhard Schölkopf, and Manuel Gomez-Rodriguez. 2018. Leveraging the crowd to detect and reduce the spread of fake news and misinformation. In Proceedings of the eleventh ACM international conference on web search and data mining . 324–332

  43. [52]

    I’m Not Sure, But

    Sunnie SY Kim, Q Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Wortman Vaughan. 2024. "I’m Not Sure, But... ": Examining the Impact of Large Language Models’ Uncertainty Expression on User Reliance and Trust. In The 2024 ACM Conference on Fairness, Accountabil...

  44. [53]

    Aniket Kittur, Susheel Khamkar, Paul André, and Robert Kraut. 2012. CrowdWeaver: visually managing complex crowd work. In Proceedings of the ACM 2012 Conference on Computer Supported Cooperative Work . 1033–1036

  45. [54]

    Aniket Kittur, Boris Smus, Susheel Khamkar, and Robert E Kraut. 2011. Crowdforge: Crowdsourcing complex work. In Proceedings of the 24th annual ACM symposium on User interface software and technology . 43–52

  46. [55]

    Moritz Körber. 2019. Theoretical considerations and development of a questionnaire to measure trust in automation. In Proceedings of the 20th Congress of the International Ergonomics Association (IEA 2018) Volume VI: Transport Ergonomics and Human Factors (TEHF), Aerospace Hum...

  47. [56]

    Joshua A Kroll. 2021. Outlining traceability: A principle for operationalizing accountability in computing systems. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . 758–771

  48. [57]

    Vera Liao, and Chenhao Tan

    Vivian Lai, Chacha Chen, Alison Smith-Renner, Q. Vera Liao, and Chenhao Tan. 2023. Towards a Science of Human-AI Decision Making: An Overview of Design Space in Empirical Human-Subject Studies. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transpar...

  49. [58]

    Why is ’Chicago’ deceptive?

    Vivian Lai, Han Liu, and Chenhao Tan. 2020. "Why is ’Chicago’ deceptive?" Towards Building Model-Driven Tutorials for Humans. In CHI ’20: CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, April 25-30, 2020 , Regina Bernhaupt, Florian ’Floyd’ Mueller, Dav...

  50. [59]

    Yunshi Lan, Gaole He, Jinhao Jiang, Jing Jiang, Wayne Xin Zhao, and Ji-Rong Wen. 2022. Complex knowledge base question answering: A survey. IEEE Transactions on Knowledge and Data Engineering (2022)

  51. [60]

    John D Lee and Neville Moray. 1994. Trust, self-confidence, and operators’ adaptation to automation. International journal of human-computer studies 40, 1 (1994), 153–184

  52. [61]

    John D Lee and Katrina A See. 2004. Trust in automation: Designing for appropriate reliance. Human factors 46, 1 (2004), 50–80

  53. [62]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing S...

  54. [63]

    Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. 2023. Trustworthy AI: From principles to practices. Comput. Surveys 55, 9 (2023), 1–46

  55. [64]

    Miaoran Li, Baolin Peng, Michel Galley, Jianfeng Gao, and Zhu Zhang. 2024. Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models. In Findings of the Association for Computational Linguistics: NAACL

  56. [65]

    Yugang Li, Baizhou Wu, Yuqi Huang, and Shenghua Luan. 2024. Developing trustworthy artificial intelligence: insights from research on interpersonal, human-automation, and human-AI trust. Frontiers in Psychology 15 (2024), 1382693

  57. [66]

    Q Vera Liao and Kush R Varshney. 2021. Human-Centered Explainable AI (XAI): From Algorithms to User Experiences. arXiv preprint arXiv:2110.10790 (2021)

  58. [67]

    Vera Liao and Jennifer Wortman Vaughan

    Q. Vera Liao and Jennifer Wortman Vaughan. 2024. AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap. Harvard Data Science Review Special Issue 5 (may 31 2024). https://hdsr.mitpress.mit.edu/pub/aelql9qy

  59. [68]

    Greg Little, Lydia B Chilton, Max Goldman, and Robert C Miller. 2010. Exploring iterative and parallel human computation processes. In Proceedings of the ACM SIGKDD workshop on human computation . 68–76

  60. [69]

    Zhuoran Lu, Dakuo Wang, and Ming Yin. 2024. Does more advice help? the effects of second opinions in AI-assisted decision making. Proceedings of the ACM on Human-Computer Interaction 8, CSCW1 (2024), 1–31

  61. [70]

    Zhuoran Lu and Ming Yin. 2021. Human Reliance on Machine Learning Models When Performance Feedback is Limited: Heuristics and Risks. In CHI ’21: CHI Conference on Human Factors in Computing Systems, Virtual Event / Yokohama, Japan, May 8-13, 2021, Yoshifumi Kitamura, Aaron Qui...

  62. [71]

    Potsawee Manakul, Adian Liusie, and Mark Gales. 2023. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and K...

  63. [72]

    David Mascharka, Philip Tran, Ryan Soklaski, and Arjun Majumdar. 2018. Transparency by design: Closing the gap between performance and interpretability in visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognition. 4942–4950

  64. [73]

    Joseph E Mercado, Michael A Rupp, Jessie YC Chen, Michael J Barnes, Daniel Barber, and Katelyn Procci. 2016. Intelligent agent transparency in human–agent teaming for Multi-UxV management. Human factors 58, 3 (2016), 401–415

  65. [74]

    Christopher A Miller. 2021. Trust, transparency, explanation, and planning: Why we need a lifecycle perspective on human-automation interaction. In Trust in human-robot interaction . Elsevier, 233–257

  66. [75]

    Brent Mittelstadt, Chris Russell, and Sandra Wachter. 2019. Explaining explanations in AI. In Proceedings of the conference on fairness, accountability, and transparency . 279–288

  67. [76]

    Mainack Mondal, Leandro Araújo Silva, and Fabrício Benevenuto. 2017. A measurement study of hate speech in social media. In Proceedings of the 28th ACM conference on hypertext and social media . 85–94

  68. [77]

    An Nguyen, Aditya Kharosekar, Matthew Lease, and Byron Wallace. 2018. An interpretable joint graphical model for fact-checking from crowds. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32

  69. [78]

    An T Nguyen, Aditya Kharosekar, Saumyaa Krishnan, Siddhesh Krishnan, Elizabeth Tate, Byron C Wallace, and Matthew Lease. 2018. Believe it or not: designing a human-ai partnership for mixed-initiative fact-checking. In Proceedings of the 31st annual ACM symposium on user interf...

  70. [79]

    Mahsan Nourani, Chiradeep Roy, Jeremy E Block, Donald R Honeycutt, Tahrima Rahman, Eric Ragan, and Vibhav Gogate. 2021. Anchoring Bias Affects Mental Model Formation and User Reliance in Explainable AI Systems. In 26th International Conference on Intelligent User Interfaces . 340–350

  71. [80]

    Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, and Preslav Nakov

  72. [81]

    Sihang Qiu, Ujwal Gadiraju, and Alessandro Bozzon. 2020. Improving worker engagement through conversational microtask crowdsourcing. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–12

  73. [82]

    Amy Rechkemmer and Ming Yin. 2022. When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning Models. In CHI ’22: CHI Conference on Human Factors in Computing Systems, New Orleans, LA, USA, 29 April 2022 - 5 May 2022 ,...

  74. [83]

    Daniela Retelny, Michael S Bernstein, and Melissa A Valentine. 2017. No workflow can ever be enough: How crowdsourcing workflows constrain complex work. Proceedings of the ACM on Human-Computer Interaction 1, CSCW (2017), 1–23

  75. [84]

    Why should i trust you?

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " Why should i trust you?" Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 1135–1144

  76. [85]

    Vincent Robbemond, Oana Inel, and Ujwal Gadiraju. 2022. Understanding the Role of Explanation Modality in AI- assisted Decision-making. In Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization . 223–233

  77. [86]

    Stephen E Robertson and Steve Walker. 1994. Some simple effective approximations to the 2-poisson model for probabilistic weighted retrieval. InSIGIR’94: Proceedings of the Seventeenth Annual International ACM-SIGIR Conference on Research and Development in Information Retriev...

  78. [87]

    Kevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina, Stefano Mizzaro, and Gianluca Demartini. 2020. Can the crowd identify misinformation objectively? The effects of judgment scale and assessor’s background. In Proceedings of the 43rd international ACM SIGIR conference...

  79. [88]

    Jon Roozenbeek and Sander Van der Linden. 2024. The psychology of misinformation . Cambridge University Press

  80. [89]

    Mohammed Saeed, Nicolas Traub, Maelle Nicolas, Gianluca Demartini, and Paolo Papotti. 2022. Crowdsourced fact-checking at Twitter: How does the crowd compare with experts?. In Proceedings of the 31st ACM international conference on information & knowledge management . 1736–1746

  81. [90]

    Waddah Saeed and Christian Omlin. 2023. Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities. Knowledge-Based Systems 263 (2023), 110273

  82. [91]

    Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2023. A Missing Piece in the Puzzle: Considering the Role of Task Complexity in Human-AI Decision Making. In Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization. 215–227

  83. [92]

    Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2024. Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision Making. InProceedings of the CHI Conference on Human Factors in Computing Systems . 1–17. , ...

  84. [93]

    Max Schemmer, Patrick Hemmer, Niklas Kühl, Carina Benz, and Gerhard Satzger. 2022. Should I Follow AI-based Advice? Measuring Appropriate Reliance in Human-AI Decision-Making. In ACM Conference on Human Factors in Computing Systems (CHI’22), Workshop on Trust and Reliance in A...

  85. [94]

    Max Schemmer, Patrick Hemmer, Maximilian Nitsche, Niklas Kühl, and Michael Vössing. 2022. A meta-analysis of the utility of explainable artificial intelligence in human-AI decision-making. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society . 617–626

  86. [95]

    Max Schemmer, Niklas Kühl, Carina Benz, and Gerhard Satzger. 2022. On the Influence of Explainable AI on Automation Bias. In Proceedings of the 30th European Conference on Information Systems (ECIS), Timis , oara, RO, June 18-24, 2022

  87. [96]

    The human body is a black box

    Mark Sendak, Madeleine Clare Elish, Michael Gao, Joseph Futoma, William Ratliff, Marshall Nichols, Armando Bedoya, Suresh Balu, and Cara O’Brien. 2020. " The human body is a black box" supporting clinical decision-making with deep learning. In Proceedings of the 2020 conferenc...

  88. [97]

    Hua Shen, Chieh-Yang Huang, Tongshuang Wu, and Ting-Hao Kenneth Huang. 2023. ConvXAI: Delivering heteroge- neous AI explanations via conversations to support human-AI scientific writing. In Companion Publication of the 2023 Conference on Computer Supported Cooperative Work and...

  89. [98]

    Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu. 2017. Fake news detection on social media: A data mining perspective. ACM SIGKDD explorations newsletter 19, 1 (2017), 22–36

  90. [99]

    Dylan Slack, Satyapriya Krishna, Himabindu Lakkaraju, and Sameer Singh. 2023. Explaining machine learning models with interactive natural language conversations using TalkToModel. Nature Machine Intelligence 5, 8 (2023), 873–883

  91. [100]

    Jessie J Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wortman Vaughan. 2022. Real ml: Recognizing, exploring, and articulating limitations of machine learning research. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparenc...

  92. [101]

    James Thorne and Andreas Vlachos. 2018. Automated Fact Checking: Task Formulations, Methods and Future Directions. In Proceedings of the 27th International Conference on Computational Linguistics . 3346–3359

  93. [102]

    Richard Tomsett, Alun Preece, Dave Braines, Federico Cerutti, Supriyo Chakraborty, Mani Srivastava, Gavin Pearson, and Lance Kaplan. 2020. Rapid trust calibration through interpretable and uncertainty-aware AI. Patterns 1, 4 (2020), 100049

  94. [103]

    Svitlana Volkova and Jin Yea Jang. 2018. Misleading or falsification: Inferring deceptive strategies and types in online news and social media. In Companion Proceedings of the The Web Conference 2018 . 575–583

  95. [104]

    Svitlana Volkova, Kyle Shaffer, Jin Yea Jang, and Nathan Hodas. 2017. Separating facts from fiction: Linguistic models to classify suspicious and trusted news posts on twitter. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (volume 2...

  96. [105]

    Michael Vössing, Niklas Kühl, Matteo Lind, and Gerhard Satzger. 2022. Designing transparency for effective human-AI collaboration. Information Systems Frontiers 24, 3 (2022), 877–895

  97. [106]

    Xinru Wang and Ming Yin. 2021. Are Explanations Helpful? A Comparative Study of the Effects of Explanations in AI-Assisted Decision-Making. In 26th International Conference on Intelligent User Interfaces . 318–328

  98. [107]

    Tongshuang Wu, Ellen Jiang, Aaron Donsbach, Jeff Gray, Alejandra Molina, Michael Terry, and Carrie J Cai. 2022. Promptchainer: Chaining large language model prompts through visual programming. In CHI Conference on Human Factors in Computing Systems Extended Abstracts . 1–10

  99. [108]

    Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In Proceedings of the 2022 CHI conference on human factors in computing systems. 1–22

  100. [109]

    Tongshuang Wu, Haiyi Zhu, Maya Albayrak, Alexis Axon, Amanda Bertsch, Wenxing Deng, Ziqi Ding, Bill Guo, Sireesh Gururaja, Tzu-Sheng Kuo, et al. 2023. Llms as workers in human-computational algorithms? replicating crowdsourcing pipelines with llms. arXiv preprint arXiv:2307.10...

  101. [110]

    Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. 2019. Understanding the effect of accuracy on trust in machine learning models. In Proceedings of the 2019 chi conference on human factors in computing systems . 1–12

  102. [111]

    Lifan Yuan, Yangyi Chen, Ganqu Cui, Hongcheng Gao, Fangyuan Zou, Xingyi Cheng, Heng Ji, Zhiyuan Liu, and Maosong Sun. 2023. Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and LLMs evaluations. Advances in Neural Information Processing Systems 36 (2023)...

  103. [112]

    Daniel Zhang, Yang Zhang, Qi Li, Thomas Plummer, and Dong Wang. 2019. Crowdlearn: A crowd-ai hybrid system for deep learning-based damage assessment applications. In 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS). IEEE, 1221–1232

  104. [113]

    Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay. 2019. Deep learning based recommender system: A survey and new perspectives. ACM computing surveys (CSUR) 52, 1 (2019), 1–38

  105. [114]

    Vera Liao, and Rachel K

    Yunfeng Zhang, Q. Vera Liao, and Rachel K. E. Bellamy. 2020. Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. In FAT* ’20: Conference on Fairness, Accountability, and Transparency, , Vol. 1, No. 1, Article . Publication dat...

  106. [2017]

    In Proceedings of the 2017 ACM conference on computer supported cooperative work and social computing

    A closer look at the self-correcting crowd: Examining corrections in online rumors. In Proceedings of the 2017 ACM conference on computer supported cooperative work and social computing . 155–168

  107. [2020]

    International journal of social robotics 12, 2 (2020), 459–478

    Towards a theory of longitudinal trust calibration in human–robot teams. International journal of social robotics 12, 2 (2020), 459–478

  108. [2023]

    In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Fact-Checking Complex Claims with Program-Guided Reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 6981–7004

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.