REVIEW 3 major objections 5 minor 118 references
Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read A 248-person study finds that letting users edit an LLM assistant's plan and approve each action does not calibrate their trust; users trust plausible-sounding plans even when they are wrong.
desk verdict Well-powered, honest HCI study with a useful null result, but the post-hoc filtering of execution-stage analyses needs a robustness check before the trust findings can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying objects are the plan-then-execute LLM agent, which first outputs a hierarchical step-wise plan and then converts each primary step into one tool call; a 2x2 factorial manipulation of user involvement (automatic vs user-involved planning, automatic vs user-involved execution); and two calibrated-trust measures, $CT_p$ and $CT_e$, computed as the frequency with which a user's binary 'do you trust it' answer agrees with ground truth (author-annotated plan quality for planning, execution accuracy for execution). These measures let the authors separate whether involvement changes what users actually do from whether it changes how accurately they gauge the agent.
What would settle it
A replication that scores plan quality with blind independent annotators and measures trust with a validated multi-item scale and behavioral reliance (e.g., frequency of overrides) would settle the null result: if user involvement then produces significant calibrated-trust differences, the paper's conclusion is an artifact; if not, the plausibility explanation stands. A second test is to make wrong plans visibly inconsistent with the task instructions (e.g., omitting a required step) and check whether involvement recovers calibration.
Extended reading notes
Core claim
The central claim is that in a plan-then-execute LLM agent workflow, user involvement in planning and execution is a double-edged sword: it can improve task performance when the agent's plan is imperfect or its execution drifts, but it fails to calibrate user trust because LLM-generated plans are convincingly plausible even when wrong. Across four experimental conditions (automatic vs user-involved planning crossed with automatic vs user-involved execution) and six daily tasks, the authors found no significant improvement in calibrated trust in the plan ($CT_p$) or in the execution ($CT_e$) from user involvement; the only significant trust effect was negative, with user-involved planning lowering calibrated trust on a simple alarm task where the agent's original plan was already correct. Task performance results were more mixed: user planning edits fixed a grammar-error plan in a currency transaction task, and user execution oversight corrected a wrong itinerary selection in a trip-planning task, but on tasks with initially high-quality plans user edits often reduced plan quality. The authors interpret the pattern as miscalibration caused by plan plausibility, arguing that users treat well-structured, readable plans as trustworthy regardless of their actual correctness.
Load-bearing premise
The study's measurement of calibrated trust rests on a single binary trust question combined with the authors' own ratings of plan quality and execution accuracy; if that measure is too coarse or biased, the null effect of user involvement could be a measurement artifact instead of a real psychological result.
Editorial extensions
If this is right
- If correct, adding plan-editing and action-approval controls to an LLM assistant is not by itself enough to make users trust it appropriately; the assistant must also surface the actual quality of its plans (e.g., by flagging uncertain steps) to prevent over-trust in plausible wrong plans.
- User involvement in the execution stage yields more consistent performance gains than involvement in planning, so designers should prioritize lightweight step-level verification over heavy plan-editing when resources are constrained.
- Because user edits can reduce the quality of already-good plans, an adaptive scheme that restricts editing when the plan is likely correct could preserve performance while lowering user effort.
- The strong positive correlation between plan quality and calibrated trust implies that improving the quality of the initial LLM plan (or making its flaws more visible) is the most direct route to better trust calibration and task success.
- High-risk tasks with imperfect initial plans show very low calibrated trust, meaning users tend to trust wrong high-stakes plans; risk flags or independent plan checks could be needed for finance-type tasks.
Reading between the lines
- The observed null effect may be partly a measurement artifact: a single binary trust item collapsed over six heterogeneous tasks cannot capture the multi-faceted trust (reliability, predictability, intention) that the paper's own questionnaire measures, so richer trust measures or behavioral reliance (e.g., how often users override) might show calibration effects the paper misses.
- If plausibility is the driver, a concrete testable extension is to systematically degrade the visual/logical structure of a wrong plan (e.g., inserting contradictions or omitting one required action) and measure whether calibrated trust rises; the paper's own task-1 (a grammar error) hints that users can fix such flaws, but only when they notice them.
- The paper's 'simulate first, then involve' suggestion points to a design where the LLM runs a full plan-and-execute pass, the user reviews the outcome, and only then does the assistant act for real; this could cut cognitive load while keeping the performance benefit of execution oversight, and could be tested directly against the fixed user-involved-planning/user-involved-execution condition.
- Because the study is built on a simulated API environment, the transferability of the trust-calibration findings to real daily tasks (with real money, real bookings, real consequences) is untested; a field deployment comparing trust and performance on the same six tasks would clarify whether the plausibility effect survives real stakes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a 2×2 factorial between-subjects experiment (N=248) in which participants collaborated with a plan-then-execute LLM agent on six daily tasks varying in risk and initial plan quality. The planning factor contrasts automatic vs. user-editable plans; the execution factor contrasts automatic vs. user-approve/modify action execution. The main dependent variables are calibrated trust in planning (CT_p) and in execution (CT_e), plan quality, action-sequence accuracy (ACC_s), and execution accuracy (ACC_e). The authors test four hypotheses about user involvement improving trust calibration and task performance. Results show that user involvement in planning or execution does not significantly increase calibrated trust, and performance effects are inconsistent: user involvement helps in some tasks (e.g., fixing grammar errors or correcting wrong itinerary selection) but harms plan quality in tasks with initially correct plans. The authors interpret the findings as evidence that LLM agents are a double-edged sword, with plausible but incorrect plans misleading user trust.
Significance. If the results are robust, the paper is a valuable empirical contribution to the growing literature on human-AI collaboration with LLM agents. It reports a well-powered sample with a documented sample-size calculation, honest presentation of null results, a task set with varying stakes, and openly available code and analysis repositories. The finding that user involvement does not calibrate trust despite high subjective trust is important and has clear design implications. However, the post-hoc filtering used in the execution-stage analyses introduces a risk of selection bias, and the reliance on author-annotated plan quality without inter-rater reliability weakens the measurement foundation. These issues need to be addressed before the central conclusions can be fully accepted.
major comments (3)
- [5.2.3 and 5.2.4 (Tables 5 and 6)] The decision to 'filter out the tasks where plan quality decreased after user-involved planning' before testing H3 and H4 is a post-hoc, outcome-dependent exclusion. Plan quality is not a pre-treatment covariate; it is itself affected by the planning manipulation, and Table 8 shows that plan quality is significantly correlated with CT_e (r=0.221, p<0.001), ACC_s (r=0.446, p<0.001), and ACC_e (r=0.400, p<0.001). Excluding exactly the observations in which user-involved planning was harmful removes the cases most likely to show lower trust calibration and lower execution accuracy, and because the probability of exclusion differs by condition, the comparison in Tables 5 and 6 is potentially biased in favor of user-involved execution. The paper does not report the corresponding analyses on the full data, nor does it specify the exact filtering rule (e.g., whether the exclusion removes the task instance from all four conditions or only from the UP conditions). Because H3 and H4 are load-bearing for the central claim that user involvement fails to calibrate trust and does not consistently improve performance, the authors should re-analyze the full data and show that the conclusions are unchanged, or justify the filtering with a pre-specified, outcome-independent criterion.
- [5.1 and 4.3 (plan quality rating)] The plan quality ratings used both for CT_p (Eq. 1) and for the filtering in Sections 5.2.3-5.2.4 are author-annotated, and the paper reports no inter-rater reliability. Section 5.1 states that 'all edited plans in user-involved planning conditions are evaluated by the authors,' but no agreement statistic (e.g., Cohen's kappa) or a second coder is mentioned. Given that the 5-point rubric includes qualitative distinctions (e.g., 'grammar errors' vs. 'action intent mismatch'), the absence of reliability evidence makes it difficult to rule out idiosyncratic rating as an alternative explanation for the task-specific patterns in Table 3. The authors should report reliability on a subset of plans, or use a more objective scoring procedure (e.g., automatic comparison against the ground-truth plan structure).
- [4.3, Eqs. (1)-(2)] The calibrated trust measures combine a single binary trust item with a dichotomized version of a 5-point plan quality rating and of execution accuracy. This coarse measurement may be insensitive to the trust miscalibration the paper claims to rule out: for example, Eq. (1) codes a participant as calibrated if they distrust any plan with quality <5, regardless of whether the plan is quality 4 (near-correct) or quality 1 (seriously flawed). The paper does not report whether the null H1/H3 results persist when the multi-item Trust-in-Automation subscales are used to construct a graded calibration score, or when plan quality is treated ordinally. Without such robustness evidence, the null effects could reflect measurement ceiling/floor rather than a genuine absence of calibration.
minor comments (5)
- [6.3] The section title 'Limitations and Potentail Biases' contains a typo; 'Potentail' should be 'Potential'.
- [Table 1] The 'Notes' column classifies tasks as 'imperfect plan' or 'correct plan', but the basis for this classification is not defined in the table; please add a brief explanation or refer the reader to the appendix for the distinction.
- [Figure 4] The asterisks indicating significant post-hoc Tukey HSD comparisons are not fully explained in the caption; please specify which pairwise comparisons are shown as significant.
- [5.2.2] The sentence 'In most tasks, condition UP-UE achieved better or compatible performance as other conditions' uses 'compatible' where 'comparable' is intended; this should be corrected.
- [4.4] The phrase 'The reserved 248 participants' is unclear; please reword, for example, to 'The final sample comprised 248 participants'. Additionally, the exclusion of 4 participants identified as outliers based on plan quality is an outcome-dependent exclusion; please report the main analyses with and without these participants.
Circularity Check
No significant circularity: the trust and performance measures are independent composites of user responses and annotated outcomes; self-citations are background only.
full rationale
The paper is an empirical study, not a derivation. The calibrated-trust measures in Eqs. 1-2 combine a separately elicited binary trust judgment with independently annotated plan quality or execution accuracy; neither term is fitted from the other, and the hypotheses compare these measures across experimental conditions rather than predicting them from their own definitions. The plan-quality annotations are author-made but are not the output being predicted, and the claim that user involvement fails to calibrate trust is a statistical comparison, not a quantity that reduces by construction to its inputs. Some related-work citations are authored by the present authors (e.g., [28], [29], [79], [80]), but they are used only as background motivation for hypotheses and for task-complexity choices; the central empirical claims do not rest on any self-citation chain or imported uniqueness theorem. The post-hoc filtering described in Sections 5.2.3 and 5.2.4 removes tasks where plan quality decreased before testing execution-stage hypotheses; this is a legitimate robustness and validity concern about selective exclusion, but it is not circularity because the filter is not equivalent to the outcome measure and no fitted parameter is being renamed as a prediction. No circular step can be exhibited from the paper's equations or citations, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- standard math Standard statistical assumptions for ANOVA (normality, homoscedasticity, independence) apply to the calibrated trust and cognitive load measures.
- domain assumption The author-annotated plan quality scores faithfully reflect true plan quality.
- domain assumption The simulation environment and the one-primary-step-to-one-action mapping represent realistic LLM agent execution.
- domain assumption The binary trust question measures the construct of trust adequately for calibration.
Cite this review
Pith. "Pith review of Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant." pith.science (2026). https://pith.science/paper/LWQD5OF6
@misc{pith2026250201390,
author = {Pith},
title = {Pith review of: Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant},
year = {2026},
howpublished = {\url{https://pith.science/paper/LWQD5OF6}},
note = {Machine review of arXiv:2502.01390}
}
read the original abstract
Since the explosion in popularity of ChatGPT, large language models (LLMs) have continued to impact our everyday lives. Equipped with external tools that are designed for a specific purpose (e.g., for flight booking or an alarm clock), LLM agents exercise an increasing capability to assist humans in their daily work. Although LLM agents have shown a promising blueprint as daily assistants, there is a limited understanding of how they can provide daily assistance based on planning and sequential decision making capabilities. We draw inspiration from recent work that has highlighted the value of 'LLM-modulo' setups in conjunction with humans-in-the-loop for planning tasks. We conducted an empirical study (N = 248) of LLM agents as daily assistants in six commonly occurring tasks with different levels of risk typically associated with them (e.g., flight ticket booking and credit card payments). To ensure user agency and control over the LLM agent, we adopted LLM agents in a plan-then-execute manner, wherein the agents conducted step-wise planning and step-by-step execution in a simulation environment. We analyzed how user involvement at each stage affects their trust and collaborative team performance. Our findings demonstrate that LLM agents can be a double-edged sword -- (1) they can work well when a high-quality plan and necessary user involvement in execution are available, and (2) users can easily mistrust the LLM agents with plans that seem plausible. We synthesized key insights for using LLM agents as daily assistants to calibrate user trust and achieve better overall task outcomes. Our work has important implications for the future design of daily assistants and human-AI collaboration with LLM agents.
Figures
Reference graph
Works this paper leans on
-
[1]
Stefano V Albrecht and Peter Stone. 2018. Autonomous agents modelling other agents: A comprehensive survey and open problems. Artificial Intelligence 258 (2018), 66–95
2018
-
[2]
Nikola Banovic, Zhuoran Yang, Aditya Ramesh, and Alice Liu. 2023. Being trustworthy is not enough: How untrustworthy artificial intelligence (AI) can deceive the end-users and gain their trust. Proceedings of the ACM on Human- Computer Interaction 7, CSCW1 (2023), 1–17
2023
-
[3]
Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz. 2019. Beyond accuracy: The role of mental models in human-AI team performance. InProceedings of the AAAI Conference on Human Computation and Crowdsourcing, Vol. 7. 2–11
2019
-
[4]
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021)
arXiv 2021
-
[5]
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1 (2021), 1–21
2021
-
[6]
Chun-Wei Chiang and Ming Yin. 2022. Exploring the Effects of Machine Learn- ing Literacy Interventions on Laypeople’s Reliance on Machine Learning Models. In IUI 2022: 27th International Conference on Intelligent User Interfaces, Helsinki, An Empirical Study of User Trust and Team Performance with LLM Agents As A Daily Assistant CHI ’25, April 26-May 1, 2...
2022
-
[7]
Chun-Wei Chiang and Ming Yin. 2021. You’d better stop! Understanding human reliance on machine learning models under covariate shift. In Proceedings of the 13th ACM Web Science Conference 2021 . 120–129
2021
-
[8]
Michael Chromik, Malin Eiband, Felicitas Buchner, Adrian Krüger, and Andreas Butz. 2021. I think i get your point, AI! the illusion of explanatory depth in explainable AI. In Proceedings of the 26th International Conference on Intelligent User Interfaces. 307–317
2021
Show all 118 references
-
[9]
Lacey Colligan, Henry WW Potts, Chelsea T Finn, and Robert A Sinkin. 2015. Cognitive workload changes for nurses transitioning from a legacy system with paper documentation to a commercial electronic health record. International journal of medical informatics 84, 7 (2015), 469–476
2015
-
[10]
Jason P Davis, Kathleen M Eisenhardt, and Christopher B Bingham. 2007. De- veloping theory through simulation methods. Academy of management review 32, 2 (2007), 480–499
2007
-
[11]
Berkeley J Dietvorst, Joseph P Simmons, and Cade Massey. 2015. Algorithm aversion: people erroneously avoid algorithms after seeing them err. Journal of experimental psychology: General 144, 1 (2015), 114
2015
-
[12]
Berkeley J Dietvorst, Joseph P Simmons, and Cade Massey. 2018. Overcoming algorithm aversion: People will use imperfect algorithms if they can (even slightly) modify them. Management science 64, 3 (2018), 1155–1170
2018
-
[13]
Shi Dong, Ping Wang, and Khushnood Abbas. 2021. A survey on deep learning and its applications. Computer Science Review 40 (2021), 100379
2021
-
[14]
Finale Doshi-Velez and Been Kim. 2017. Towards a rigorous science of inter- pretable machine learning. arXiv preprint arXiv:1702.08608 (2017)
2017 arXiv
-
[15]
Mary T Dzindolet, Scott A Peterson, Regina A Pomranky, Linda G Pierce, and Hall P Beck. 2003. The role of trust in automation reliance. International journal of human-computer studies 58, 6 (2003), 697–718
2003
-
[16]
Sven Eckhardt, Niklas Kühl, Mateusz Dolata, and Gerhard Schwabe. 2024. A Survey of AI Reliance. arXiv preprint arXiv:2408.03948 (2024)
2024 arXiv
-
[17]
Alexander Erlei, Abhinav Sharma, and Ujwal Gadiraju. 2024. Understanding Choice Independence and Error Types in Human-AI Collaboration. In Proceed- ings of the CHI Conference on Human Factors in Computing Systems . 1–19
2024
-
[18]
Shaoyang Fan, Ujwal Gadiraju, Alessandro Checco, and Gianluca Demartini
-
[19]
Franz Faul, Edgar Erdfelder, Axel Buchner, and Albert-Georg Lang. 2009. Statis- tical power analyses using G* Power 3.1: Tests for correlation and regression analyses. Behavior research methods 41, 4 (2009), 1149–1160
2009
-
[20]
Tharindu Fernando, Harshala Gammulle, Simon Denman, Sridha Sridharan, and Clinton Fookes. 2021. Deep learning for medical anomaly detection–a survey. ACM Computing Surveys (CSUR) 54, 7 (2021), 1–37
2021
-
[21]
Riccardo Fogliato, Alexandra Chouldechova, and Zachary Lipton. 2021. The impact of algorithmic risk assessments on human predictions and its analysis via crowdsourcing studies. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–24
2021
-
[22]
Andreas Fügener, Jörn Grahl, Alok Gupta, and Wolfgang Ketter. 2022. Cognitive challenges in human–artificial intelligence collaboration: Investigating the path toward productive delegation. Information Systems Research 33, 2 (2022), 678– 696
2022
-
[23]
Ujwal Gadiraju, Jie Yang, and Alessandro Bozzon. 2017. Clarity is a worthwhile quality: On the role of task clarity in microtask crowdsourcing. In Proceedings of the 28th ACM conference on hypertext and social media . 5–14
2017
-
[24]
Florian Geissler, Karsten Roscher, and Mario Trapp. 2024. Concept-Guided LLM Agents for Human-AI Safety Codesign. In Proceedings of the AAAI Symposium Series, Vol. 3. 100–104
2024
-
[25]
Zoubin Ghahramani. 2015. Probabilistic machine learning and artificial intelli- gence. Nature 521, 7553 (2015), 452–459
2015
-
[26]
Ben Green and Yiling Chen. 2021. Algorithmic risk assessments can alter human decision-making processes in high-stakes government contexts. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–33
2021
-
[27]
Gaole He, Nilay Aishwarya, and Ujwal Gadiraju. 2025. Is Conversational XAI All You Need? Human-AI Decision Making With a Conversational XAI Assistant. In Proceedings of the 30th International Conference on Intelligent User Interfaces
2025
-
[28]
Gaole He, Abri Bharos, and Ujwal Gadiraju. 2024. To Err Is AI! Debugging as an Intervention to Facilitate Appropriate Reliance on AI Systems. In Proceedings of the 35th ACM Conference on Hypertext and Social Media . 98–105
2024
-
[29]
Gaole He, Lucie Kuiper, and Ujwal Gadiraju. 2023. Knowing About Knowing: An Illusion of Human Competence Can Hinder Appropriate Reliance on AI Systems. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–18
2023
-
[30]
Zeyu He, Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Shaurya Rohatgi, and Ting-Hao Kenneth Huang. 2024. If in a Crowdsourced Data Annotation Pipeline, a GPT-4. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–25
2024
-
[31]
Patrick Hemmer, Sebastian Schellhammer, Michael Vössing, Johannes Jakubik, and Gerhard Satzger. 2022. Forming Effective Human-AI Teams: Building Machine Learning Models that Complement the Capabilities of Multiple Experts. In Proceedings of the Thirty-First International Joint...
2022 doi
-
[32]
Patrick Hemmer, Max Schemmer, Niklas Kühl, Michael Vössing, and Gerhard Satzger. 2024. Complementarity in Human-AI Collaboration: Concept, Sources, and Evidence. arXiv preprint arXiv:2404.00029 (2024)
2024 arXiv
-
[33]
Patrick Hemmer, Max Schemmer, Michael Vössing, and Niklas Kühl. 2021. Human-AI Complementarity in Hybrid Intelligence Systems: A Structured Lit- erature Review. PACIS (2021), 78
2021
-
[34]
Patrick Hemmer, Monika Westphal, Max Schemmer, Sebastian Vetter, Michael Vössing, and Gerhard Satzger. 2023. Human-AI collaboration: the effect of AI delegation on human task performance and task satisfaction. In Proceedings of the 28th International Conference on Intelligent ...
2023
-
[35]
Yoyo Tsung-Yu Hou and Malte F Jung. 2021. Who is the expert? Reconciling algorithm aversion and algorithm appreciation in AI-supported decision making. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–25
2021
-
[36]
Shijue Huang, Wanjun Zhong, Jianqiao Lu, Qi Zhu, Jiahui Gao, Weiwen Liu, Yutai Hou, Xingshan Zeng, Yasheng Wang, Lifeng Shang, Xin Jiang, Ruifeng Xu, and Qun Liu. 2024. Planning, Creation, Usage: Benchmarking LLMs for Comprehen- sive Tool Utilization in Real-World Complex Scen...
2024 arXiv
-
[37]
Alon Jacovi, Ana Marasović, Tim Miller, and Yoav Goldberg. 2021. Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. 624–635
2021
-
[38]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. Comput. Surveys 55, 12 (2023), 1–38
2023
-
[39]
Robin Jia and Percy Liang. 2017. Adversarial Examples for Evaluating Reading Comprehension Systems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . 2021–2031
2017
-
[40]
Jialun Aaron Jiang, Kandrea Wade, Casey Fiesler, and Jed R Brubaker. 2021. Sup- porting serendipity: Opportunities and challenges for Human-AI Collaboration in qualitative analysis. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1 (2021), 1–23
2021
-
[41]
Patricia K Kahr, Gerrit Rooks, Martijn C Willemsen, and Chris CP Snijders. 2024. Understanding Trust and Reliance Development in AI Advice: Assessing Model Accuracy, Model Explanations, and Experiences from Previous Interactions. ACM Transactions on Interactive Intelligent Sys...
2024
-
[42]
Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Kaya Stechly, Mudit Verma, Siddhant Bhambri, Lucas Saldyt, and Anil Murthy. 2024. LLMs Can’t Plan, But Can Help Planning in LLM-Modulo Frameworks. arXiv preprint arXiv:2402.01817 (2024)
2024 arXiv
-
[43]
Davinder Kaur, Suleyman Uslu, Kaley J Rittichier, and Arjan Durresi. 2022. Trustworthy artificial intelligence: a review. ACM computing surveys (CSUR) 55, 2 (2022), 1–38
2022
-
[44]
Veton Kepuska and Gamal Bohouta. 2018. Next-generation of virtual personal assistants (microsoft cortana, apple siri, amazon alexa and google home). In 2018 IEEE 8th annual computing and communication workshop and conference (CCWC). IEEE, 99–103
2018
-
[45]
I’m Not Sure, But
Sunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Wortman Vaughan. 2024. "I’m Not Sure, But... ": Examining the Impact of Large Language Models’ Uncertainty Expression on User Reliance and Trust. In Proceedings of the 2024 ACM Conference on Fa...
2024
-
[46]
Moritz Körber. 2019. Theoretical considerations and development of a ques- tionnaire to measure trust in automation. In Proceedings of the 20th Congress of the International Ergonomics Association (IEA 2018) Volume VI: Transport Er- gonomics and Human Factors (TEHF), Aerospace...
2019
-
[47]
Vivian Lai, Samuel Carton, Rajat Bhatnagar, Q Vera Liao, Yunfeng Zhang, and Chenhao Tan. 2022. Human-ai collaboration via conditional delegation: A case study of content moderation. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–18
2022
-
[48]
Vera Liao, and Chenhao Tan
Vivian Lai, Chacha Chen, Alison Smith-Renner, Q. Vera Liao, and Chenhao Tan. 2023. Towards a Science of Human-AI Decision Making: An Overview of Design Space in Empirical Human-Subject Studies. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transpar...
2023
-
[49]
Why is’ Chicago’deceptive?
Vivian Lai, Han Liu, and Chenhao Tan. 2020. " Why is’ Chicago’deceptive?" Towards Building Model-Driven Tutorials for Humans. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13
2020
-
[50]
John Lee and Neville Moray. 1992. Trust, control strategies and allocation of function in human-machine systems. Ergonomics 35, 10 (1992), 1243–1270. CHI ’25, April 26-May 1, 2025, Yokohama, Japan Gaole He, Gianluca Demartini, and Ujwal Gadiraju
1992
-
[51]
John D Lee and Katrina A See. 2004. Trust in automation: Designing for appro- priate reliance. Human factors 46, 1 (2004), 50–80
2004
-
[52]
Zhuoyan Li and Ming Yin. [n. d.]. Utilizing Human Behavior Modeling to Manipulate Explanations in AI-Assisted Decision Making: The Good, the Bad, and the Scary. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
-
[53]
Q Vera Liao and S Shyam Sundar. 2022. Designing for responsible trust in AI systems: A communication perspective. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency . 1257–1268
2022
-
[54]
Q Vera Liao and Jennifer Wortman Vaughan. 2023. AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap.arXiv preprint arXiv:2306.01941 (2023)
2023 arXiv
-
[55]
Gabriel Lima, Nina Grgić-Hlača, and Meeyoung Cha. 2021. Human perceptions on moral responsibility of AI: A case study in AI-assisted bail decision-making. In Proceedings of the 2021 CHI conference on human factors in computing systems . 1–17
2021
-
[56]
Jessy Lin, Nicholas Tomlin, Jacob Andreas, and Jason Eisner. 2024. Decision- oriented dialogue for human-ai collaboration. Transactions of the Association for Computational Linguistics 12 (2024), 892–911
2024
-
[57]
Han Liu, Vivian Lai, and Chenhao Tan. 2021. Understanding the effect of out- of-distribution examples and interactive explanations on human-ai decision making. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–45
2021
-
[58]
Jennifer M Logg, Julia A Minson, and Don A Moore. 2019. Algorithm apprecia- tion: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes 151 (2019), 90–103
2019
-
[59]
Zhuoran Lu, Dakuo Wang, and Ming Yin. 2024. Does more advice help? the effects of second opinions in AI-assisted decision making. Proceedings of the ACM on Human-Computer Interaction 8, CSCW1 (2024), 1–31
2024
-
[60]
Zhuoran Lu and Ming Yin. 2021. Human Reliance on Machine Learning Models When Performance Feedback is Limited: Heuristics and Risks. In CHI ’21: CHI Conference on Human Factors in Computing Systems, Virtual Event / Yokohama, Japan, May 8-13, 2021 , Yoshifumi Kitamura, Aaron Qu...
2021
-
[61]
Brian Lubars and Chenhao Tan. 2019. Ask not what AI can do, but what AI should do: Towards a framework of task delegability. Advances in neural information processing systems 32 (2019)
2019
-
[62]
Shuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng, Zhenhui Peng, Ming Yin, and Xiaojuan Ma. 2024. Towards human-ai deliberation: Design and evaluation of llm-empowered deliberative ai for ai-assisted decision-making. arXiv preprint arXiv:2403.16812 (2024)
2024 arXiv
-
[63]
Shuai Ma, Ying Lei, Xinru Wang, Chengbo Zheng, Chuhan Shi, Ming Yin, and Xiaojuan Ma. 2023. Who should i trust: Ai or myself? leveraging human and ai correctness likelihood to promote appropriate trust in ai-assisted decision- making. In Proceedings of the 2023 CHI Conference ...
2023
-
[64]
Are You Really Sure?
Shuai Ma, Xinru Wang, Ying Lei, Chuhan Shi, Ming Yin, and Xiaojuan Ma. 2024. “Are You Really Sure?” Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making. InProceedings of the CHI Conference on Human Factors in Computing Systems . 1–20
2024
-
[65]
David Madras, Toni Pitassi, and Richard Zemel. 2018. Predict responsibly: improving fairness and accuracy by learning to defer. Advances in neural information processing systems 31 (2018)
2018
-
[66]
Hasan Mahmud, AKM Najmul Islam, Syed Ishtiaque Ahmed, and Kari Smolander
-
[67]
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, F...
2019
-
[68]
Siddharth Mehrotra, Chadha Degachi, Oleksandra Vereschak, Catholijn M Jonker, and Myrthe L Tielman. 2024. A Systematic Review on Fostering Appro- priate Trust in Human-AI Interaction: Trends, Opportunities and Challenges. ACM Journal on Responsible Computing 1, 4 (2024), 1–45
2024
-
[69]
George A Miller. 1956. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological review 63, 2 (1956), 81
1956
-
[70]
Harikrishna Narasimhan, Wittawat Jitkrittum, Aditya K Menon, Ankit Rawat, and Sanjiv Kumar. 2022. Post-hoc estimators for learning to defer to an expert. Advances in Neural Information Processing Systems 35 (2022), 29292–29304
2022
-
[71]
Judith S Olson and Wendy A Kellogg. 2014. Ways of Knowing in HCI . Vol. 2. Springer
2014
-
[72]
Behrooz Omidvar Tehrani and Anmol Anubhai. 2024. Evaluating Human-AI Partnership for LLM-based Code Migration. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–8
2024
-
[73]
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology . 1–22
2023
-
[74]
Niccolò Pescetelli and Nicholas Yeung. 2021. The role of decision confidence in advice-taking and trust formation. Journal of Experimental Psychology: General 150, 3 (2021), 507
2021
-
[75]
Marc Pinski, Martin Adam, and Alexander Benlian. 2023. AI knowledge: Im- proving AI delegation through human enablement. In Proceedings of the 2023 CHI conference on human factors in computing systems . 1–17
2023
-
[76]
Samira Pouyanfar, Saad Sadiq, Yilin Yan, Haiman Tian, Yudong Tao, Maria Presa Reyes, Mei-Ling Shyu, Shu-Ching Chen, and Sundaraja S Iyengar. 2018. A survey on deep learning: Algorithms, techniques, and applications. ACM computing surveys (CSUR) 51, 5 (2018), 1–36
2018
-
[77]
Charvi Rastogi, Yunfeng Zhang, Dennis Wei, Kush R Varshney, Amit Dhurand- har, and Richard Tomsett. 2022. Deciding fast and slow: The role of cognitive bi- ases in ai-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction 6, CSCW1 (2022), 1–22
2022
-
[78]
Amy Rechkemmer and Ming Yin. 2022. When confidence meets accuracy: Exploring the effects of multiple performance indicators on trust in machine learning models. In Proceedings of the 2022 chi conference on human factors in computing systems. 1–14
2022
-
[79]
Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2023. A Missing Piece in the Puzzle: Considering the Role of Task Complexity in Human-AI Decision Making. In Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization. 215–227
2023
-
[80]
Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2024. Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision Making. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–17
2024
-
[81]
Max Schemmer, Patrick Hemmer, Niklas Kühl, Carina Benz, and Gerhard Satzger
-
[82]
Max Schemmer, Niklas Kuehl, Carina Benz, Andrea Bartos, and Gerhard Satzger
-
[83]
Ashish Sharma, Inna W Lin, Adam S Miner, David C Atkins, and Tim Althoff
-
[84]
InACM Conference on Human Factors in Computing Systems (CHI’22), Workshop on Trust and Reliance in AI-Human Teams (trAIt)
Should I Follow AI-based Advice? Measuring Appropriate Reliance in Human-AI Decision-Making. InACM Conference on Human Factors in Computing Systems (CHI’22), Workshop on Trust and Reliance in AI-Human Teams (trAIt)
-
[85]
Chenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, and Jordan Boyd-Graber. 2024. Large Language Models Help Humans Verify Truthfulness–Except When They Are Convincingly Wrong. In Proceedings of the 2024 Conference of the North American Chapter of the...
2024
-
[86]
Dylan Slack, Satyapriya Krishna, Himabindu Lakkaraju, and Sameer Singh
-
[87]
Suzanne Tolmeijer, Markus Christen, Serhiy Kandul, Markus Kneer, and Abra- ham Bernstein. 2022. Capable but amoral? Comparing AI and human expert collaboration in ethical decision making. In Proceedings of the 2022 CHI Confer- ence on Human Factors in Computing Systems . 1–17
2022
-
[88]
Nature Machine Intelligence 5, 1 (2023), 46–57
Human–AI collaboration enables more empathic conversations in text- based peer-to-peer mental health support. Nature Machine Intelligence 5, 1 (2023), 46–57
2023
-
[89]
Hua Shen, Chieh-Yang Huang, Tongshuang Wu, and Ting-Hao Kenneth Huang
-
[90]
In Companion Publication of the 2023 Conference on Computer Supported Cooperative Work and Social Computing
ConvXAI: Delivering heterogeneous AI explanations via conversations to support human-AI scientific writing. In Companion Publication of the 2023 Conference on Computer Supported Cooperative Work and Social Computing . 384–387
2023
-
[91]
Karthik Valmeekam, Matthew Marquez, Sarath Sreedharan, and Subbarao Kamb- hampati. 2023. On the planning abilities of large language models-a critical investigation. Advances in Neural Information Processing Systems 36 (2023), 75993–76005
2023
-
[92]
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S Bernstein, and Ranjay Krishna. 2023. Explanations can An Empirical Study of User Trust and Team Performance with LLM Agents As A Daily Assistant CHI ’25, April 26-May 1, 2025, Yokoham...
2023
-
[93]
Nature Machine Intelligence 5, 8 (2023), 873– 883
Explaining machine learning models with interactive natural language conversations using TalkToModel. Nature Machine Intelligence 5, 8 (2023), 873– 883
2023
-
[94]
Michael Vössing, Niklas Kühl, Matteo Lind, and Gerhard Satzger. 2022. Design- ing transparency for effective human-AI collaboration. Information Systems Frontiers 24, 3 (2022), 877–895
2022
-
[95]
Suzanne Tolmeijer, Ujwal Gadiraju, Ramya Ghantasala, Akshit Gupta, and Abra- ham Bernstein. 2021. Second chance for a first impression? Trust development in intelligent system interaction. In Proceedings of the 29th ACM Conference on user modeling, adaptation and personalizati...
2021
-
[96]
Richard Tomsett, Alun Preece, Dave Braines, Federico Cerutti, Supriyo Chakraborty, Mani Srivastava, Gavin Pearson, and Lance Kaplan. 2020. Rapid trust calibration through interpretable and uncertainty-aware AI. Patterns 1, 4 (2020), 100049
2020
-
[97]
Amos Tversky and Daniel Kahneman. 1991. Loss aversion in riskless choice: A reference-dependent model. The quarterly journal of economics 106, 4 (1991), 1039–1061
1991
-
[98]
Xinru Wang and Ming Yin. 2021. Are Explanations Helpful? A Comparative Study of the Effects of Explanations in AI-Assisted Decision-Making. In 26th International Conference on Intelligent User Interfaces . 318–328
2021
-
[99]
Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian Ma, Yitao Liang, and Team CraftJarvis. 2023. Describe, explain, plan and select: interactive planning with large language models enables open-world multi-task agents. In Proceedings of the 37th International Conference...
2023
-
[100]
Oleksandra Vereschak, Gilles Bailly, and Baptiste Caramiaux. 2021. How to eval- uate trust in AI-assisted decision making? A survey of empirical methodologies. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–39
2021
-
[101]
Robert E Wood. 1986. Task complexity: Definition of the construct. Organiza- tional behavior and human decision processes 37, 1 (1986), 60–82
1986
-
[102]
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science 18, 6 (2024), 186345
2024
-
[103]
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023. Plan-and-Solve Prompting: Improving Zero-Shot Chain- of-Thought Reasoning by Large Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Li...
2023
-
[104]
Xinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra, and Zhengjie Miao
-
[105]
Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. 2019. Understanding the effect of accuracy on trust in machine learning models. In Proceedings of the 2019 chi conference on human factors in computing systems . 1–12
2019
-
[106]
Vera Liao, and Rachel K
Yunfeng Zhang, Q. Vera Liao, and Rachel K. E. Bellamy. 2020. Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. In FAT* ’20: Conference on Fairness, Accountability, and Transparency, Barcelona, Spain, January 27-30, 2020 , Mi...
2020
-
[107]
It’s a Fair Game
Zhiping Zhang, Michelle Jia, Hao-Ping Lee, Bingsheng Yao, Sauvik Das, Ada Lerner, Dakuo Wang, and Tianshi Li. 2024. “It’s a Fair Game”, or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents. In Proceedings of the CHI Co...
2024
-
[108]
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al . 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682 (2022)
2022 arXiv
-
[109]
Qingxiao Zheng, Zhongwei Xu, Abhinav Choudhary, Yuting Chen, Yongming Li, and Yun Huang. 2023. Synergizing human-AI agency: a guide of 23 heuristics for service co-creation with LLM-based agents. arXiv preprint arXiv:2310.15065 (2023)
2023 arXiv
-
[110]
Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. Ai chains: Transpar- ent and controllable human-ai interaction by chaining large language model prompts. In Proceedings of the 2022 CHI conference on human factors in computing systems. 1–22
2022
-
[111]
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. 2023. The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864 (2023)
2023 arXiv
-
[112]
Yanming Yang, Xin Xia, David Lo, and John Grundy. 2022. A survey on deep learning for software engineering. ACM Computing Surveys (CSUR) 54, 10s (2022), 1–73
2022
-
[116]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023)
2023 arXiv
-
[118]
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. 2024. Toolqa: A dataset for llm question answering with external tools. Advances in Neural Information Processing Systems 36 (2024). CHI ’25, April 26-May 1, 2025, Yokohama, Japan Gaole He, Gianluca Demartini, and ...
2024
-
[2020]
Proceedings of the ACM on Human-Computer Interaction 4, CSCW2 (2020), 1–24
Crowdco-op: Sharing risks and rewards in crowdsourcing. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2 (2020), 1–24
2020
-
[2022]
Technological Forecasting and Social Change 175 (2022), 121390
What influences algorithmic decision-making? A systematic literature review on algorithm aversion. Technological Forecasting and Social Change 175 (2022), 121390
2022
-
[2023]
In Proceedings of the 28th International Conference on Intelligent User Interfaces
Appropriate reliance on AI advice: Conceptualization and the effect of explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces. 410–422
-
[2024]
In Proceedings of the CHI Conference on Human Factors in Computing Systems
Human-LLM collaborative annotation through effective verification of LLM labels. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–21
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.