Pith. sign in

REVIEW 3 major objections 80 references

Rule-guided mismatch cues and previews of edit impact help people catch AI failures and refine models more safely in clinical assessment.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 22:32 UTC pith:RZCKP43S

load-bearing objection Solid HCI systems paper: mismatch cues cleanly lift team accuracy and reliance; Phase-2 local gains are real under their hybrid protocol but rest on a frozen-NN proxy, so treat the adaptation claim carefully. the 3 major comments →

arxiv 2606.00011 v1 pith:RZCKP43S submitted 2026-04-12 cs.HC cs.AIcs.LG

RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview

classification cs.HC cs.AIcs.LG
keywords human-AI collaborationmodel editingfailure detectionrule-based feedbackprospective previewstroke rehabilitationcalibrated reliancelocal-global tradeoff
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that people working with AI still lack practical tools to spot when the model is likely wrong on a given case and to see what a proposed fix would do before they commit it. The authors present RuleEdit, a system that flags disagreements between a neural network’s prediction and simple clinical rules, and that lets users rewrite those rules while previewing projected accuracy change and how cases would move in embedding space. In a stroke rehabilitation assessment study with therapists and students, the mismatch cues raised team accuracy by about 14 percentage points, increased rejection of wrong AI advice, and cut harmful decision flips. When users could also inspect embedding-shift previews, the quality of their local rule edits rose sharply, with local performance gains rising from roughly 12% to 36%. The same work shows a clear boundary: edits that help one patient can hurt when they are pooled into a global update, so failure-aware editing needs both local control and safeguards against unsafe transfer.

Core claim

Mismatch signals from compact clinical rule tables, shown at decision time, calibrate human reliance on AI and improve Human+AI team accuracy in stroke rehabilitation assessment; prospective previews of performance and embedding shifts then improve the quality of user-authored local rule edits, while exposing that locally helpful edits can degrade performance when transferred globally.

What carries the argument

RuleEdit: an interactive pipeline that surfaces case-level mismatches between a frozen neural network and Top-3 rule-table verdicts, then lets users author rule or label feedback while previewing projected Δ performance and supervised-UMAP embedding shifts (neighborhoods and class centroids) before commit.

Load-bearing premise

The offline hybrid setup—frozen network plus patient-specific rules or a kNN proxy on supervised embeddings—faithfully reflects the quality and safety of edits that would appear in live, ongoing clinical use.

What would settle it

Deploy the same rule-edit and preview interface in a live multi-session rehabilitation clinic and test whether decision-time mismatch cues still raise team accuracy and reduce ChangedToWrong rates, and whether embedding previews still produce large local gains without global regressions when therapists’ edits accumulate over real patients.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Decision interfaces that only show explanations or confidence are incomplete; adding interpretable consistency checks can raise team accuracy and cut over- and under-reliance.
  • Model-editing tools should show pre-commit previews of both accuracy and representation change, not only post-hoc results after an update is applied.
  • Local personalization of rules can deliver large gains for a single patient while the same rules, when pooled globally, can reduce overall performance—so transfer must be gated.
  • Safety metrics such as AI-harm rate, Regret_best, and ChangedToWrong become practical design targets alongside raw accuracy in high-stakes human-AI work.
  • Failure-aware systems can move users from passive interpretation of AI outputs to active inspection and controllable refinement of model behavior.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Similar mismatch-plus-preview patterns could transfer to other structured clinical scores (imaging triage, risk scores) where domain rules already exist and local personalization is valued.
  • The local-global tradeoff implies that production systems may need per-patient or per-task rule layers with explicit conflict detection rather than single global fine-tunes of user feedback.
  • Embedding-shift previews may act as a teaching tool: they force non-ML users to reason about how feedback reshapes neighborhoods, which could reduce trial-and-error editing over time.
  • If mismatch AUPRC stays higher than confidence or ensemble disagreement across domains, rule tables become a general complementary failure signal rather than a domain-specific trick.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper presents RuleEdit, an interactive system for failure-aware human–AI model editing in stroke rehabilitation assessment. It combines (i) rule-table mismatch cues that flag inconsistencies between a neural classifier and Top-3 interpretable rules, and (ii) user-authored label/rule feedback with pre-commit previews of projected performance and supervised-UMAP embedding shifts (class centroids and neighborhoods) while keeping the base NN frozen. In a two-phase within-subjects study (n=21 experts and novices), Phase 1 finds that mismatch cues raise Human–AI team accuracy by 14.16% (p<0.001), improve reject-on-wrong and accept-on-correct, and reduce AI-harm and ChangedToWrong rates; an AUPRC comparison (0.64 vs 0.46–0.52 for uncertainty/ensemble baselines) supports the cue’s discriminative value. Phase 2 finds that adding embedding-shift previews raises post-update local ΔPerf from 11.50% to 36.38% (p<0.001) under offline hybrid/kNN evaluation, while pooled global updates can regress—highlighting a local–global tradeoff.

Significance. If the results hold under tighter measurement of adaptation, the work is a solid HCI contribution: it moves human–AI decision support from post-hoc explanation toward decision-time failure inspection and anticipatory, user-authored editing with pre-commit impact preview. Phase 1 is particularly valuable—interpretable consistency checks as lightweight guardrails for calibrated reliance, with safety-oriented metrics (AI-harm, ChangedToWrong, Regret_best) and subgroup trends for experts and novices. The honest local–global tradeoff finding is useful for design of controllable editing systems. Strengths include a counterbalanced within-subjects protocol, normality-checked paired tests, LOSO model evaluation, and an explicit AUPRC comparison against alternative failure signals. The main significance risk is that Phase 2’s headline local gains are measured under a frozen-NN hybrid and kNN-on-UMAP proxy rather than live model updates.

major comments (3)
  1. Phase 2’s central claim (embedding previews raise local ΔPerf from 11.50% to 36.38%, p<0.001; Abstract; Table 4; §4.2) is evaluated under the offline protocol of §3.2.6 and §3.3.5/A.5: the NN stays frozen; rule feedback is a patient-specific inference-time layer (probability aggregation with Top-3 RuleTree rules); label feedback is scored via lightweight kNN on supervised UMAP of penultimate embeddings (and separately via offline fine-tuning for re-labels). Table 4 therefore reports proxy ΔPerf, not performance of a model that would be updated and re-deployed in clinic. The causal claim that previews “improved participants’ feedback for model adaptation” is only established inside this proxy world. Please either (a) validate that the same user edits yield comparable local gains under true fine-tuning / online hybrid inference, or (b) reframe Abstract/§1/§6 claims as “edit quality under t
  2. Abstract and §4.2 attribute the 11.50→36.38 local lift specifically to “users’ rule-based feedback,” but Phase 2 also includes re-label feedback and offline fine-tuning (§3.3.5, A.5). Table 4 reports only aggregate “overall” local/global ΔPerf. Without a feedback-type breakdown (rules-only vs labels-only vs combined) and without stating which evaluation path produced the headline numbers, it is hard to attribute the preview effect to rule editing as claimed. Please disaggregate Table 4 (and Appendix Table 11) by feedback mechanism and clarify which path underlies the abstract claim.
  3. Phase 1 case curation fixed 10 AI-correct and 4 AI-incorrect trials per condition to match model accuracy (§3.3.3). That is reasonable for power on failure detection, but Accept-on-wrong / Reject-on-wrong rates and AI-harm then depend on an artificial error base rate and on which failures were selected. Please report sensitivity of reliance metrics to case mix (e.g., leave-one-error-case-out or natural LOSO error incidence), or qualify that calibrated-reliance gains are conditional on this curated error distribution.

Circularity Check

0 steps flagged

No circularity: claims are empirical user-study outcomes, not derivations that reduce to fitted inputs or self-defined quantities.

full rationale

RuleEdit’s load-bearing claims (Phase-1 Human+AI accuracy +14.16%, reliance calibration, ChangedToWrong reduction; Phase-2 local ΔPerf 11.50%→36.38% with embedding preview) are measured outcomes of a within-subjects user study and offline re-evaluation on held-out data, not first-principles predictions. Rule thresholds and Top-3 RuleTree selection are standard LOSO design choices validated on held-out folds; hybrid inference and kNN/UMAP previews are fixed evaluation protocols, not quantities defined as the performance deltas they report. Self-citations (e.g. prior rehab datasets and feature work) supply domain setup and are not uniqueness theorems or load-bearing premises that force the reported effect sizes. Measurement-proxy concerns (frozen NN + patient-specific rule layer vs live fine-tuning) are validity issues, not circular reductions of claim to input. No self-definitional loop, fitted-input-as-prediction, or ansatz-via-citation chain is present.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

Empirical HCI/ML systems paper. Free parameters are design and hyper-parameter choices (Top-K, network sizes, learning rates, UMAP settings). Axioms are standard domain and methodological assumptions of clinical assessment and leave-one-subject-out evaluation. Invented entities are the system components themselves rather than new physical or mathematical objects.

free parameters (4)
  • Top-K rules (K=3) = 3
    Number of rules retained for mismatch cues and hybrid inference; selected by validation performance (Appendix Table 5).
  • NN architecture and learning rates = ROM 3×64 / 1e-2; COMP 32-32-64 / 5e-3
    ROM: 3×64 units, lr=1e-2; COMP: 32-32-64, lr=5e-3; chosen by grid search under LOSO.
  • Rule thresholding strategy = RuleTree
    RuleTree (max_depth=2) chosen over Median/Avg/Percentile after LOSO comparison.
  • Hybrid aggregation method = probability aggregation
    Probability averaging of NN and rule soft outputs used in the user study.
axioms (4)
  • domain assumption Kinematic features extracted from Kinect skeletons (joint angles, distances, compensatory displacements) are sufficient to assess ROM and compensation quality.
    Inherited from prior rehab-assessment work [46,47] and used throughout model and rule construction (Section 3.2.1–3.2.2).
  • domain assumption Leave-one-subject-out cross-validation yields a realistic estimate of generalization to new patients.
    Standard for subject-dependent clinical data; used for both NN and rule baselines (Section 3.2).
  • ad hoc to paper A compact Top-3 rule set can serve as an interpretable mismatch signal complementary to model uncertainty.
    Justified by AUPRC comparison (0.64 vs 0.46–0.52) but remains a design choice specific to this system (Section 4.1).
  • ad hoc to paper Supervised UMAP of penultimate-layer embeddings plus class-centroid medians provides a faithful preview of how rule/label feedback would reorganize the representation.
    Core of the prospective-preview mechanism (Section 3.2.6); not independently validated against actual retraining trajectories.
invented entities (2)
  • Rule-guided mismatch cue (rule-table disagreement signal) no independent evidence
    purpose: Surface case-level contradictions between NN prediction and Top-3 rule verdicts to support accept/override decisions.
    System component introduced and evaluated in Phase 1; independent evidence limited to the reported AUPRC and user-study deltas.
  • Prospective embedding-shift preview (supervised UMAP + centroids under frozen NN) no independent evidence
    purpose: Let users inspect neighborhood and class-centroid movement before committing rule or label feedback.
    Central interaction mechanism of Phase 2; its causal contribution is inferred from the Condition A vs B comparison, not from external validation.

pith-pipeline@v1.1.0-grok45 · 28897 in / 3265 out tokens · 42099 ms · 2026-07-12T22:32:24.685993+00:00 · methodology

0 comments
read the original abstract

Despite the promise of AI to assist complex decisions, practitioners still lack ways to detect likely failures and inspect the consequences of model edits before committing them. We present RuleEdit, an interactive, rule-guided human-AI model editing system that (i) surfaces likely failures through interpretable mismatch signals from rule tables and (ii) supports user-authored rule feedback with prospective previews of projected performance changes and embedding shifts. We instantiate RuleEdit in stroke rehabilitation assessment and evaluate it with health professionals and students. Rule-guided failure detection significantly increased Human + AI performance by 14.16\% ($p<0.001$) while improving rejection of incorrect AI and reducing both over- and under- reliance as well as ChangedToWrong decisions. In addition, presenting prospective embedding previews improved participants' feedback for model adaptation, increasing post-update local performance gains from 11.50\% to 36.38\% after incorporating users' rule-based feedback ($p<0.001$). Our findings show that mismatch-based failure cues and prospective impact previews can support failure-aware human-AI model editing, while also revealing a local-global tradeoff: edits that help a specific case can degrade performance when transferred globally. We discuss implications of designing failure-aware and controllable human-AI systems.

Figures

Figures reproduced from arXiv: 2606.00011 by Justin Yu Feng Teo, Min Hun Lee.

Figure 1
Figure 1. Figure 1: RuleEdit supports failure inspection, user-authored editing, and prospective preview in AI-assisted decision-making. Given [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The system interface includes the video of a post-stroke survivor and the corresponding AI output (prediction, confidence [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Two-phase within-subject study. In Phase 1 (AI-assisted decisions), participants complete tasks under Condition A (baseline) [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: User Interface of Phase 1 [PITH_FULL_IMAGE:figures/full_fig_p022_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

80 extracted references · 6 linked inside Pith

  1. [1]

    Ashraf Abdul, Jo Vermeulen, Danding Wang, Brian Y Lim, and Mohan Kankanhalli. 2018. Trends and trajectories for explainable, accountable and intelligible systems: An hci research agenda. InProceedings of the 2018 CHI conference on human factors in computing systems. 1–18

  2. [2]

    Saleema Amershi, Maya Cakmak, William Bradley Knox, and Todd Kulesza. 2014. Power to the people: The role of humans in interactive machine learning.AI magazine35, 4 (2014), 105–120

  3. [3]

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. InProceedings of the 2019 chi conference on human factors in computing systems. 1–13

  4. [4]

    Vijay Arya, Rachel KE Bellamy, Pin-Yu Chen, Amit Dhurandhar, Michael Hind, Samuel C Hoffman, Stephanie Houde, Q Vera Liao, Ronny Luss, Aleksandra Mojsilović, et al . 2019. One explanation does not fit all: A toolkit and taxonomy of ai explainability techniques.arXiv preprint arXiv:1909.03012(2019)

  5. [5]

    Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. InProceedings of the 2021 CHI conference on human factors in computing systems. 1–16

  6. [6]

    Etienne Becht, Leland McInnes, John Healy, Charles-Antoine Dutertre, Immanuel WH Kwok, Lai Guan Ng, Florent Ginhoux, and Evan W Newell

  7. [7]

    Dimensionality reduction for visualizing single-cell data using UMAP.Nature biotechnology37, 1 (2019), 38–44

  8. [8]

    Emma Beede, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M Vardoulakis. 2020. A human- centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. InProceedings of the 2020 CHI conference on human factors in computing systems. 1–12

  9. [9]

    Yoshua Bengio, Stephen Clare, Carina Prunkl, Maksym Andriushchenko, Ben Bucknall, Philip Fox, Nestor Maslej, Conor McGlynn, Malcolm Murray, Shalaleh Rismani, et al. 2025. International ai safety report 2025: Second key update: Technical safeguards and risk management.arXiv preprint arXiv:2511.19863(2025)

  10. [10]

    Angie Boggust, Brandon Carter, and Arvind Satyanarayan. 2022. Embedding comparator: Visualizing differences in global structure and local neighborhoods via small multiples. In27th international conference on intelligent user interfaces. 746–766

  11. [11]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making.Proceedings of the ACM on Human-computer Interaction5, CSCW1 (2021), 1–21

  12. [12]

    Adrian Bussone, Simone Stumpf, and Dympna O’Sullivan. 2015. The role of explanations on trust and reliance in clinical decision support systems. In2015 international conference on healthcare informatics. IEEE, 160–169

  13. [13]

    Ángel Alexander Cabrera, Abraham J Druck, Jason I Hong, and Adam Perer. 2021. Discovering and validating ai errors with crowdsourced failure reports.Proceedings of the ACM on Human-Computer Interaction5, CSCW2 (2021), 1–22

  14. [14]

    Carrie J Cai, Emily Reif, Narayan Hegde, Jason Hipp, Been Kim, Daniel Smilkov, Martin Wattenberg, Fernanda Viegas, Greg S Corrado, Martin C Stumpe, et al. 2019. Human-centered tools for coping with imperfect algorithms during medical decision-making. InProceedings of the 2019 chi conference on human factors in computing systems. 1–14

  15. [15]

    Hello AI

    Carrie J Cai, Samantha Winter, David Steiner, Lauren Wilcox, and Michael Terry. 2019. " Hello AI": uncovering the onboarding needs of medical practitioners for human-AI collaborative decision-making.Proceedings of the ACM on Human-computer Interaction3, CSCW (2019), 1–24

  16. [16]

    Valerie Chen, Q Vera Liao, Jennifer Wortman Vaughan, and Gagan Bansal. 2023. Understanding the role of human intuition on reliance in human-AI decision-making with explanations.Proceedings of the ACM on Human-computer Interaction7, CSCW2 (2023), 1–32

  17. [17]

    Hao-Fei Cheng, Ruotong Wang, Zheng Zhang, Fiona O’Connell, Terrance Gray, F Maxwell Harper, and Haiyi Zhu. 2019. Explaining decision-making algorithms through UI: Strategies to help non-expert stakeholders. InProceedings of the 2019 chi conference on human factors in computing systems. 1–12

  18. [18]

    Jack Cuzick. 1985. A Wilcoxon-type test for trend.Statistics in medicine4, 1 (1985), 87–90

  19. [19]

    Maria De-Arteaga, Riccardo Fogliato, and Alexandra Chouldechova. 2020. A case for humans-in-the-loop: Decisions in the presence of erroneous algorithmic scores. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems. 1–12

  20. [20]

    John J Dudley and Per Ola Kristensson. 2018. A review of user interface design for interactive machine learning.ACM Transactions on Interactive Intelligent Systems (TiiS)8, 2 (2018), 1–37

  21. [21]

    Alan M Frisch, Christopher Jefferson, Bernadette Martínez Hernández, and Ian Miguel. 2005. The rules of constraint modelling. InIJCAI. 109–116

  22. [22]

    Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, et al. 2023. A survey of uncertainty in deep neural networks.Artificial intelligence review56, Suppl 1 (2023), 1513–1589

  23. [23]

    David J Gladstone, Cynthia J Danells, and Sandra E Black. 2002. The Fugl-Meyer assessment of motor recovery after stroke: a critical review of its measurement properties.Neurorehabilitation and neural repair16, 3 (2002), 232–240

  24. [24]

    Ben Green and Yiling Chen. 2019. The principles and limits of algorithm-in-the-loop decision making.Proceedings of the ACM on human-computer interaction3, CSCW (2019), 1–24

  25. [25]

    O’Reilly Media, Inc

    Miguel Grinberg. 2018.Flask web development: developing web applications with python. " O’Reilly Media, Inc. "

  26. [26]

    Lijie Guo, Elizabeth M Daly, Oznur Alkan, Massimiliano Mattetti, Owen Cornec, and Bart Knijnenburg. 2022. Building trust in interactive machine learning via user contributed interpretable rules. InProceedings of the 27th international conference on intelligent user interfaces. 537–548. RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impac...

  27. [27]

    Ziyang Guo, Yifan Wu, Jason D Hartline, and Jessica Hullman. 2024. A decision theoretic framework for measuring AI reliance. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. 221–236

  28. [28]

    Kenneth Holstein, Maria De-Arteaga, Lakshmi Tumati, and Yanghuidi Cheng. 2023. Toward supporting perceptual complementarity in human-AI collaboration via reflection on unobservables.Proceedings of the ACM on Human-Computer Interaction7, CSCW1 (2023), 1–20

  29. [29]

    Min Hun Lee, Daniel P Siewiorek, Asim Smailagic, Alexandre Bernardino, and Sergi Bermudez i Badia. 2023. Design, development, and evaluation of an interactive personalized social robot to monitor and coach post-stroke rehabilitation exercises.User Modeling and User-Adapted Interaction33, 2 (2023), 545–569

  30. [30]

    Ashish Kapoor, Bongshin Lee, Desney Tan, and Eric Horvitz. 2010. Interactive optimization for steering machine classification. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems. 1343–1352

  31. [31]

    Anna Kawakami, Venkatesh Sivaraman, Hao-Fei Cheng, Logan Stapleton, Yanghuidi Cheng, Diana Qing, Adam Perer, Zhiwei Steven Wu, Haiyi Zhu, and Kenneth Holstein. 2022. Improving human-AI partnerships in child welfare: understanding worker practices, challenges, and desires for algorithmic decision support. InProceedings of the 2022 CHI Conference on Human F...

  32. [32]

    Dajung Kim, Niko Vegt, Valentijn Visch, and Marina Bos-De Vos. 2024. How Much Decision Power Should (A) I Have?: Investigating Patients’ Preferences Towards AI Autonomy in Healthcare Decision Making. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–17

  33. [33]

    I’m Not Sure, But

    Sunnie SY Kim, Q Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Wortman Vaughan. 2024. " I’m Not Sure, But... ": Examining the Impact of Large Language Models’ Uncertainty Expression on User Reliance and Trust. InProceedings of the 2024 ACM conference on fairness, accountability, and transparency. 822–835

  34. [34]

    Sunnie SY Kim, Jennifer Wortman Vaughan, Q Vera Liao, Tania Lombrozo, and Olga Russakovsky. 2025. Fostering appropriate reliance on large language models: The role of explanations, sources, and inconsistencies. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–19

  35. [35]

    Todd Kulesza, Margaret Burnett, Weng-Keen Wong, and Simone Stumpf. 2015. Principles of explanatory debugging to personalize interactive machine learning. InProceedings of the 20th international conference on intelligent user interfaces. 126–137

  36. [36]

    Todd Kulesza, Simone Stumpf, Weng-Keen Wong, Margaret M Burnett, Stephen Perona, Amy J Ko, and Ian Oberst. 2011. Why-oriented end-user debugging of naive Bayes text classification.ACM Transactions on Interactive Intelligent Systems (TiiS)1, 1 (2011), 1–31

  37. [37]

    Tzu-Sheng Kuo, Hong Shen, Jisoo Geum, Nev Jones, Jason I Hong, Haiyi Zhu, and Kenneth Holstein. 2023. Understanding Frontline Workers’ and Unhoused Individuals’ Perspectives on AI Used in Homeless Services. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17

  38. [38]

    Vivian Lai, Chacha Chen, Q Vera Liao, Alison Smith-Renner, and Chenhao Tan. 2021. Towards a science of human-ai decision making: a survey of empirical studies.arXiv preprint arXiv:2112.11471(2021)

  39. [39]

    Vivian Lai and Chenhao Tan. 2019. On human predictions with explanations and predictions of machine learning models: A case study on deception detection. InProceedings of the conference on fairness, accountability, and transparency. 29–38

  40. [40]

    Himabindu Lakkaraju, Julius Adebayo, and Sameer Singh. 2020. Explaining machine learning predictions: State-of-the-art, challenges, and opportunities.NeurIPS Tutorial(2020)

  41. [41]

    Himabindu Lakkaraju, Stephen H Bach, and Jure Leskovec. 2016. Interpretable decision sets: A joint framework for description and prediction. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 1675–1684

  42. [42]

    Kimin Lee, Laura Smith, and Pieter Abbeel. 2021. Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training.arXiv preprint arXiv:2106.05091(2021)

  43. [43]

    Min Hun Lee. 2026. From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making.arXiv preprint arXiv:2603.18895(2026)

  44. [44]

    Min Hun Lee and Chong Jun Chew. 2023. Understanding the Effect of Counterfactual Explanations on Trust and Reliance on AI for Human-AI Collaborative Clinical Decision Making.Proceedings of the ACM on Human-Computer Interaction7, CSCW2 (2023), 1–22

  45. [45]

    Min Hun Lee, Silvana Xin Yi Choo, Shamala D Thilarajah, et al. 2024. Improving Health Professionals’ Onboarding with AI and XAI for Trustworthy Human-AI Collaborative Decision Making.arXiv preprint arXiv:2405.16424(2024)

  46. [46]

    Min Hun Lee, Renee Bao Xuan Ng, Silvana Xinyi Choo, and Shamala Thilarajah. 2024. Interactive Example-based Explanations to Improve Health Professionals’ Onboarding with AI for Human-AI Collaborative Decision Making.arXiv preprint arXiv:2409.15814(2024)

  47. [47]

    Min Hun Lee, Daniel P Siewiorek, Asim Smailagic, Alexandre Bernardino, and Sergi Bermúdez i Badia. 2019. Learning to assess the quality of stroke rehabilitation exercises. InProceedings of the 24th International Conference on Intelligent User Interfaces. 218–228

  48. [48]

    Min Hun Lee, Daniel P Siewiorek, Asim Smailagic, Alexandre Bernardino, and Sergi Bermúdez i Badia. 2020. Co-design and evaluation of an intelligent decision support system for stroke rehabilitation assessment.Proceedings of the ACM on Human-Computer Interaction4, CSCW2 (2020), 1–27

  49. [49]

    Min Hun Lee, Daniel P Siewiorek, Asim Smailagic, Alexandre Bernardino, and Sergi Bermúdez i Badia. 2021. A human-ai collaborative approach for clinical decision making on rehabilitation assessment. InProceedings of the 2021 CHI conference on human factors in computing systems. 1–14

  50. [50]

    Min Hun Lee and Martyn Zhe Yu Tok. 2025. Towards Uncertainty Aware Task Delegation and Human-AI Collaborative Decision-Making. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. 2274–2289

  51. [51]

    Jingshu Li, Yitian Yang, Q Vera Liao, Junti Zhang, and Yi-Chieh Lee. 2025. As Confidence Aligns: Understanding the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–16. 18 Lee and Teo

  52. [52]

    Gabriel Lima, Nina Grgić-Hlača, and Meeyoung Cha. 2021. Human perceptions on moral responsibility of AI: A case study in AI-assisted bail decision-making. InProceedings of the 2021 CHI conference on human factors in computing systems. 1–17

  53. [53]

    Scott M Lundberg and Su-In Lee. 2017. A Unified Approach to Interpreting Model Predictions. InAdvances in Neural Information Processing Systems 30, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.). Curran Associates, Inc., 4765–4774. http://papers.nips.cc/paper/7062-a-unified-approach-to-interpreting-model-...

  54. [54]

    Frank J Massey Jr. 1951. The Kolmogorov-Smirnov test for goodness of fit.Journal of the American statistical Association46, 253 (1951), 68–78

  55. [55]

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. InProceedings of the conference on fairness, accountability, and transparency. 220–229

  56. [56]

    Claudio Novelli, Mariarosaria Taddeo, and Luciano Floridi. 2024. Accountability in artificial intelligence: What it is and how it works.Ai & Society 39, 4 (2024), 1871–1882

  57. [57]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library.Advances in neural information processing systems32 (2019)

  58. [58]

    Snehal Prabhudesai, Leyao Yang, Sumit Asthana, Xun Huan, Q Vera Liao, and Nikola Banovic. 2023. Understanding uncertainty: how lay decision- makers perceive and interpret uncertainty in human-AI decision making. InProceedings of the 28th international conference on intelligent user interfaces. 379–396

  59. [59]

    Alun Preece. 2018. Asking ‘Why’in AI: Explainability of intelligent systems–perspectives and challenges.Intelligent Systems in Accounting, Finance and Management25, 2 (2018), 63–72

  60. [60]

    Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio. 2019. Transfusion: Understanding transfer learning for medical imaging.Advances in neural information processing systems32 (2019)

  61. [61]

    Rahul Rahaman et al. 2021. Uncertainty quantification and deep ensembles.Advances in neural information processing systems34 (2021), 20063–20075

  62. [62]

    Pranav Rajpurkar, Emma Chen, Oishi Banerjee, and Eric J Topol. 2022. AI in health and medicine.Nature medicine28, 1 (2022), 31–38

  63. [63]

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018. Anchors: High-precision model-agnostic explanations. InProceedings of the AAAI conference on artificial intelligence, Vol. 32

  64. [64]

    Cynthia Rudin. 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.Nature machine intelligence1, 5 (2019), 206–215

  65. [65]

    Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong. 2022. Interpretable machine learning: Fundamental principles and 10 grand challenges.Statistic Surveys16 (2022), 1–85

  66. [66]

    Tim Sainburg, Leland McInnes, and Timothy Q Gentner. 2021. Parametric UMAP embeddings for representation and semisupervised learning. Neural Computation33, 11 (2021), 2881–2907

  67. [67]

    Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2024. Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision Making. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–17

  68. [68]

    Edward H Shortliffe. 1986. Medical expert systems—knowledge tools for physicians.Western Journal of Medicine145, 6 (1986), 830

  69. [69]

    Logan Stapleton, Min Hun Lee, Diana Qing, Marya Wright, Alexandra Chouldechova, Ken Holstein, Zhiwei Steven Wu, and Haiyi Zhu. 2022. Imagining new futures beyond predictive systems in child welfare: A qualitative study with impacted stakeholders. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 1162–1177

  70. [70]

    Ningjing Tang, Jiayin Zhi, Tzu-Sheng Kuo, Calla Kainaroi, Jeremy J Northup, Kenneth Holstein, Haiyi Zhu, Hoda Heidari, and Hong Shen. 2024. AI failure cards: Understanding and supporting grassroots efforts to mitigate AI failures in homeless services. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. 713–732

  71. [71]

    Stefano Teso and Kristian Kersting. 2019. Explanatory interactive machine learning. InProceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. 239–245

  72. [72]

    Ayush K Varshney and Vicenç Torra. 2023. Literature review of the recent trends and applications in various fuzzy rule-based systems.International Journal of Fuzzy Systems25, 6 (2023), 2163–2186

  73. [73]

    Brilliant AI doctor

    Dakuo Wang, Liuping Wang, Zhan Zhang, Ding Wang, Haiyi Zhu, Yvonne Gao, Xiangmin Fan, and Feng Tian. 2021. “Brilliant AI doctor” in rural clinics: Challenges in AI-powered clinical decision support system deployment. InProceedings of the 2021 CHI conference on human factors in computing systems. 1–18

  74. [74]

    Danding Wang, Qian Yang, Ashraf Abdul, and Brian Y Lim. 2019. Designing theory-driven user-centric explainable AI. InProceedings of the 2019 CHI conference on human factors in computing systems. 1–15

  75. [75]

    Xinru Wang and Ming Yin. 2021. Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making. In 26th international conference on intelligent user interfaces. 318–328

  76. [76]

    Bryan Wilder, Eric Horvitz, and Ece Kamar. 2020. Learning to complement humans.arXiv preprint arXiv:2005.00582(2020)

  77. [77]

    Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld. 2019. Errudite: Scalable, reproducible, and testable error analysis. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 747–763

  78. [78]

    Qian Yang, Jina Suh, Nan-Chen Chen, and Gonzalo Ramos. 2018. Grounding interactive machine learning tool design in how non-experts actually build models. InProceedings of the 2018 designing interactive systems conference. 573–584. RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview 19 Table 5. Top-𝐾rule performance across rule ...

  79. [79]

    Wencan Zhang and Brian Y Lim. 2022. Towards relatable explainable AI with the perceptual process. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–24

  80. [80]

    Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. 2020. Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. InProceedings of the 2020 conference on fairness, accountability, and transparency. 295–305. A System Implementation Details A.1 Feature Extractions We extend prior rehabilitation assessment work [...