REVIEW 3 major objections 5 minor 1 cited by
Supporting AI-Augmented Meta-Decision Making with InDecision
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper argues that AI tools should help people decide how to decide, by generating concrete decision options that provoke reflection on the criteria they are using.
desk verdict A well-scoped design paper with honest limits: the D1-D7 goals and the InDecision prototype are worth reading, but the 11-person pilot is too thin to establish that LLM provocations improve real decision criteria. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is InDecision's iterative "option provocation" loop, in which an LLM generates eight concrete alternatives, the user is required to narrow them to three, and the system infers up to six criteria that could explain those choices. Users then prioritize and refine the definitions of those criteria, and the cycle repeats with eight new options generated to align with or challenge the refined criteria. Three theoretical ingredients make the loop work: Variation Theory, which says comparing diverse instances helps people notice dimensions of difference; the known phenomenon of criteria drift, in which stated criteria diverge from what actually excites people as they compare cases; and the design stance that the human remains the judge while the AI acts as provocateur. The mechanism's purpose is to generate productive dissonance between what users say they value and what their choices reveal they value.
What would settle it
Run a controlled study in which participants use InDecision to refine criteria and then evaluate a separate set of real options: if their refined criteria predict their choices no better than their initial stated criteria, the paper's central claim is not supported.
Extended reading notes
Core claim
The central claim is that "deciding how to decide"—meta-decision making—can be augmented by AI that generates concrete, systematically varied, and deliberately provocative decision options. Because people often know their values only when they see concrete cases, comparing such cases is itself a mode of learning; AI's role is to provoke reflection rather than to recommend. The paper translates this into seven design goals (D1–D7) and a working case study, InDecision, whose loop elicits a user's description of a decision, generates eight options, forces a narrowing to the three that most resonate, infers candidate criteria, and lets users prioritize and redefine those criteria before repeating the cycle. Pilot observations with eleven users indicate that users generate new criteria while evaluating options, want to represent relationships among criteria, and sometimes discover that their decision had more than the two or three options they initially imagined. The paper stops short of claiming that these reflections lead to objectively better decisions; the claim is that the process of criteria development can be made iterative and reflective with AI support.
Load-bearing premise
The whole loop depends on LLM-generated options being plausible, relevant stand-ins for the user's real alternatives, so that what users discover while evaluating simulated cases carries over to the decisions they actually face.
Editorial extensions
If this is right
- Tools built on the same loop could let individuals and groups prototype their values before committing to a final rubric, much as designers prototype artifacts.
- If reflection on concrete options is enough to surface tacit criteria, then AI tools could be paired with, or even replaced by, structured human comparison exercises in low-resource settings.
- Group decision processes could use AI-generated edge cases to surface productive dissensus before teams converge prematurely.
- Because all elements of the decision can evolve, future tools should support fluid, side-by-side iteration of options and criteria rather than a fixed sequence.
Reading between the lines
- Editorial extension: the paper does not test transfer; the decisive next experiment would compare how well users' refined criteria predict actual choices against their initial stated criteria.
- Editorial extension: a controlled arm that swaps LLM edge cases for randomly selected cases could isolate whether the provocation itself, rather than the act of comparing options, drives criteria revision.
- Editorial extension: the same loop could scale to institutional rubrics, stress-testing hiring or admissions criteria against synthetic candidate profiles before they are used on real people.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This short paper argues that AI tools have a rich opportunity space to support human "meta-decision making," the process of deciding how to decide, by helping people iteratively prototype and refine decision criteria. The authors derive seven design goals (D1-D7) from prior literature in HCI, psychology, learning sciences, and decision science, focusing on positioning humans as judges and AI as provocateur, engaging users with concrete options, systematically varying examples, generating edge cases, enabling iteration, preserving process, and mitigating social pressure. They instantiate these goals in InDecision, a mixed-initiative LLM-based tool that elicits a decision description, generates eight concrete options for users to narrow to three, infers and lets users refine up to six criteria, and then loops. The paper reports early insights from designing and piloting InDecision with 11 participants, highlighting needs for fluid transitions between options and criteria, expressive representations of inter-criteria structure, expansion of the considered option space, and clarity about whether the system works for the user or the LLM. The authors conclude by raising open questions for future research on AI-augmented meta-decision making.
Significance. The paper's central strength is its careful, literature-grounded articulation of an under-explored design space: rather than helping users make a specific decision, AI tools can help them understand what they value before committing to criteria. The design goals D1-D7 are plausible and traceable to established concepts such as criteria drift, variation theory, reflective practice, and cognitive dissonance; the InDecision prototype makes the idea concrete and provides a starting point for empirical work. The authors are appropriately modest, labeling the pilot results as "early insights" rather than demonstrated efficacy. If the design goals are taken up by the community, this paper could meaningfully influence future tools for decision preparation in high-stakes contexts. The main risks are the lack of systematic empirical support and the unresolved reliance on LLM-generated options faithfully representing a user's real decision space.
major comments (3)
- [3.1] The pilot study is reported without essential methodological details: recruitment and participant demographics, the specific decisions participants worked on, the think-aloud and study protocol, how data were recorded and analyzed, and whether the thematic insights in Sections 3.1.1-3.1.4 were derived through any coding or consensus process. Because these insights are presented as findings that motivate future directions, the absence of this information prevents readers from assessing their validity. Please add an appendix with the protocol and a systematic analysis (e.g., theme definitions, representative quotes, and frequencies), or explicitly reframe these observations as informal pilot notes whose evidentiary weight is illustrative only.
- [3, Steps 2-3; Figure 1] The LLM-driven generation process is not described in enough detail to be reproduced or evaluated. Step 2 states the system "generates a diverse starter set of eight concrete decision options," and Step 3 states the system generates "a set of up to six inferred criteria," but the paper does not report the model, prompts, sampling strategy, or how diversity and edge-case-ness were operationalized. Without this information, claims that the options are "diverse" or "challenging" cannot be scrutinized. Please include the prompts or a detailed generation strategy, and ideally release the code/prompts as supplementary material.
- [3.1.3] The observation that users "were quick to reject options that seemed unrealistic" directly challenges the tool's core assumption that LLM-generated options simulate the user's real decision space. This tension is central to the paper's contribution: if many generated options are discarded as irrelevant, the reflection loop may fail to produce criteria that transfer to real decisions. The paper should (a) report how often options were rejected in the pilots, (b) discuss how rejections affected the criteria-refinement step, and (c) propose concrete mitigations, such as allowing users to edit generated options or gathering richer context before generation. As written, the paper leaves an unresolved threat to the validity of the entire approach.
minor comments (5)
- [2] The transition from cited literature to each design goal is sometimes implicit; a mapping table or explicit sentences linking each goal to the supporting references would improve traceability.
- [3.1.4] The finding expressed as "Is this for me, or the LLM?" is intriguing but under-developed. Please clarify whether this perception affected users' engagement or their willingness to iterate on criteria.
- [4] The paper lacks a dedicated Limitations section. A short paragraph acknowledging the exploratory nature of the pilot, the small sample, and the dependence on LLM output quality would strengthen the manuscript.
- [References] Several citations are to arXiv preprints that may not have been peer-reviewed (e.g., [16], [19], [20], [21]). Please mark preprints as such and indicate whether they have been accepted or published.
- [Figure 1] The figure caption does not explain the numbered steps in enough detail; consider referencing the corresponding subsections (1)-(4) in the caption to help readers map the figure to the text.
Circularity Check
No significant circularity: InDecision's design goals and pilot insights are presented as hypotheses and future directions, not as predictions derived from the tool's own outputs.
full rationale
This paper does not derive a result from first principles or make predictions that are forced by its own definitions. It proposes design goals D1-D7 grounded in prior literature, describes a prototype called InDecision that instantiates those goals, and reports initial qualitative pilot observations as 'early insights' and directions for future research. The central claim is an opportunity-space proposal: that AI tools can support meta-decision making by helping people prototype and stress-test decision criteria. The pilot observations are explicitly framed as exploratory, not as validation: for example, Section 3.1 says 'Below are some of our early insights from these pilot studies, which point to directions for future research and design.' No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no design choice is justified solely by a self-citation. The few self-references (e.g., prior work on participatory design and AI-supported deliberation) are contextual citations rather than load-bearing arguments. The paper even acknowledges open questions and limitations, such as whether LLM-generated provocations transfer to real decisions, which further corroborates the absence of circular reasoning. Accordingly, the paper is self-contained as a design-research contribution and merits a circularity score of 0.
Assumptions & free parameters
free parameters (3)
- starter set size =
8
- selection target =
3
- inferred criteria cap =
6
assumptions (4)
- domain assumption People have limited conscious access to their own values and learn them through comparing concrete options ('know it when they see it').
- domain assumption Criteria drift is a genuine and costly phenomenon in real decision processes.
- domain assumption Variation Theory from concept learning generalizes to preference elicitation, so systematically varying options helps users discover dimensions.
- ad hoc to paper LLM-generated options can simulate realistic, decision-relevant cases.
Cite this review
Pith. "Pith review of Supporting AI-Augmented Meta-Decision Making with InDecision." pith.science (2026). https://pith.science/paper/Y2E4LH4A
@misc{pith2026250412433,
author = {Pith},
title = {Pith review of: Supporting AI-Augmented Meta-Decision Making with InDecision},
year = {2026},
howpublished = {\url{https://pith.science/paper/Y2E4LH4A}},
note = {Machine review of arXiv:2504.12433}
}
read the original abstract
From school admissions to hiring and investment decisions, the first step behind many high-stakes decision-making processes is "deciding how to decide." Formulating effective criteria to guide decision-making requires an iterative process of exploration, reflection, and discovery. Yet, this process remains under-supported in practice. In this short paper, we outline an opportunity space for AI-driven tools that augment human meta-decision making. We draw upon prior literature to propose a set of design goals for future AI tools aimed at supporting human meta-decision making. We then illustrate these ideas through InDecision, a mixed-initiative tool designed to support the iterative development of decision criteria. Based on initial findings from designing and piloting InDecision with users, we discuss future directions for AI-augmented meta-decision making.
Figures
Forward citations
Cited by 1 Pith paper
-
Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop
A synthesis of the CHI 2025 workshop maps research and design opportunities for understanding, protecting, and augmenting human cognition with generative AI.
Reference graph
Works this paper leans on
-
[1]
Y-Lan Boureau, Peter Sokol-Hessner, and Nathaniel D Daw. 2015. Deciding how to decide: Self-control and meta-decision making. Trends in cognitive sciences 19, 11 (2015), 700–710
work page 2015
-
[2]
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-computer Interaction 5, CSCW1 (2021), 1–21
2021
-
[3]
Alice Cai, Ian Arawjo, and Elena L Glassman. 2024. Antagonistic AI. arXiv preprint arXiv:2402.07350 (2024)
arXiv 2024
-
[4]
Quan Ze Chen and Amy X Zhang. 2023. Judgment Sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (2023), 1–26
work page 2023
-
[5]
Kate Compton and Michael Mateas. 2015. Casual Creators. In ICCC. 228–235
work page 2015
-
[6]
Dawn Culpepper, Damani White-Lewis, KerryAnn O’Meara, Lindsey Templeton, and Julia Anderson. 2023. Do rubrics live up to their promise? Examining how rubrics mitigate bias in faculty hiring. The Journal of Higher Education 94, 7 (2023), 823–850
work page 2023
-
[7]
Avedis Donabedian. 1981. Advantages and limitations of explicit criteria for assessing the quality of health care.The Milbank Memorial Fund Quarterly. Health and Society 59, 1 (1981), 99–106
work page 1981
-
[8]
Ian Drosos, Advait Sarkar, Neil Toronto, et al. 2025. “It makes you think”: Provo- cations Help Restore Critical Thinking to AI-Assisted Knowledge Work. arXiv preprint arXiv:2501.17247 (2025)
arXiv 2025
Show all 36 references
-
[9]
Salvatore Greco, Jose Figueira, and Matthias Ehrgott. 2016. Multiple criteria decision analysis. Vol. 37. Springer
2016
-
[10]
The Bridgespan Group. 2023. Developing decision criteria: Achieving strate- gic clarity. https://www.bridgespan.org/getmedia/f978e90c-f37d-4042-bf11- 6cedfcbc2a97/ASC-DevelopingDecisionCriteria-Complete.pdf
2023
-
[11]
Eddie Harmon-Jones and Judson Mills. 2019. An introduction to cognitive disso- nance theory and an overview of current perspectives on the theory. InCognitive dissonance: Reexamining a pivotal theory in psychology, 2nd ed . American Psycho- logical Association, Washington, DC,...
2019
-
[12]
Daniel Kahneman, Andrew M Rosenfield, Linnea Gandhi, and Tom Blaser. 2016. Noise: How to overcome the high, hidden cost of inconsistent decision making. Harvard Business Review 94, 10 (2016), 38–46
2016
-
[13]
Anna Kawakami, Venkatesh Sivaraman, Hao-Fei Cheng, Logan Stapleton, Yanghuidi Cheng, Diana Qing, Adam Perer, Zhiwei Steven Wu, Haiyi Zhu, and Kenneth Holstein. 2022. Improving human-AI partnerships in child welfare: understanding worker practices, challenges, and desires for a...
2022
-
[14]
Mary Beth Kery, Marissa Radensky, Mahima Arya, Bonnie E John, and Brad A Myers. 2018. The story in the notebook: Exploratory data science using a literate programming tool. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. 1–11
2018
-
[15]
Vinay Koshy, Frederick Choi, Yi-Shyuan Chiang, Hari Sundaram, Eshwar Chan- drasekharan, and Karrie Karahalios. 2024. Venire: A Machine Learning-Guided Panel Review System for Community Content Moderation. arXiv preprint arXiv:2410.23448 (2024)
2024 arXiv
-
[16]
Tzu-Sheng Kuo, Quan Ze Chen, Amy X Zhang, Jane Hsieh, Haiyi Zhu*, and Ken- neth Holstein*. 2024. PolicyCraft: Supporting Collaborative and Participatory Pol- icy Design through Case-Grounded Deliberation. arXiv preprint arXiv:2409.15644 (2024)
2024 arXiv
-
[17]
Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein*, and Haiyi Zhu*. 2024. Wikibench: Community-driven data curation for AI evaluation on Wikipedia. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–24
2024
-
[18]
Vivian Lai, Chacha Chen, Alison Smith-Renner, Q Vera Liao, and Chenhao Tan
-
[19]
Michelle S Lam, Fred Hohman, Dominik Moritz, Jeffrey P Bigham, Kenneth Holstein*, and Mary Beth Kery*. 2024. AI Policy Projector: Grounding LLM Policy Design in Iterative Mapmaking. arXiv preprint arXiv:2409.18203 (2024)
2024 arXiv
-
[20]
Michael Xieyang Liu, Tongshuang Wu, Tianying Chen, Franklin Mingzhe Li, Aniket Kittur, and Brad A Myers. 2023. Selenite: Scaffolding decision making with comprehensive overviews elicited from large language models. arXiv preprint arXiv:2310.02161 (2023)
2023 arXiv
-
[21]
Shuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng, Zhenhui Peng, Ming Yin, and Xiaojuan Ma. 2024. Towards human-ai deliberation: Design and evaluation of llm-empowered deliberative ai for ai-assisted decision-making. arXiv preprint arXiv:2403.16812 (2024)
2024 arXiv
-
[22]
Ference Marton. 2014. Necessary conditions of learning . Routledge
2014
-
[23]
April McGrath. 2017. Dealing with Dissonance: A Review of Cognitive Dis- sonance Reduction. Social and Personality Psychology Compass 11, 12 (2017), e12362
2017
-
[24]
John McMackin and Paul Slovic. 2000. When does explicit justification impair decision making? Applied Cognitive Psychology: The Official Journal of the Society for Applied Research in Memory and Cognition 14, 6 (2000), 527–541
2000
-
[25]
unstructured
Henry Mintzberg, Duru Raisinghani, and Andre Theoret. 1976. The structure of" unstructured" decision processes. Administrative Science Quarterly (1976), 246–275
1976
-
[26]
Ben R Newell and David R Shanks. 2014. Unconscious influences on decision making: A critical review. Behavioral and brain sciences 37, 1 (2014), 1–19
2014
-
[27]
Peter Pirolli and Stuart Card. 2005. The sensemaking process and leverage points for analyst technology as identified through cognitive task analysis. In Proceedings of International Conference on Intelligence Analysis , Vol. 5. McLean, VA, USA, 2–4
2005
-
[28]
William E Remus and Jeffrey E Kottemann. 1986. Toward intelligent decision support systems: An artificially intelligent statistician. Mis Quarterly (1986), 403–418
1986
-
[29]
Donald A Schön. 1992. Designing as reflective conversation with the materials of a design situation. Knowledge-based systems 5, 1 (1992), 3–14
1992
-
[30]
Shreya Shankar, JD Zamfirescu-Pereira, Björn Hartmann, Aditya Parameswaran, and Ian Arawjo. 2024. Who validates the validators? aligning llm-assisted evalu- ation of llm outputs with human preferences. In Proceedings of the 37th Annual ACM Symposium on User Interface Software ...
2024
-
[31]
Sarah Sterman, Molly Jane Nicholas, and Eric Paulos. 2022. Towards Creative Version Control. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 336 (Nov. 2022), 25 pages
2022
-
[32]
Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024. Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation. InProceedings of the CHI Conference on Human Factors in Computing Systems . 1–26
2024
-
[33]
Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: En- abling multilevel exploration and sensemaking with large language models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–18
2023
-
[34]
Qian Yang, Aaron Steinfeld, and John Zimmerman. 2019. Unremarkable AI: Fitting intelligent decision support into critical, clinical decision-making processes. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . 1–11
2019
-
[35]
Angie Zhang, Olympia Walker, Kaci Nguyen, Jiajun Dai, Anqing Chen, and Min Kyung Lee. 2023. Deliberating with AI: improving decision-making for the future through participatory AI design and stakeholder deliberation. Proceedings of the ACM on Human-Computer Interaction 7, CSCW...
2023
-
[2023]
In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency
Towards a science of human-AI decision making: An overview of de- sign space in empirical human-subject studies. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency . 1369–1385
2023
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.