Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Supported designers scored 3.5 vs 1.0 in GenAI CAD feasibility study

desk verdict A transparent exploratory WoZ study that opens a useful design space for metacognitive support in GenAI CAD, but its headline outcome gap lacks an attention control and should be reframed as a formative result. read the letter →

arxiv 2506.12879 v1 pith:JBOOGWSS submitted 2025-06-15 cs.HC cs.AI

classification cs.HCcs.AI
keywords human-AIco-creationmetacognitivesupportgenerativedesignWizardofOzthink-aloudagentsCAD
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a voice-based metacognitive support agent—software that asks reflective questions, prompts planning and sketching, or lets an expert guide the user—can help people work with generative AI design tools that are otherwise hard to use. In a Wizard of Oz study with 20 mechanical designers, participants who received agent support produced designs that met more of the engineering brief's feasibility criteria (average 3.5 out of 5) than unsupported participants (1.0 out of 5, with every unsupported design failing structural analysis). The paper's aim is not to crown a single best strategy, but to open a design space: different support styles helped with different parts of the design process, and each came with tradeoffs. The central insight is that prompting reflection during AI-assisted design may matter as much as giving answers.

What carries the argument

The carrying mechanism is the metacognitive support agent probe: a voice-based conversational agent that hears the designer's think-aloud speech and sees their screen in real time, operated behind the scenes by a human wizard. Three probes instantiate different support strategies: SocratAIs answers only with reflective questions; HephAIstus offers a project planning sheet, a free-body diagram sketching board, and proactive suggestions; Expert-Freeform lets invited experts choose their own tactics. The mechanism targets the three cognitive bottlenecks the authors identify in GenAI workflows—intent formulation, problem exploration, and outcome evaluation—prompting designers to externalize and reflect on assumptions before and while the black-box solver runs.

What would settle it

A decisive check would be to run the same 20-designer task with an additional attention-control arm in which a voice agent gives friendly, design-irrelevant encouragement at the same frequency as the support agents; if that group also averages near 3.5 on the five-point feasibility score while no-support averages 1.0, the metacognitive content is not the cause. Alternatively, pre-register an automated version of SocratAIs that asks the same questions without a human wizard; if the feasibility gap disappears, the human facilitator's expertise is the active ingredient.

Watch

Extended reading notes

Core claim

The central discovery is that metacognitive support agents can shift outcomes in GenAI-based CAD work from uniformly infeasible to mostly feasible. Using an engine-bracket generative design task, the study compares five designers supported by SocratAIs (questions only), five by HephAIstus (planning sheet, free-body diagram sketching, and suggestions), five by external generative-design experts acting freeform, and five with no support. Supported groups averaged 3.5 of 5 feasibility criteria versus 1.0 for unsupported; all five unsupported participants set up the bracket's load case incorrectly and none passed finite element analysis. The authors report qualitative differences across strategies: reflective questions helped with intent formulation and problem exploration, sketching anchored conversations about loads, and direct suggestions helped software operation but were less effective at overturning solidified misconceptions.

Load-bearing premise

The study's central comparison assumes the outcome gap between supported and unsupported designers is caused by the metacognitive support strategies, but it lacks an attention-control condition, the first author enacted two of the three agent types, and the expert-freeform group contained more experienced professionals; if the gap came from generic attention or from participant backgrounds, the central claim would not hold.

Editorial extensions

If this is right

  • If the result holds, adding a reflective-prompting layer to GenAI design tools could reduce the common failure mode in which novices submit infeasible parts.
  • Question-asking alone can support intent formulation, but the paper finds it can also entrench wrong assumptions or sow doubt in correct ones, so support must be adaptive.
  • Planning and sketching activities helped designers reason about loads even when few individual agent messages were coded as impactful, suggesting structured externalization carries part of the benefit.
  • None of the three strategies dominated, so the paper's own conclusion is that future systems should blend strategies and let users choose support style based on experience level.
  • Voice interaction and screen annotations were generally appreciated and reduced context-switching, motivating multimodal support interfaces for CAD and other visual tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not include an attention-control condition, so a plausible reading—one the authors acknowledge—is that any engaged human attention could produce part of the gain; a direct test would be an agent that speaks generic encouragement instead of design-specific prompts.
  • If the effect scales beyond the Wizard of Oz setting, the natural product implication is a 'reflection layer' in CAD tools that detects transitions between subtasks and inserts checkpoints or design reviews, as the paper suggests.
  • The identified bottlenecks—specifying criteria upfront and evaluating generated outputs—are not unique to mechanical engineering, so the findings may transfer to other GenAI co-creation tasks such as prompt-based image or circuit generation.
  • Replacing the human wizard with an automated large-language-model agent would test whether the metacognitive behaviors themselves, rather than the human facilitator's expertise, carry the outcome; the paper leaves this open.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper reports a Wizard-of-Oz study with 20 mechanical designers comparing three metacognitive support agents (SocratAIs, HephAIstus, and an Expert-Freeform condition) against a No Support baseline on a realistic Fusion 360 Generative Design task. The authors measure design outcome feasibility with a five-criterion rubric, analyze think-aloud video for message impact, and conduct thematic analysis of post-task interviews. The paper's central claim, stated in the abstract and Section 6.1.1, is that agent-supported users created more feasible designs than non-supported users (M = 3.5 vs. M = 1.0), with different support strategies showing different process-level effects. The paper positions the work as an exploratory prototyping study and derives design considerations for future metacognitive support agents.

Significance. If the central claim were established, this would be a meaningful contribution to HCI for AI-assisted design: it opens a design space for metacognitive support agents and provides qualitative insights about reflective questioning, planning, sketching, and expert freeform support. The study has notable strengths: a realistic and professionally relevant task, a transparent and reproducible five-point outcome rubric, the convergence of the No Support baseline with prior unsupported-task results in [42], and rich time-series process visualizations. The paper also ships a large appendix with wizard guidelines and interview protocols, which supports replication. However, the current evidence is exploratory: the headline outcome gap conflates metacognitive support with the presence of a live attentive wizard, the per-condition sample size is five, no inferential statistics are reported, and the video-impact coding lacks inter-rater reliability metrics. These issues prevent the paper from supporting its causal claims as stated, though the qualitative and design-space contributions remain valuable.

major comments (4)
  1. [Section 6.1.1 (Table 2, Figure 3)] The headline comparison between supported and unsupported users confounds the independent variable of interest with the presence of a live, attentive wizard. In all three support conditions, participants were told an AI agent would support them, heard a synthetic voice, saw an 'agent processing' idle cue, and received real-time personalized messages. In the No Support condition there was no such attentive presence, no expectation of help, and no comparable social interaction. The observed outcome gap could therefore be due to generic human attention, accountability, or demand characteristics rather than to metacognitive support strategies. This threat is load-bearing because the abstract and Section 7.1 attribute the outcome difference to metacognitive support. The paper does not include an attention-control arm, and Section 7.3's acknowledgment of wizard-dependent results does not repair the missing control. I recommend either adding an attention-control condition or substantially reframing the central claim as an exploratory demonstration that real-time human-like agent support can help, without attributing the effect specifically to metacognitive strategies.
  2. [Section 5.5.2 and Section 6.1.2] The video interaction analysis that produces the 'triggered new considerations' counts is central to RQ1 and to the between-condition comparisons in Section 6.1.2, but the paper reports no inter-rater reliability or agreement metric. The text says the three coders 'met frequently to discuss edge cases and ensure consistency,' but it does not quantify consistency. Without an agreement measure, the reader cannot assess whether the message-impact codes are reliable or idiosyncratic. This is a specific, fixable weakness: the authors should report a reliability statistic (e.g., Cohen's kappa or Krippendorff's alpha) for a subset of coded sessions, or explicitly present the counts as a single coder's interpretive reading rather than as a reliable measurement.
  3. [Section 7.3 and Table 4] The small per-condition sample (n = 5) and the demographic imbalance between conditions are acknowledged but not adequately handled. Section 7.3 states that most of the more experienced professionals were in the Expert-Freeform group and says 'we disregarded this potential bias since the observed behaviors were similar across all supported groups.' This is not a valid justification: similarity of observed behavior does not remove the confound between support strategy and participant experience for the outcome scores. In addition, Section 6.1.1 reports a 'consistent gap' between supported and unsupported users, but no inferential statistics are provided, and the No Support group's constant score of 1.0 (SD = 0.0) suggests a floor effect. The authors should either provide appropriate statistical analysis (or a clear statement that no inferential claims are made) and a sensitivity analysis for the experience confound, or temper the causal language throughout the abstract and discussion.
  4. [Section 6.1.2 and Figure 5] The message-impact analysis is used to argue that SocratAIs' questions were more effective than HephAIstus' suggestions for intent formulation and problem exploration, but this comparison is itself confounded by agent response policy and session length. HephAIstus and Expert-Freeform also responded to user-initiated queries, and the normalized frequency metric divides agent-initiated messages by session duration without controlling for the number or timing of user messages. Figure 5's y-axis label, 'Number of Impactful Messages / Total Messages,' is ambiguous about whether the bar height is a count, a percentage, or both. The qualitative examples (e.g., S5, H4) are informative, but the quantitative contrast should be presented with more caution or with a clearer model of the messaging process.
minor comments (6)
  1. [Author affiliations] The affiliation for Ye Wang contains a typo: 'San Franscisco' should be 'San Francisco.'
  2. [Section 6.1.4 D] The text says 'only a fraction (0.06%) directly helped overcome cognitive design challenges (16/196 messages)'; 16/196 is approximately 8.2%, not 0.06%. Please correct this percentage.
  3. [Table 2] Several rows in Table 2 are difficult to parse: the 'Feasible Part Size' row mixes checkmarks, blank cells, and 'AVG' entries, and the 'Normalized Message Frequency' row appears to contain stray values that do not align cleanly with the per-participant columns. Please reformat the table so each cell corresponds unambiguously to a participant and metric.
  4. [Section 5.5.2] The description of how the three coders were 'equally distributed' across sessions and how 'edge cases' were resolved would benefit from more detail about whether each session was coded by one researcher or more than one, and whether any dual-coding was used for agreement assessment.
  5. [Figure 5] The caption and legend of Figure 5 should define the difference between 'total messages' and 'impactful messages' explicitly, and clarify whether the saturated areas represent percentages of the total messages per category or percentages per agent group.
  6. [Section 5.2] The piloting of the task is described in one sentence ('we verified the suitability of the task...'). Please provide a sentence or two on the number of pilots, their outcomes, and any resulting changes to the protocol, as this bears on the validity of the task.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the central supported-vs-unsupported contrast is an empirical comparison computed from the paper's own study data.

full rationale

This paper is an empirical Wizard-of-Oz study rather than a derivation chain. Its central claim, that agent-supported users produced more feasible designs than unsupported users, rests on outcome scores in Table 2 and Figure 3 (supported M=3.5, SD=1.4 vs No Support M=1.0, SD=0.0), which are coded from submitted CAD models against five independent feasibility criteria described in Section 5.5.1. No parameter is fitted and then renamed as a prediction; the outcome measure is not defined in terms of agent-message counts or agent strategy. The only self-citations are to the authors' prior CHI 2023 study [42], used to adopt the engine-bracket task and to note that the unsupported baseline is consistent with earlier results. Those citations are corroborative context, not the evidence for the supported-vs-unsupported contrast; that contrast is computed from the current study's own No Support and supported groups. No uniqueness theorem, ansatz, or definitional identity is imported from the authors' prior work. The acknowledged confounds (wizard identity, lack of attention control, demographic imbalance) are internal-validity limitations, not circular-reasoning steps. Therefore, no circularity is present.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim relies on qualitative research assumptions rather than fitted numerical parameters. The main assumptions concern think-aloud validity, Wizard of Oz generalizability, the outcome rubric, and the absence of an attention control.

assumptions (4)
  • domain assumption Designers' concurrent think-aloud verbalizations reliably reveal their cognitive processes and task-relevant knowledge.
    This is used throughout the procedure to give the wizard material for agent messages, as described in Sections 5.1 and 5.4.2.
  • domain assumption A human wizard following strategy guidelines approximates the behavior of a future automated metacognitive support agent.
    This is the standard Wizard of Oz assumption, presented in Section 5.4, but it is unverified for a deployed system.
  • domain assumption Design outcome feasibility can be captured by the five binary criteria in Section 5.5.1.
    The entire supported versus unsupported comparison rests on this rubric, and no external validation of the rubric against independent expert ratings is reported.
  • domain assumption The observed benefit of agent support is not entirely a generic attention or demand-characteristics effect.
    No active-control condition was included, so the paper assumes the benefit comes from support content rather than mere human presence; this is untested in Sections 5.3 and 7.3.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation." pith.science (2026). https://pith.science/paper/JBOOGWSS

@misc{pith2026250612879,
  author       = {Pith},
  title        = {Pith review of: Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JBOOGWSS}},
  note         = {Machine review of arXiv:2506.12879}
}
read the original abstract

Despite the potential of generative AI (GenAI) design tools to enhance design processes, professionals often struggle to integrate AI into their workflows. Fundamental cognitive challenges include the need to specify all design criteria as distinct parameters upfront (intent formulation) and designers' reduced cognitive involvement in the design process due to cognitive offloading, which can lead to insufficient problem exploration, underspecification, and limited ability to evaluate outcomes. Motivated by these challenges, we envision novel metacognitive support agents that assist designers in working more reflectively with GenAI. To explore this vision, we conducted exploratory prototyping through a Wizard of Oz elicitation study with 20 mechanical designers probing multiple metacognitive support strategies. We found that agent-supported users created more feasible designs than non-supported users, with differing impacts between support strategies. Based on these findings, we discuss opportunities and tradeoffs of metacognitive support agents and considerations for future AI-based design tools.

Figures

Figures reproduced from arXiv: 2506.12879 by the authors.

Figure 1
Figure 1. Overview of the Fusion 360 design task (A-E), workflow (F-G), common user mistakes (H), and cognitive challenges (I). [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Process diagram of Wizard of Oz setup. The remotely located wizard (right) followed the designer’s actions (left) by [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Plot showing the design outcome scores between [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (4 more)
Figure 5
Figure 5. Figure 5: Plot illustrating the total number of messages per [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Timeline excerpts visualizing participant and agent interactions throughout the design task; timelines are divided [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: Timeline plots visualizing participant and agent interactions throughout the design task; timelines are divided into [PITH_FULL_IMAGE:figures/full_fig_p022_7.png]
Figure 8
Figure 8. Figure 8: Screenshot of the sketching board HephAIstus sent to users (with scribbles from H3 on it). A.1.5 Expert-Freeform Agent Introduction. Expert Agent: Hey! I am a voice agent here to support you during the design task. I can hear what you are saying, and I can see your scr…

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stop Writing for Me: Generative Refusal in AI Tools for Thought

    cs.HC 2026-05 conditional novelty 5.0 of 10

    AI tools that deliberately refuse to generate text and ask context-aware questions can reduce cognitive burden in creative journaling and leave users with internalized questioning habits.

Reference graph

Works this paper leans on

127 extracted references · 37 canonical work pages · cited by 1 Pith paper

  1. [42]

    Frederic Gmeiner, Humphrey Yang, Lining Yao, Kenneth Holstein, and Nikolas Martelaro. 2023. Exploring Challenges and Opportunities to Support Designers in Learning to Co-create with AI-based Manufacturing Design Tools. In Pro- ceedings of the 2023 CHI Conference on Human Factors in Computing Systems . ACM, Hamburg Germany, 1–20. https://doi.org/10.1145/35...

  2. [1]

    Mark Abdelshiheed, John Wesley Hostetter, Tiffany Barnes, and Min Chi. 2023. Leveraging Deep Reinforcement Learning for Metacognitive Interventions Across Intelligent Tutoring Systems. In Artificial Intelligence in Education , Ning Wang, Genaro Rebolledo-Mendez, Noboru Matsuda, Olga C. Santos, and Vania Dimitrova (Eds.). Vol. 13916. Springer Nature Switze...

  3. [2]

    Erfan Al-Hossami, Razvan Bunescu, Ryan Teehan, Laurel Powell, Khyati Maha- jan, and Mohsen Dorodchi. 2023. Socratic Questioning of Novice Debuggers: A Benchmark Dataset and Preliminary Evaluations. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023). Association for Computational Linguistics, Toron...

  4. [3]

    Tamang, and V

    Zeyad Alshaikh, L. Tamang, and V. Rus. 2020. Experiments with a Socratic Intelligent Tutoring System for Source Code Understanding. In The Florida AI Research Society

  5. [4]

    Marco Aurisicchio, Rob H Bracewell, Ken M Wallace, et al. 2007. Characterising Design Questions That Involve Reasoning. In DS 42: Proceedings of ICED 2007, the 16th International Conference on Engineering Design, Paris, France, 28.-31.07. 2007

  6. [5]

    Autodesk. 2020. Fusion 360 Generative Design. https://www.autodesk.com/solutions/generative-design/manufacturing

  7. [6]

    Paul Ayres and John Sweller. 2005. The Split-Attention Principle in Multimedia Learning. In The Cambridge Handbook of Multimedia Learning (1 ed.), Richard Mayer (Ed.). Cambridge University Press, 135–146. https://doi.org/10.1017/ CBO9780511816819.009

  8. [7]

    Maral Babapour. 2015. Roles of Externalisation Activities in the Design Process. Swedish Design Research Journal 2014 (May 2015). https://doi.org/10.3384/svid. 2000-964x.14134

Show all 127 references
  1. [8]

    Ball and Bo T

    Linden J. Ball and Bo T. Christensen. 2009. Analogical Reasoning and Mental Simulation in Design: Two Strategies Linked to Uncertainty Resolution. Design Studies 30, 2 (March 2009), 169–186. https://doi.org/10.1016/j.destud.2008.12.005

  2. [9]

    Ball and Bo T

    Linden J. Ball and Bo T. Christensen. 2019. Advancing an Understanding of Design Cognition and Design Metacognition: Progress and Prospects. Design Studies 65 (Nov. 2019), 35–59. https://doi.org/10.1016/j.destud.2019.10.003

  3. [10]

    Maria Bannert, Peter Reimann, and Christoph Sonnenberg. 2014. Process Mining Techniques for Analysing Patterns and Strategies in Students’ Self-Regulated Learning. Metacognition and Learning 9, 2 (Aug. 2014), 161–185. https://doi. org/10.1007/s11409-013-9107-6

  4. [11]

    Kicking the Tires

    Eric P.S. Baumer and Bill Tomlinson. 2011. Comparing Activity Theory with Distributed Cognition for Video Analysis: Beyond "Kicking the Tires". In Pro- ceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’11). Association for Computing Machinery, New Y...

  5. [12]

    Warren Berger. 2014. A More Beautiful Question: The Power of Inquiry to Spark Breakthrough Ideas. Bloomsbury USA, New York, NY

  6. [13]

    Berry and Donald E

    Dianne C. Berry and Donald E. Broadbent. 1984. On the Relationship be- tween Task Performance and Associated Verbalizable Knowledge. The Quar- terly Journal of Experimental Psychology Section A 36, 2 (May 1984), 209–231. https://doi.org/10.1080/14640748408402156

  7. [14]

    Kirsten Boehner, Janet Vertesi, Phoebe Sengers, and Paul Dourish. 2007. How HCI Interprets the Probes. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’07) . Association for Computing Machinery, New York, NY, USA, 1077–1086. https://doi.org/1...

  8. [15]

    Baker, and Vincent Aleven

    Conrad Borchers, Jiayi Zhang, Ryan S. Baker, and Vincent Aleven. 2024. Us- ing Think-Aloud Data to Understand Relations between Self-Regulation Cycle Characteristics and Student Performance in Intelligent Tutoring Systems. In Proceedings of the 14th Learning Analytics and Know...

  9. [16]

    Virginia Braun and Victoria Clarke. 2019. Reflecting on Reflexive Thematic Analysis. Qualitative Research in Sport, Exercise and Health 11, 4 (Aug. 2019), 589–597. https://doi.org/10.1080/2159676X.2019.1628806

  10. [17]

    Sahan Bulathwela, Hamze Muse, and Emine Yilmaz. 2023. Scalable Educational Question Generation with Pre-trained Language Models. InArtificial Intelligence in Education, Ning Wang, Genaro Rebolledo-Mendez, Noboru Matsuda, Olga C. Santos, and Vania Dimitrova (Eds.). Vol. 13916. ...

  11. [18]

    Çetin Tünger, Çetin Tünger, Şule Taşlı Pektaş, and Sule Tasli Pektas. 2020. A Comparison of the Cognitive Actions of Designers in Geometry-Based and Parametric Design Environments. Open House International 45 (June 2020), 87–101. https://doi.org/10.1108/ohi-04-2020-0008

  12. [19]

    Carlos Cardoso, Ozgur Eris, Petra Badke-Schaub, and Marco Aurisicchio. 2014. Question Asking in Design Reviews: How Does Inquiry Facilitate the Learning Interaction?. In Proceedings of the 10th Design Thinking Research Symposium (DTRS). Purdue University, 18

  13. [20]

    Castro-Alonso and John Sweller

    Juan C. Castro-Alonso and John Sweller. 2020. The Modality Effect of Cognitive Load Theory. In Advances in Human Factors in Training, Education, and Learning Sciences, Waldemar Karwowski, Tareq Ahram, and Salman Nazir (Eds.). Springer International Publishing, Cham, 75–84

  14. [21]

    Chase, Jenna Marks, Deena Bernett, Melissa Bradley, and Vincent Aleven

    Catherine C. Chase, Jenna Marks, Deena Bernett, Melissa Bradley, and Vincent Aleven. 2015. Towards the Development of the Invention Coach: A Naturalistic Study of Teacher Guidance for an Exploratory Learning Task. In Artificial Intelligence in Education , Cristina Conati, Neil...

  15. [22]

    Rimika Chaudhury, Taha Liaqat, and Parmit K. Chilana. 2023. Exploring the Needs of Informal Learners of Computational Skills: Probe-Based Elicitation for the Design of Self-Monitoring Interventions

  16. [23]

    Xiang ’Anthony’ Chen, Ye Tao, Guanyun Wang, Runchang Kang, Tovi Grossman, Stelian Coros, and Scott E. Hudson. 2018. Forte: User-Driven Generative Design. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. ACM, Montreal QC Canada, 1–12. https://doi...

  17. [24]

    Michelene T. H. Chi. 2006. Laboratory Methods for Assessing Experts’ and Novices’ Knowledge. In The Cambridge Handbook of Expertise and Expert Performance, K. Anders Ericsson, Neil Charness, Paul J. Feltovich, and Robert R. Hoffman (Eds.). Cambridge University Press, Cambridge...

  18. [25]

    Chilana, Nathaniel Hudson, Srinjita Bhaduri, Prashant Shashiku- mar, and Shaun K

    Parmit K. Chilana, Nathaniel Hudson, Srinjita Bhaduri, Prashant Shashiku- mar, and Shaun K. Kane. 2018. Supporting Remote Real-Time Expert Help: Opportunities and Challenges for Novice 3D Modelers. (Oct. 2018), 157–166. https://doi.org/10.1109/vlhcc.2018.8506568

  19. [26]

    Carlos Coimbra Cardoso, Petra Badke-Schaub, and Ozgur Eris. 2016. Inflection Moments in Design Discourse: How Questions Drive Problem Framing during Idea Generation. Design Studies 46 (Sept. 2016), 59–78. https://doi.org/10.1016/ j.destud.2016.07.002

  20. [27]

    Craig, Jeremiah Sullins, Amy Witherspoon, and Barry Gholson

    Scotty D. Craig, Jeremiah Sullins, Amy Witherspoon, and Barry Gholson. 2006. The Deep-Level-Reasoning-Question Effect: The Role of Dialogue and Deep- Level-Reasoning Questions During Vicarious Learning.Cognition and Instruction 24, 4 (2006), 565–591. jstor:27739846

  21. [28]

    Nigel Cross. 2006. Designerly Ways of Knowing. Springer, Berlin London

  22. [29]

    Nils Dahlbäck, Arne Jönsson, and Lars Ahrenberg. 1993. Wizard of Oz Studies — Why and How. Knowledge-Based Systems 6, 4 (Dec. 1993), 258–266. https: //doi.org/10.1016/0950-7051(93)90017-N

  23. [30]

    Valdemar Danry, Pat Pataranutaporn, Yaoli Mao, and Pattie Maes. 2023. Don’t Just Tell Me, Ask Me: AI Systems That Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI Ex- planations. In Proceedings of the 2023 CHI Conference o...

  24. [31]

    Nicholas Davis, Chih-PIn Hsiao, Kunwar Yashraj Singh, Lisa Li, Sanat Moningi, and Brian Magerko. 2015. Drawing Apprentice: An Enactive Co-Creative Agent for Artistic Collaboration. In Proceedings of the 2015 ACM SIGCHI Conference Exploring the Potential of Metacognitive Suppor...

  25. [32]

    Flood, Ana Asuncion, Dor Abraham- son, Noel Enyedy, and Francis Steen

    David DeLiema, Maggie Dahn, Virginia J. Flood, Ana Asuncion, Dor Abraham- son, Noel Enyedy, and Francis Steen. 2019. Debugging as a Context for Fostering Reflection on Critical Thinking and Emotion (1 ed.). Routledge, London, 209–228. https://doi.org/10.4324/9780429323058-13

  26. [33]

    J. T. Dillon. 1984. The Classification of Research Questions.Review of Educational Research 54, 3 (Sept. 1984), 327–361. https://doi.org/10.3102/00346543054003327

  27. [34]

    Kees Dorst and Nigel Cross. 2001. Creativity in the Design Process: Co-Evolution of Problem–Solution. Design Studies 22, 5 (Sept. 2001), 425–437. https://doi. org/10.1016/S0142-694X(01)00009-6

  28. [35]

    Graham Dove, Nicolai Brodersen Hansen, and Kim Halskov. 2016. An Argu- ment For Design Space Reflection. In Proceedings of the 9th Nordic Conference on Human-Computer Interaction (NordiCHI ’16) . Association for Computing Machinery, New York, NY, USA, 1–10. https://doi.org/10....

  29. [36]

    It’s like a Rubber Duck That Talks Back

    Ian Drosos, Advait Sarkar, Xiaotong Xu, Carina Negreanu, Sean Rintel, and Lev Tankelevitch. 2024. "It’s like a Rubber Duck That Talks Back": Understand- ing Generative AI-Assisted Data Analysis Workflows through a Participatory Prompting Study. In Proceedings of the 3rd Annual...

  30. [37]

    E, Yu-Chun Grace Yen, Isabelle Yan Pan, Grace Lin, Mingyi Li, Hy- oungwook Jin, Mengyi Chen, Haijun Xia, and Steven P

    Jane L. E, Yu-Chun Grace Yen, Isabelle Yan Pan, Grace Lin, Mingyi Li, Hy- oungwook Jin, Mengyi Chen, Haijun Xia, and Steven P. Dow. 2024. When to Give Feedback: Exploring Tradeoffs in the Timing of Design Feedback. In Proceedings of the 16th Conference on Creativity & Cognitio...

  31. [38]

    Linda Elder and Richard Paul. 2016. The Thinker’s Guide to The Art of Socratic Questioning. Foundation for Critical Thinking Press

  32. [39]

    Ozgur Eris. 2004. Effective Inquiry for Innovative Engineering Design . Springer US, Boston, MA. https://doi.org/10.1007/978-1-4419-8943-7

  33. [40]

    John H. Flavell. 1979. Metacognition and Cognitive Monitoring: A New Area of Cognitive–Developmental Inquiry. American Psychologist 34, 10 (Oct. 1979), 906–911. https://doi.org/10.1037/0003-066X.34.10.906

  34. [41]

    formlabs. 2020. Generative Design 101. https://formlabs.com/blog/generative- design/

  35. [43]

    Gabriela Goldschmidt, Hagay Hochman, and Itay Dafni. 2010. The Design Studio “Crit”: Teacher–Student Communication. Artificial Intelligence for Engineering Design, Analysis and Manufacturing 24, 3 (Aug. 2010), 285–302. https://doi.org/ 10.1017/S089006041000020X

  36. [44]

    Graesser and Natalie K

    Arthur C. Graesser and Natalie K. Person. 1994. Question Asking During Tutoring. American Educational Research Journal 31, 1 (March 1994), 104–137. https://doi.org/10.3102/00028312031001104

  37. [45]

    Hacker (Ed.)

    Douglas J. Hacker (Ed.). 1998. Metacognition in Educational Theory and Practice . Erlbaum, Mahwah, NJ

  38. [46]

    Robert Hausmann and Kurt Vanlehn. 2010. The Effect of Self-Explaining on Robust Learning. I. J. Artificial Intelligence in Education 20 (Jan. 2010), 303–332. https://doi.org/10.3233/JAI-2010-010

  39. [47]

    Sofie Heirweg, Mona De Smul, Emmelien Merchie, Geert Devos, and Hilde Keer

  40. [48]

    Kingma, Ben Poole, Mohammad Norouzi, David J

    Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. 2022. Imagen Video: High Definition Video Generation with Diffusion Models. https://arxiv.org/abs/2210.02303v1

  41. [49]

    Yueh-Ren Ho, Bao-Yu Chen, and Chien-Ming Li. 2023. Thinking More Wisely: Using the Socratic Method to Develop Critical Thinking Skills amongst Health- care Students. BMC Medical Education 23, 1 (March 2023), 173. https: //doi.org/10.1186/s12909-023-04134-2

  42. [50]

    Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxuan Zhang, Juanzi Li, Bin Xu, Yuxiao Dong, Ming Ding, and Jie Tang. 2023. CogAgent: A Visual Language Model for GUI Agents. arXiv:2312.08914

  43. [51]

    Nespoli, and John S

    Ada Hurst, Shirley Lin, Claire Treacy, Oscar G. Nespoli, and John S. Gero. 2023. Comparing Academics and Practitioners Q&A Tutoring in the Engineering Design Studio. Proceedings of the Design Society 3 (July 2023), 997–1006. https: //doi.org/10.1017/pds.2023.100

  44. [52]

    Nikhita Joshi, Justin Matejka, Fraser Anderson, Tovi Grossman, and George Fitzmaurice. 2020. MicroMentor: Peer-to-Peer Software Help Sessions in Three Minutes or Less. (April 2020), 1–13. https://doi.org/10.1145/3313831.3376230

  45. [53]

    Miller, and Patricia A

    Shabnam Kavousi, Patrick A. Miller, and Patricia A. Alexander. 2020. The Role of Metacognition in the First-Year Design Lab. Educational Technology Research and Development 68, 6 (Dec. 2020), 3471–3494. https://doi.org/10.1007/s11423- 020-09848-4

  46. [54]

    Rubaiat Habib Kazi, Tovi Grossman, Hyunmin Cheong, Ali Hashemi, and George Fitzmaurice. 2017. DreamSketch: Early Stage 3D Design Explorations with Sketching and Generative Design. In Proceedings of the 30th Annual ACM Sym- posium on User Interface Software and Technology. ACM,...

  47. [55]

    Sumbul Khan and Bige Tunçer. 2019. Gesture and Speech Elicitation for 3D CAD Modeling in Conceptual Design. Automation in Construction 106 (Oct. 2019), 102847. https://doi.org/10.1016/j.autcon.2019.102847

  48. [56]

    Hannah Kim, Jaegul Choo, Haesun Park, and Alex Endert. 2016. InterAxis: Steering Scatterplot Axes via Observation-Level Interaction. IEEE Transactions on Visualization and Computer Graphics 22, 1 (Jan. 2016), 131–140. https: //doi.org/10.1109/TVCG.2015.2467615

  49. [57]

    Ko and Brad A

    Amy J. Ko and Brad A. Myers. 2004. Designing the Whyline: A Debugging Interface for Asking Questions about Program Behavior. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . ACM, Vienna Austria, 151–158. https://doi.org/10.1145/985692.985712

  50. [58]

    Jon Kolko. 2010. Abductive Thinking and Sensemaking: The Drivers of Design Synthesis. Design Issues 26, 1 (Jan. 2010), 15–28. https://doi.org/10.1162/desi. 2010.26.1.15

  51. [59]

    Lasecki, Tovi Grossman, and George Fitzmaurice

    Rebecca Krosnick, Fraser Anderson, Justin Matejka, Steve Oney, Walter S. Lasecki, Tovi Grossman, and George Fitzmaurice. 2021. Think-Aloud Com- puting: Supporting Rich and Low-Effort Knowledge Capture. International Conference on Human Factors in Computing Systems (May 2021). ...

  52. [60]

    Kelly Y. L. Ku and Irene T. Ho. 2010. Metacognitive Strategies That Enhance Critical Thinking. Metacognition and Learning 5, 3 (Dec. 2010), 251–267. https: //doi.org/10.1007/s11409-010-9060-6

  53. [61]

    Todd Kulesza, Margaret Burnett, Weng-Keen Wong, and Simone Stumpf. 2015. Principles of Explanatory Debugging to Personalize Interactive Machine Learn- ing. In Proceedings of the 20th International Conference on Intelligent User Inter- faces. ACM, Atlanta Georgia USA, 126–137. ...

  54. [62]

    Mustafa Kurt and Sevinc Kurt. 2017. Improving Design Understandings and Skills through Enhanced Metacognition: Reflective Design Journals. Inter- national Journal of Art & Design Education 36, 2 (2017), 226–238. https: //doi.org/10.1111/jade.12094

  55. [63]

    Hannu Kuusela and Pallab Paul. 2000. A Comparison of Concurrent and Retro- spective Verbal Protocol Analysis. The American Journal of Psychology 113, 3 (2000), 387. https://doi.org/10.2307/1423365 jstor:1423365

  56. [64]

    J. D. Lee and K. A. See. 2004. Trust in Automation: Designing for Appropriate Reliance. Human Factors: The Journal of the Human Factors and Ergonomics Society 46, 1 (Jan. 2004), 50–80. https://doi.org/10.1518/hfes.46.1.50_30392

  57. [65]

    Wendy G. Lehnert. 1978. The Process of Question Answering: A Computer Simulation of Cognition (1 ed.). Routledge, London. https://doi.org/10.4324/ 9781003316817

  58. [66]

    Justin Matejka, Michael Glueck, Erin Bradner, Ali Hashemi, Tovi Grossman, and George Fitzmaurice. 2018. Dream Lens: Exploration and Visualization of Large- Scale Generative Design Datasets. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . ACM, ...

  59. [67]

    Justin Matejka, Tovi Grossman, and George Fitzmaurice. 2011. Ambient Help. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’11). Association for Computing Machinery, New York, NY, USA, 2751–2760. https://doi.org/10.1145/1978942.1979349

  60. [68]

    Ambuj Mehrish, Navonil Majumder, Rishabh Bhardwaj, Rada Mihalcea, and Soujanya Poria. 2023. A Review of Deep Learning Techniques for Speech Processing. arXiv:2305.00359

  61. [69]

    Achim Menges and Sean Ahlquist (Eds.). 2011. Computational Design Thinking (1. publ ed.). Wiley, Chichester

  62. [70]

    Hongwei Niu, Cees Van Leeuwen, Jia Hao, Guoxin Wang, and Thomas Lach- mann. 2022. Multimodal Natural Human–Computer Interfaces for Computer- Aided Design: A Review Paper. Applied Sciences 12, 13 (June 2022), 6510. https://doi.org/10.3390/app12136510

  63. [71]

    Runliang Niu, Jindong Li, Shiqi Wang, Yali Fu, Xiyu Hu, Xueyuan Leng, He Kong, Yi Chang, and Qi Wang. 2024. ScreenAgent: A Vision Language Model-driven Computer Control Agent. (2024). https://doi.org/10.48550/ARXIV.2402.07945

  64. [72]

    Gross, and Ellen Yi-Luen Do

    Yeonjoo Oh, Suguru Ishizaki, Mark D. Gross, and Ellen Yi-Luen Do. 2013. A Theoretical Framework of Design Critiquing in Architecture Studios. Design Studies 34, 3 (May 2013), 302–325. https://doi.org/10.1016/j.destud.2012.08.004

  65. [73]

    Ernesto Panadero. 2017. A Review of Self-regulated Learning: Six Models and Four Directions for Research. Frontiers in Psychology 8 (April 2017), 422. https://doi.org/10.3389/fpsyg.2017.00422

  66. [74]

    Soya Park, Hari Subramonyam, and Chinmay Kulkarni. 2024. Thinking As- sistants: LLM-Based Conversational Assistants That Help Users Think By DIS ’25, July 5–9, 2025, Funchal, Portugal Gmeiner et al. Asking Rather than Answering. https://doi.org/10.48550/arXiv.2312.06024 arXiv:...

  67. [75]

    Maria Teresa Parreira, Sarah Gillet, and Iolanda Leite. 2023. Robot Duck De- bugging: Can Attentive Listening Improve Problem Solving?. In Proceedings of the 25th International Conference on Multimodal Interaction (ICMI ’23) . As- sociation for Computing Machinery, New York, N...

  68. [76]

    Carolyn Plumb, Rose Marra, Douglas Hacker, and John Dunlosky. 2018. Mea- suring Engineering Students’ Metacognition with a Think-Aloud Protocol. In 2018 ASEE Annual Conference & Exposition Proceedings . ASEE Conferences, Salt Lake City, Utah, 30796. https://doi.org/10.18260/1-2--30796

  69. [77]

    Raquel Plumed, Carmen González-Lluch, David Pérez-López, Manuel Contero, and Jorge D Camba. 2021. A Voice-Based Annotation System for Collaborative Computer-Aided Design. Journal of Computational Design and Engineering 8, 2 (April 2021), 536–546. https://doi.org/10.1093/jcde/qwaa092

  70. [78]

    Pontificia Universidad Javeriana, Juanita Tobón, Fabio Tellez, and Oscar Alzate

  71. [79]

    Rebecca Anne Price and Peter Lloyd. 2022. Asking Effective Questions: Aware- ness of Bias in Designerly Thinking. In Handbook of Engineering Systems Design, Anja Maier, Josef Oehmen, and Pieter E. Vermaas (Eds.). Springer International Publishing, Cham, 1–16. https://doi.org/1...

  72. [80]

    Xiangshi Ren, Gao Zhang, and Guozhong Dai. 2000. An Experimental Study of Input Modes for Multimodal Human-Computer Interaction. In Advances in Multimodal Interfaces — ICMI 2000 (Lecture Notes in Computer Science) , Tieniu Tan, Yuanchun Shi, and Wen Gao (Eds.). Springer, Berli...

  73. [81]

    Risko and Sam J

    Evan F. Risko and Sam J. Gilbert. 2016. Cognitive Offloading.Trends in Cognitive Sciences 20, 9 (Sept. 2016), 676–688. https://doi.org/10.1016/j.tics.2016.07.002

  74. [82]

    Argüelles, Nicolás F

    Debrina Roy, Nicole Calpin, Kathy Cheng, Alison Olechowski, Andrea P. Argüelles, Nicolás F. Soria Zurita, and Jessica Menold. 2024. Designing To- gether: Exploring Collaborative Dynamics of Multi-Objective Design Problems in Virtual Environments. Journal of Mechanical Design 1...

  75. [83]

    Marta Royo, Elena Mulet, Vicente Chulvi, and Francisco Felip. 2021. Guiding Questions for Increasing the Generation of Product Ideas to Meet Changing Needs (QuChaNe). Research in Engineering Design 32, 3 (July 2021), 411–430. https://doi.org/10.1007/s00163-021-00364-x

  76. [84]

    Sara Mah- davi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mah- davi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. 2022. Photorealistic Text-to-Image Diff...

  77. [85]

    Donald A. Schön. 1983. The Reflective Practitioner: How Professionals Think in Action. Basic Books, New York

  78. [86]

    I Reflect to Improve My Design

    Moushumi Sharmin and Brian P. Bailey. 2011. "I Reflect to Improve My Design": Investigating the Role and Process of Reflection in Creative Design. In Pro- ceedings of the 8th ACM Conference on Creativity and Cognition (C&C ’11). Association for Computing Machinery, New Yor...

  79. [87]

    Kumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha, Manu Kapur, and Mrinmaya Sachan. 2022. Automatic Generation of Socratic Subques- tions for Teaching Math Word Problems. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing ,...

  80. [88]

    Keith Stenning, Alexander Schmoelz, Heather Wren, Elias Stouraitis, Theodore Scaltsas, Constantine Alexopoulos, and Amelie Aichhorn. 2016. Socratic Dia- logue as a Teaching and Research Method for Co-Creativity? Digital Culture and Education 8, 2 (2016), 13

  81. [89]

    Hari Subramonyam, Roy Pea, Christopher Pondoc, Maneesh Agrawala, and Colleen Seifert. 2024. Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMs. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24) ...

  82. [90]

    Lasang Jimba Tamang, Zeyad Alshaikh, Nisrine Ait Khayi, Priti Oli, and Vasile Rus. 2021. A Comparative Study of Free Self-Explanations and Socratic Tutoring Explanations for Source Code Comprehension. Proceedings of the 52nd ACM Technical Symposium on Computer Science Educatio...

  83. [91]

    Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2023. The Metacognitive Demands and Opportunities of Generative AI. https://doi.org/10.48550/arXiv.2312.10893 arXiv:2312.10893 [cs]

  84. [92]

    Tawfik, Arthur Graesser, Jessica Gatewood, and Jaclyn Gishbaugher

    Andrew A. Tawfik, Arthur Graesser, Jessica Gatewood, and Jaclyn Gishbaugher

  85. [93]

    Suzanne Tolmeijer, Naim Zierau, Andreas Janson, Jalil Sebastian Wahdatehagh, Jan Marco Marco Leimeister, and Abraham Bernstein. 2021. Female by Default? – Exploring the Effect of Voice Assistant Gender and Pitch on Trait and Trust Attribution. In Extended Abstracts of the 2021...

  86. [94]

    Schuller, Gökçe İymen, Metin Sezgin, Xiangheng He, Zijiang Yang, Panagiotis Tzirakis, Shuo Liu, Silvan Mertes, Elisa- beth André, Ruibo Fu, and Jianhua Tao

    Andreas Triantafyllopoulos, Björn W. Schuller, Gökçe İymen, Metin Sezgin, Xiangheng He, Zijiang Yang, Panagiotis Tzirakis, Shuo Liu, Silvan Mertes, Elisa- beth André, Ruibo Fu, and Jianhua Tao. 2023. An Overview of Affective Speech Synthesis and Conversion in the Deep Learning...

  87. [95]

    Educational Technology Research and Development 68, 2 (April 2020), 653–678

    Role of Questions in Inquiry-Based Instruction: Towards a Design Taxonomy for Question-Asking and Implications for Design. Educational Technology Research and Development 68, 2 (April 2020), 653–678. https: //doi.org/10.1007/s11423-020-09738-9

  88. [96]

    Jones, and Michelene T.H

    Kurt VanLehn, Randolph M. Jones, and Michelene T.H. Chi. 1992. A Model of the Self-Explanation Effect. Journal of the Learning Sciences 2, 1 (Jan. 1992), 1–59. https://doi.org/10.1207/s15327809jls0201_1

  89. [97]

    Geoffrey Vaughan. 2022. Metacognition and Self-Regulated Learning: Recent Perspectives for an International Context. Opus et Educatio 9, 2 (Aug. 2022). https://doi.org/10.3311/ope.501

  90. [98]

    Maaike Van Den Haak, Menno De Jong, and Peter Jan Schellens. 2003. Ret- rospective vs. Concurrent Think-Aloud Protocols: Testing the Usability of an Online Library Catalogue. Behaviour & Information Technology 22, 5 (Sept. 2003), 339–351. https://doi.org/10.1080/0044929031000

  91. [99]

    Min Wang and Yong Zeng. 2009. Asking the Right Questions to Elicit Product Requirements. International Journal of Computer Integrated Manufacturing 22, 4 (April 2009), 283–298. https://doi.org/10.1080/09511920802232902

  92. [100]

    Judith D. Wilson. 1987. A Socratic Approach to Helping Novice Programmers Debug Programs. In Proceedings of the Eighteenth SIGCSE Technical Symposium on Computer Science Education - SIGCSE ’87 . ACM Press, St. Louis, Missouri, United States, 179–182. https://doi.org/10.1145/31...

  93. [101]

    Tom Veuskens, Danny Leen, and Raf Ramakers. 2022. Identifying Opportunities to Reimagine Parametric Modeling for Makers. (2022)

  94. [102]

    Humphrey Yang, Kuanren Qian, Haolin Liu, Yuxuan Yu, Jianzhe Gu, Matthew McGehee, Yongjie Jessica Zhang, and Lining Yao. 2020. SimuLearn: Fast and Accurate Simulator to Support Morphing Materials Design and Workflows. In Proceedings of the 33rd Annual ACM Symposium on User Inte...

  95. [103]

    Kirsty Young. 2009. Direct from the Source: The Value of ’think-Aloud’ Data in Understanding Learning. The Journal of Educational Enquiry 6 (2009)

  96. [104]

    Robert Woodbury. 2010. Elements of Parametric Design . Routledge, London

  97. [105]

    Zamfirescu-Pereira, David Sirkin, David Goedicke, Ray LC, Natalie Fried- man, Ilan Mandel, Nikolas Martelaro, and Wendy Ju

    J.D. Zamfirescu-Pereira, David Sirkin, David Goedicke, Ray LC, Natalie Fried- man, Ilan Mandel, Nikolas Martelaro, and Wendy Ju. 2021. Fake It to Make It: Exploratory Prototyping in HRI. In Companion of the 2021 ACM/IEEE In- ternational Conference on Human-Robot Interaction (H...

  98. [106]

    Zamfirescu-Pereira, Richmond Y

    J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang

  99. [107]

    Loutfouz Zaman, Wolfgang Stuerzlinger, Christian Neugebauer, Rob Woodbury, Maher Elkhaldi, Naghmi Shireen, and Michael Terry. 2015. GEM-NI: A System for Creating and Managing Alternatives In Generative Design. In Proceedings of the 33rd Annual ACM Conference on Human Factors i...

  100. [108]

    Jiayi Zhang, Conrad Borchers, Vincent Aleven, and Ryan Shaun Baker. 2024. Using Large Language Models to Detect Self-Regulated Learning in Think-Aloud Protocols. https://doi.org/10.35542/osf.io/hrtz6 A Additional Materials Table 4: Overview of study participants. ID Agent Grou...

  101. [111]

    Guanglu Zhang, Ayush Raina, Jonathan Cagan, and Christopher McComb. 2021. A Cautionary Tale about the Impact of AI on Human Design Teams. Design Studies 72 (Jan. 2021), 100990. https://doi.org/10.1016/j.destud.2021.100990

  102. [113]

    Follow the designer’s verbalizations and screen actions and pay close attention to the task-specific design steps and challenges as outlined in Section 3, such as specifying the bracket’s load cases (forces and structural constraints), modeling appropriate geometry for keeping...

  103. [114]

    Pay close attention to inconsistencies between the requirements stated in the design brief and the input parameters set by the designer. Such requirements could be explicit (e.g., the force the bracket needs to hold) or implicit features, such as bolt clearances, which were no...

  104. [115]

    Never directly tell the participant what to do, but rather provide supportive questions, hints, or suggestions (depending on the enacted agent type)

  105. [116]

    I am unsure if

    You are free to send messages whenever and how often you consider it helpful to the designer. However, pay special attention to moments in which designers transition between design sub-tasks (such as from specifying obstacle geometry to specifying loads), as well as when desig...

  106. [117]

    A.1.2 Guidelines for the Expert-Freeform wizards

    You are free to formulate the messages in a way you consider to be most helpful, while adhering to the agent’s support strategy (e.g., only asking questions). A.1.2 Guidelines for the Expert-Freeform wizards . The Expert-Freeform wizards (external experts not part of the resea...

  107. [118]

    • Understand the operational conditions of the ship engine

    Project Scope and Requirements • Define the objectives of the bracket design. • Understand the operational conditions of the ship engine. • Identify load types (static, dynamic, thermal) and magnitudes. • Clarify space constraints and installation considerations

  108. [119]

    • Consider material properties such as strength, weight, corrosion resistance, and cost

    Material Selection • Discuss different material options (metal alloys, composites, etc.). • Consider material properties such as strength, weight, corrosion resistance, and cost. • Review the material performance under extreme marine conditions

  109. [120]

    • Evaluate the pros and cons of each method concerning the design objectives

    Manufacturing Method • Determine feasible manufacturing methods (casting, machining, additive manufacturing, etc.). • Evaluate the pros and cons of each method concerning the design objectives. • Discuss generative design constraints for each manufacturing process

  110. [121]

    • Define the design space and apply necessary constraints and conditions

    Generative Design Parameters • Set up load cases and boundary conditions in Fusion 360. • Define the design space and apply necessary constraints and conditions. • Choose the resolution of the generative design mesh

  111. [122]

    • Define requirements for vibration dampening

    Design Constraints and Criteria • Set criteria for minimum safety factors. • Define requirements for vibration dampening. • Consider access for maintenance and installation

  112. [123]

    • Analyze stress distribution, deformation, and fatigue life

    Simulation and Analysis • Plan for simulations to predict performance under various loads. • Analyze stress distribution, deformation, and fatigue life. • Review thermal and fluid flow analysis if necessary

  113. [124]

    • Discuss trade-offs between different optimization objectives

    Optimization Objectives • Establish the optimization goals, such as weight reduction, strength optimization, cost efficiency, etc. • Discuss trade-offs between different optimization objectives

  114. [125]

    • Consider classification society requirements and certifications

    Compliance and Standards • Ensure the design meets marine industry standards and regulatory compliance. • Consider classification society requirements and certifications

  115. [126]

    • Plan for interfaces with other systems and parts

    Integration with Existing Systems • Discuss how the bracket will integrate with the ship's engine and surrounding structures. • Plan for interfaces with other systems and parts

  116. [127]

    • Maintenance

    Lifecycle Considerations • Consider the lifecycle impacts, such as ease of manufacture, sustainability, recyclability, and end-of-life disposal. • Maintenance. Free-body Diagram Sketching Activity: The agent can suggest that the designer sketch out load case-relevant forces an...

  117. [2019]

    In Insider Knowledge - Proceedings of the Design Research Society Learn X Design Con- ference, 2019

    Metacognition in the Wild: Metacognitive Studies in Design Education. In Insider Knowledge - Proceedings of the Design Research Society Learn X Design Con- ference, 2019. Design Research Society. https://doi.org/10.21606/learnxdesign. 2019.09128

  118. [2020]

    Instructional Science 48 (Aug

    Mine the Process: Investigating the Cyclical Nature of Upper Primary School Students’ Self-Regulated Learning. Instructional Science 48 (Aug. 2020). https://doi.org/10.1007/s11251-020-09519-0

  119. [2023]

    In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems

    Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. ACM, Hamburg Germany, 1–21. https://doi.org/10.1145/ 3544548.3581388

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.