REVIEW 3 major objections 5 minor 130 references
Controlling Context: Generative AI at Work in Integrated Circuit Design and Other High-Precision Domains
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that for integrated circuit designers, generative AI's accuracy failures are secondary to context failures: engineers are more troubled by generic outputs and misparsed documents than by wrong answers, and the biggest…
desk verdict Worth a careful read for HCI/CSCW folks: a useful 'trouble' framing and early empirical data, but the heavy-user sampling undercuts the general claim that accuracy is secondary. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The analytical engine of the paper is the concept of 'trouble': any difficulty that forces a user to re-prompt, edit, or abandon an output. Trouble is treated as orthogonal to accuracy and is mapped onto three elements of a generative AI sociotechnical system — the pipeline (training data, tokenization, fine-tuning, and retrieval-augmented generation), the features (interface, metaprompts, conversation threading), and grounding (the inability of text tokens to carry situated meaning without human context). This mapping does the work of converting scattered interview complaints into design targets and recommendations.
What would settle it
A comparable interview or survey study sampling light and non-users of generative AI in IC design that found accuracy failures—for example wrong simulation code slipping past review—were the most commonly cited reason for limiting or abandoning the tools would directly contradict the claim.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that concerns about accuracy are eclipsed by other problems 'without exception' among the engineers interviewed. Only seven of 17 commented on accuracy at all, and none treated it as a barrier to use. The most common troubles were systems using supplied documents out of context (n=11), outputs too generic to be useful (n=9), and failed numerical operations (n=6). These troubles are not user error or a need for training; they are manifestations of the gap between the general-purpose design of generative AI and the situated, organization-specific context of chip design. Engineers repair the gap by atomizing tasks, iterating, making context explicit in prompts, and relying on existing code review and verification workflows, and the paper concludes that the repair burden should be shifted back to the system through interactive control of context.
Load-bearing premise
All interview evidence comes from 17 intensive users (over 500,000 tokens in 30 days) at one firm using internal tools only about six months old; if this early-adopter, heavy-use group tolerates inaccuracy better than typical engineers, the paper's 'accuracy is secondary' claim may not hold for chip designers broadly.
Editorial extensions
If this is right
- Improving generative AI for engineering should focus on interactive context control—letting users constrain documents, style, and conversation state—rather than on pushing benchmark accuracy scores higher.
- Existing verification, code review, and unit testing remain load-bearing for safe GenAI use and may need to be strengthened, since engineers rely on them to absorb inaccuracies.
- The most frequent trouble, document parsing, points to a concrete fix: RAG and file-input components should be made layout-aware so tables, footnotes, and adjacent cells are not misread.
- Transparency features—showing metaprompts, persistent context, and uncertainty estimates—could reduce the guesswork engineers currently put into prompt repair.
- Organizations should preserve pathways for novice engineers to build the judgment senior engineers used to supervise outputs, since most interviewees doubted novices could reliably catch wrong outputs.
Reading between the lines
- The trouble taxonomy likely transfers to other high-precision document-heavy domains such as medicine, law, and scientific research, where layout and local conventions carry much of the meaning; a comparative study could test this directly.
- Productivity claims for generative AI should subtract repair labor: measured speed gains from tool use may be offset by the time spent re-prompting, editing tone, and re-establishing context, so token or task counts alone overstate value.
- If context control becomes a first-class feature, it may reduce the advantage of expert prompt crafters, flattening the skill gradient among users and shifting skill demands toward domain verification rather than prompting.
- A testable extension would log prompt-repair sequences in a RAG-based engineering tool and measure whether layout-aware document chunking reduces the rate of re-prompts compared with current chunking.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a qualitative interview study of 17 hardware and software engineers at a multinational integrated circuit (IC) design firm, all of whom were intensive users of internally developed generative AI (GenAI) tools, defined as consuming more than 500,000 tokens in 30 days. The authors argue that accuracy concerns were secondary to other 'troubles' engineers encountered, most notably difficulties with document parsing, overly generic outputs, and failed numerical operations. They introduce 'trouble' as a category of interaction difficulty orthogonal to accuracy, map the identified forms of trouble onto three elements of GenAI systems (pipeline, features, and grounding), and conclude that controlling the context of human-GenAI interaction is one of the largest challenges in high-precision engineering work. The paper closes with organizational, transparency, and technical recommendations for mitigating context-related trouble.
Significance. If the central finding holds, the paper provides a valuable counterpoint to the prevailing research focus on GenAI accuracy and hallucinations, suggesting that for professional users in high-precision domains, context-level usability may matter more than raw output accuracy. The study is one of the few qualitative investigations of GenAI use in IC design, a domain of clear practical importance. The paper's strengths include a transparently reported interview protocol (Appendix A), detailed thematic counts in Tables 1 and 2, and a conceptually grounded link to Ackerman's socio-technical gap. The authors also appropriately note that generalization to other high-precision domains requires further comparative analysis. The contribution is significant for CSCW and HCI audiences concerned with the real-world deployment of GenAI in professional work.
major comments (3)
- [Section 3] The recruitment criterion of 'intensive users' (more than 500,000 tokens in 30 days) introduces a selection effect that directly bears on the paper's central claim about accuracy. By construction, the sample consists of engineers who already found the internal GenAI tools sufficiently useful to sustain heavy use over a 30-day period. Engineers for whom inaccuracies or other troubles were severe enough to reduce or abandon use are excluded. Consequently, the finding that accuracy was not a barrier to use (Section 4.1) and the conclusion that context troubles are the largest challenge (Section 5) are, at most, claims about heavy adopters, not about IC designers generally. The paper should either explicitly restrict its conclusions to this subpopulation or include a comparison group of light users, non-adopters, or trial-then-abandon users to support population-level claims. This issue is load-bearing for the paper's main contribution and needs to be addressed in revision.
- [Section 4.1] The statement that 'engineers’ concerns about other issues eclipsed those about accuracy, without exception' overstates the evidence presented. The paper reports that only 7 of 17 interviewees expressed opinions about the accuracy or inaccuracy of GenAI, and that none of these 7 saw accuracy as a barrier. That supports a claim that accuracy was not a dominant concern among those who mentioned it, but it does not demonstrate that all 17 participants ranked other issues above accuracy. The interview protocol (Appendix A, Q3) included explicit accuracy questions, so it would be possible to report how many participants, when prompted, discussed accuracy versus other troubles. Without such a systematic comparison, 'without exception' is an unsupported generalization. I recommend revising this sentence to reflect the actual distribution of responses, for example by reporting the number of participants who spontaneously raised accuracy relative to other trouble categories.
- [Table 3 and Section 5] The mapping of trouble types onto pipeline, features, and grounding elements in Table 3 is a purely analytic construction. While the paper acknowledges this in Section 5, the map is then used as a foundation for the conclusion that 'controlling the context of interactions is one of the largest challenges' and for the technical and organizational recommendations in Section 6. No inter-rater reliability, member checking, or independent validation of this mapping is reported, so it remains an interpretive hypothesis rather than an empirical finding. The paper should present Table 3 more explicitly as an analytical framework for future research, and temper claims that depend on the mapping's validity. Additionally, the counts in Table 2 are treated as indicators of the prevalence of trouble types; given the small sample (n=17) and the subjective thematic coding, these counts should not be read as a quantitative ranking. The text would benefit from a clearer statement of how the authors intend these counts to be interpreted by readers.
minor comments (5)
- [Section 4, introductory paragraph] The phrase 'needing repair and covery' appears to be a typographical error for 'needing repair and recovery.'
- [Section 1] The sentence 'indiscriminate use of GenAI tools in software development teams can increase the number of software flaws as it increase software programmers’ coding speed' contains a grammatical error ('increase' should be 'increases').
- [Section 3] The word 'useage' should be 'usage' in the description of company-specific safety features.
- [Section 5.1] In discussing RAG, the text says 'the orginal GPT would not have had access to'; 'orginal' should be 'original.'
- [Section 5.3] The footnote reference to [100] is clear, but the inline citation 'Banerjee et al. point out' does not appear in the reference list as an author name in the text; consider consistent author-date formatting.
Circularity Check
No load-bearing circularity: the empirical finding rests on interview coding, not on a definitional identity or self-citation chain.
full rationale
The paper's central claim—that context trouble outweighs accuracy concerns—is an empirical result from 17 interviews, coded into categories in Tables 1 and 2. The 'trouble' construct is broad ('any difficulty or inconvenience users experience ... such that they ... must edit or alter the output'), but the definition precedes the coding, and the reported counts (parsing tables n=11, too generic n=9, numerical operations n=6) are not entailed by that definition. The conclusion that most trouble is context-related is an analytic summary of observed codes, not a tautology. The treatment of inaccuracy as a separate category is an analytic choice, yet the paper still reports accuracy comments (n=7 engineers) and does not hide them. The heavy-user sampling criterion (>500,000 tokens/30 days) creates a genuine generalizability limitation: the conclusion about accuracy being non-blocking may not extend to non-adopters, and this is an external-validity concern, not circularity. The paper itself flags scope limits in Section 1 and Section 5. The one self-citation (Moss 2021, reference [61]) supports a background claim about AI evaluation and is not load-bearing. No step in the derivation reduces by construction to its own input.
Assumptions & free parameters
free parameters (2)
- Usage threshold for 'intensive user' =
>500,000 tokens in 30 days
- Sample size (n=17) =
17 interviews
assumptions (3)
- domain assumption All GenAI tools studied are described as 'comparable to off-the-shelf GenAI implementations' despite detailed technical specs being withheld.
- domain assumption Thematic counts from interviews provide meaningful evidence about the relative prevalence of troubles.
- domain assumption Engineers' self-reports accurately represent their actual interactions with the tools.
invented entities (1)
-
'Trouble' as a defined category of GenAI interaction difficulty
Cite this review
Pith. "Pith review of Controlling Context: Generative AI at Work in Integrated Circuit Design and Other High-Precision Domains." pith.science (2026). https://pith.science/paper/ZUPEYXA3
@misc{pith2026250614567,
author = {Pith},
title = {Pith review of: Controlling Context: Generative AI at Work in Integrated Circuit Design and Other High-Precision Domains},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZUPEYXA3}},
note = {Machine review of arXiv:2506.14567}
}
read the original abstract
Generative AI tools have become more prevalent in engineering workflows, particularly through chatbots and code assistants. As the perceived accuracy of these tools improves, questions arise about whether and how those who work in high-precision domains might maintain vigilance for errors, and what other aspects of using such tools might trouble their work. This paper analyzes interviews with hardware and software engineers, and their collaborators, who work in integrated circuit design to identify the role accuracy plays in their use of generative AI tools and what other forms of trouble they face in using such tools. The paper inventories these forms of trouble, which are then mapped to elements of generative AI systems, to conclude that controlling the context of interactions between engineers and the generative AI tools is one of the largest challenges they face. The paper concludes with recommendations for mitigating this form of trouble by increasing the ability to control context interactively.
Reference graph
Works this paper leans on
-
[1]
Introduction to VLSI systems
Carver Mead and Lynn Conway. Introduction to VLSI systems . 1980
1980
-
[2]
Accessing lexical ambiguities during sentence comprehension: Effects of frequency of meaning and contextual bias
William Onifer and David A Swinney. “Accessing lexical ambiguities during sentence comprehension: Effects of frequency of meaning and contextual bias”. In: Memory & Cognition 9 (1981), pp. 225–236
1981
-
[3]
User approaches to computer-supported teams
Robert Johansen. User approaches to computer-supported teams . Tech. rep. CISR WP. 155. Cambridge, MA: MIT Institute for the Future, 1987. url: https://dspace.mit.edu/bitstream/handle/1721.1/49359/userapproachesto00joha.pdf
1987
-
[4]
Plans and situated actions: The problem of human-machine communication
Lucille Alice Suchman. Plans and situated actions: The problem of human-machine communication . Cambridge university press, 1987
1987
-
[5]
The articulation of project work: An organizational process
Anselm Strauss. “The articulation of project work: An organizational process”. In: Sociological Quarterly 29.2 (1988), pp. 163–178
1988
-
[6]
The symbol grounding problem
Stevan Harnad. “The symbol grounding problem”. In: Physica D: Nonlinear Phenomena 42.1-3 (1990), pp. 335–346
1990
-
[7]
Engineering in history
Richard Shelton Kirby. Engineering in history. Courier Corporation, 1990
1990
-
[8]
What engineers know and how they know it
WG Vincenti. What engineers know and how they know it . 1990
1990
Show all 130 references
-
[9]
Saussure: signs, system and arbitrariness
David Holdcroft. Saussure: signs, system and arbitrariness . Cambridge University Press, 1991
1991
-
[10]
Awareness and coordination in shared workspaces
Paul Dourish and Victoria Bellotti. “Awareness and coordination in shared workspaces”. In: Proceedings of the 1992 ACM conference on Computer- supported cooperative work. 1992, pp. 107–114
1992
-
[11]
Taking CSCW seriously: Supporting articulation work
Kjeld Schmidt and Liam Bannon. “Taking CSCW seriously: Supporting articulation work”. In: Computer supported cooperative work (CSCW) 1 (1992), pp. 7–40
1992
-
[12]
Characteristics and models of iteration in engineering design
Robert P Smith and Steven P Eppinger. Characteristics and models of iteration in engineering design . Citeseer, 1993
1993
-
[13]
Integrated circuit, hybrid, and multichip module package design guidelines: a focus on reliability
Michael Pecht. Integrated circuit, hybrid, and multichip module package design guidelines: a focus on reliability . John Wiley & Sons, 1994
1994
-
[14]
Shaping electronic communication: The metastructuring of technology in the context of use
Wanda J Orlikowski et al. “Shaping electronic communication: The metastructuring of technology in the context of use”. In: Organization science 6.4 (1995), pp. 423–444
1995
-
[15]
Handbook of software reliability engineering
Michael R Lyu et al. Handbook of software reliability engineering . Vol. 222. IEEE computer society press Los Alamitos, 1996
1996
-
[16]
Of bicycles, bakelites, and bulbs: Toward a theory of sociotechnical change
Wiebe E Bijker. Of bicycles, bakelites, and bulbs: Toward a theory of sociotechnical change . MIT press, 1997
1997
-
[17]
The Audit Society: Rituals of Verification
M Power. The Audit Society: Rituals of Verification . Oxford University Press, 1997
1997
-
[18]
On" technomethodology
Paul Dourish and Graham Button. “On" technomethodology": Foundational relationships between ethnomethodology and system design”. In: Human-computer interaction 13.4 (1998), pp. 395–432
1998
-
[19]
Actor-network theory—the market test
Michel Callon. “Actor-network theory—the market test”. In: The sociological review 47.S1 (1999), pp. 181–195
1999
-
[20]
The intellectual challenge of CSCW: the gap between social requirements and technical feasibility
Mark S Ackerman. “The intellectual challenge of CSCW: the gap between social requirements and technical feasibility”. In: Human–Computer Interaction 15.2-3 (2000), pp. 179–203
2000
-
[21]
The structure and value of modularity in software design
Kevin J Sullivan et al. “The structure and value of modularity in software design”. In: ACM SIGSOFT Software Engineering Notes 26.5 (2001), pp. 99–108
2001
-
[22]
Analyzing interview data: The development and evolution of a coding system
Cynthia Weston et al. “Analyzing interview data: The development and evolution of a coding system”. In: Qualitative sociology 24 (2001), pp. 381–400
2001
-
[23]
Decomposition of interdependent task group for concurrent engineering
Shi-Jie Gary Chen and Li Lin. “Decomposition of interdependent task group for concurrent engineering”. In: Computers & Industrial Engineering 44.3 (2003), pp. 435–459
2003
-
[24]
Using prosody to avoid ambiguity: Effects of speaker awareness and referential context
Jesse Snedeker and John Trueswell. “Using prosody to avoid ambiguity: Effects of speaker awareness and referential context”. In: Journal of Memory and language 48.1 (2003), pp. 103–130
2003
-
[25]
Bringing context to CSCW
MRS Borges et al. “Bringing context to CSCW”. In: 8th International Conference on Computer Supported Cooperative Work in Design . Vol. 2. IEEE. 2004, pp. 161–166
2004
-
[26]
What we talk about when we talk about context
Paul Dourish. “What we talk about when we talk about context”. In: Personal and ubiquitous computing 8 (2004), pp. 19–30. 20 Moss et al
2004
-
[27]
Advanced formal verification
Rolf Drechsler. Advanced formal verification. Springer, 2004
2004
-
[28]
Agile and iterative development: a manager’s guide
Craig Larman. Agile and iterative development: a manager’s guide . Addison-Wesley Professional, 2004
2004
-
[29]
Ambiguity resolution in sentence processing: The role of lexical and contextual information
Despina Papadopoulou and Harald Clahsen. “Ambiguity resolution in sentence processing: The role of lexical and contextual information”. In: Journal of Linguistics 42.1 (2006), pp. 109–138
2006
-
[30]
Thematic coding and categorizing
Graham R Gibbs. “Thematic coding and categorizing”. In: Analyzing qualitative data 703.38-56 (2007)
2007
-
[31]
Reassembling the social: An introduction to actor-network-theory
Bruno Latour. Reassembling the social: An introduction to actor-network-theory . Oup Oxford, 2007
2007
-
[32]
System software reliability
Hoang Pham. System software reliability. Springer Science & Business Media, 2007
2007
-
[33]
The Verilog® hardware description language
Donald Thomas and Philip Moorby. The Verilog® hardware description language. Springer Science & Business Media, 2008
2008
-
[34]
Nature’s queer performativity
Karen Barad. “Nature’s queer performativity”. In: Qui Parle: Critical Humanities and Social Sciences 19.2 (2011), pp. 121–158
2011
-
[35]
SoC HW/SW verification and validation
Chung-Yang Huang et al. “SoC HW/SW verification and validation”. In: 16th Asia and South Pacific Design Automation Conference (ASP-DAC 2011). IEEE. 2011, pp. 297–300
2011
-
[36]
Research methods in anthropology
H Russell Bernard. “Research methods in anthropology”. In: AltaMira (2012)
2012
-
[37]
Frames of meaning: The social construction of extraordinary science
Harry M Collins and Trevor J Pinch. Frames of meaning: The social construction of extraordinary science . Routledge, 2013
2013
-
[38]
What is qualitative interviewing? Bloomsbury Academic, 2013
Rosalind Edwards and Janet Holland. What is qualitative interviewing? Bloomsbury Academic, 2013
2013
-
[39]
Organizational social structures for software engineering
Damian A Tamburri, Patricia Lago, and Hans van Vliet. “Organizational social structures for software engineering”. In: ACM Computing Surveys (CSUR) 46.1 (2013), pp. 1–35
2013
-
[40]
Rethinking Repair
Steven J Jackson. “Rethinking Repair”. en. In: Media Technologies: Essays on Communication, Materiality, and Society . Ed. by Tarleton Gillespie, Pablo J. Boczkowski, and Kirsten A. Foot. Cambridge, MA: MIT Press, 2014
2014
-
[41]
How to see values in social computing: methods for studying values dimensions
Katie Shilton, Jes A Koepfler, and Kenneth R Fleischmann. “How to see values in social computing: methods for studying values dimensions”. In: Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing . 2014, pp. 426–435
2014
-
[42]
Expert vs. novice: Problem decomposition/recomposition in engineering design
Ting Song and Kurt Becker. “Expert vs. novice: Problem decomposition/recomposition in engineering design”. In: 2014 International Conference on interactive collaborative learning (ICL) . IEEE. 2014, pp. 181–190
2014
-
[43]
Anthropocene, capitalocene, plantationocene, chthulucene: Making kin
Donna Haraway. “Anthropocene, capitalocene, plantationocene, chthulucene: Making kin”. In: Environmental humanities 6.1 (2015), pp. 159–165
2015
-
[44]
CSCW and social computing: the past and the future
Michael Koch, Gerhard Schwabe, and Robert O Briggs. CSCW and social computing: the past and the future . 2015
2015
-
[45]
Qualitative research in digital environments: A research toolkit
Alessandro Caliandro and Alessandro Gandini. Qualitative research in digital environments: A research toolkit . Routledge, 2016
2016
-
[46]
Staying with the trouble: Making kin in the Chthulucene
Donna J Haraway. “Staying with the trouble: Making kin in the Chthulucene”. In: Staying with the Trouble. Duke University Press, 2016
2016
-
[47]
Deep reinforcement learning from human preferences
Paul F Christiano et al. “Deep reinforcement learning from human preferences”. In: Advances in neural information processing systems 30 (2017)
2017
-
[48]
Attention Is All You Need
Ashish Vaswani et al. Attention Is All You Need . en. arXiv:1706.03762 [cs]. Dec. 2017. url: http://arxiv.org/abs/1706.03762 (visited on 07/15/2022)
2017 arXiv
-
[49]
Performance metrics (error measures) in machine learning regression, forecasting and prognostics: Properties and typology
Alexei Botchkarev. “Performance metrics (error measures) in machine learning regression, forecasting and prognostics: Properties and typology”. In: arXiv preprint arXiv:1809.03006 (2018)
2018 arXiv
-
[50]
Software development and CSCW: Standardization and flexibility in large-scale agile development
Helena Tendedez, Maria Angela MAF Ferrario, and Jon Whittle. “Software development and CSCW: Standardization and flexibility in large-scale agile development”. In: Proceedings of the ACM on Human-Computer Interaction 2.CSCW (2018), pp. 1–23
2018
-
[51]
Reliability seeking virtual organizations: Challenges for high reliability organizations and resilience engineering
Martha Grabowski and Karlene H Roberts. “Reliability seeking virtual organizations: Challenges for high reliability organizations and resilience engineering”. In: Safety science 117 (2019), pp. 512–522
2019
-
[52]
Ways of knowing when research subjects care
Dorothy Howard and Lilly Irani. “Ways of knowing when research subjects care”. In: Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 2019, pp. 1–16
2019
-
[53]
Climbing towards NLU: On meaning, form, and understanding in the age of data
Emily M Bender and Alexander Koller. “Climbing towards NLU: On meaning, form, and understanding in the age of data”. In: Proceedings of the 58th annual meeting of the association for computational linguistics . 2020, pp. 5185–5198
2020
-
[54]
Two bits: The cultural significance of free software
Christopher M Kelty. Two bits: The cultural significance of free software . Duke University Press, 2020
2020
-
[55]
Rising with the machines: A sociotechnical framework for bringing artificial intelligence into the organization
Erin E Makarius et al. “Rising with the machines: A sociotechnical framework for bringing artificial intelligence into the organization”. In: Journal of business research 120 (2020), pp. 262–273
2020
-
[56]
The who, what, how of software engineering research: a socio-technical framework
Margaret-Anne Storey et al. “The who, what, how of software engineering research: a socio-technical framework”. In: Empirical Software Engineering 25 (2020), pp. 4097–4129
2020
-
[57]
Adaptive folk theorization as a path to algorithmic literacy on changing platforms
Michael Ann DeVito. “Adaptive folk theorization as a path to algorithmic literacy on changing platforms”. In: Proceedings of the ACM on Human-Computer Interaction 5.CSCW2 (2021), pp. 1–38
2021
-
[58]
The labor of maintaining and scaling free and open-source software projects
R Stuart Geiger, Dorothy Howard, and Lilly Irani. “The labor of maintaining and scaling free and open-source software projects”. In: Proceedings of the ACM on human-computer interaction 5.CSCW1 (2021), pp. 1–28
2021
-
[59]
Cognition, context, and learning: A social semiotic perspective
Jay L Lemke. “Cognition, context, and learning: A social semiotic perspective”. In: Situated cognition. Routledge, 2021, pp. 37–55
2021
-
[60]
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.11401 [cs]. Apr. 2021. url: http://arxiv.org/abs/ 2005.11401 (visited on 10/30/2023)
2005 arXiv
-
[61]
The objective function: Science and society in the age of machine intelligence
Emanuel D Moss. “The objective function: Science and society in the age of machine intelligence”. PhD thesis. City University of New York, 2021
2021
-
[62]
Artificial intelligence and the world of work, a co-constitutive relationship
Carsten Østerlund et al. “Artificial intelligence and the world of work, a co-constitutive relationship”. In:Journal of the Association for Information Science and Technology 72.1 (2021), pp. 128–135
2021
-
[63]
Retrieval augmentation reduces hallucination in conversation
Kurt Shuster et al. “Retrieval augmentation reduces hallucination in conversation”. In: arXiv preprint arXiv:2104.07567 (2021)
2021 arXiv
-
[64]
Comparative analysis of image classification algorithms based on traditional machine learning and deep learning
Pin Wang, En Fan, and Peng Wang. “Comparative analysis of image classification algorithms based on traditional machine learning and deep learning”. In: Pattern recognition letters 141 (2021), pp. 61–67. Controlling Context: Generative AI at Work in Integrated Circuit Design an...
2021
-
[65]
A primer on AI in/from the Majority World: An Empirical Site and a Standpoint
Sareeta Amrute, Ranjit Singh, and Rigoberto Lara Guzmán. “A primer on AI in/from the Majority World: An Empirical Site and a Standpoint”. In: A vailable at SSRN 4199467 (2022)
2022
-
[66]
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai et al. “Constitutional ai: Harmlessness from ai feedback”. In: arXiv preprint arXiv:2212.08073 (2022)
2022 arXiv
-
[67]
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai et al. “Training a helpful and harmless assistant with reinforcement learning from human feedback”. In:arXiv preprint arXiv:2204.05862 (2022)
2022 arXiv
-
[68]
Recent Advances in Retrieval-Augmented Text Generation
Deng Cai et al. “Recent Advances in Retrieval-Augmented Text Generation”. en. In:Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . Madrid Spain: ACM, July 2022, pp. 3417–3419. isbn: 978-1-4503-8732-3. doi: 10.1145...
2022
-
[69]
AI-based conversational agents: a scoping review from technologies to future directions
Sheetal Kusal et al. “AI-based conversational agents: a scoping review from technologies to future directions”. In:IEEE Access 10 (2022), pp. 92337– 92356
2022
-
[70]
Introducing ChatGPT
OpenAI. Introducing ChatGPT. en-US. Nov. 2022. url: https://openai.com/index/chatgpt/ (visited on 08/06/2024)
2022
-
[71]
Reliance on metrics is a fundamental challenge for AI
Rachel L Thomas and David Uminsky. “Reliance on metrics is a fundamental challenge for AI”. In: Patterns 3.5 (2022)
2022
-
[72]
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
Yuxin Xiao et al. “Uncertainty quantification with pre-trained language models: A large-scale empirical analysis”. In:arXiv preprint arXiv:2210.04714 (2022)
2022 arXiv
-
[74]
AI/ML algorithms and applications in VLSI design and technology
Deepthi Amuru et al. “AI/ML algorithms and applications in VLSI design and technology”. In: Integration 93 (2023), p. 102048
2023
-
[75]
Erik Brynjolfsson, Danielle Li, and Lindsey Raymond.Generative AI at Work. en. Tech. rep. w31161. Cambridge, MA: National Bureau of Economic Research, Apr. 2023, w31161. doi: 10.3386/w31161. url: http://www.nber.org/papers/w31161.pdf (visited on 10/29/2024)
2023 doi
-
[76]
Exploring the intersection of Generative AI and Software Development
Filipe Calegario et al. “Exploring the intersection of Generative AI and Software Development”. In: arXiv preprint arXiv:2312.14262 (2023)
2023 arXiv
-
[77]
Unleashing the potential of prompt engineering in Large Language Models: a comprehensive review
Banghao Chen et al. “Unleashing the potential of prompt engineering in Large Language Models: a comprehensive review”. In: arXiv preprint arXiv:2310.14735 (2023)
2023 arXiv
-
[78]
Generative AI for software practitioners
Christof Ebert and Panos Louridas. “Generative AI for software practitioners”. In: IEEE Software 40.4 (2023), pp. 30–38
2023
-
[79]
AI Is Reshaping Chip Design. But Where Will It End?
Karl Freund. “AI Is Reshaping Chip Design. But Where Will It End?” In: Forbes (Dec. 2023). url: https://www.forbes.com/sites/karlfreund/2023/ 12/19/ai-is-reshaping-chip-design-but-where-will-it-end/ (visited on 08/12/2024)
2023
-
[80]
Handling bias in toxic speech detection: A survey
Tanmay Garg et al. “Handling bias in toxic speech detection: A survey”. In: ACM Computing Surveys 55.13s (2023), pp. 1–32
2023
-
[81]
Look before you leap: An exploratory study of uncertainty measurement for large language models
Yuheng Huang et al. “Look before you leap: An exploratory study of uncertainty measurement for large language models”. In: arXiv preprint arXiv:2307.10236 (2023)
2023 arXiv
-
[82]
Survey of hallucination in natural language generation
Ziwei Ji et al. “Survey of hallucination in natural language generation”. In: ACM Computing Surveys 55.12 (2023), pp. 1–38
2023
-
[83]
Active Retrieval Augmented Generation
Zhengbao Jiang et al. Active Retrieval Augmented Generation. en. arXiv:2305.06983 [cs]. Oct. 2023. url: http://arxiv.org/abs/2305.06983 (visited on 01/23/2024)
2023 arXiv
-
[84]
Assessing the accuracy and reliability of AI-generated medical responses: an evaluation of the Chat-GPT model
Douglas Johnson et al. “Assessing the accuracy and reliability of AI-generated medical responses: an evaluation of the Chat-GPT model”. In: Research square (2023)
2023
-
[85]
Holistic Evaluation of Language Models
Percy Liang et al. Holistic Evaluation of Language Models . en. arXiv:2211.09110 [cs]. Oct. 2023. url: http://arxiv.org/abs/2211.09110 (visited on 01/23/2024)
2023 arXiv
-
[86]
CircuitOps: An ML Infrastructure Enabling Generative AI for VLSI Circuit Optimization
Rongjian Liang et al. “CircuitOps: An ML Infrastructure Enabling Generative AI for VLSI Circuit Optimization”. In: 2023 IEEE/ACM International Conference on Computer Aided Design (ICCAD) . IEEE. 2023, pp. 1–6
2023
-
[87]
In Race for AI Chips, Google DeepMind Uses AI to Design Specialized Semiconductors - WSJ
Belle Lin. “In Race for AI Chips, Google DeepMind Uses AI to Design Specialized Semiconductors - WSJ”. In: Wall Street Journal (July 2023). url: https://www.wsj.com/articles/in-race-for-ai-chips-google-deepmind-uses-ai-to-design-specialized-semiconductors-dcd78967 (visited on ...
2023
-
[88]
Chipnemo: Domain-adapted llms for chip design
Mingjie Liu et al. “Chipnemo: Domain-adapted llms for chip design”. In: arXiv preprint arXiv:2311.00176 (2023)
2023 arXiv
-
[89]
A review of evaluation metrics in machine learning algorithms
Gireen Naidu, Tranos Zuva, and Elias Mmbongeni Sibanda. “A review of evaluation metrics in machine learning algorithms”. In: Computer Science On-line Conference. Springer. 2023, pp. 15–25
2023
-
[90]
Experimental evidence on the productivity effects of generative artificial intelligence
Shakked Noy and Whitney Zhang. “Experimental evidence on the productivity effects of generative artificial intelligence”. In: Science 381.6654 (2023), pp. 187–192
2023
-
[91]
The potential of generative artificial intelligence across disciplines: Perspectives and future directions
Keng-Boon Ooi et al. “The potential of generative artificial intelligence across disciplines: Perspectives and future directions”. In: Journal of Computer Information Systems (2023), pp. 1–32
2023
-
[92]
Do users write more insecure code with AI assistants?
Neil Perry et al. “Do users write more insecure code with AI assistants?” In: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. 2023, pp. 2785–2799
2023
-
[93]
A Qualitative Study That Explores the Implementation of Artificial Intelligence in Integrated Circuit Design
Danny Rittman. “A Qualitative Study That Explores the Implementation of Artificial Intelligence in Integrated Circuit Design”. PhD thesis. Colorado Technical University, 2023
2023
-
[94]
ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?
Jürgen Rudolph, Samson Tan, and Shannon Tan. “ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?” In: Journal of applied learning and teaching 6.1 (2023), pp. 342–363
2023
-
[95]
Instruction tuning for large language models: A survey
Shengyu Zhang et al. “Instruction tuning for large language models: A survey”. In: arXiv preprint arXiv:2308.10792 (2023)
2023
-
[96]
To believe or not to believe your llm: Iterative prompting for estimating epistemic uncertainty
Yasin Abbasi Yadkori et al. “To believe or not to believe your llm: Iterative prompting for estimating epistemic uncertainty”. In:Advances in Neural Information Processing Systems 37 (2024), pp. 58077–58117. 22 Moss et al
2024
-
[97]
A Critical Analysis of the Largest Source for Generative AI Training Data: Common Crawl
Stefan Baack. “A Critical Analysis of the Largest Source for Generative AI Training Data: Common Crawl”. en. In: The 2024 ACM Conference on Fairness, Accountability, and Transparency. Rio de Janeiro Brazil: ACM, June 2024, pp. 2199–2208. isbn: 9798400704505. doi: 10.1145/36301...
2024
-
[98]
A Critical Analysis of the Largest Source for Generative AI Training Data: Common Crawl
Stefan Baack. “A Critical Analysis of the Largest Source for Generative AI Training Data: Common Crawl”. In: The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2024, pp. 2199–2208
2024
-
[99]
Integrating Generative AI for Advancing Agile Software Development and Mitigating Project Management Challenges
Anas Bahi, Jihane GHARI, and Youssef Gahi. “Integrating Generative AI for Advancing Agile Software Development and Mitigating Project Management Challenges.” In: International Journal of Advanced Computer Science & Applications 15.3 (2024)
2024
-
[100]
LLMs Will Always Hallucinate, and We Need to Live With This
Sourav Banerjee, Ayushi Agarwal, and Saloni Singla. “LLMs Will Always Hallucinate, and We Need to Live With This”. In: arXiv preprint arXiv:2409.05746 (2024)
2024 arXiv
-
[101]
Middle Tech: Software Work and the Culture of Good Enough
Paula Bialski. Middle Tech: Software Work and the Culture of Good Enough . Vol. 36. Princeton University Press, 2024
2024
-
[102]
Middle Tech: Software Work and the Culture of Good Enough
Paula Bialski. Middle Tech: Software Work and the Culture of Good Enough . en. Google-Books-ID: 3h3lEAAAQBAJ. Princeton University Press, May 2024. isbn: 978-0-691-25716-7
2024
-
[103]
Evaluating LLMs for Hardware Design and Test
Jason Blocklove et al. “Evaluating LLMs for Hardware Design and Test”. In: arXiv preprint arXiv:2405.02326 (2024)
2024 arXiv
-
[104]
Foundation Model Transparency Reports
Rishi Bommasani et al. “Foundation Model Transparency Reports”. In: arXiv preprint arXiv:2402.16268 (2024)
2024 arXiv
-
[105]
Design principles for collaborative generative AI systems in software development
Johannes Chen and Jan Zacharias. “Design principles for collaborative generative AI systems in software development”. In: International Conference on Design Science Research in Information Systems and Technology . Springer. 2024, pp. 341–354
2024
-
[106]
Uncertainty in Visual Generative AI
Kara Combs, Adam Moyer, and Trevor J Bihl. “Uncertainty in Visual Generative AI”. In: Algorithms 17.4 (2024), p. 136
2024
-
[107]
The Role of Generative AI in Software Development Productivity: A Pilot Case Study
Mariana Coutinho et al. “The Role of Generative AI in Software Development Productivity: A Pilot Case Study”. In: Proceedings of the 1st ACM International Conference on AI-Powered Software . 2024, pp. 131–138
2024
-
[108]
Generative AI enhances individual creativity but reduces the collective diversity of novel content
Anil R Doshi and Oliver P Hauser. “Generative AI enhances individual creativity but reduces the collective diversity of novel content”. In:Science Advances 10.28 (2024), eadn5290
2024
-
[109]
Generative ai
Stefan Feuerriegel et al. “Generative ai”. In: Business & Information Systems Engineering 66.1 (2024), pp. 111–126
2024
-
[110]
The importance of generalizability in machine learning for systems
Varun Gohil et al. “The importance of generalizability in machine learning for systems”. In: IEEE Computer Architecture Letters (2024)
2024
-
[111]
When teams embrace AI: human collaboration strategies in generative prompting in a creative design task
Yuanning Han et al. “When teams embrace AI: human collaboration strategies in generative prompting in a creative design task”. In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . 2024, pp. 1–14
2024
-
[112]
ChatGPT is bullshit
Michael Townsen Hicks, James Humphries, and Joe Slater. “ChatGPT is bullshit”. In: Ethics and Information Technology 26.2 (2024), p. 38
2024
-
[113]
A survey of uncertainty estimation in llms: Theory meets practice
Hsiu-Yuan Huang et al. “A survey of uncertainty estimation in llms: Theory meets practice”. In: arXiv preprint arXiv:2410.15326 (2024)
2024 arXiv
-
[114]
Adaptive Uncertainty Quantification for Generative AI
Jungeum Kim, Sean O’Hagan, and Veronika Rockova. “Adaptive Uncertainty Quantification for Generative AI”. In:arXiv preprint arXiv:2408.08990 (2024)
2024 arXiv
-
[115]
" I’m Not Sure, But
Sunnie SY Kim et al. “" I’m Not Sure, But... ": Examining the Impact of Large Language Models’ Uncertainty Expression on User Reliance and Trust”. In: The 2024 ACM Conference on Fairness, Accountability, and Transparency . 2024, pp. 822–835
2024
-
[116]
Inadequacies of large language model benchmarks in the era of generative artificial intelligence
Timothy R McIntosh et al. “Inadequacies of large language model benchmarks in the era of generative artificial intelligence”. In: arXiv preprint arXiv:2402.09880 (2024)
2024 arXiv
-
[117]
Reading between the lines: Modeling user behavior and costs in AI-assisted programming
Hussein Mozannar et al. “Reading between the lines: Modeling user behavior and costs in AI-assisted programming”. In: Proceedings of the CHI Conference on Human Factors in Computing Systems . 2024, pp. 1–16
2024
-
[118]
GPT-4 Technical Report
OpenAI et al. GPT-4 Technical Report. arXiv:2303.08774 [cs]. Mar. 2024. url: http://arxiv.org/abs/2303.08774 (visited on 03/18/2024)
2024 arXiv
-
[119]
Navigating the complexity of generative ai adoption in software engineering
Daniel Russo. “Navigating the complexity of generative ai adoption in software engineering”. In: ACM Transactions on Software Engineering and Methodology (2024)
2024
-
[120]
SPADE: Synthesizing Data Quality Assertions for Large Language Model Pipelines
Shreya Shankar et al. SPADE: Synthesizing Data Quality Assertions for Large Language Model Pipelines . en. arXiv:2401.03038 [cs]. Mar. 2024. url: http://arxiv.org/abs/2401.03038 (visited on 07/16/2024)
2024 arXiv
-
[121]
How generative AI can boost highly skilled workers’ productivity
Meredith Somers. “How generative AI can boost highly skilled workers’ productivity”. In: Ideas Made to Matter - MIT Sloan School of Management (Oct. 2024). url: https://mitsloan.mit.edu/ideas-made-to-matter/how-generative-ai-can-boost-highly-skilled-workers-productivity (visit...
2024
-
[122]
Content-Centric Prototyping of Generative AI Applications: Emerging Approaches and Challenges in Collaborative Software Teams
Hari Subramonyam et al. “Content-Centric Prototyping of Generative AI Applications: Emerging Approaches and Challenges in Collaborative Software Teams”. In: arXiv preprint arXiv:2402.17721 (2024)
2024 arXiv
-
[123]
Meta-prompting: Enhancing language models with task-agnostic scaffolding
Mirac Suzgun and Adam Tauman Kalai. “Meta-prompting: Enhancing language models with task-agnostic scaffolding”. In: arXiv preprint arXiv:2401.12954 (2024)
2024 arXiv
-
[124]
Transforming software development with generative AI: empirical insights on collaboration and workflow
Rasmus Ulfsnes et al. “Transforming software development with generative AI: empirical insights on collaboration and workflow”. In: Generative AI for effective software development . Springer, 2024, pp. 219–234
2024
-
[125]
Fine-tuning GPT-3 for machine learning electronic and functional properties of organic molecules
Zikai Xie et al. “Fine-tuning GPT-3 for machine learning electronic and functional properties of organic molecules”. In: Chemical science 15.2 (2024), pp. 500–510
2024
-
[126]
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Ce Zhou et al. “A comprehensive survey on pretrained foundation models: A history from bert to chatgpt”. In: International Journal of Machine Learning and Cybernetics (2024), pp. 1–65
2024
-
[127]
Toolqa: A dataset for llm question answering with external tools
Yuchen Zhuang et al. “Toolqa: A dataset for llm question answering with external tools”. In: Advances in Neural Information Processing Systems 36 (2024). Controlling Context: Generative AI at Work in Integrated Circuit Design and Other High-Precision Domains 23
2024
-
[128]
The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers
Hao-Ping (Hank) Lee et al. “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers”. In: Proceedings of the ACM CHI Conference on Human Factors in Computing Systems . ACM, Apr
-
[129]
FIT-RAG: Black-Box RAG with Factual Information and Token Reduction
Yuren Mao et al. “FIT-RAG: Black-Box RAG with Factual Information and Token Reduction”. In: ACM Trans. Inf. Syst. 43.2 (Jan. 2025). issn: 1046-8188. doi: 10.1145/3676957. url: https://doi.org/10.1145/3676957
2025 doi
-
[130]
Reflection, Confabulation, and Reasoning
Jennifer Nagel. “Reflection, Confabulation, and Reasoning”. In:Kornblith and His Critics. Ed. by Luis Oliveira and Joshua DiPaolo. Wiley-Blackwell, forthcoming. A Interview Protocol (1) Context-setting questions: • How would you describe your role at the company? • What softwa...
-
[2025]
url: https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions- in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.