{"id":"50194bc9-4cd1-4c80-b6f5-774b0a59def1","arxiv_id":"2506.14567","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Qualitative interviews with 17 IC-design engineers show that the main difficulty with generative AI tools is not output accuracy but context control, supporting a shift toward interactive context-steering features.","lead":"This paper interviews 17 hardware and software engineers at a large chip company who use internal generative AI chatbots, and finds that accuracy is not their biggest concern; a broader category of 'trouble' from mismatched context matters more. The paper maps these troubles onto parts of AI systems and argues that engineers need more interactive control over the context the AI works in.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sampling only heavy users may make accuracy's absence as a barrier an artifact: engineers who already tolerate GenAI inaccuracy are the ones studied.","rationale":"The reader's weakest_assumption identified sample generalizability, and I agree this is the load-bearing issue. The paper is a careful qualitative study with a transparent protocol and thoughtful analysis; the trouble taxonomy is useful. However, the central empirical claim that accuracy is secondary depends entirely on who was sampled. The 500,000-token threshold is an extreme inclusion criterion that selects for users who have already overcome any accuracy-related adoption barrier. The claim is thus at risk of circularity: heavy use is both the selection criterion and the evidence that accuracy does not block use. The paper acknowledges limits on domain generalization but does not address this within-domain selection effect. A comparison group of light or non-users would settle whether accuracy is genuinely secondary for IC designers or merely for those who already tolerate inaccuracy. Other potential concerns—such as absence of inter-rater reliability or the use of mention counts as a proxy for importance—are secondary and would not change the conditional verdict. Therefore I recommend no change to the reader's CONDITIONAL verdict.","tokens_in":22352,"tokens_out":5517,"duration_ms":57747,"concrete_test":"Conduct the same interview protocol with two additional groups at the same firm: (a) engineers who used the internal GenAI tools but stopped or use them rarely (e.g., under 10,000 tokens in 30 days), and (b) engineers who have not used the tools. If accuracy emerges as a primary stated barrier in these groups, the heavy-user sample artifact is confirmed. If not, the generalizability concern is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that accuracy concerns are eclipsed by context troubles and that controlling context is the largest challenge—rests on interviews with 17 'intensive users' defined as consuming more than 500,000 tokens in 30 days (Section 3). This is a survivorship-biased sample: by construction, these engineers have already adopted the tool heavily, so they are precisely the people for whom accuracy has not been a sufficient barrier to stop use. The finding that none of them sees accuracy as a barrier (Section 4.1) is therefore unsurprising and cannot support the general statement that accuracy is not a significant source of trouble for IC designers. The paper acknowledges that extrapolation to other high-precision domains needs comparative analysis (Section 1), but it does not acknowledge that the recruitment criterion itself excludes the very users for whom accuracy might be decisive—light users, trial-then-abandon users, or non-adopters. Without such a comparison group, the ordering of 'trouble' categories may reflect the repair practices these heavy users have developed (atomizing, iteration, making explicit, Section 4.3) rather than the population-level importance of accuracy versus context.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a qualitative interview study of 17 hardware and software engineers at a multinational integrated circuit (IC) design firm, all of whom were intensive users of internally developed generative AI (GenAI) tools, defined as consuming more than 500,000 tokens in 30 days. The authors argue that accuracy concerns were secondary to other 'troubles' engineers encountered, most notably difficulties with document parsing, overly generic outputs, and failed numerical operations. They introduce 'trouble' as a category of interaction difficulty orthogonal to accuracy, map the identified forms of trouble onto three elements of GenAI systems (pipeline, features, and grounding), and conclude that controlling the context of human-GenAI interaction is one of the largest challenges in high-precision engineering work. The paper closes with organizational, transparency, and technical recommendations for mitigating context-related trouble.","tokens_in":22532,"tokens_out":3942,"duration_ms":43418,"significance":"If the central finding holds, the paper provides a valuable counterpoint to the prevailing research focus on GenAI accuracy and hallucinations, suggesting that for professional users in high-precision domains, context-level usability may matter more than raw output accuracy. The study is one of the few qualitative investigations of GenAI use in IC design, a domain of clear practical importance. The paper's strengths include a transparently reported interview protocol (Appendix A), detailed thematic counts in Tables 1 and 2, and a conceptually grounded link to Ackerman's socio-technical gap. The authors also appropriately note that generalization to other high-precision domains requires further comparative analysis. The contribution is significant for CSCW and HCI audiences concerned with the real-world deployment of GenAI in professional work.","major_comments":[{"comment":"The recruitment criterion of 'intensive users' (more than 500,000 tokens in 30 days) introduces a selection effect that directly bears on the paper's central claim about accuracy. By construction, the sample consists of engineers who already found the internal GenAI tools sufficiently useful to sustain heavy use over a 30-day period. Engineers for whom inaccuracies or other troubles were severe enough to reduce or abandon use are excluded. Consequently, the finding that accuracy was not a barrier to use (Section 4.1) and the conclusion that context troubles are the largest challenge (Section 5) are, at most, claims about heavy adopters, not about IC designers generally. The paper should either explicitly restrict its conclusions to this subpopulation or include a comparison group of light users, non-adopters, or trial-then-abandon users to support population-level claims. This issue is load-bearing for the paper's main contribution and needs to be addressed in revision.","section":"Section 3"},{"comment":"The statement that 'engineers’ concerns about other issues eclipsed those about accuracy, without exception' overstates the evidence presented. The paper reports that only 7 of 17 interviewees expressed opinions about the accuracy or inaccuracy of GenAI, and that none of these 7 saw accuracy as a barrier. That supports a claim that accuracy was not a dominant concern among those who mentioned it, but it does not demonstrate that all 17 participants ranked other issues above accuracy. The interview protocol (Appendix A, Q3) included explicit accuracy questions, so it would be possible to report how many participants, when prompted, discussed accuracy versus other troubles. Without such a systematic comparison, 'without exception' is an unsupported generalization. I recommend revising this sentence to reflect the actual distribution of responses, for example by reporting the number of participants who spontaneously raised accuracy relative to other trouble categories.","section":"Section 4.1"},{"comment":"The mapping of trouble types onto pipeline, features, and grounding elements in Table 3 is a purely analytic construction. While the paper acknowledges this in Section 5, the map is then used as a foundation for the conclusion that 'controlling the context of interactions is one of the largest challenges' and for the technical and organizational recommendations in Section 6. No inter-rater reliability, member checking, or independent validation of this mapping is reported, so it remains an interpretive hypothesis rather than an empirical finding. The paper should present Table 3 more explicitly as an analytical framework for future research, and temper claims that depend on the mapping's validity. Additionally, the counts in Table 2 are treated as indicators of the prevalence of trouble types; given the small sample (n=17) and the subjective thematic coding, these counts should not be read as a quantitative ranking. The text would benefit from a clearer statement of how the authors intend these counts to be interpreted by readers.","section":"Table 3 and Section 5"}],"minor_comments":[{"comment":"The phrase 'needing repair and covery' appears to be a typographical error for 'needing repair and recovery.'","section":"Section 4, introductory paragraph"},{"comment":"The sentence 'indiscriminate use of GenAI tools in software development teams can increase the number of software flaws as it increase software programmers’ coding speed' contains a grammatical error ('increase' should be 'increases').","section":"Section 1"},{"comment":"The word 'useage' should be 'usage' in the description of company-specific safety features.","section":"Section 3"},{"comment":"In discussing RAG, the text says 'the orginal GPT would not have had access to'; 'orginal' should be 'original.'","section":"Section 5.1"},{"comment":"The footnote reference to [100] is clear, but the inline citation 'Banerjee et al. point out' does not appear in the reference list as an author name in the text; consider consistent author-date formatting.","section":"Section 5.3"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and practically important topic, and the qualitative data seem rich. The main concern is external validity: the heavy-user sampling and the overstatement of 'without exception' undermine the general conclusion about the relative unimportance of accuracy. These issues are fixable through careful reframing and additional transparency about the limits of the sample. I would encourage the editors to seek a revision rather than reject, as the empirical material is valuable and the conceptual framing of 'trouble' could be a useful contribution to CSCW once its scope is appropriately bounded."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful part: this is one of the first empirical studies of GenAI use in IC design, and the 'trouble' inventory is a genuinely useful counterweight to the field's accuracy obsession. The counts are presented as thematic indicators, not statistics, and the interview protocol is included. Mapping trouble onto pipeline, features, and grounding (Table 3) is a nice analytic device, even if it is a reasoning artifact. The paper also earns credit for acknowledging that extension to other domains needs comparative analysis.\n\nThe soft spots are real but not fatal. The stress-test concern holds: recruiting only intensive users (over 500k tokens in 30 days) selects for people who have already decided inaccuracy is tolerable. So the finding that accuracy is not a barrier says something about heavy adopters, not about IC designers at large. The paper's conclusion overstates slightly when it says accuracy 'is not a significant source of trouble' for designers who use such tools; a more precise statement would limit it to the studied subpopulation. A comparison group of light users or non-adopters would have strengthened the design. The 'trouble' construct also has a whiff of circularity—defined as anything needing repair, then coded as present—but in qualitative research with an explicit protocol, that is a minor issue. The mapping in Table 3 is plausible, but no inter-rater reliability or member checking is reported, so treat it as a hypothesis. The recommendations about uncertainty communication and RLHF are speculative, and the paper mostly says so.\n\nOverall, this is a solid qualitative contribution. It should go to peer review, not be desk-rejected. I would recommend the authors moderate the generalization in the title and conclusion, and add a sentence acknowledging the heavy-user selection effect. For HCI/CSCW readers interested in GenAI at work, it is a worthwhile read.","headline":"Worth a careful read for HCI/CSCW folks: a useful 'trouble' framing and early empirical data, but the heavy-user sampling undercuts the general claim that accuracy is secondary.","tokens_in":23099,"tokens_out":2357,"would_cite":true,"duration_ms":24311,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that for integrated circuit designers, generative AI's accuracy failures are secondary to context failures: engineers are more troubled by generic outputs and misparsed documents than by wrong answers, and the biggest…","keywords":["generative AI","integrated circuit design","qualitative interviews","human-computer interaction","sociotechnical systems","context control","accuracy","repair work"],"falsifier":"A comparable interview or survey study sampling light and non-users of generative AI in IC design that found accuracy failures—for example wrong simulation code slipping past review—were the most commonly cited reason for limiting or abandoning the tools would directly contradict the claim.","tokens_in":22122,"feed_emoji":"⚙️","tokens_out":6271,"duration_ms":62024,"temperature":0.7,"pith_summary":"This paper examines how 17 hardware and software engineers at one large integrated circuit design firm actually use internal generative AI chatbots. It tries to establish that accuracy is not the primary barrier to using these tools: engineers work inside verification and review systems that already catch errors, so they judge outputs 'good enough' and move on. What genuinely troubles them, the paper argues, is context — outputs that are too generic, documents parsed from the wrong table cell or footnote, company-specific terms mangled, and numerical failures. The central claim is that controlling the context of interactions between engineers and generative AI is one of the largest challenges they face, and that this challenge, not accuracy, should organize design and organizational response.","feed_headline":"For chip designers, context beats accuracy as AI's biggest hurdle","feed_subtitle":"Interviews with 17 engineers show the real rework comes from generic outputs and misparsed documents, not wrong answers.","key_machinery":"The analytical engine of the paper is the concept of 'trouble': any difficulty that forces a user to re-prompt, edit, or abandon an output. Trouble is treated as orthogonal to accuracy and is mapped onto three elements of a generative AI sociotechnical system — the pipeline (training data, tokenization, fine-tuning, and retrieval-augmented generation), the features (interface, metaprompts, conversation threading), and grounding (the inability of text tokens to carry situated meaning without human context). This mapping does the work of converting scattered interview complaints into design targets and recommendations.","core_discovery":"On the paper's own terms, the discovery is that concerns about accuracy are eclipsed by other problems 'without exception' among the engineers interviewed. Only seven of 17 commented on accuracy at all, and none treated it as a barrier to use. The most common troubles were systems using supplied documents out of context (n=11), outputs too generic to be useful (n=9), and failed numerical operations (n=6). These troubles are not user error or a need for training; they are manifestations of the gap between the general-purpose design of generative AI and the situated, organization-specific context of chip design. Engineers repair the gap by atomizing tasks, iterating, making context explicit in prompts, and relying on existing code review and verification workflows, and the paper concludes that the repair burden should be shifted back to the system through interactive control of context.","pith_inferences":["The trouble taxonomy likely transfers to other high-precision document-heavy domains such as medicine, law, and scientific research, where layout and local conventions carry much of the meaning; a comparative study could test this directly.","Productivity claims for generative AI should subtract repair labor: measured speed gains from tool use may be offset by the time spent re-prompting, editing tone, and re-establishing context, so token or task counts alone overstate value.","If context control becomes a first-class feature, it may reduce the advantage of expert prompt crafters, flattening the skill gradient among users and shifting skill demands toward domain verification rather than prompting.","A testable extension would log prompt-repair sequences in a RAG-based engineering tool and measure whether layout-aware document chunking reduces the rate of re-prompts compared with current chunking."],"forward_implications":["Improving generative AI for engineering should focus on interactive context control—letting users constrain documents, style, and conversation state—rather than on pushing benchmark accuracy scores higher.","Existing verification, code review, and unit testing remain load-bearing for safe GenAI use and may need to be strengthened, since engineers rely on them to absorb inaccuracies.","The most frequent trouble, document parsing, points to a concrete fix: RAG and file-input components should be made layout-aware so tables, footnotes, and adjacent cells are not misread.","Transparency features—showing metaprompts, persistent context, and uncertainty estimates—could reduce the guesswork engineers currently put into prompt repair.","Organizations should preserve pathways for novice engineers to build the judgment senior engineers used to supervise outputs, since most interviewees doubted novices could reliably catch wrong outputs."],"supporting_citations":[{"why":"Supplies the notion of a socio-technical gap that the paper uses to frame trouble as eruptions of mismatch between social requirements and technical feasibility.","marker":"[20]"},{"why":"Supplies 'articulation work,' which grounds the repair and recovery practices engineers use to re-situate GenAI outputs in their local work context.","marker":"[11]"},{"why":"Defines retrieval-augmented generation, the technique whose document-parsing failures account for the most common trouble reported.","marker":"[60]"},{"why":"Provides the survey treatment of hallucination that the paper draws on for one trouble category and the accuracy discourse it argues is secondary.","marker":"[82]"},{"why":"Marks the arrival of general-purpose GPT chatbots, the starting point for the general-purpose-versus-context mismatch motivating the study.","marker":"[70]"},{"why":"Formulates the symbol grounding problem, which underpins the 'grounding' element in the paper's mapping of trouble.","marker":"[6]"},{"why":"Describes reinforcement learning from human feedback, cited both as a pipeline source of generic preference shaping and as a target for domain-specific intervention.","marker":"[67]"},{"why":"Supplies the 'good enough' characterization of software culture that supports the claim engineers need usable outputs, not perfectly accurate ones.","marker":"[102]"}],"fun_headline_variants":["Chip design's AI headache: context loss, not accuracy gaps","Generic AI replies, not wrong ones, stall chip engineers","AI context errors beat accuracy as chip design's top foe","For chip engineers, AI's context blindness outweighs errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All interview evidence comes from 17 intensive users (over 500,000 tokens in 30 days) at one firm using internal tools only about six months old; if this early-adopter, heavy-use group tolerates inaccuracy better than typical engineers, the paper's 'accuracy is secondary' claim may not hold for chip designers broadly.","fun_headline_variants_meta":{"raw":{"variants":["Chip design's AI headache: context loss, not accuracy gaps","Generic AI replies, not wrong ones, stall chip engineers","AI context errors beat accuracy as chip design's top foe","For chip engineers, AI's context blindness outweighs errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000291,"raw_usage":{"total_tokens":1660,"prompt_tokens":865,"completion_tokens":795,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":726}},"tokens_in":481,"tokens_out":795,"duration_ms":8277,"temperature":1.0,"reasoning_tokens":726,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:50:08.432124+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A comparable interview or survey study sampling light and non-users of generative AI in IC design that found accuracy failures—for example wrong simulation code slipping past review—were the most commonly cited reason for limiting or abandoning the tools would directly contradict the claim.","supporting_citations":[{"cited_title":"Survey of hallucination in natural language generation","cited_arxiv_id":null,"evidence_quote":"Provides the survey treatment of hallucination that the paper draws on for one trouble category and the accuracy discourse it argues is secondary."},{"cited_title":"Introducing ChatGPT","cited_arxiv_id":null,"evidence_quote":"Marks the arrival of general-purpose GPT chatbots, the starting point for the general-purpose-versus-context mismatch motivating the study."},{"cited_title":"Middle Tech: Software Work and the Culture of Good Enough","cited_arxiv_id":null,"evidence_quote":"Supplies the 'good enough' characterization of software culture that supports the claim engineers need usable outputs, not perfectly accurate ones."}],"review_version":2}