{"id":"ae23de57-e5e0-4cbe-9763-3043bba4f27c","arxiv_id":"2412.15444","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Through interviews and co-design with genetics professionals, the paper identifies sensemaking during whole genome sequencing analysis as a key target for generative AI support, with two prioritized tasks: flagging cases for reanalysis and synthesizing gene and variant information.","lead":"This paper studies how genetic professionals who analyze whole genome sequences would want a generative AI assistant to help them. It finds they prioritize AI that flags cases for reanalysis and summarizes scientific literature, and it offers design considerations for such tools.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the central claim is appropriately qualified, and the main limitation—the single-institution Phase II sample—is explicitly acknowledged in §7.4.","rationale":"The reader correctly identified generalizability as the weakest assumption, but the paper's central claim is about what these participants envisioned and prioritized, not about the full population of genetic professionals. The two prioritized tasks were selected by all five walk-through participants, with explicit rationales, and the authors repeatedly qualify their contributions as insights that 'could' transfer. Section 7.4 openly states the single-institution and voluntary-participation limitations. For a qualitative HCI study, this sample size and setting are within normal practice, and the internal validity of the reported themes is not compromised. The only concrete issue I found is a numerical inconsistency in the Phase I role breakdown in §4.2, which appears to be a typo rather than a substantive flaw. Because the central claim is descriptive and appropriately hedged, no verdict change is warranted.","tokens_in":25702,"tokens_out":5368,"duration_ms":50474,"concrete_test":"Recruit a new Phase II co-design panel of 6–8 genetic professionals from at least two other institutions, including laboratory directors and clinicians, without requiring prior familiarity with seqr; repeat the group workshop and individual prioritization. If the same two tasks fail to emerge among the top three, the generalizability of the design considerations weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper does not overreach: it reports what its six Phase II participants prioritized, not a statistically generalizable population statement. The two top tasks were selected by all five walk-through participants with stated rationales, and the design considerations are presented as tentative rather than definitive. The weakest point is external validity: all Phase II participants were from the Broad Institute, familiar with seqr, and mostly variant analysts (5 of 6), so the prioritized tasks may reflect one institution's workflow and tool rather than the broader population of genetic professionals. Section 7.4 acknowledges this, and the paper hedges its language (e.g., 'could provide insight'). This is a limitation, not an internal inconsistency, and it does not undermine the reported findings. One minor reporting inconsistency exists: §4.2's parenthetical role breakdown sums to 15, while Table 1 sums to 18 for the combined unique sample; the Phase I clinician count appears to be 4 rather than 2. This should be corrected but does not affect the central argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a two-phase qualitative study with genetic professionals involved in whole-genome sequencing (WGS) analysis for rare disease diagnosis. Phase I comprised semi-structured interviews with 17 professionals and identified three challenges: aggregating and synthesizing gene/variant information, sharing findings with colleagues, and prioritizing cases for reanalysis. Phase II comprised co-design sessions, including a group workshop and individual design walk-throughs with six professionals from the Broad Institute, which led to a prototype of an AI assistant embedded in the seqr platform. Participants prioritized two AI tasks—flagging cases for reanalysis based on new scientific findings, and aggregating/synthesizing key gene/variant information from publications—and their feedback produced two design themes: balancing comprehensive and selective evidence, and collaboratively interpreting/verifying AI-generated information. The paper frames these findings through sensemaking theory and proposes three design considerations for generative AI support of sensemaking in knowledge work.","tokens_in":25880,"tokens_out":6679,"duration_ms":60966,"significance":"If the findings hold, the paper makes a useful empirical contribution to HCI research on human-centered generative AI. It documents how one group of domain experts envisions delegating sensemaking tasks and grounds design considerations in participant-generated prototype feedback. The study is methodologically transparent about recruitment, data analysis, and limitations; it includes participant quotes and figures showing the prototype and workflow; and it explicitly acknowledges the single-institution, voluntary-participation limitations of Phase II. The central claims are descriptive and appropriately qualified, with no overreach to population-level generalizations. The paper does not claim that the prototype was evaluated for effectiveness; it is clearly framed as a design probe. These strengths make the work appropriately scoped for publication.","major_comments":[],"minor_comments":[{"comment":"The role counts in the text do not match Table 1: the text reports 17 interviewees as seven variant analysts, two laboratory directors, two clinicians, two methods developers, and two program managers, which sums to 15, whereas Table 1 lists eight variant analysts and four clinicians in the combined unique sample of 18. Please reconcile the Phase I role breakdown and clarify the unique participant count.","section":"Section 4.2 / Table 1"},{"comment":"The subsection numbering is duplicated: both 'Individual Design Walk-Through Sessions: Protocol' and 'Individual Design Walk-Through Sessions: Data Analysis' are labeled 4.4.3. The latter should be renumbered.","section":"Section 4.4.3"},{"comment":"The caption says 'WGS (whole gene sequencing)' but WGS stands for whole genome sequencing; please correct this.","section":"Figure 3 caption"},{"comment":"There is a typo: 'analysists' should be 'analysts', and 't o' should be 'to'.","section":"Section 7.2.2"},{"comment":"Please clarify whether the prototype content (summaries, tables, chat responses) was produced by an actual LLM or hand-crafted for the design probe; this affects how readers interpret the walk-through feedback on 'AI-generated' artifacts.","section":"Section 6.2.1"},{"comment":"Consider adding to the limitations that Phase II participants were all familiar with seqr by design, which may make their design ideas incremental relative to that tool; the current limitation paragraph covers institution and voluntariness but not tool familiarity.","section":"Section 7.4"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is suitable for publication after minor revisions. The main correction is the participant-count inconsistency in Section 4.2; the remaining issues are presentation-level. I do not see a need for additional data collection or for changes to the central claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, squarely-reported qualitative HCI study. The new contribution is not the method but the application: a careful needs-elicitation (17 interviews) and co-design (6 participants, workshop plus walkthroughs) with genetic professionals, resulting in two AI tasks they prioritize—flagging cases for reanalysis based on new literature, and aggregating/synthesizing evidence about gene-variant pairs—and three sensible design considerations for AI-supported sensemaking.\n\nThe paper does what qualitative work should do: it grounds claims in quotes, describes the workshop protocol, uses a clickable prototype as a design probe, and labels findings as tentative. I think the central claim holds: sensemaking is a real bottleneck in WGS analysis, and the analysts want generative AI in a human-in-the-loop form, not full automation. Credit is also due for distinguishing interview-phase themes from co-design priorities, and for noting the gap between this human-in-the-loop vision and the fully automated reanalysis tools in the literature.\n\nSoft spots, in proportion. The Phase II sample is small, from one institution (Broad), mostly variant analysts, and all familiar with seqr. That is acknowledged in §7.4, and the claims are worded so the limitation does not sink the paper. The bigger issue is a numbers mismatch: §4.2 says 17 Phase I participants and lists roles that sum to 15, while Table 1 shows 18 unique participants with 4 clinicians and 8 variant analysts. The Phase II count also needs a careful pass. It is a minor reporting error, but it needs fixing before publication. Also, no raw interview data is provided, which is normal for this venue but does mean the analysis is not independently verifiable beyond the quoted excerpts.\n\nBottom line: this is a well-executed, unpretentious empirical study with a moderate contribution. It deserves serious peer review and will be useful to HCI researchers working on AI for knowledge work and to clinical-genomics tool builders. I would send it out; with the participant counts corrected, I would accept.","headline":"Solid qualitative study of AI support for WGS analysis; worth refereeing, but fix the participant-count inconsistency before publication.","tokens_in":26390,"tokens_out":2994,"would_cite":true,"duration_ms":24385,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that genetic professionals want a generative AI assistant that flags cases for reanalysis and synthesizes gene and variant evidence, with humans verifying the AI's output.","keywords":["whole genome sequencing","generative AI","large language models","knowledge work","sensemaking","co-design","rare disease","human-AI interaction"],"falsifier":"Deploy a working version of the assistant with a larger set of genetic professionals across multiple institutions and measure whether the two prioritized tasks actually change practice: if analysts outside the study site do not rank reanalysis flagging and gene-and-variant evidence synthesis as top AI tasks, or if task-based testing shows no reduction in per-case analysis time or no increase in diagnostic yield, the paper's central claim about the high-value AI sensemaking tasks would be undermined.","tokens_in":25524,"feed_emoji":"🧬","tokens_out":9490,"duration_ms":72948,"temperature":0.7,"pith_summary":"This paper argues that the hard part of whole-genome sequencing (WGS) for rare disease diagnosis is sensemaking, and that genetic professionals want a generative AI assistant aimed at that bottleneck rather than at full automation. Through interviews with 17 genetics professionals and co-design sessions with six, the study identifies two AI tasks that every co-design participant prioritized: flagging unsolved cases for reanalysis when new scientific findings appear, and aggregating and synthesizing key information about genes and variants from scientific publications. The paper further claims that the assistant should be human-in-the-loop: AI drafts tables, summaries, notes, and presentations, while analysts edit, verify, tag, and share those artifacts. From this, the authors derive three design considerations for AI-enhanced sensemaking: facilitate distributed sensemaking, support both initial sensemaking and re-sensemaking, and combine multiple modalities of evidence. If the paper is right, these results give concrete targets for generative AI in clinical genomics and a template for designing such tools in other knowledge-work domains.","feed_headline":"Give AI two jobs: flag reanalysis cases, synthesize gene evidence","feed_subtitle":"Co-design with analysts points AI at reanalysis triage and literature synthesis, keeping humans in the loop.","key_machinery":"The carrying mechanism is the co-design loop: interviews surface current challenges, a group workshop generates candidate AI tasks and interaction sketches, those sketches are turned into a clickable prototype of a generative-AI assistant embedded in the web-based genome analysis platform used at the study's partner institution, and individual design walk-throughs refine the resulting design considerations. Conceptually, the argument is organized around the named sensemaking model that distinguishes foraging (searching, filtering, and synthesizing information) from sensemaking proper (building, refining, and presenting models of information). The prototype's three features—reanalysis flagging, an AI-generated evidence table for gene and variant interpretation, and AI-drafted presentation slides—are the concrete objects through which the paper maps AI tasks onto that model.","core_discovery":"The paper's central claim is that sensemaking—the foraging and model-building work of finding, synthesizing, and interpreting information—is both the main bottleneck in WGS analysis and the process generative AI can most usefully support. Analysts struggle to aggregate information about genes and variants scattered across publications and databases, to share findings with colleagues, and to decide which unsolved cases deserve reanalysis as new papers appear. Asked to design an assistant, all six co-design participants prioritized the same two capabilities: flagging cases for reanalysis based on new scientific findings, and aggregating and synthesizing key gene-and-variant information from publications. The paper does not conclude that AI should replace the analyst; instead, participants envisioned a human-in-the-loop assistant that produces evidence tables, summaries, notes, and presentation drafts that analysts then edit, verify, and share, turning individual sensemaking into collaborative and distributed sensemaking. Three design considerations follow: facilitate distributed sensemaking, support initial sensemaking and re-sensemaking, and combine evidence from multiple modalities.","pith_inferences":["A testable extension of the paper's design considerations is that shared, verified AI evidence tables will measurably reduce duplicate literature searching in a laboratory, shrinking per-case interpretation time for genes already curated by colleagues.","The same human-in-the-loop artifact model likely transfers to other knowledge-work domains where professionals track a growing literature and revisit prior cases, such as diagnostic radiology, pathology, or legal research; the paper only gestures at this generality.","If continuous reanalysis becomes practical, the reimbursement models and patient- or clinician-initiated reanalysis triggers the paper mentions may become binding constraints, so real deployments would need policy-aware designs rather than pure automation.","Because the co-design participants all came from one institution and were mostly variant analysts, the two prioritized tasks are best treated as hypotheses about the broader profession until a multi-site participatory study confirms them."],"forward_implications":["A generative AI assistant that flags cases for reanalysis based on new scientific findings could turn reanalysis from a periodic, manually triggered event into a more continuous process driven by new evidence.","If analysts adopt the AI-generated evidence table, the time spent foraging across databases and publications per case could drop, shifting effort from searching to verifying and editing AI output.","Sharing verified, editable AI-generated evidence tables and notes across analysts could reduce duplicated sensemaking work and support distributed sensemaking within and between institutions.","Systems should show verification status, edits, and note authorship so readers can calibrate trust in AI-generated artifacts without hiding the model's raw inaccuracies.","Designers need to balance comprehensive evidence for gene and variant review against selective alerts for reanalysis flags, because analysts reject both over-filtering and too much noise."],"supporting_citations":[{"why":"Supplies the standard WGS interpretation workflow, per-case time estimates, and reanalysis barriers that frame the study.","marker":"[2]"},{"why":"Defines sensemaking as foraging plus model building, the conceptual lens the paper applies to variant interpretation.","marker":"[66]"},{"why":"Supports the premise that fewer than half of rare disease cases are solved on initial WGS analysis, motivating reanalysis and AI support.","marker":"[83]"},{"why":"Establishes that periodic reanalysis can increase diagnostic yield and frames AI tools as a key opportunity.","marker":"[18]"},{"why":"Documents the annual volume of new disease and gene-disease publications that makes manual flagging of cases for reanalysis infeasible.","marker":"[22]"},{"why":"Describes the open-source genome analysis platform into which the AI assistant prototype was embedded.","marker":"[61]"},{"why":"Introduces distributed sensemaking, the concept underlying sharing AI-generated evidence artifacts among analysts.","marker":"[25]"},{"why":"Provides the framework for when users reuse others' sensemaking artifacts, informing the verification and trust design considerations.","marker":"[48]"},{"why":"Supplies the task-delegation categories used to characterize the desired human-in-the-loop level of AI involvement.","marker":"[52]"}],"fun_headline_variants":["AI helps genetics experts spot reanalysis cases and synthesize evidence","Co-designed AI assistant: flag reanalysis, synthesize gene evidence","Human-in-the-loop AI for genetic diagnosis: triage and synthesis","Why geneticists want an AI that flags new clues and writes summaries","AI as sensemaking sidekick: reanalysis prompts and evidence digests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design considerations rest on the assumption that six self-selected genetic professionals from a single institution, mostly variant analysts, are representative enough of the profession that their prioritized AI tasks and interaction preferences can support general design guidance.","fun_headline_variants_meta":{"raw":{"variants":["AI helps genetics experts spot reanalysis cases and synthesize evidence","Co-designed AI assistant: flag reanalysis, synthesize gene evidence","Human-in-the-loop AI for genetic diagnosis: triage and synthesis","Why geneticists want an AI that flags new clues and writes summaries","AI as sensemaking sidekick: reanalysis prompts and evidence digests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000219,"raw_usage":{"total_tokens":1484,"prompt_tokens":1030,"completion_tokens":454,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":364}},"tokens_in":646,"tokens_out":454,"duration_ms":4484,"temperature":1.0,"reasoning_tokens":364,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:24:46.120732+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Deploy a working version of the assistant with a larger set of genetic professionals across multiple institutions and measure whether the two prioritized tasks actually change practice: if analysts outside the study site do not rank reanalysis flagging and gene-and-variant evidence synthesis as top AI tasks, or if task-based testing shows no reduction in per-case analysis time or no increase in diagnostic yield, the paper's central claim about the high-value AI sensemaking tasks would be undermined.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the standard WGS interpretation workflow, per-case time estimates, and reanalysis barriers that frame the study."},{"cited_title":"and Reddy, M.C","cited_arxiv_id":null,"evidence_quote":"Defines sensemaking as foraging plus model building, the conceptual lens the paper applies to variant interpretation."},{"cited_title":"and Lindstrand, A","cited_arxiv_id":null,"evidence_quote":"Supports the premise that fewer than half of rare disease cases are solved on initial WGS analysis, motivating reanalysis and AI support."},{"cited_title":"and Phan, T.G","cited_arxiv_id":null,"evidence_quote":"Establishes that periodic reanalysis can increase diagnostic yield and frames AI tools as a key opportunity."},{"cited_title":"and Evelo, C.T","cited_arxiv_id":null,"evidence_quote":"Documents the annual volume of new disease and gene-disease publications that makes manual flagging of cases for reanalysis infeasible."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the open-source genome analysis platform into which the AI assistant prototype was embedded."},{"cited_title":"and Myers, B.A","cited_arxiv_id":null,"evidence_quote":"Provides the framework for when users reuse others' sensemaking artifacts, informing the verification and trust design considerations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the task-delegation categories used to characterize the desired human-in-the-loop level of AI involvement."}],"review_version":1}