{"id":"e56c9f7b-dfd6-4e1b-bb68-c2c7440057a0","arxiv_id":"2506.18727","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A knowledge-graph framework for digital nuclear control rooms maps procedure steps to interface paths and automates execution, reducing task completion time in a small simulator study.","lead":"AutoGraph is a knowledge-graph framework that links digital nuclear control room interfaces to procedure text, letting a computer automate clicks and navigation. The paper reports that the automated system completes parameter-checking tasks faster than human operators in a simulator test.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'fully automated execution' claim is not yet supported: mapping relies on a hand-built partial IE-KG, and no component is shown to read or verify parameter values, so the human comparison mainly measures scripted menu navigation.","rationale":"The reader identified the manual IE-KG as the weakest assumption; I agree it is serious. I would go one step further: even with an automatic KG, the paper does not establish the 'check' part of the procedure, so the strongest claim is under-supported on two linked grounds. The framework's own text admits the KG is manually constructed and partial, and the 100% accuracy statement in Sec. 4.2 is a tautology because the path is defined by the graph nodes it traverses. I am not arguing the system is useless: a scripted navigation engine for parameter lookup is plausible, and the tracker plus demonstration are real contributions. But the headline claim 'fully automated execution' and the human comparison rest on the unshown equivalence between replaying clicks and completing a verification procedure. The concrete test above would settle whether the execution engine performs the verification step or only navigation. If it fails, the verdict should remain conditional with the automation claim explicitly scoped to navigation, which is why I keep the reader's conditional verdict unchanged.","tokens_in":11630,"tokens_out":4351,"duration_ms":50137,"concrete_test":"Take one Table 1 verification step (e.g., 2LBA10CP801C = 13.86 MPa), set the simulator to display a different value, and feed AutoGraph the step text with no additional human annotation. Record whether it (a) maps the text to the correct interface path and (b) reports a mismatch rather than completing silently. A pass on both would support the verification claim; failure on either would narrow the claim to scripted navigation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Central claim: AutoGraph enables fully automated procedure execution with minimal operator input and outperforms humans. The evidence does not yet support this. Phase II (Sec. 3.3) says the IE-KG is 'manually constructed based on the tracker data'; Sec. 4.1 says it is a 'partial' graph for selected panels. Phase III (Sec. 3.4) 'searches within the constructed IE-KG' for a path, but no algorithm is specified for resolving arbitrary procedure text to graph nodes, so 'automatic mapping' is at best a lookup in a hand-authored index. Sec. 4.2 calls any multi-node path a multi-action step and claims 100% accuracy; that is definitional, not a measured detection result. More important, Table 1 steps are verification tasks ('check whether the value ... is X'), yet Sec. 4.3 only describes reproducing navigation/interaction sequences and 'collecting' parameter values. No module is described that reads the displayed value, compares it to the expected value, or flags a mismatch. Sec. 5.1 thus compares humans doing navigate+read+verify against a system that may only navigate. The all-points-below-y=x plot and p<0.001 Mann-Whitney U on six graduate students are then expected, not evidence of full procedure automation. The manual KG and missing verification step directly undermine the central automation claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes AutoGraph, a layered framework for modeling and automating procedure execution on digital nuclear control room interfaces. It combines an interaction tracker (HTRPM), a manually constructed Interface Element Knowledge Graph (IE-KG), a procedure-to-path mapping that labels multi-node paths as multi-action steps, and an execution engine that replays click sequences. The evaluation uses six graduate students performing five parameter-check tasks in an HTGRSim full-scope simulator; the paper reports that automated execution was faster in all scenarios (Mann-Whitney U p<0.001) and illustrates integration with the COGMIF dynamic HRA framework and the DRIF decision-support framework.","tokens_in":11936,"tokens_out":4913,"duration_ms":51830,"significance":"If the framework were to achieve the claimed automatic mapping and full verification, it would be a useful contribution to procedure automation and human reliability analysis. Credit is due for the concrete implementation on a real simulator, the direct human-automation timing comparison, and the explicit integration demonstrations with ACT-R/COGMIF. However, the central claims go beyond the current evidence: the knowledge graph is hand-built, the multi-action detection accuracy is definitional, and the automated system does not perform the verification step that the tested procedures require. The paper is a promising proof-of-concept but not a validated demonstration of fully automated procedure execution.","major_comments":[{"comment":"The claim of 'automatic mapping from textual procedures to executable interface paths' is not supported by the described pipeline. Section 3.3 states that the IE-KG is manually constructed from tracker data, and Section 3.4 says the system 'searches within the constructed IE-KG' for a path, without specifying any algorithm that resolves arbitrary procedure text to graph nodes. As written, mapping is a lookup in a hand-authored index over a partial set of panels (Section 4.1), so the scalability and automation claims are not established. To support the claim, the paper should either provide the mapping algorithm and report its performance on procedures not used in KG construction, or explicitly restrict the contribution to a manually curated demonstration.","section":"3.3, 3.4"},{"comment":"The '100%' multi-action detection accuracy is a tautology, not a measured result. The text defines a multi-action step as one whose mapped path contains multiple sequential interface nodes and then reports that this classification achieves 100% accuracy; the ground truth and the detection rule are identical by construction. The claim that this constitutes 'dynamic detection of human error traps' therefore needs independent validation, for example human annotation of step complexity, eye-tracking or workload data, or comparison with an existing HRA taxonomy.","section":"4.2"},{"comment":"The time comparison does not cover the full procedure content. The tasks in Table 1 are verification tasks of the form 'Check whether the value of parameter ... is X', but Section 4.3 describes the execution module only as reproducing navigation and interaction sequences and collecting parameter values. No component is described that reads the displayed value, compares it to the expected value, or flags a mismatch. Consequently, the comparison in Section 5.1 contrasts humans who navigate, read, and verify against an automation that may only navigate; the reported time advantage is therefore not evidence for the 'fully automated execution' claim. The paper should either implement and evaluate the verification step or restate the claim as automated navigation/parameter retrieval.","section":"4.3, 5.1"},{"comment":"The statistical support is thinner than presented. The evaluation uses six graduate students and five scenarios with no reported error bars, effect sizes, or per-scenario distributions; the Mann-Whitney U p-value on this sample, combined with the all-points-below the y=x line, is consistent with a scripted macro being faster than manual navigation but does not establish performance on realistic operator tasks. Please report the full distributions, effect sizes, and the amount of human supervision or setup time required by the automated runs, and treat the result as a demonstration rather than a general claim.","section":"5.1"}],"minor_comments":[{"comment":"Phase numbering is inconsistent: Section 3.1 describes Phase I as IE-KG construction and Phase II as semantic matching, while Sections 3.2-3.5 label tracker development as Phase I and IE-KG construction as Phase II.","section":"3.1-3.5"},{"comment":"The abstract lists contributions (3) and (4) as the same capability ('automatic mapping from textual procedures to executable interface paths' and 'an execution engine that maps textual procedures to executable interface paths'); reword to distinguish mapping from execution.","section":"Abstract"},{"comment":"The Author contribution statement names Peng Chen, Shunshun Liu, and Qianqian Jia, none of whom appear in the author list; clarify the authorship/acknowledgment arrangement.","section":"Author contribution"},{"comment":"Several figure references are loose: Section 4.2 and Section 4.3 refer to 'Figure 6' for the navigation path and demo scenarios, but the numbering in the text does not align with the captioned figures 1-9; please check all cross-references.","section":"Figures"},{"comment":"The COGMIF/DRIF integration results in Sections 5.2 and 5.3 should be labeled as illustrative demonstrations; the single reported T_reqd and HEP value has no validation against observed operator performance.","section":"5.2, 5.3"}],"recommendation":"major_revision","confidential_remarks":"The paper has a contribution statement that includes names not on the title page; this is an editorial integrity matter. The related-work and integration sections rely heavily on the authors' own recent preprints (e.g., [27]); that is acceptable but should be disclosed clearly. The manuscript's central claim is currently overstated relative to the evidence; major revision is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper builds a knowledge graph of interface elements in an HTR-PM simulator, tracks clicks, maps procedure text to navigation paths, and uses those paths to drive automated clicks. That combination is new for nuclear control room interfaces, and the prototype is real: end-to-end click navigation to parameter displays, plus a direct time comparison against six graduate students that is statistically significant (p<0.001, Mann-Whitney U). The paper is also honest in places—it explicitly says the IE-KG is manually constructed and partial, which is more than many papers do.\n\nThe central problem is the gap between the abstract's 'fully automated execution' and what the execution module actually does. All five tasks in Table 1 are 'check whether the value ... is X'. Section 4.3 says the module reproduces navigation and interaction sequences and collects parameter values, but no component is described that reads the displayed value, compares it to the expected value, or flags a mismatch. There is a passing mention of 'dynamic field recognition' and 'condition-based value confirmation' in a demo description, but no algorithm, no accuracy, and no evaluation. So the time comparison in Section 5.1 likely measures humans doing navigate+read+verify against a system that only navigates. That makes the all-points-below-y=x plot expected, not evidence of full procedure automation. The central claim is therefore not yet supported; the paper actually demonstrates automated navigation for parameter search, which is a useful but narrower result.\n\nThe other soft spots are proportionate. The IE-KG is manually built and partial, so scalability is not demonstrated—the authors admit this. The multi-action '100% detection accuracy' is definitional, but they also call it 'straightforward', so it is not a hidden flaw. Six graduate students with an incentive to be fast is a weak proxy for licensed operators; the p-value is not backed by effect sizes or error bars. The COGMIF and DRIF integration sections are illustrative, not validated.\n\nThat said, the core idea is sound and the prototype is concrete enough that a serious referee can engage. It deserves peer review, but the abstract and Section 5.1 need to be rewritten to say 'automated navigation and data collection' until value verification is implemented and measured. I would send it out.","headline":"A concrete click-automation prototype for parameter search in a nuclear simulator, but the 'fully automated execution' claim overreaches because no component is shown to read or verify displayed values.","tokens_in":12417,"tokens_out":3354,"would_cite":false,"duration_ms":36228,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A knowledge graph of interface elements lets software execute nuclear control-room procedures automatically and faster than human operators.","keywords":["Knowledge Graph","Digital Control Room","Procedure Automation","Human Reliability Analysis","Nuclear Power Plant","Human-System Interface","Multi-action Step","Parameter Search"],"falsifier":"Give AutoGraph a digital control-room simulator it has never been configured for, build the knowledge graph automatically from raw screen content or an accessibility dump rather than from curated tracker data, and run a parameter-check procedure: if the system cannot locate the target elements or execute the path without manual graph edits, the automation claim is refuted. A sharper challenge is a procedure with conditional branching, where the current path search has no explicit decision logic.","tokens_in":11446,"feed_emoji":"⚛️","tokens_out":8666,"duration_ms":84302,"temperature":0.7,"pith_summary":"The paper tries to establish that the gap between written nuclear procedures and the screens operators click on can be closed with a knowledge graph of interface elements, and that once closed, procedures can be executed automatically. It argues that such a graph lets a computer navigate multiple interface layers, perform parameter checks, and do so faster than human operators in simulator tests. The motivation is human error: roughly half of reportable U.S. nuclear events are attributed to human factors, and current computer-based procedures lack the semantic hooks needed to automate complex multi-action steps. If the claim holds, the same graph can flag cognitively demanding steps, feed timing data into human reliability models, and compare operator behavior against an expected interaction path.","feed_headline":"Software runs nuclear control-room procedures faster than operators","feed_subtitle":"A graph of clickable interface elements lets software execute multi-step control-room tasks in less time.","key_machinery":"The interface element knowledge graph (IE-KG). It is a directed labeled graph $G = (V, E)$, where each node is an interactive GUI element carrying a human-readable name and two-dimensional screen coordinates, and each edge encodes a hierarchy such as a system panel containing a specific component. It is the load-bearing representation: procedure-to-interface mapping searches the graph, multi-action steps are recognized from mapped path length, and the execution engine replays node-to-node navigation as simulated clicks. A tracking module supplies the element positions and interaction traces from which the graph is built.","core_discovery":"AutoGraph's central claim is that a machine-readable interface element knowledge graph — a directed graph whose nodes are labeled, coordinate-bearing interface elements and whose edges encode parent-child nesting — makes procedural text executable in a digital control room. The paper reports that in all five tested scenarios on a full-scope simulator, automated execution completed every parameter-check task faster than human operators did, with a nonparametric rank-sum test giving $p < 0.001$, and that any mapped path longer than one step can be classified as a multi-action step with 100% accuracy. It further claims the framework integrates with existing dynamic human reliability analysis and real-time decision support systems by automatically decomposing high-level tasks into concrete interface-linked operations.","pith_inferences":["If graph construction were automated from accessibility trees or screen captures, the same machinery could transfer to other digital control rooms or safety-critical GUIs; the paper's manual, partial graph is the current bottleneck.","The mapped shortest paths suggest a real-time operator-monitoring oracle: any deviation from the expected path could be flagged immediately, extending the paper's retrospective risk-analysis integration into live error detection.","Procedure-design review could be inverted: procedures whose mapped paths are deep or long would empirically be the ones worth simplifying, and the framework quantifies that directly.","Cognitive-simulation integration points toward generating synthetic operator timing data at scale, which could supplement scarce real incident data for training data-driven human reliability models."],"forward_implications":["Textual procedures can be turned into executable click-level paths without modifying the underlying simulator, so parameter checks and similar fixed tasks can run unattended.","Multi-action steps are automatically identifiable from mapped path length, giving human reliability analysts a concrete way to flag cognitively demanding procedure segments.","The framework can feed task-completion times into cognitive operator models, and one demonstrated integration yields an estimated error probability of $8.2 \\times 10^{-3}$ for a specific step.","Integrated with a diagnostic decision support system, AutoGraph can provide a reference path against which real operator actions are compared, supporting data collection for dynamic human reliability analysis.","Because a direct timing comparison shows automation faster in every tested scenario, the framework offers a repeatable baseline for evaluating procedure efficiency."],"supporting_citations":[{"why":"Supplies the statistic that about 48% of reportable U.S. nuclear events are attributed to human error, motivating the need for automation.","marker":"[1]"},{"why":"Identifies multi-action steps as a limitation of current computer-based procedures, the specific gap AutoGraph targets.","marker":"[2]"},{"why":"Provides human-factors engineering guidance and documents cognitive-overload risks from digitalization that motivate semantic interface modeling.","marker":"[7]"},{"why":"Provides the real-time decision support architecture into which AutoGraph plugs its procedure-to-path mapping.","marker":"[9]"},{"why":"Establishes the absence of machine-interpretable task models in computer-based procedures, the barrier AutoGraph is designed to remove.","marker":"[13]"},{"why":"Describes prior automated work-package capabilities whose semantic depth is limited by current procedures, the baseline AutoGraph extends.","marker":"[24]"},{"why":"Supplies the nonparametric rank-sum test used to show the automated-versus-human time difference is statistically significant.","marker":"[25]"},{"why":"Provides the cognitive-mechanistic human reliability framework used to demonstrate automatic task decomposition into modellable operator actions.","marker":"[27]"}],"fun_headline_variants":["AutoGraph: knowledge-graph control-room automation cuts task time","AutoGraph automates nuclear procedure execution via interface graph","Knowledge graph turns nuclear procedures into executable tasks","AutoGraph: faster control-room procedures via automated execution","Graph-based automation speeds up nuclear control-room task completion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a hand-built, partial interface-element knowledge graph can stand in for the full interface; if a new control room needs manual graph construction for every panel, the claimed automatic mapping and scalability do not yet follow.","fun_headline_variants_meta":{"raw":{"variants":["AutoGraph: knowledge-graph control-room automation cuts task time","AutoGraph automates nuclear procedure execution via interface graph","Knowledge graph turns nuclear procedures into executable tasks","AutoGraph: faster control-room procedures via automated execution","Graph-based automation speeds up nuclear control-room task completion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000976,"raw_usage":{"total_tokens":4136,"prompt_tokens":927,"completion_tokens":3209,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":3133}},"tokens_in":543,"tokens_out":3209,"duration_ms":20593,"temperature":1.0,"reasoning_tokens":3133,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:00:58.216955+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give AutoGraph a digital control-room simulator it has never been configured for, build the knowledge graph automatically from raw screen content or an accessibility dump rather than from curated tracker data, and run a parameter-check procedure: if the system cannot locate the target elements or execute the path without manual graph edits, the automation claim is refuted. A sharper challenge is a procedure with conditional branching, where the current path search has no explicit decision logic.","supporting_citations":[{"cited_title":"Nuclear Engineering and Technology, 103687 (2025)","cited_arxiv_id":null,"evidence_quote":"Supplies the statistic that about 48% of reportable U.S. nuclear events are attributed to human error, motivating the need for automation."},{"cited_title":"Technical report, Idaho National Lab.(INL), Idaho Falls, ID (United States) (2012)","cited_arxiv_id":null,"evidence_quote":"Identifies multi-action steps as a limitation of current computer-based procedures, the specific gap AutoGraph targets."},{"cited_title":"US Nuclear Regulatory Commission, Washington, DC, USA (2008)","cited_arxiv_id":null,"evidence_quote":"Provides human-factors engineering guidance and documents cognitive-overload risks from digitalization that motivate semantic interface modeling."},{"cited_title":"In: Computing in Civil Engineering 2023, pp","cited_arxiv_id":null,"evidence_quote":"Establishes the absence of machine-interpretable task models in computer-based procedures, the barrier AutoGraph is designed to remove."},{"cited_title":"Nuclear Technology 202(2-3), 201–209 (2018)","cited_arxiv_id":null,"evidence_quote":"Describes prior automated work-package capabilities whose semantic depth is limited by current procedures, the baseline AutoGraph extends."},{"cited_title":"The Corsini encyclopedia of psychology, 1–1 (2010)","cited_arxiv_id":null,"evidence_quote":"Supplies the nonparametric rank-sum test used to show the automated-versus-human time difference is statistically significant."}],"review_version":1}