{"id":"125e3dd9-1863-4f5a-a0a3-008062860270","arxiv_id":"2507.10817","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Model risk in safety-critical ML applications can be evaluated under an ALARP test by combining decision analysis, Bayesian classification uncertainty, and value-of-information analysis, demonstrated for weld radiograph inspection.","lead":"This paper proposes a framework for deciding when a complex machine learning model is safe enough to use in high-consequence engineering decisions, using expected costs and value-of-information analysis to test whether model risk is As Low As Reasonably Practicable. It demonstrates the approach on a weld radiograph classification task, comparing manual inspection, full automation, and a hybrid strategy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.4.2 equates ALARP with a cost-benefit break-even (verification cost > VoI); UK ALARP requires stopping only when further risk reduction is grossly disproportionate, so the paper's criterion can declare ALARP too early.","rationale":"I considered whether the unvalidated cost inputs and perfect-manual benchmark in Table 2 are the weakest point. That is a real limitation of the illustrative numbers, and sensitivity analysis would improve confidence. But the paper explicitly says the framework is compatible with alternative cost models, so a bad input can be fixed without changing the method. The ALARP equivalence in Section 3.4.2 is different: if it is wrong, the method's output is mislabeled even with perfect inputs. I therefore focus there. The reader's verdict is CONDITIONAL and already flags the legal and expected-value interpretation in its rationale, but the stated weakest_assumption is the numerical inputs, so my emphasis is only partially aligned. A gross-disproportion check is a concrete, low-cost way to settle the concern: if the stopping rule changes when 'disproportionate' is interpreted as 'grossly disproportionate', then the paper should be revised to either adopt the legal standard or explicitly present a cost-benefit operationalization that is not called ALARP. Other issues, such as the figure and text discrepancy in the VoPI values and missing uncertainty bounds, are real but secondary.","tokens_in":11100,"tokens_out":11205,"duration_ms":142150,"concrete_test":"Use the Section 3.4.2 VoPI values. For the porosity scenario (Figure 11: £51.07 per radiograph), take a project of 100 welds, giving VoPI = £5,107. Choose a hypothetical verification cost C_v between £5,107 and 2 x £5,107, for example £6,000. Under the paper's rule, C_v > VoI, so further verification is declared not cost-effective and 'ALARP reached'. Recompute with a gross-disproportion rule C_v > g x VoI for g = 2: £6,000 < £10,214, so verification is still required. If the conclusion flips for any plausible g > 1, the break-even criterion in Section 3.4.2 is not the ALARP test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 3.4.2 is that ALARP is reached 'i.e. the point where further verification costs exceed their benefits to reduce risk'. This is not the UK ALARP test. Under HSE practice, further risk reduction is required unless the cost, including time and trouble, is grossly disproportionate to the safety benefit; a simple expected-value inequality C_v > VoI is a cost-benefit break-even, not a gross-disproportion test. With the paper's rule, if verification costs only slightly exceed expected benefit, the template tells an auditor to stop, whereas ALARP would still require the reduction. Because this equation is the step that converts a VoI calculation into an ALARP determination, the main conclusion does not follow even if every cost input and confusion matrix is correct. The statistical decision analysis itself is coherent; the problem is the label attached to the stopping rule. This is a correctness risk in the central claim, not a matter of tuning the example's numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes using the UK ALARP principle to decide whether the model risk introduced by deploying a complex machine-learning model in a safety-critical setting has been reduced to an acceptable level. It defines model risk as expected adverse impact, develops a decision-analysis workflow (influence diagram, cost matrix, Dirichlet-multinomial posterior over a confusion matrix, rank-ordering of manual, automated, and hybrid strategies, and value-of-information analysis for further verification), and applies this workflow to automated weld radiograph classification. The numerical results indicate that manual inspection is risk-optimal for defective welds, a hybrid strategy is best for no-anomaly radiographs, and further model verification is not cost-effective for most scenarios. The paper frames these results as demonstrating when model risk is ALARP.","tokens_in":11287,"tokens_out":16445,"duration_ms":180256,"significance":"The framework is a constructive step: it connects model reliability uncertainty to downstream decisions, and the value-of-information willingness-to-pay for verification is a useful concept. The Bayesian treatment of classification uncertainty is standard and appropriate, and the authors are transparent about assumptions such as perfect manual inspection. The main weakness is interpretive: the paper's operational definition of ALARP as a cost-benefit break-even is not the UK ALARP test, which requires stopping only when further risk reduction is grossly disproportionate. Because that identification is the paper's central claim, the contribution in its current form is not yet suitable for publication; the decision-analytic core is sound and can likely be reframed. The paper also leaves important numerical uncertainties unreported, which limits the strength of the demonstrated example.","major_comments":[{"comment":"The statement that model risk has reached ALARP \"i.e. the point where further verification costs exceed their benefits to reduce risk\" equates ALARP with a simple expected-value break-even. UK ALARP practice, as applied by HSE, requires further risk reduction unless the cost, including time and trouble, is grossly disproportionate to the safety benefit; it does not permit stopping merely because verification costs are slightly larger than expected benefits. With the paper's rule, an auditor would be instructed to stop precisely in cases where ALARP would still require the reduction. Since this sentence is the step that converts the VoI calculation into an ALARP determination, the main conclusion does not follow even if every cost input and confusion matrix is correct. Please reframe the contribution as a cost-benefit break-even for verification, or incorporate a gross-disproportion element (for example, a multiplier or a regulatory test) and justify it against HSE guidance.","section":"§3.4.2"},{"comment":"The expected costs in Table 4 and the VoI values in Figure 11 are computed by sampling from the Dirichlet posterior, but the paper reports neither the Monte Carlo sample size nor standard errors or credible intervals. This matters for the demonstrated ranking: in the porosity row, the automated strategy (£2,424.17) and hybrid strategy (£2,282.77) differ by only £141, a small margin relative to the variability induced by the £9,750 expected failure penalty that occurs with posterior probability around 0.16 under the stated Dirichlet(1,1,1,1) prior; the reported VoI for porosity is only £0.33 per radiograph. Without uncertainty quantification on these Monte Carlo outputs, the risk-optimal ordering and the conclusion that further verification is not cost-effective are not established at the precision claimed. No code or data are provided to reproduce the numbers, so the sampling details should be stated explicitly.","section":"§3.3.2–§3.3.3, Table 4, Figure 11, Algorithm 1"},{"comment":"The numerical demonstration rests on cost inputs that are stated without data or standard references: unit inspection and repair costs, the failure-cost mixture parameters (including the Gamma parameters), the 1, 1/2, and 1/10 failure-cost multipliers, and the assumption of error-free manual inspection. The paper notes that the framework is \"equally compatible with alternative input models,\" but it does not test how the Table 4 ranking or the Section 3.4.2 ALARP threshold respond to plausible variations in these inputs. Because the example is the only demonstration of the framework, please add a sensitivity analysis, or at least a stated range of inputs over which the conclusions are unchanged.","section":"§3.3.1, Table 2, Eqs. (4)–(6)"}],"minor_comments":[{"comment":"The text says \"Section 3.3 presets a Bayesian analysis\"; \"presets\" should be \"presents\".","section":"§4, step 2"},{"comment":"There is a missing space in \"threshold foracceptable risk\" in the ALARP definition.","section":"§1.1"},{"comment":"The symbol z in Eq. (13) is used for prospective verification data but is not defined in the nomenclature or at first use; please define it explicitly and clarify the relationship between VoPI and the more general z-based VoI.","section":"§3.4.2, Eq. (13)"},{"comment":"Reference [10] attributes \"Artificial intelligence concepts and terminology\" to BS ISO/IEC 42001, but that standard is an AI management system standard; the concepts-and-terminology standard is ISO/IEC 22989. Please correct the citation or the title.","section":"Reference [10]"},{"comment":"The ordering of the bar labels and the numerical values in Figure 11 is confusing; please align the labels (no anomaly, porosity, cracking, lack of penetration) with the corresponding bars and values.","section":"Figure 11"}],"recommendation":"major_revision","confidential_remarks":"The paper is best viewed as a decision-analysis template with an ALARP label attached; the template itself is coherent and potentially useful, but the ALARP interpretation is a load-bearing step that needs correction. Since the author list includes HSE-affiliated researchers, a revision that engages directly with HSE's published ALARP guidance (e.g., gross disproportion) would be particularly appropriate. I do not see evidence of circular derivation or data fabrication; the concerns are interpretive and about robustness reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a competent applied-statistics paper that packages standard tools—expected-risk minimization, Dirichlet-multinomial posteriors, value of information—into a practical template for deciding whether a machine learning model is acceptable for a safety-critical use. The worked weld-radiograph example is genuinely instructive: Table 4 is internally consistent, the VoI calculation correctly identifies that verification should target scenarios where it changes decisions, and the authors are honest that the perfect-manual benchmark is the most challenging case.\n\nWhat's new is the integration, not the ingredients. The decision analysis is straightforward and the posterior calculations are sound. The explainability section is a bit tacked on, but it is short and does not hurt. The paper's real contribution is showing safety auditors a coherent route from confusion matrices to verification investment decisions.\n\nThe main problem is in Section 3.4.2. The authors define ALARP as the point where 'further verification costs exceed their benefits to reduce risk.' That is a cost-benefit break-even. UK ALARP, as HSE practises it, does not stop at break-even: further risk reduction is required unless the cost is grossly disproportionate to the safety benefit. The paper's rule can therefore declare model risk ALARP too early. This is not a tuning issue; it is the load-bearing step that converts a VoI calculation into a regulatory determination. The authors should either adopt a gross-disproportion threshold or reframe the claim as 'verification is not cost-effective under expected-value analysis,' rather than 'ALARP has been reached.'\n\nThe numerical example also rests on unvalidated placeholders: repair costs, the failure-cost mixture, and the perfect-manual assumption. The authors flag the last one, but the cost inputs themselves could flip the risk-optimal ranking without a sensitivity analysis. That limits the example to an illustration, not a deployable benchmark.\n\nNone of this is fatal. The framework is coherent; the problem is the label and the missing sensitivity work. A careful referee would catch this and push for a proper ALARP criterion.\n\nThis paper is for safety scientists and engineering firms who want a quantitative template. It deserves a serious referee, but with heavy revision.\n\nRecommendation: send to peer review.","headline":"A coherent, instructive template for evaluating ML model risk against ALARP, but the central operationalization of ALARP as a cost-benefit break-even is a regulatory overreach that needs fixing.","tokens_in":11827,"tokens_out":2109,"would_cite":false,"duration_ms":24504,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a machine-learning model's risk is ALARP precisely when the value of further verification no longer exceeds its cost, and demonstrates the threshold with a weld-radiograph case study.","keywords":["model risk","ALARP","value of information","decision analysis","weld radiograph classification","Bayesian uncertainty quantification","model verification","safety-critical machine learning"],"falsifier":"Recompute the expected costs using real defect-repair and inspection cost data and measured, non-perfect manual inspector error rates from a weld-inspection organization; if manual evaluation no longer has the lowest expected cost for cracking, porosity, or lack of penetration, or if the value of perfect verification exceeds quoted verification costs, the case-study conclusion fails. A simpler direct test is to compare actual verification quotes with the computed £51.07-per-radiograph value for lack of penetration: if verification can be bought for less than that, the ALARP claim for that scenario is overturned.","tokens_in":10869,"feed_emoji":"🔍","tokens_out":6892,"duration_ms":76343,"temperature":0.7,"pith_summary":"The paper argues that a proposed machine-learning model for a safety-critical job should be accepted only when its model risk, the expected harm of wrong or unhelpful outputs, is reduced to a level that is As Low As Reasonably Practicable (ALARP), and it gives a quantitative way to find that point. It defines ALARP operationally as the moment where the value of further model verification no longer exceeds what that verification costs. The demonstration uses automated classification of weld radiographs: a Bayesian analysis of test-set performance feeds a decision analysis comparing fully manual, fully automated, and hybrid inspection strategies. In the case study, manual inspection is the cheapest option whenever a real defect is present, while the hybrid strategy wins for clean radiographs, and value-of-information analysis says further verification is not worth its cost for most scenarios. The framework is presented as general guidance for any high-consequence decision informed by a complex model.","feed_headline":"Model risk is acceptable when more verification stops paying","feed_subtitle":"Weld-inspection case study shows how to price the safety threshold 'as low as reasonably practicable'.","key_machinery":"The engine of the argument is the model-risk integral R_m = ∫_S ∫_{MO} Pr(s) Pr(mo|s) I(s,mo) dmo ds, which measures expected adverse impact of using model m. The classification model's reliability enters through Pr(mo|s), estimated from a Dirichlet-multinomial posterior over confusion-matrix rows, and consequences enter through scenario-specific costs including a mixture model for failure cost, Cfail. The ALARP stopping rule is computed by value of information: compare the prior optimal expected cost with the pre-posterior optimal expected cost after hypothetical perfect verification; the difference is the maximum worth of further verification, and when actual verification costs exceed it, model risk is declared ALARP.","core_discovery":"The central claim is that model risk should be treated as a decision-analytic quantity, computed as expected impact over scenarios and model outputs, and that the ALARP principle gives the stopping rule for model verification. In the weld radiograph example, ranking the three strategies by expected cost per radiograph shows manual evaluation is risk-optimal for cracking, porosity, and lack of penetration; the hybrid strategy is risk-optimal only for no-anomaly radiographs; and value-of-information analysis of perfect verification shows the largest per-radiograph benefit is £51.07 when lack of penetration is present, while the no-anomaly scenario gains essentially nothing from further verification. The authors read this as evidence that ALARP can be identified in practice, and that verification effort should follow decision value rather than raw uncertainty.","pith_inferences":["The paper leaves implicit that the same value-of-information calculation could set contract or procurement thresholds, telling buyers how much verification a model vendor should be required to fund.","Because the case study assumes a perfect manual inspector, a natural extension is to replace that benchmark with measured inspector error rates; the framework is compatible with that, and the optimal rankings could shift.","The break-even prevalence thresholds (around 84.5% no-anomaly radiographs for an equal defect mix) suggest a deployment rule for sites with known defect density, a testable extension the paper does not pursue.","Extending the cost inputs from pounds to multi-attribute consequences such as injury or environmental harm would require a utility model, but the decision-analytic structure would remain unchanged."],"forward_implications":["Model risk becomes a documented, quantitative part of a safety case rather than a vague concern, with an explicit budget for verification.","Verification spending should be allocated to scenarios where value of information is highest, not where model uncertainty looks largest.","In the weld-inspection setting, the risk-optimal deployment depends on defect prevalence: the hybrid strategy is preferable only when most radiographs are clean.","The same decision-analytic template can be applied to any complex model that informs high-consequence decisions, such as predictive maintenance or clinical decision support.","ALARP determinations can be made concrete: further testing is justified only while its cost stays below the calculated value of information."],"supporting_citations":[{"why":"Supplies the value-of-information comparison (prior versus pre-posterior) that defines the ALARP stopping rule.","marker":"[38]"},{"why":"Provides the decision-analysis framing and pre-posterior terminology used to compute expected verification value.","marker":"[39]"},{"why":"Supports the asymptotic approach to value of perfect information and risk-optimal data requirements for digital twins.","marker":"[40]"},{"why":"Provides the dataset of weld radiographs used to train and test the classification model.","marker":"[36]"},{"why":"Establishes the feasibility of convolutional neural network classification for weld defects, which the case study builds on.","marker":"[37]"},{"why":"Supplies the engineering standard that specifies what proportion of welds must be radiographically inspected.","marker":"[31]"}],"fun_headline_variants":["ALARP sets the stop sign for model risk verification","Model risk: verify until extra checking doesn't pay","Weld example: ALARP governs when to stop verifying AI","Model risk ALARP: decision value, not uncertainty, sets stop"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The numerical rankings depend on the assumed costs in Table 2, including repair costs, the fixed manual-inspection cost, the hand-specified failure-cost mixture, and the perfect-manual-inspector benchmark, so if those values are wrong for a real deployment, the risk-optimal strategy and the ALARP threshold can flip.","fun_headline_variants_meta":{"raw":{"variants":["ALARP sets the stop sign for model risk verification","Model risk: verify until extra checking doesn't pay","Weld example: ALARP governs when to stop verifying AI","Model risk ALARP: decision value, not uncertainty, sets stop"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000263,"raw_usage":{"total_tokens":1549,"prompt_tokens":844,"completion_tokens":705,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":637}},"tokens_in":460,"tokens_out":705,"duration_ms":9207,"temperature":1.0,"reasoning_tokens":637,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:25:05.881457+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the expected costs using real defect-repair and inspection cost data and measured, non-perfect manual inspector error rates from a weld-inspection organization; if manual evaluation no longer has the lowest expected cost for cracking, porosity, or lack of penetration, or if the value of perfect verification exceeds quoted verification costs, the case-study conclusion fails. A simpler direct test is to compare actual verification quotes with the computed £51.07-per-radiograph value for lack of penetration: if verification can be bought for less than that, the ALARP claim for that scenario is overturned.","supporting_citations":[{"cited_title":"Raiffa, Information Value Theory, IEEE Transactions on System Science and Cybernetics SSC-2 (1966) 22–34","cited_arxiv_id":null,"evidence_quote":"Supplies the value-of-information comparison (prior versus pre-posterior) that defines the ALARP stopping rule."},{"cited_title":"Jordaan, Decisions under Uncertainty, Cambridge University Press,","cited_arxiv_id":null,"evidence_quote":"Provides the decision-analysis framing and pre-posterior terminology used to compute expected verification value."},{"cited_title":"System Effects in Identifying Risk-Optimal Data Requirements for Digital Twins of Structures","cited_arxiv_id":"2309.07695","evidence_quote":"Supports the asymptotic approach to value of perfect information and risk-optimal data requirements for digital twins."},{"cited_title":"Totino, F","cited_arxiv_id":null,"evidence_quote":"Provides the dataset of weld radiographs used to train and test the classification model."},{"cited_title":"Perri, F","cited_arxiv_id":null,"evidence_quote":"Establishes the feasibility of convolutional neural network classification for weld defects, which the case study builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the engineering standard that specifies what proportion of welds must be radiographically inspected."}],"review_version":1}