{"id":"9a4aa849-0389-4136-a904-18897c59b8e6","arxiv_id":"1908.06540","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A conservative Bayesian framework that incorporates rare failures and prior knowledge is derived, and software reliability growth models are shown to provide useful but non-definitive forecasts of AV disengagement rates.","lead":"This paper extends Conservative Bayesian Inference to handle rare failures in autonomous vehicle road testing, and applies software reliability growth models to Waymo's disengagement data. The CBI results show how strong prior beliefs can reduce required test miles, while the SRGM analysis illustrates how forecast accuracy can be evaluated and improved.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SRGM results rest on an unreported synthetic preprocessing step; the empirical half of the paper is not independently reproducible.","rationale":"I read the paper in good faith and focused on the central technical claim: Theorem 1's worst-case posterior confidence formula, Eq. (3), and its use in the CBI comparisons. The proof in Appendix A is careful: the three-point reduction via the integral mean-value theorem is valid, the partial-derivative argument that the beta mass should be zero is correct, and the case analysis in Stage 3 correctly identifies the extreme points of x^k(1-x)^(n-k) on the relevant intervals. I independently re-derived Eq. (10) and the values in Table I and Figures 2-4; they are consistent with the theorem. Thus the CBI contribution does not present a soundness problem. The real vulnerability is exactly the one the reader identified: the SRGM empirical section depends on turning monthly aggregate disengagement data into a sequence of inter-failure mileages by random generation within each month, and the sensitivity check is mentioned but not reported. This is a reproducibility and potential artifact concern, not a contradiction or an internal inconsistency. It affects the secondary empirical contribution but does not undermine the theorem, so the appropriate disposition remains conditional: the CBI material can be accepted, while the SRGM empirical claims need either the missing sensitivity evidence or a clear statement that they are illustrative. The reader's weakest_assumption matches this concern, and no stronger objection emerged from my review.","tokens_in":17806,"tokens_out":18214,"duration_ms":186530,"concrete_test":"Re-run the Section IV-A preprocessing with, say, 100 independent random seeds (or a documented deterministic alternative such as placing events at equal intervals or at monthly-constant rates), and report the resulting MMTD forecasts, u-plot statistics, PLR rankings, and the 7,000-8,000 mile estimate. If the MMTD estimate or the ranked ordering of SRGMs changes materially across seeds, the empirical SRGM conclusions should be treated as unverified; if they are stable, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The CBI theorem and its numerical illustrations are internally consistent: the proof in Appendix A is sound, and the values in Figures 2-4 and Table I reproduce from Eq. (3) and Eq. (10) with the stated parameters. The load-bearing weakness is in the SRGM contribution (Section IV-A). Waymo's public data are monthly disengagement counts and monthly mileages, but PETERS requires a sequence of exact inter-failure mileages. The paper says it preprocessed the raw data by generating random points in a Poisson process within each month, and that it repeated this to check sensitivity, but it reports neither the sensitivity results nor the random seed. Consequently, all SRGM-dependent outputs--Figure 5A-F, the u-plots, the PLR-plots, the recalibration comparisons, and the headline current MMTD estimate of 7,000-8,000 miles--depend on an unreported arbitrary synthetic sequence. No code or data artifacts are provided. The paper honestly disclaims SRGMs for safety claims, but the empirical demonstration that SRGMs are useful planning aids is not reproducible from the manuscript as written, and the claimed MMTD value could be an artifact of the synthetic sequence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper addresses two related problems in assessing autonomous vehicle safety from road testing. First, it develops a Conservative Bayesian Inference (CBI) framework for claims about a constant per-mile failure probability (pfm/pcm) when testing produces a small number of failures. Theorem 1 identifies the prior distribution, subject to partial prior knowledge Pr(X≤ε)=θ and Pr(X>pl)=1, that minimizes the posterior confidence in any bound p; the result generalizes existing CBI from failure-free testing to arbitrary failure counts. The authors compare the required number of test miles against the classical RAND approach and standard Bayesian priors, and derive asymptotic properties for the number of extra miles needed to compensate for a fatality (Q3). Second, the paper applies Software Reliability Growth Models (SRGMs) to Waymo's public disengagement data over 51 months, with u-plot and PLR-plot accuracy assessment and recalibration, concluding that SRGMs can be a useful planning aid. The abstract and introduction frame this as a contribution to safety and reliability assessment of autonomous vehicles.","tokens_in":18039,"tokens_out":6360,"duration_ms":62013,"significance":"The CBI theorem is a rigorous and useful extension: its proof in Appendix A is constructive and self-contained, and the numerical comparisons in Figures 2–4 and Table I follow from the stated formulas. If the SRGM empirical analysis were fully reproducible, the paper would provide a valuable demonstration of how forecast-accuracy assessment and recalibration can make SRGMs usable for planning AV disengagement trends. The paper explicitly avoids overclaiming SRGMs for safety certification, which is appropriate. The main significance is the new conservative Bayesian method for combining partial prior knowledge with sparse failure evidence; the SRGM application is a secondary, methodological contribution.","major_comments":[{"comment":"The preprocessing of Waymo's monthly data into a sequence of inter-failure mileages is described only in a footnote as 'generating random points in a Poisson Process for each month, repeating to check sensitivity of the results to this manipulation,' with no details of the procedure, the number of replications, the random seed, or the results of the sensitivity checks. Consequently, the MMTD predictions, u-plots, PLR-plots, and the headline MMTD estimate of 7,000–8,000 miles are not independently reproducible and may depend on the particular synthetic sequence generated. This is a load-bearing deficiency because the paper's second contribution is an empirical demonstration that SRGMs with recalibration are useful planning aids. The authors should report the full preprocessing protocol, the sensitivity results (e.g., the range of MMTD predictions across replications), and make the code and data available, or explicitly reframe the results as conditional on a specific synthetic realization.","section":"Section IV-A, footnote 8"},{"comment":"The statement 'the best estimate of current MMTD is thus about 7-8000 miles' is a point estimate with no uncertainty quantification. Given that the SRGMs in Figure 5A produce widely different MMTD predictions, and given the stochastic preprocessing, the paper should provide confidence intervals or a sensitivity range for this estimate, for example by showing the distribution of recalibrated MMTD predictions over the random replications. Without this, the claim that SRGMs can provide practical planning forecasts is not quantitatively supported.","section":"Section IV-A, Figure 5C"}],"minor_comments":[{"comment":"The abbreviations GO, MO, LV, and Li are used in the text and figure without a complete mapping in the main body; consider placing the model names in the caption of Figure 5 or in a separate table.","section":"Section IV-A"},{"comment":"Equation (10) appears to involve a floor or ceiling function, but this notation is not defined in the main text; please clarify the operation used.","section":"Appendix B"},{"comment":"The paper refers to the PETERS toolset without providing a citation or a description of its provenance; a reference or a brief explanation would help a reader locate the tool.","section":"Section IV-A"},{"comment":"The phrase 'we thus postulate an observed number of fatalities' should explicitly state that the k=43 value is a hypothetical expectation used for comparison, not an actual observation, to avoid confusion.","section":"Section III-B, Q2"},{"comment":"The text consistently uses 'A Vs' with a space; this should be standardized to 'AVs' for stylistic consistency.","section":"Throughout"},{"comment":"The sentence 'In Fig. 5A,C,E,F, the 528 failures...' is confusing because the panels have different roles; consider rephrasing to describe the content of each panel explicitly.","section":"Section IV-A, Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The CBI part of the paper is a sound and publishable contribution. The SRGM empirical section is the main obstacle: the unreported synthetic preprocessing means the results cannot be checked. If the authors provide the sensitivity analysis and the data/code in the revision, I would support acceptance. The paper might also be considered by a reliability-engineering audience; the AV framing is a natural application."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know about this paper: the theoretical half is genuinely good, and the empirical half is less solid than it looks. The CBI result generalizes existing conservative Bayesian inference from failure-free testing to cases with k > 0 failures, and the proof in Appendix A is constructive and checks out. The numerical illustrations (Figures 2-4, Table I) reproduce from the stated equations and parameters, and the comparison with the RAND study is fair. That part deserves credit.\n\nThe soft spot is the SRGM contribution. The authors take Waymo's monthly disengagement counts and miles, then generate exact inter-failure mileages by randomly placing points inside each month under a Poisson assumption. They mention repeating this to check sensitivity, but they do not report the sensitivity results or a random seed. That means the u-plots, PLR-plots, recalibration curves, and the headline 7,000-8,000 mile MMTD estimate all depend on an arbitrary synthetic sequence that no one can reproduce from the manuscript. No code or data artifacts are provided either. This is not a minor omission; it is the load-bearing step for the empirical half. The authors are honest that SRGMs are not suitable for safety claims, but they do claim SRGMs are useful as planning aids, and that claim is currently supported only by a non-reproducible preprocessing step.\n\nTo be fair, the CBI theorem is exactly what the paper says it is: a self-contained extension of prior work, and the self-citations to earlier CBI papers are appropriate since this is a direct generalization. The SRGM methodology is standard, and the paper makes no pretense of inventing it. The weakness is not in the theory but in the execution of the empirical demonstration.\n\nI would send this to a serious referee. The theoretical result is worth publishing, and the SRGM application, if the authors can either provide the synthetic sequences or analyze the sensitivity transparently, could be made reproducible. The paper is written clearly and the authors know their limitations, which is more than many in this space. If I were refereeing, I would ask for the sensitivity analysis to be shown or the random seed released, and for confidence intervals on the MMTD predictions. With that, the paper would be a solid contribution to the AV safety literature.\n\nFor a reading group, it is a reasonable choice if the group cares about conservative Bayesian methods; the SRGM part is more of a cautionary tale about preprocessing.","headline":"The new CBI theorem for non-zero failure counts is sound and worth knowing, but the SRGM half of the paper rests on an unreported synthetic preprocessing step that makes its empirical results hard to trust as presented.","tokens_in":18549,"tokens_out":1128,"would_cite":true,"duration_ms":13705,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"With only two quantile constraints on an AV fatality rate, a worst-case Bayesian prior yields conservative posterior confidence and can cut required road-test mileage by about three quarters when prior confidence is strong.","keywords":["autonomous vehicles","conservative Bayesian inference","reliability assessment","prior knowledge","software reliability growth models","disengagement data","recalibration","road testing"],"falsifier":"Take a concrete instance of the theorem, say $\\epsilon=1.09\\times10^{-10}$, $pl=10^{-15}$, $p=1.09\\times10^{-8}$, $\\theta=0.9$, $k=0$, $n=69\\times10^6$, and numerically minimize the posterior confidence in Eq. (1) over a dense family of feasible priors; any feasible prior giving posterior confidence below the value from Eq. (3) would refute the claimed conservatism.","tokens_in":17609,"feed_emoji":"🚗","tokens_out":17789,"duration_ms":157705,"temperature":0.7,"pith_summary":"The paper targets a practical impasse: demonstrating the very low fatality rates expected of autonomous vehicles from road testing alone would require hundreds of millions of miles, while most published analyses ignore evidence available before testing. It proposes a variant of Conservative Bayesian Inference in which the assessor needs only two numbers about the unknown fatality rate $X$: a prior confidence $\\theta$ that $X$ is below the engineering goal $\\epsilon$, and certainty that $X$ is above a physical lower bound $pl$. Theorem 1 shows that, among all priors satisfying those constraints, the one minimizing posterior confidence in any claimed bound $p$ is a two-point distribution, with the worst-case confidence given explicitly by Eq. (3) for any observed failure count $k$. This generalizes CBI from failure-free testing to rare failures, and the paper uses it to compare required test miles with classical and standard Bayesian approaches. For disengagements, it fits several software reliability growth models to 51 months of public road-test records, checks their forecast accuracy, and shows recalibration brings their median-miles-to-next-disengagement forecasts to around 7,000--8,000 miles.","feed_headline":"Worst-case prior cuts AV test miles for a 95% safety claim","feed_subtitle":"A worst-case prior lets engineering knowledge count without optimistic bias; recalibrated forecasts track disengagement trends.","key_machinery":"The load-bearing object is the conservative two-point prior: a prior distribution placing probability $\\theta$ at a lower point $x_1$ between the physical lower bound $pl$ and the engineering goal $\\epsilon$, and probability $1-\\theta$ at an upper point $x_3$ above $\\epsilon$. Theorem 1 proves this is the feasible prior that minimizes the posterior probability of the safety bound $p$. The argument works in three stages: first any feasible prior is replaced by an equivalent three-point prior; then the middle mass is shown to increase the posterior, so it is set to zero; finally $x_1$ and $x_3$ are chosen by monotonicity of the likelihood factor $x^k(1-x)^{n-k}$. That factor carries the whole computation, turning an infinite-dimensional optimization over priors into a two-point placement from which mileage requirements follow directly from Eq. (3).","core_discovery":"On its own terms, the discovery is that conservatism and prior knowledge are compatible: a Bayesian assessor does not need a full prior distribution to obtain a sound posterior bound. For a Bernoulli per-mile failure process with unknown rate $X$, given constraints $\\Pr(X\\le\\epsilon)=\\theta$ and $\\Pr(X>pl)=1$, the worst-case posterior confidence $\\Pr(X\\le p \\mid k \\text{ failures in } n \\text{ miles})$ is attained by a two-point prior with mass $\\theta$ at $x_1\\in[pl,\\epsilon]$ and mass $1-\\theta$ at $x_3>\\epsilon$. The minimized value is the rational expression in Eq. (3), which is zero for $p\\le\\epsilon$ and otherwise depends on $x_1^k(1-x_1)^{n-k}\\theta$ relative to $x_3^k(1-x_3)^{n-k}(1-\\theta)$. The proof collapses any feasible prior to a three-point prior, shows the middle mass can be removed, and then places the two remaining points by monotonicity of $x^k(1-x)^{n-k}$, so the required number of test miles can be solved directly. A direct corollary is that under these partial prior constraints, no amount of failure-free testing supports a safety bound better than the engineering goal $\\epsilon$.","pith_inferences":["A consequence the paper leaves implicit is that the same worst-case-prior construction applies to any safety-critical machine-learning system with a verifiable non-ML backstop: whenever an assessor can state $\\epsilon$, $\\theta$ and $pl$, Eq. (3) determines the operational evidence required, so the paper's mileage results can be read as generic evidence requirements.","A testable extension is to derive $\\theta$ from simulation, formal verification, or component testing results rather than eliciting it as a judgement; the paper lists these as possible sources but does not develop a calibration procedure that turns such evidence into the prior confidence value.","The same forecast-recalibration pipeline could be applied to other fleets' monthly disengagement reports; because those reports are monthly aggregates, exact event dates would make third-party reliability-growth forecasts fully reproducible, whereas reconstructed event times leave the forecasts dependent on the reconstruction procedure."],"forward_implications":["If an assessor can justify a strong prior confidence $\\theta=0.9$ in the engineering goal, CBI supports a 95% claim on a human-comparable fatality bound with about 69 million fatality-free miles, versus 275 million under the classical confidence approach; a weak prior $\\theta=0.1$ raises the requirement to about 476 million miles.","Under the stated partial prior constraints, no amount of failure-free road testing can support any fatality-rate claim better than the engineering goal $\\epsilon$; CBI makes this limit explicit rather than hidden in a prior distribution.","When observed failures are included, CBI demands substantially more miles than classical or uniform-prior approaches (roughly 79 billion versus 5.0 billion miles in the 43-fatality scenario), reflecting its avoidance of optimistic prior assumptions.","After one observed fatality, the additional fatality-free miles needed to restore a given claim has a ceiling of $1/\\epsilon$ once the initial test mileage is large, and this ceiling is independent of the confidence levels $c$ and $\\theta$.","Recalibrated software reliability growth models applied to 51 months of disengagement records converge near a current median-miles-to-disengagement estimate of 7,000--8,000 miles, making them a practical test-planning aid rather than a tool for demonstrating ultra-high safety."],"supporting_citations":[{"why":"Introduces the conservative Bayesian inference method for failure-free operation that the paper generalizes to observed failures.","marker":"[23]"},{"why":"Supplies the classical confidence-interval mileage benchmark that the paper's CBI mileage requirements are compared against.","marker":"[9]"},{"why":"Establishes the infeasibility of demonstrating ultra-high dependability from operational testing alone, motivating the use of prior knowledge.","marker":"[13]"},{"why":"Provides the exponential order-statistic software reliability growth model family used for disengagement forecasting.","marker":"[27]"},{"why":"Introduces the u-plot and prediction-accuracy comparison techniques used to evaluate forecast bias and accuracy.","marker":"[43]"},{"why":"Supplies the recalibration procedure applied to improve the SRGMs' predictive accuracy.","marker":"[48]"},{"why":"Documents the prediction-analysis and recalibration techniques used in the forecast-accuracy workflow.","marker":"[28]"}],"fun_headline_variants":["Worst-case prior slashes AV road-test miles needed","Bayesian prior cuts AV test miles, stays conservative","Conservative prior shrinks AV test miles for safety claim","Prior knowledge reduces AV test miles without optimistic bias"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that an assessor can state a trustworthy prior confidence $\\theta$ in the engineering goal and a physical lower bound $pl$, and that the per-mile failure probability is constant over the assessment period; if those prior numbers are not defensible, the mileage reductions the method offers are not defensible either.","fun_headline_variants_meta":{"raw":{"variants":["Worst-case prior slashes AV road-test miles needed","Bayesian prior cuts AV test miles, stays conservative","Conservative prior shrinks AV test miles for safety claim","Prior knowledge reduces AV test miles without optimistic bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000241,"raw_usage":{"total_tokens":1552,"prompt_tokens":1008,"completion_tokens":544,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":481}},"tokens_in":624,"tokens_out":544,"duration_ms":5191,"temperature":1.0,"reasoning_tokens":481,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:41:27.029619+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a concrete instance of the theorem, say $\\epsilon=1.09\\times10^{-10}$, $pl=10^{-15}$, $p=1.09\\times10^{-8}$, $\\theta=0.9$, $k=0$, $n=69\\times10^6$, and numerically minimize the posterior confidence in Eq. (1) over a dense family of feasible priors; any feasible prior giving posterior confidence below the value from Eq. (3) would refute the claimed conservatism.","supporting_citations":[{"cited_title":"Toward a formalism for conservative claims about the dependability of software-based systems,","cited_arxiv_id":null,"evidence_quote":"Introduces the conservative Bayesian inference method for failure-free operation that the paper generalizes to observed failures."},{"cited_title":"Driving to safety: How many miles of driving would it take to demonstrate autonomous vehicle reliability?","cited_arxiv_id":null,"evidence_quote":"Supplies the classical confidence-interval mileage benchmark that the paper's CBI mileage requirements are compared against."},{"cited_title":"Validation of ultra-high dependability for software-based systems,","cited_arxiv_id":null,"evidence_quote":"Establishes the infeasibility of demonstrating ultra-high dependability from operational testing alone, motivating the use of prior knowledge."},{"cited_title":"Exponential order statistic models of software reliability growth,","cited_arxiv_id":null,"evidence_quote":"Provides the exponential order-statistic software reliability growth model family used for disengagement forecasting."},{"cited_title":"New ways to get accurate reliability measures,","cited_arxiv_id":null,"evidence_quote":"Introduces the u-plot and prediction-accuracy comparison techniques used to evaluate forecast bias and accuracy."},{"cited_title":"Recalibrating software reliability models,","cited_arxiv_id":null,"evidence_quote":"Supplies the recalibration procedure applied to improve the SRGMs' predictive accuracy."},{"cited_title":"Techniques for prediction analysis and recalibration,","cited_arxiv_id":null,"evidence_quote":"Documents the prediction-analysis and recalibration techniques used in the forecast-accuracy workflow."}],"review_version":1}