{"id":"c51479f4-efe8-44ee-b5d2-3e0d2fa21237","arxiv_id":"2608.10025","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Coarse failure data without false-positive vs false-negative labels can make conservative Bayesian reliability claims for AV software dangerously optimistic, sometimes infinitely so after a single failure.","lead":"This paper extends conservative Bayesian inference to show how coarse operational data, which records failures but not their types, can support dangerously optimistic reliability claims for safety-critical autonomous vehicle software. It provides the first conservative bounds on the size of this optimism gap and guidance for logging failure-type labels.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The infinite-gap conclusion depends on Theorem 3's zero-infimum result, which is stated without proof; the omitted proof is the least secure load-bearing step.","rationale":"The reader's verdict is CONDITIONAL primarily because Theorems 1-3 are unproved. My stress-test agrees that the central claim is plausible and well-supported by Theorem 4's proof and inequality (9), but identifies Theorem 3's zero-infimum case as the hinge of the strongest quantitative conclusion. Unlike Theorem 4, this case depends on an exactly-zero likelihood at a boundary point of S2, so it is not covered by the 'analogous steps' remark. A full proof or a counterexample would settle the concern. The ground-truth-label limitation noted by the reader is real but is a scope condition the paper explicitly acknowledges (Section III-B, Section VII-C); it affects practical applicability, not the internal validity of the mathematical comparison. I therefore keep the CONDITIONAL verdict: the concern does not by itself refute the paper, but it identifies a load-bearing step that must be verified before the infinite-gap claim can be accepted. Agreement is partial because the reader's weakest assumption was the gold-standard typed benchmark, whereas I focus on the unproved theorem that produces the benchmark's most extreme value.","tokens_in":23906,"tokens_out":16063,"duration_ms":181438,"concrete_test":"Independently verify Theorem 3. For k1>=1, k2=0, exhibit a prior in D satisfying PK1/PK2 whose posterior confidence in (7) is 0, e.g. mass a at (0,1-b1) in S2 and mass 1-a at a point in S1∩{p+q>b} with positive likelihood; check this point satisfies P+Q>=l. For k1=k2=0, solve the reduced one-dimensional problem over r=p+q and confirm the stated two-point formula. If the zero case admits no feasible exact prior, determine whether the infimum is 0 as a limit and whether the 'no finite n2' claim still holds under the paper's limit-point convention.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline result — that after a single untyped failure no finite amount of additional failure-free evidence can restore 95% confidence, while an assessor aware of failure types is reduced to zero posterior confidence — is a direct consequence of Theorem 3's assertion that the typed-data infimum is 0 whenever k1>=1 or k2>=1. Section V states only that 'Theorem 4 is proved in Appendix A; the other theorems are proved using analogous steps,' but the zero case is not analogous to Theorem 4: PK3's positive lower bounds l1,l2 keep the typed likelihood positive on S2, whereas under PK1 the 'good' support point can have p=0 or q=0, making the likelihood of an observed type-specific failure exactly zero. The proof must establish both the zero-infimum claim and the claimed two-point form for k1=k2=0; if the zero case fails (e.g., because observed counts force positive likelihood on every admissible S2 point once boundary conventions are fixed), the infinite gap collapses to the finite gap already visible in Theorem 4. This is the least secure step in the central argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses how coarse operational data, i.e. failure counts without type labels, affects conservative Bayesian reliability claims for safety-critical AV classifier and monitor software. It extends conservative Bayesian inference (CBI) to a trinomial model with FP and FN failure modes, and derives worst-case posterior confidence bounds on the probability of failure per classification under two partial-prior specifications (PK1+PK2 and PK2+PK3), in both untyped and typed versions (Theorems 1-4). The paper applies these bounds to AV safety assessment: after failure-free operation the four theorems give identical evidence requirements, while after a single failure the untyped analyses (Theorems 1-2) permit finite recovery, the typed analysis under PK3 (Theorem 4) requires substantially more evidence, and the typed analysis under PK1 (Theorem 3) is claimed to yield zero posterior confidence so that no finite amount of additional failure-free evidence can restore confidence. Inequality (9) is offered as a formal comparison between untyped and typed infima, and sensitivity analyses examine prior confidence and the i.i.d. assumption.","tokens_in":24085,"tokens_out":8915,"duration_ms":100115,"significance":"If the results are fully substantiated, the paper makes a useful contribution: it gives closed-form conservative posterior bounds for a two-failure-mode CBI problem, identifies a concrete scenario in which coarser data used with conservative intent is optimistic relative to typed data, and connects the analysis to AV safety standards and lifecycle processes. The supplied proof of Theorem 4 in Appendix A and the proof of the infimum interchange in Appendix C are valuable, and the application figures make the practical stakes clear. The variational formulation of the problem and the explicit two-point worst-case priors are also useful for practitioners. However, three of the four theorems are stated without proofs, and the most dramatic conclusion, the infinite recovery gap, depends on the unproved zero-infimum case of Theorem 3; the contribution will be fully credible only when those proofs are supplied or delegated with sufficient detail.","major_comments":[{"comment":"The sentence \"Theorem 4 is proved in Appendix A; the other theorems are proved using analogous steps\" is not adequate for the load-bearing zero-infimum claim. For k1>=1 or k2>=1, the zero result is not analogous to Theorem 4: under PK1 the 'good' region S2 touches the coordinate axes, so the typed likelihood p^{k1} q^{k2} (1-p-q)^{n-k1-k2} can vanish on S2, whereas PK3's positive lower bounds l1,l2 prevent this. The manuscript should supply the argument explicitly, including the boundary/limit-point convention and the positivity of the denominator. For example, when k1=1 and k2=0, a prior with mass a at (0,l) in S2 and mass 1-a at a point in S1 with p>0 gives numerator zero and denominator positive, yielding value 0 in the closure; the proof should state this and treat the k1=k2=0 case separately. This is the least secure step in the central argument because the infinite-gap conclusion in Section VI depends directly on it.","section":"Section V, Theorem 3"},{"comment":"Theorems 1 and 2, which drive the untyped curves in Figures 6 and 7, are also stated without proof. A reader cannot verify the reduction to two-point priors or the claimed extrema L2* and L1* for the untyped likelihood (p+q)^k(1-p-q)^{n-k} from the statement 'proved using analogous steps.' Since these theorems are used in the application section to quantify the evidence needed for 95% confidence after an untyped failure, the paper should include at least a proof sketch or an appendix treatment, or explicitly state where the complete proof appears.","section":"Section V, Theorems 1 and 2"},{"comment":"Inequality (9) is formally correct as an interchange of infima over a finite set and D, as shown in Appendix C, but as a 'guarantee of conservatism' it is uninformative for k>=1. Because Theorem 3's infimum is zero whenever k1>=1 or k2>=1, the left-hand minimum over k1 is 0 for every k>=1, so the inequality only says the untyped confidence is at least 0. The text should not suggest that (9) quantitatively bounds the optimism gap; the quantitative content is in the comparison of Theorems 1-2 with Theorems 3-4, not in (9) itself.","section":"Section VII-A, inequality (9)"}],"minor_comments":[{"comment":"The formula \"Φ* = ... 11−b1≤b\" appears to be a rendering error for the indicator 1_{1-b1 <= b}; it should be typeset as a subscripted indicator throughout Theorems 1-4 to avoid confusion with the number 11.","section":"Theorems 1-4, notation"},{"comment":"The paper is transparent that typed data requires reliable ground truth and counterfactual reasoning, but given that the headline conclusion is about the danger of untyped data, the conclusion should restate this caveat prominently: the comparison to the 'gold-standard' typed analysis holds only when FP/FN labels are available and reliable.","section":"Section III-B and Section VII-C"},{"comment":"The sentence \"using (9) ensures the CBI-based assessment remains conservative\" should be qualified: for k>=1 the bound is vacuous, and the conservatism guarantee that remains meaningful is the definitional one that holds only when the same evidence and likelihood are used in the typed analysis.","section":"Section VII-A, inequality (9)"},{"comment":"The phrase \"The integer asymptotic supremum on n2 this implies is ceil(b1/(1-b1))\" is unclear: the limit n2 -> b1/(1-b1) is a real limit, and the ceiling appears without justification; please clarify what exactly is being claimed about the integer-valued n2.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The central issue is the missing proof of Theorem 3; the result is likely correct, and a constructive boundary-support argument can probably fill the gap, but the current text does not provide it. The paper also relies on companion work [59] for the fixed-point machinery used in the only supplied proof, so the editor may wish to confirm that the companion results are available or submitted for review. The paper's contribution is otherwise useful and well positioned for a software-reliability or AV-safety venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The genuinely new material is the two-dimensional CBI analysis of typed (FP/FN) failure data: Theorems 3–4 and inequality (9), which lower-bounds the coarse-data confidence by the typed infima. That inequality is proved, Theorem 4 is proved, and Appendix B gives a clean asymptotic derivation for the hump. The second thing: the paper's headline claim — that after a single untyped failure no finite amount of clean operation restores confidence, while typed-data analysis collapses to zero under PK1 — is Theorem 3, and Theorem 3 is not proved. The text says the other theorems are 'proved using analogous steps,' but the zero case is not analogous to Theorem 4. Theorem 4 works because PK3 keeps P and Q bounded away from zero; PK1 allows degenerate points on the axes, so the typed likelihood of an observed FP or FN can vanish. The proof needs to pin down the boundary conventions and establish both the zero infimum and the two-point support form. If that step fails, the infinite gap becomes the finite gap of Theorem 4. So I would not build on the infinite-gap result until the proof exists.\n\nWhat the paper does well: it is honest about its assumptions — i.i.d. demands, reliable ground truth for FP/FN labeling, failure-centric focus, no calibrated mapping from pfc to end-to-end safety metrics; it situates itself in the AV safety standards landscape without getting lost; and it credits Popov for the qualitative 'low-fidelity data can be non-conservative' message, positioning its contribution as quantification rather than discovery. The self-citations to prior CBI work are legitimate: Theorems 1–2 collapse to univariate results in the limit, and the new comparisons are derivations.\n\nSoft spots beyond the missing proof: the 'first conservative estimates' claim is asserted without a systematic literature comparison; the reader's concern about noisy or missing failure-type labels is real and acknowledged in the text, so the typed benchmark is conditional on data that may not exist in practice; the i.i.d. assumption is relaxed only in a sensitivity analysis, not in the main theorems. None of these sinks the paper, but together they mean the contribution is narrower than the abstract suggests.\n\nMy recommendation: send it to serious peer review. It deserves referee time. The referee should ask for proofs of Theorems 1–3 or explicit pointers to where they appear; without Theorem 3 the central shock value is not established. If those proofs materialize, this is a solid, citable result. As it stands, I would cite inequality (9) and Theorem 4 cautiously, but not the infinite gap.","headline":"A useful but partially unproved extension of CBI to typed failure data; the headline infinite-gap result rests on Theorem 3, whose zero-infimum proof is omitted.","tokens_in":24663,"tokens_out":3044,"would_cite":true,"duration_ms":32119,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Coarse operational logs can justify more optimistic safety bounds for AV software than typed failure data, and after a single failure the gap can be infinite.","keywords":["conservative Bayesian inference","software reliability assessment","operational data fidelity","autonomous vehicle safety","binary classifier failure modes","false positive and false negative","probability of failure per classification","worst-case prior"],"falsifier":"Take any concrete instance of the model, say $n=1000$, $k=1$, $b=10^{-2}$, and fixed PK parameters, and compute the infimum in Theorem 1 together with the minimum over $k_1=0,1$ of the Theorem 3 infima. If the untyped infimum is smaller than that minimum, inequality (9) is contradicted and the paper's central claim fails; a simulation over many parameter values can search for such a counterexample. Alternatively, in field data, find an AV safety monitor with logged failure types and compare the coarse-data confidence curve with the typed-data curve after the first failure: the paper predicts the coarse curve stays above the typed curve, and under PK1/PK2 the typed curve drops to zero.","tokens_in":23667,"feed_emoji":"🚗","tokens_out":10736,"duration_ms":103134,"temperature":0.7,"pith_summary":"This paper asks whether a safety assessor who only knows how many times an autonomous-vehicle classifier failed, but not whether each failure was a false positive or a false negative, can still make conservative reliability claims. It extends conservative Bayesian inference (CBI), a worst-case Bayesian method that works with a set of priors consistent with stated evidence, to a binary classifier with two failure modes. The central finding is that coarse, untyped operational data can produce posterior confidence in a bound on the probability of failure per classification that is larger than the confidence justified by typed data, and in one scenario the gap is infinite after a single failure. The paper proves four theorems giving worst-case posterior confidence bounds and derives inequality (9), which lets an assessor correct a coarse-data assessment by taking the minimum over all possible typed decompositions of the observed failures. If the paper is right, safety cases for AV software built from aggregate intervention logs can be dangerously optimistic, and logging systems should record failure types or apply the paper's correction.","feed_headline":"Coarse failure logs can make AV software look safer than it is","feed_subtitle":"Untyped logs hide failure types; conservative Bayesian bounds show the confidence gap can be infinite.","key_machinery":"The engine of the analysis is a two-parameter categorical model of classifier outcomes: unknown probabilities $P$ for false positives and $Q$ for false negatives, with success probability $\\Theta = 1-P-Q$, so the probability of failure per classification is $P+Q$. The assessor specifies only partial prior knowledge, namely PK1, PK2, or PK3, which defines a set $\\mathcal{D}$ of admissible priors over the feasible triangle $\\Omega$ in the unit square. The four theorems compute the infimum posterior confidence in $P+Q\\le b$ by solving a fractional optimization problem whose extremal solution is always a two-point discrete prior: mass $1-a$ placed at the point in the \"not good enough\" region where the likelihood is largest, and mass $a$ at the point in the \"good enough\" region where the likelihood is smallest. The untyped likelihood $(p+q)^k(1-p-q)^{n-k}$ is a convex combination of the typed likelihoods $p^{k_1}q^{k_2}(1-p-q)^{n-k_1-k_2}$, which is the mechanism behind inequality (9): the coarse-data confidence is always at least the minimum typed-data confidence over splits of the failure count.","core_discovery":"The paper's central claim is that ignoring failure-type labels is not automatically conservative: replacing typed counts $(k_1,k_2)$ of false positives and false negatives with a single total $k$ can raise the assessor's conservative posterior confidence in the statement that the probability of failure per classification $P+Q$ is at most $b$. Under prior knowledge PK1, which says the classifier is imperfect so $P+Q\\ge l$, and PK2, which says the assessor has prior confidence $a$ that accuracy exceeds $1-b_1$, Theorem 3 says that any observed false positive or false negative takes the typed-data posterior confidence to zero, so after one failure no finite amount of failure-free operation can restore a 95% confidence target. Theorems 1 and 2, which use only untyped counts, give a positive confidence bound that can be recovered with finitely many further successes. Under the stronger prior PK3, which places positive lower limits on both $P$ and $Q$, Theorem 4 gives a finite recovery requirement that can still be orders of magnitude larger than the untyped requirements. Inequality (9) expresses the general structure of the gap: the untyped-data confidence infimum is bounded below by the minimum, over all ways of splitting $k$ into $k_1$ and $k_2$, of the typed-data confidence infimum. The paper presents this as a demonstration that attempts to use low-fidelity data conservatively can be naive, and as a first conservative estimate of the impact of data fidelity on AV software assessments.","pith_inferences":["Beyond the paper, the same convex-mixture argument should extend to any reliability metric that is a convex or linear function of typed failure counts; metrics that weight false positives and false negatives differently could show an even larger coarse-versus-typed gap.","Beyond the paper, a practical design consequence is that AV safety logging should record failure-type labels at collection time, since aggregate counters cannot be re-derived once counterfactual ground truth is lost.","Beyond the paper, the theorems can be tested on shadow-mode logs where the monitor does not intervene and ground truth is directly observable; the prediction is that coarse summaries will look more optimistic than type-aware summaries as soon as failures appear."],"forward_implications":["Safety evidence from AV operational logs that record only aggregate failure counts can overstate confidence in a classifier's probability of failure per classification, compared with the conservative bound the same evidence would justify if failure types were logged.","Under prior knowledge PK1 and PK2, a single logged failure of known type collapses the conservative posterior confidence to zero, so no finite amount of subsequent failure-free operation can restore a target confidence level.","With the more detailed prior PK3, recovery after one failure is finite but can require orders of magnitude more failure-free classifications than an untyped analysis suggests.","When failure types are unknown, inequality (9) provides a way to keep the assessment conservative: take the minimum typed-data confidence over all possible decompositions of the observed failures.","During failure-free operation all four theorems coincide, so data fidelity only starts to matter once failures are observed."],"supporting_citations":[{"why":"Supplies the road-test CBI framework and the univariate no-failure and single-failure scenarios that Theorems 1-4 extend to two failure modes.","marker":"[17]"},{"why":"Defines conservative Bayesian inference and its guarantee that a prior consistent with stated evidence cannot give a more conservative posterior than the worst case.","marker":"[12]"},{"why":"Provides the white-box/black-box comparison that motivates the paper's point that CBI's conservatism is conditional on the evidence used.","marker":"[60]"},{"why":"Supplies the sensitivity-analysis method for relaxing the i.i.d. assumption, used in Section VII to show that non-i.i.d. models can be even less optimistic.","marker":"[19]"},{"why":"Supplies the epsilon-contamination robust-Bayesian result used in the proof of Theorem 4 to restrict the optimization to two-point extremal priors.","marker":"[28]"},{"why":"Supplies the fixed-point characterization of extremal distributions used to justify worst-case priors as limits of feasible priors in the proofs.","marker":"[59]"}],"fun_headline_variants":["Untyped AV failure logs hide a dangerous optimism gap","One failure, zero recovery: how coarse logs break AV safety claims","Low-fidelity AV logs can't be trusted even with conservative priors","Why conservative use of coarse AV logs is still optimistic","First quantitative proof: failure-type detail is crucial for AV safety"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison assumes every past outcome can be reliably labeled as a success, a false positive, or a false negative using ground truth and counterfactual reasoning; if those labels are noisy or impossible to obtain, the typed-data benchmark that the paper treats as the conservative gold standard may itself be unavailable.","fun_headline_variants_meta":{"raw":{"variants":["Untyped AV failure logs hide a dangerous optimism gap","One failure, zero recovery: how coarse logs break AV safety claims","Low-fidelity AV logs can't be trusted even with conservative priors","Why conservative use of coarse AV logs is still optimistic","First quantitative proof: failure-type detail is crucial for AV safety"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000195,"raw_usage":{"total_tokens":1419,"prompt_tokens":1069,"completion_tokens":350,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":685,"completion_tokens_details":{"reasoning_tokens":264}},"tokens_in":685,"tokens_out":350,"duration_ms":4456,"temperature":1.0,"reasoning_tokens":264,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:23:45.914752+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any concrete instance of the model, say $n=1000$, $k=1$, $b=10^{-2}$, and fixed PK parameters, and compute the infimum in Theorem 1 together with the minimum over $k_1=0,1$ of the Theorem 3 infima. If the untyped infimum is smaller than that minimum, inequality (9) is contradicted and the paper's central claim fails; a simulation over many parameter values can search for such a counterexample. Alternatively, in field data, find an AV safety monitor with logged failure types and compare the coarse-data confidence curve with the typed-data curve after the first failure: the paper predicts the coarse curve stays above the typed curve, and under PK1/PK2 the typed curve drops to zero.","supporting_citations":[{"cited_title":"Assessing the safety and reliability of autonomous vehicles from road testing,","cited_arxiv_id":null,"evidence_quote":"Supplies the road-test CBI framework and the univariate no-failure and single-failure scenarios that Theorems 1-4 extend to two failure modes."},{"cited_title":"Toward a formalism for conservative claims about the dependability of software-based systems,","cited_arxiv_id":null,"evidence_quote":"Defines conservative Bayesian inference and its guarantee that a prior consistent with stated evidence cannot give a more conservative posterior than the worst case."},{"cited_title":"Why black-box bayesian safety assessment of autonomous vehicles is problematic and what can be done about it?","cited_arxiv_id":null,"evidence_quote":"Provides the white-box/black-box comparison that motivates the paper's point that CBI's conservatism is conditional on the evidence used."},{"cited_title":"The unnecessity of assuming statistically independent tests in bayesian software reliability assessments,","cited_arxiv_id":null,"evidence_quote":"Supplies the sensitivity-analysis method for relaxing the i.i.d. assumption, used in Section VII to show that non-i.i.d. models can be even less optimistic."},{"cited_title":"Robust bayesian analysis withϵ- contaminations partially known,","cited_arxiv_id":null,"evidence_quote":"Supplies the epsilon-contamination robust-Bayesian result used in the proof of Theorem 4 to restrict the optimization to two-point extremal priors."},{"cited_title":"Fixed-Point Characterisations of Extremal Distributions under Partial Distributional Constraints","cited_arxiv_id":"2608.04315","evidence_quote":"Supplies the fixed-point characterization of extremal distributions used to justify worst-case priors as limits of feasible priors in the proofs."}],"review_version":1}