{"id":"c0a3a4df-19c7-41bc-b2df-83fbe146e7af","arxiv_id":"2507.13661","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Most current ADS test methods generate impossible scenarios and rely on assumptions of rationality and determinacy that eight open autopilots do not satisfy.","lead":"This paper argues that most current methods for testing self-driving cars create situations the car cannot handle, and that standard testing ideas like worst-case analysis and coverage fail unless the autopilot behaves consistently. It proposes a framework and reports that most of eight open autopilots violate the two consistency properties it identifies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.1 admits that the critical boundary values x̂_a and x̂_f are imprecise, yet Table 2's rationality/failure classifications are computed relative to those estimated A/D functions; without a sensitivity analysis, the 'five of eight autopilots are irrational' result and the impossibility…","rationale":"The paper's main conceptual contribution—the criticality framework and the determinacy measurements—is valuable and largely independent of the A/D estimation issue. The determinacy observations in Figures 10–12 are direct behavioral measurements and would likely survive re-testing. However, the rationality claim, which is essential to the headline conclusion that autopilots must be designed to be rational and determinate, depends on classifying test cases as more or less critical using x̂_a and x̂_f computed from B, T_A, and V_A. The paper itself acknowledges in Section 3.1 that these values cannot be estimated precisely, and the absence of any sensitivity analysis or independent validation means that the Table 2 frequencies and the 'five of eight' count could shift if the A/D estimates are inaccurate. This is a genuine load-bearing concern rather than a manufactured one: the central empirical assertion about widespread irrationality is not yet supported to the precision the conclusion requires. The concern does not refute the paper; it makes the central empirical claim conditional on validation that is currently missing. The reader's CONDITIONAL verdict is therefore the right disposition, and no verdict adjustment is needed. I agree with the reader that the A/D estimation assumption is the weakest point, while noting that the determinacy evidence is more robust than the rationality evidence.","tokens_in":29416,"tokens_out":15827,"duration_ms":192706,"concrete_test":"For at least one modular autopilot (e.g., Apollo) and one end-to-end autopilot (e.g., Carla), independently measure the A/D functions B(v), T_A(x,v), V_A(x,v) from high-fidelity simulator logs over a dense grid of speeds and distances. Recompute x̂_a and x̂_f from the measured functions, re-run the Table 2 exploration, and compare the IS/OF classifications under the estimated versus measured functions, including a perturbation study that varies x̂_a and x̂_f by ±20%. If the set of autopilots labeled irrational or any IS frequency changes materially (e.g., by more than a few percentage points), the Table 2 evidence for the rationality claim is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 explicitly says (p. 21) that 'it is hard to estimate precise values of x̂_a and x̂_f because the A/D functions are not known exactly or due to other measurement errors,' but Table 2's failure taxonomy—especially the irrational-safety (IS) frequencies and the overall-failure (OF) zones—is computed relative to those estimated boundaries. The A/D functions B(v), T_A(x,v), V_A(x,v) are described in §2.3.3 as 'either mathematically defined by autopilots developers or estimated experimentally [20]', yet the paper gives no error bars, no independent validation, and no sensitivity analysis. If the estimates are off, the location of the most critical test case shifts, and test cases labeled 'less critical' can actually be more critical (or vice versa) under the true dynamics. The dominance order in §2.3.3 makes the existence of a matching pass/fail pair robust once such a pair is found, but the reported frequencies ('IS 3.9%', 'IS 1.7–7.3%') and the claim that five of eight autopilots lack rationality depend on which test cases are counted inside/outside the computed safe-progress region. The determinacy measurements in Figures 10–12 are more direct and less sensitive to A/D estimates, but the rationality pillar of the central 'impossibility' claim rests on an uncalibrated estimate. The paper flags this limitation in passing but does not quantify its impact on the main empirical conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for assessing the effectiveness and validity of autonomous driving system (ADS) test methods, distinguishing false acceptance and false rejection. It derives, for elementary adverse intersection/merge scenarios, an optimal criticality order based on inclusion of safe control policies, and identifies the most critical test cases in terms of acceleration/deceleration (A/D) functions. The authors then report experiments on eight open autopilots, claiming that most lack rationality (passing a critical test does not imply passing less critical tests) and determinacy (behavior from visited states is inconsistent with behavior from the initial state), and conclude that current ADS testing cannot provide required safety guarantees unless autopilots are designed to be rational and determinate.","tokens_in":29838,"tokens_out":5252,"duration_ms":60212,"significance":"If the central claims hold, the paper makes a substantial contribution by providing a principled, partly formal basis for comparing ADS test methods and by identifying two concrete autopilot design properties that would make testing tractable. The kinematic derivation of most-critical test cases in Section 2.3.3 is clean under the stated assumptions, and the experimental study across eight open autopilots is a useful concrete illustration. The paper also offers an actionable test methodology (Section 3.4) that combines rationality, determinacy, and monotonicity to reduce the number of required test cases. However, the empirical pillar of the paper—especially the reported frequencies of irrationality and the five-of-eight claim—rests on estimated A/D functions whose precision is admitted to be limited, and no sensitivity analysis is provided. The monotonicity assumptions in Section 3.3 are also not validated for the tested autopilots. These issues must be addressed before the main conclusions can be regarded as quantitatively grounded.","major_comments":[{"comment":"The manuscript states in Section 3.1 that 'it is hard to estimate precise values of x̂_a and x̂_f because the A/D functions are not known exactly or due to other measurement errors,' yet Table 2 reports specific irrational-safety (IS) frequencies (e.g., 1.7%–7.3%) and the conclusion states that five of eight autopilots lack rationality. These classifications depend on which test cases lie inside or outside the computed safe-progress region, which in turn depends on the estimated A/D functions B(v), T_A(x,v), and V_A(x,v). If the estimates are biased, test cases labeled IS could instead be transition failures or irrelevant cases, and the number of autopilots classified as irrational could change. Please provide a sensitivity analysis that perturbs x̂_a and x̂_f within plausible error bounds and re-derives the classifications, or report confidence intervals for the IS frequencies. Without this, the central empirical claim is not robust to the acknowledged measurement error.","section":"§3.1, Table 2"},{"comment":"The test-case partitioning technique and the coverage ratio calculation assume monotonicity of the A/D functions (V_B monotonic in v, T_A and V_A monotonic in the appropriate directions). This assumption is stated but not established for the eight tested autopilots. In fact, the non-determinacy results in Figures 10–12 show that braking and acceleration behavior can depend on the state in ways that violate simple composability, and it is not shown whether monotonicity holds in the tested data. Please validate the monotonicity assumptions on the tested autopilots, or clearly restrict the proposed methodology to autopilots for which monotonicity has been verified. As written, the coverage improvement claim in Section 3.4 is conditional on an unverified property.","section":"§3.3, Figures 13–15"},{"comment":"The derivation of the most critical test case assumes that the ego vehicle's progress policy is safe exactly when T_A(x_e,v_e) ≤ x_a/v_l and B(V_A(x_e,v_e)) ≤ x_f, with equality defining the critical boundary. This is a simplified model that neglects, for example, the finite length of the critical zone and the time spent within it, as acknowledged by the use of a distance d in the merge example in Section 2.3.1. The paper should state clearly that the derived critical values are approximations under idealized point-vehicle kinematics, and should discuss how the finite size of the conflict zone affects the boundary. This is not a fatal flaw, but it strengthens the need for the sensitivity analysis requested above.","section":"§2.3.3, Eq. for x̂_a and x̂_f"}],"minor_comments":[{"comment":"In the text accompanying Figure 6, the condition for no safe progress policy is written as 'x_a ≤ x̂_a or x̂_f ≤ x_f'; this should be 'x_f ≤ x̂_f' for consistency with the earlier definitions.","section":"§3.1, Figure 6"},{"comment":"The inequality 'T_A(v_i,x_e) < V_A(v_{i+1},x_e)' appears to be a typo; it should compare travel times, e.g., 'T_A(v_i,x_e) < T_A(v_{i+1},x_e)'.","section":"§3.3, paragraph before Figure 15"},{"comment":"The frequencies in Table 2 are reported as percentages with no denominators or confidence intervals. Please clarify how IS (and TF) percentages are computed, the number of test cases in each scenario type, and whether the percentages are over all explored (x_a,x_f) pairs or only over the safe-progress region. This would make the table interpretable and reproducible.","section":"Table 2"},{"comment":"The claim that 'all eight tested autopilots exhibit non-determinacy for basic acceleration policies' is supported only by the illustrative example in Figure 12 and summary statements. Please report quantitative measures of non-determinacy (e.g., differences in achieved speed or braking distance across state-restart experiments) for each autopilot, so that the claim is verifiable.","section":"§3.2, Conclusions"},{"comment":"The A/D functions are said to be 'either mathematically defined by autopilots developers or can be estimated experimentally [20]' in Section 2.3.3. Since reference [20] is the authors' own prior work, an independent description of the estimation procedure (e.g., number of trials, simulator settings, error metrics) would strengthen reproducibility and help readers judge the reliability of the critical boundaries.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is positioned as a broad analysis of ADS testing, and the empirical section is used to support a strong impossibility-style conclusion. The main risk is that the empirical support is more fragile than the presentation suggests: the A/D-function estimates that drive the rationality classifications are admitted to be imprecise, and no sensitivity analysis is given. The theoretical framework itself is a useful contribution and the derivations are largely sound. If the authors can supply robustness checks for the empirical claims and either validate or clearly bound the monotonicity assumptions, the paper could become a strong candidate for acceptance. I would also encourage the authors to consider whether the five-of-eight claim is necessary for the main message, or whether a weaker claim ('at least several of the tested autopilots lack rationality') would be more defensible given the acknowledged measurement uncertainties."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this paper gives the clearest statement I've seen of rationality and determinacy as testability requirements for autopilots, and it backs that with a real comparison of eight open autopilots. The determinacy measurements are direct and convincing: braking and acceleration curves that diverge depending on the starting state are hard to explain away. That part is new and worth engaging with seriously.\n\nThe kinematic derivation in Section 2.3.3 is clean under the stated assumptions. The definition of the optimal criticality order via inclusion of safe policy sets is sound, and the reduction of elementary adverse scenarios to a few parameters is helpful. The paper also does a fair job of surveying the state of the art, and the critique of NPC-based testing that ignores nominal capabilities is well placed.\n\nThe soft spot is exactly where the stress-test note lands. The rationality/failure taxonomy in Table 2 is computed relative to critical boundaries x̂_a and x̂_f, which depend on A/D functions that are either mathematically defined or estimated experimentally. Section 3.1 admits these are hard to estimate precisely, but the paper never quantifies how much error in those estimates would move the boundaries or change the reported IS frequencies. The determinacy results are much less sensitive to this, but the claim that five of eight autopilots lack rationality, and the broader impossibility conclusion, rest on that uncalibrated foundation. The monotonicity assumption in Section 3.3 is also asserted, not verified for the tested autopilots.\n\nThe conclusion that \"it seems impossible to obtain the safety guarantees\" is broader than the evidence: eight open autopilots and four scenario types do not prove impossibility, they demonstrate a pattern. That is a real gap between the strength of the claim and the support.\n\nThe paper is largely an extension of the authors' prior work, so novelty is moderate, but the explicit formalization of rationality and determinacy and the eight-autopilot comparison are genuinely new. The self-citation is not itself a problem, though the reliance on [20] for A/D estimation is exactly where the sensitivity analysis is missing.\n\nFor the right reader, this is a valuable paper: it gives the ADS testing community a vocabulary and a set of concrete properties to argue about. It deserves a serious referee. The referee should push for sensitivity analysis on the A/D estimates and a more careful scoping of the impossibility claim, but the core conceptual contribution is solid.","headline":"A useful conceptual framework and an honest but under-calibrated empirical claim: the determinacy results are robust, but the rationality frequencies and the impossibility conclusion lean on unquantified A/D estimates.","tokens_in":30298,"tokens_out":1537,"would_cite":true,"duration_ms":18904,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Passing a critical test does not imply passing easier ones for most current autopilots, so strong safety guarantees stay out of reach unless autopilots are designed to be rational and determinate.","keywords":["autonomous driving systems","safety testing","critical scenario generation","false rejection","autopilot rationality","autopilot determinacy","acceleration/deceleration functions","test coverage"],"falsifier":"Measure the actual A/D functions of a given autopilot with high precision, then run the maximally critical test case $(\\hat{x}_a,\\hat{x}_f)$ and every less critical test case in the ordered region; if any autopilot passes the critical case and all less critical cases without exception, or if restarting it from intermediate states of a braking or acceleration policy always reproduces the original trajectory, the paper's claim that these properties are systematically absent would be contradicted. A second check: if critical test methods already avoided non-ego-attributable accidents at negligible rates, the false-rejection claim would fail.","tokens_in":29241,"feed_emoji":"🚗","tokens_out":6120,"duration_ms":61763,"temperature":0.7,"pith_summary":"This paper argues that current methods for testing autonomous driving systems cannot deliver the safety guarantees they promise, because they rest on two assumptions that real autopilots violate and because many critical tests place vehicles in situations that are physically impossible to handle. It proposes a framework that separates test methods by effectiveness and validity, and defines an optimal criticality order for elementary adverse scenarios—single conflicts between the ego vehicle and one arriving or front vehicle—using acceleration and deceleration functions. Applying the framework to eight open autopilots, the paper reports that most fail to be rational (passing a critical test does not imply passing less critical tests) and that all fail to be determinate (behavior from a visited state need not match behavior from the initial state). It concludes that, as things stand, the required guarantees cannot be obtained unless autopilots are designed from the outset to be rational and determinate.","feed_headline":"Eight open autopilots fail the two assumptions safety testing rests on","feed_subtitle":"Critical-scenario tests assume rational, repeatable autopilots; tests on eight systems show both assumptions break.","key_machinery":"The load-bearing objects are the A/D functions: the braking function $B(v)$ giving the distance needed to stop from speed $v$, the acceleration time function $T_A(x,v)$ giving the time to cover distance $x$ from speed $v$, and the acceleration speed function $V_A(x,v)$ giving the speed reached over distance $x$. From these, the paper computes the most critical safe test case for an elementary adverse scenario as $\\hat{x}_a = T_A(x_e,v_e)\\cdot v_l$ and $\\hat{x}_f = B(V_A(x_e,v_e))$, where $x_e,v_e$ are the ego vehicle's distance and speed, $v_l$ is the speed limit of the arriving vehicle, and $\\hat{x}_a,\\hat{x}_f$ are the critical distances of the arriving and front vehicles. This critical boundary defines an optimal criticality order: any test case with larger $x_a,x_f$ has a superset of safe policies. Rationality and determinacy are then defined against this boundary—rationality as consistency in passing ordered test cases, determinacy as consistency of the policy when restarted from intermediate states—and the paper shows how monotonicity of the A/D functions permits partitioning the speed domain to reduce test counts.","core_discovery":"On the paper's own terms, the central discovery is that test effectiveness and validity for autonomous driving systems are not just properties of the test method; they depend on design properties of the autopilot under test. Most existing methods implicitly assume rationality—that success on a maximally critical test case implies success on all less critical ones—and determinacy—that success from an initial state implies success from every state visited along the policy. Experiments on eight open autopilots, covering four scenario types (merge with yield, lane change, intersection with yield, intersection with traffic lights), show widespread violations: transition safety failures at rates up to 71.5% in one scenario type, irrational overcaution, and nondeterminate braking for seven of eight autopilots and nondeterminate acceleration for all eight. The paper therefore maintains that critical-scenario testing as currently practiced produces both false rejections (scenarios no vehicle could handle under nominal capabilities) and false acceptances, and that strong safety guarantees are impossible under the current state of the art.","pith_inferences":["A testable implication for regulators is that acceptance criteria for autonomous driving systems should include demonstrations of rationality and determinacy on benchmark scenarios, not just accident counts.","The framework's reliance on A/D function estimation means the measured violation rates are conditional on those estimates; if better measurement shifts the critical boundary, some transition failures may be reclassified as nominal or vice versa.","The rationality requirement connects to adversarial robustness: an autopilot that is rational in this ordered sense has decisions that change monotonically with scenario difficulty, which could be checked cheaply in simulation before road testing.","For end-to-end neural autopilots, the paper's argument implies that training objectives should include consistency across scenario criticality and state restart, not just collision avoidance."],"forward_implications":["If autopilots are rational and determinate, a small number of maximally critical test cases can replace a very large test space; the paper provides a concrete partitioning methodology and a coverage-rate estimate.","Most current critical-scenario methods generate a large fraction of accidents attributable to non-ego vehicles, so reported failure rates overstate autopilot risk and induce false rejections.","End-to-end AI autopilots show more irrational safety failures than modular ones, suggesting that the rationality assumption is especially unsafe for learning-based designs.","Without determinacy, even checking a simple braking-to-a-stop property may require infinitely many tests, since success from the initial state does not transfer to states on the braking curve.","A practical design recommendation follows: autopilots should be built so that control policies are generated by iterating a transition function, which yields determinate policies.","The framework's requirement for monotonic A/D functions implies that test-space partitioning and the resulting coverage estimates hold only when the autopilot's acceleration and deceleration behavior is monotonic in speed."],"supporting_citations":[{"why":"Supplies the acceleration/deceleration functions and the prior four-autopilot test basis that this framework extends.","marker":"[20]"},{"why":"Provides the four end-to-end autopilot evaluation used for the irrationality and overall-failure measurements.","marker":"[21]"},{"why":"Doppelgänger test generation evidence that most generated accidents are attributed to non-ego vehicles, grounding the false-rejection claim.","marker":"[13]"},{"why":"Statistical demonstration that generic road testing requires over 11 billion miles, motivating critical-scenario testing.","marker":"[14]"},{"why":"Defines criticality via inclusion of safe policy sets, which the optimal criticality order builds on.","marker":"[15]"},{"why":"The fuzzing method whose generated accidents were mostly attributable to non-player vehicles in the cited manual analysis.","marker":"[22]"},{"why":"Reported that 32.3% of generated accidents were attributable to non-player vehicles, quantifying false rejection in trajectory-based methods.","marker":"[24]"},{"why":"Manual analysis attributing 151 of 192 accidents to non-player vehicles in the cited fuzzing method, load-bearing for the false-rejection claim.","marker":"[25]"}],"fun_headline_variants":["Autopilot safety tests rest on rationality and determinacy that autopilots lack","Eight autopilots break the two assumptions that make safety tests valid","Why self-driving tests reject cars unfairly: autopilots aren't rational","Critical scenarios can't validate autopilots that aren't rational or deterministic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework's computed critical boundaries and all failure classifications rest on estimates of the autopilot's acceleration and deceleration functions, which the paper admits are not known exactly, so inaccurate estimates could move the boundaries and change the measured rationality and determinacy violations.","fun_headline_variants_meta":{"raw":{"variants":["Autopilot safety tests rest on rationality and determinacy that autopilots lack","Eight autopilots break the two assumptions that make safety tests valid","Why self-driving tests reject cars unfairly: autopilots aren't rational","Critical scenarios can't validate autopilots that aren't rational or deterministic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000432,"raw_usage":{"total_tokens":2258,"prompt_tokens":1054,"completion_tokens":1204,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":670,"completion_tokens_details":{"reasoning_tokens":1120}},"tokens_in":670,"tokens_out":1204,"duration_ms":13751,"temperature":1.0,"reasoning_tokens":1120,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:19:11.307992+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the actual A/D functions of a given autopilot with high precision, then run the maximally critical test case $(\\hat{x}_a,\\hat{x}_f)$ and every less critical test case in the ordered region; if any autopilot passes the critical case and all less critical cases without exception, or if restarting it from intermediate states of a braking or acceleration policy always reproduces the original trajectory, the paper's claim that these properties are systematically absent would be contradicted. A second check: if critical test methods already avoided non-ego-attributable accidents at negligible rates, the false-rejection claim would fail.","supporting_citations":[{"cited_title":"Rigorous Simulation-based Testing for Autonomous Driving Systems -- Targeting the Achilles' Heel of Four Open Autopilots","cited_arxiv_id":"2405.16914","evidence_quote":"Supplies the acceleration/deceleration functions and the prior four-autopilot test basis that this framework extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Statistical demonstration that generic road testing requires over 11 billion miles, motivating critical-scenario testing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines criticality via inclusion of safe policy sets, which the optimal criticality order builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The fuzzing method whose generated accidents were mostly attributable to non-player vehicles in the cited manual analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reported that 32.3% of generated accidents were attributable to non-player vehicles, quantifying false rejection in trajectory-based methods."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Manual analysis attributing 151 of 192 accidents to non-player vehicles in the cited fuzzing method, load-bearing for the false-rejection claim."}],"review_version":1}