{"id":"050e1e7a-ee5f-47a2-8b38-2402411bd9ee","arxiv_id":"2507.20685","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A position paper recommending that automated driving safety assurance separate AI-specific risks from open-context uncertainties and use behavior-based analyses to bridge them.","lead":"This position paper argues that safety challenges for AI-based automated vehicles are not uniquely caused by AI; many stem from the open driving context. It proposes using behavior-based analyses to link system-level safety to AI component requirements, which could help manufacturers navigate a tangled standards landscape.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central claim depends on a lossless decomposition from behavior specifications to AI-component metrics that is asserted but not demonstrated and is explicitly conceded as future work in Sec. V-D.","rationale":"The reader's weakest assumption identifies the same gap: decomposability of behavior specifications into AI-level requirements. This is indeed the most load-bearing point because the paper's contribution is not merely that engineering rigor is necessary (a near-tautology) but that a behavior-based approach can ground the decomposition to AI metrics. The paper's own admission in Sec. V-D that the connecting metamodel is missing makes the assertion unsupported, not internally inconsistent. I agree with the CONDITIONAL verdict: the proposal is coherent and valuable, but its central claim should be read as a research agenda. The proposed concrete test would provide evidence for or against the decomposition in a small but nontrivial scenario; until such a test is done, the claim remains an unvalidated assertion.","tokens_in":12421,"tokens_out":4806,"duration_ms":55678,"concrete_test":"Work the Sec. V occluded-VRU case study end-to-end. From the safety goal 'collisions with VRUs must be prevented' and the behavioral competency 'responding to occluded VRUs,' formally derive concrete AI perception performance requirements (e.g., minimum recall of occluded-pedestrian detection at stated ranges) and state the traceability argument. Then run a closed-loop simulation of the ego vehicle in the specified ODD with a perception model that meets exactly those derived thresholds and check whether the safety goal is guaranteed across all scenario variations. If a collision occurs despite meeting the derived metrics, the decomposition is lossy and the central claim overstates what behavior-based analysis delivers; if no collision occurs, the decomposition is at least sufficient for this scenario and the concern is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that behavior-based analyses (ODD, behavior specification, behavioral competencies) provide 'solid grounds' for defining system-level requirements and their decomposition to AI-related performance metrics. The linchpin is the decomposition step: without it, the paper's answer to 'What's really different with AI?' reduces to a truism about engineering rigor. This step is not demonstrated. Section IV-B4 states that behavioral competencies 'can provide guidance' for defining pass-fail and AI performance criteria, but the only worked example (Sec. V-C1) yields a data requirement (label occluded areas), not a performance metric, and no argument that satisfying such requirements guarantees the system-level safety goal. Section V-D explicitly concedes that a comprehensive metamodel connecting behavior, ODD, and AI-specific needs 'is still yet to be established.' The risk is that the decomposition is lossy: AI components could satisfy the derived metrics while the system still exhibits hazardous behavior due to unmodeled interactions among perception, prediction, planning, and control. If so, behavior-based analysis does not close the gap claimed, and the paper's recommendation to anchor AI assurance in behavior specification would be premature. The paper is honest about the gap, but the conclusion still asserts sufficiency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that safety assurance for AI-based automated driving systems is hampered by imprecise AI definitions and by an overemphasis on AI-specific challenges. The authors propose a behavior-based perspective in which system-level safety analyses—using ODD, behavior specifications, and behavioral competencies—can provide grounds for connecting system-level requirements to AI-related performance metrics. The paper surveys definitions in the EU AI Act and ISO standards, discusses sources of uncertainty following Burton and Herd, and presents a short illustrative case study on occluded pedestrians. It concludes with recommendations for standards navigation and future research.","tokens_in":12593,"tokens_out":5210,"duration_ms":51810,"significance":"If substantiated, the proposed behavior-based framework would be a valuable integration of SOTIF, functional safety, and AI safety analyses, and would offer practical guidance for the application of ISO/PAS 8800. The paper is a useful and mostly well-grounded position statement: it correctly observes that many assurance challenges attributed to AI actually stem from open-context uncertainty, and it leverages established systems engineering concepts. The authors are transparent about the need for future work (Section V-D). Its main weakness is that the central decomposition claim is asserted rather than demonstrated; the illustrative case study stops at a dataset requirement and does not show a traceable performance metric.","major_comments":[{"comment":"The claim that behavioral competencies \"can provide guidance for a traceable definition of pass-fail criteria from system-level safety indicators, as well as the definition of meaningful performance criteria for AI components\" is not demonstrated. The case study ends with a dataset requirement (labeled occluded areas) rather than an AI performance metric, and no argument is provided that satisfying such a requirement guarantees, or even measurably advances, the system-level safety goal of preventing collisions with pedestrians. This is load-bearing because the paper's answer to \"What's really different with AI?\" rests on this decomposition. Please either extend the example to a concrete performance metric with a traceability argument, or reframe the claim as a research hypothesis that requires the metamodel identified in Section V-D.","section":"§IV-B4 and §V-C1"},{"comment":"The admission that \"a comprehensive approach that has been fully connected to AI-specific needs is still yet to be established\" is in tension with the strength of the conclusion that behavior-based analyses \"can provide solid grounds\" for the decomposition. Given the conceded gap, the conclusion should be conditioned on the future development of the metamodel, or the authors should specify conditions under which the decomposition is expected to be lossless—particularly with respect to unmodeled interactions among perception, prediction, planning, and control components. Without such qualification, the central contribution is an interesting assertion rather than a supported result.","section":"§V-D and §VI"}],"minor_comments":[{"comment":"There is a typo in the quoted definition: \"abscence\" should be \"absence\".","section":"Abstract and §III-A"},{"comment":"The discussion of differing AI definitions would benefit from a table comparing the EU AI Act, ISO/IEC 22989, ISO/IEC TR 24028, and ISO/PAS 8800 to improve readability and scannability.","section":"§II"},{"comment":"The relationship between \"risk\" and \"assurance uncertainty\" in Fig. 1a is stated but not formally defined; a brief mapping would help readers follow the rephrasing.","section":"§III-A"},{"comment":"The reference to Koopman's \"Machine Learning breaks the Vee\" and the Bayesian analogy is underdeveloped; a sentence on how the prior is updated by evidence in the assurance setting would clarify the point.","section":"§III-C"},{"comment":"The sentence \"The elements of the system context are, e.g., sources of requirements for must-have class labels\" is terse; specify how ODD elements map to input-space requirements for ML models.","section":"§IV-B2"},{"comment":"References [27] and [42] contain repeated author/consortium names; please clean them up.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a fitting position paper for IAVVC, and the topic is timely. The requested major revisions are moderate: either strengthen the illustrative example or constrain the central claims. I do not see a need for additional experimentation, but the decomposition claim should be presented as an open problem if it is not further evidenced."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a position paper, not a result paper, and it is a solid one. The genuinely new part is the framing: the authors argue that much of what gets labeled AI-specific assurance burden in automated driving is actually open-context uncertainty that would exist regardless of implementation. That separation, anchored in Burton and Herd's uncertainty model and tied to ISO/PAS 8800's gap between system-level and AI-level analyses, is well argued and practically useful. The mapping of systems engineering concepts (operational concept, ODD, behavioral competencies) onto ADS safety is clear and should help practitioners navigate the standards landscape.\n\nThe paper is also honest about its own status. The case study is explicitly hypothetical and illustrative. The authors state in Section V-D that a comprehensive metamodel connecting behavior, ODD, and AI-specific needs is still to be established. They call the behavior-based approach a starting point. That is the right register.\n\nThe soft spot is the conclusion's slide from these analyses can provide solid grounds to something close to sufficiency. The linchpin claim is that behavioral competencies and behavior specifications can be decomposed into AI-component-level performance metrics without losing the system-level safety meaning. That is asserted, not shown. The only worked example produces a dataset requirement (label occluded areas), not a performance metric, and no argument that meeting such requirements guarantees the safety goal. The authors themselves concede the traceability machinery is missing. So the central claim is a reasonable research hypothesis, not an established finding. The paper would be stronger if the conclusion said we have identified where the decomposition must be built rather than implying the decomposition is already available.\n\nThe citation pattern is fine. The authors lean on their own prior work, but the core conceptual separation is grounded in Burton and Herd, and the self-citations are relevant, not padding. This is not circular; it is a research group synthesizing its own line of work into a broader argument.\n\nWho gets value: safety engineers and researchers in ADS assurance who want a clear vocabulary for separating AI-specific and open-context risk, and who need a concise map of the standards landscape. It is a legitimate proposal that deserves serious referee attention, even though the validation is future work. I would accept it for peer review and, if I were the editor, ask the authors to temper the conclusion or add a concrete sketch of how the decomposition would be validated.","headline":"A useful, well-grounded position paper on separating AI-specific from open-context risk in ADS safety, but its central claim about decomposing behavior specifications into AI performance metrics is asserted, not demonstrated.","tokens_in":736,"tokens_out":1003,"would_cite":true,"duration_ms":27306,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper asserts that traditional engineering rigor—captured in ODD, behavior specification, and behavioral competencies—is a necessary condition for safe AI-based automated driving, and that this behavior-based layer enables traceable…","keywords":["automated driving","AI safety assurance","behavior specification","operational design domain","systems engineering","scenario-based safety analysis","ISO/PAS 8800","SOTIF"],"falsifier":"A concrete counterexample would be an AI-specific risk that cannot be expressed as a behavioral competency at the system level, such as a nondeterministic failure mode of a perception model that produces hazardous behavior in situations the behavior specification deems safe, or a documented case where a system passes all behavior-level competencies yet still exhibits an AI-induced hazardous event that no ODD or behavior artifact could have captured.","tokens_in":12216,"feed_emoji":"🚗","tokens_out":4367,"duration_ms":45911,"temperature":0.7,"pith_summary":"This position paper asks what is genuinely new, and genuinely hard, about safety assurance when automated driving systems (ADS) contain AI components. The answer it defends is that the core challenge is not the AI itself but the open context in which the vehicle must operate, and the discipline needed to handle that context is the same engineering rigor that traditional safety and systems engineering already provide. The paper claims that a behavior-based approach, defining the Operational Design Domain, specifying required behavior in scenarios, and deriving behavioral competencies, can ground system-level safety indicators and decompose them into AI-specific performance metrics. If correct, most AI-specific assurance work can attach to existing systems-engineering artifacts rather than requiring a parallel safety paradigm.","feed_headline":"Engineering rigor, not AI magic, makes self-driving safe","feed_subtitle":"A position paper shows how behavior-based analysis can trace system safety from the open road down to AI performance metrics.","key_machinery":"The load-bearing mechanism is the behavior-based safety analysis chain: stakeholder needs, use cases and abstract scenarios, ODD and system context, behavior specification, and behavioral competencies. Each step stays solution- and technology-neutral, so hazard analysis can be performed at the level of maneuvers and required capabilities before any AI-specific or implementation-specific analysis begins. The paper maps these artifacts onto systems-engineering concepts such as the operational concept and capabilities, and argues that this chain lets safety engineers trace a system-level safety indicator, such as 'collisions with vulnerable road users must be prevented', down to an AI performance metric, such as 'the dataset must contain labeled occluded areas', with behavioral competencies serving as the intermediate pass-fail criteria.","core_discovery":"The authors' central claim is that the risks introduced by AI-based components in automated driving are real but narrow: they are performance insufficiencies of AI models, captured in the inner rings of Burton and Herd's uncertainty model. Everything else that complicates assurance, such as unpredictability of the open world, incomplete knowledge, and system complexity, is shared by any ADS, whether or not it uses AI. The paper therefore argues that engineering rigor, understood as diligent problem-space analysis through operational concepts, ODD, behavior specification, and behavioral competencies, is a necessary condition for building safe AI-based systems, and that this behavior-based foundation provides the missing link that ISO/PAS 8800 calls for: traceable decomposition of system-level safety metrics into AI component performance metrics. The illustrative occlusion-pedestrian case study shows how a SOTIF functional insufficiency (inability to predict occluded pedestrians) translates first into a behavior-level competency and then into a concrete dataset requirement (labeled occluded areas).","pith_inferences":["If the behavior-based bridge works, regulators could focus on ODD, behavior, and competencies as the assurance backbone, treating AI-specific metrics as implementation details to be checked against that backbone.","The same decomposition logic could generalize beyond driving to other open-context AI systems, such as robots or drones, where an abstract behavior specification can anchor downstream AI assurance.","A testable extension would be to build the missing metamodel and run it on a set of known incidents to see whether every AI-contributed factor maps back to a behavior-level competency gap; any that do not would signal the decomposition needs revision.","The position paper implies that investment in formalizing ODD and behavior reasoning may yield higher safety leverage than investment in AI-specific verification tools alone."],"forward_implications":["Safety analyses can be performed once at the behavior level and reused for functional safety, SOTIF, and AI safety analyses, reducing duplicated effort across standards.","Standards with broad AI definitions can be scoped out of early development stages, because problem-space analyses are technology-neutral.","AI component requirements, such as dataset labels and coverage criteria, become derived artifacts from behavior-level safety goals rather than ad hoc additions.","Behavioral competencies provide a natural basis for pass-fail criteria in verification and validation, linking system-level safety indicators to test outcomes.","The identified gap in traceability between specification, test results, and field monitoring motivates further work on ontologies and metamodels for behavior specification."],"supporting_citations":[{"why":"Supplies the uncertainty model that separates open-context uncertainty from AI-specific uncertainty, which the paper uses to argue that only the inner rings are truly AI-specific.","marker":"[2]"},{"why":"ISO/PAS 8800 defines AI system and AI model in the narrow sense the paper adopts, and explicitly demands traceability between system-level and AI-specific analyses without providing methodical guidance.","marker":"[4]"},{"why":"Troubitsyna et al. emphasize that safety is a system property and call for relating AI performance metrics to system-level safety metrics, the direct research gap the paper addresses.","marker":"[20]"},{"why":"Koopman's claim that 'Machine Learning breaks the Vee' is the position the paper challenges, framing the prior-building challenge that motivates the behavior-based approach.","marker":"[24]"},{"why":"The AVSC best practice supplies the behavioral competency concept used as the intermediate artifact for decomposing system-level indicators into pass-fail criteria.","marker":"[29]"},{"why":"Provides an ontology-based approach for traceable behavior specifications, which the paper presents as a first step toward the needed metamodel.","marker":"[32]"},{"why":"The INCOSE Systems Engineering Handbook supplies the operational concept and problem-space vocabulary that the behavior-based approach instantiates.","marker":"[35]"},{"why":"ISO 21448 (SOTIF) is used in the case study to identify functional and specification insufficiencies, showing how behavior-level analysis connects to an established safety domain.","marker":"[37]"}],"fun_headline_variants":["AI risks are narrow; open-world uncertainty is the real challenge","Behavior-based analysis links AI performance to safe driving systems","For safe self-driving, separate AI model gaps from open-context risk","Trace system safety down to AI metrics with behavior-based engineering"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a behavior specification expressed through abstract scenarios and behavioral competencies can be decomposed without loss into AI-component-level requirements and performance metrics, so that system-level safety remains traceable down to AI-specific indicators; the paper itself acknowledges that the comprehensive metamodel for this is still to be established.","fun_headline_variants_meta":{"raw":{"variants":["AI risks are narrow; open-world uncertainty is the real challenge","Behavior-based analysis links AI performance to safe driving systems","For safe self-driving, separate AI model gaps from open-context risk","Trace system safety down to AI metrics with behavior-based engineering"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000398,"raw_usage":{"total_tokens":2051,"prompt_tokens":884,"completion_tokens":1167,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":1097}},"tokens_in":500,"tokens_out":1167,"duration_ms":11213,"temperature":1.0,"reasoning_tokens":1097,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T13:21:06.960919+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete counterexample would be an AI-specific risk that cannot be expressed as a behavioral competency at the system level, such as a nondeterministic failure mode of a perception model that produces hazardous behavior in situations the behavior specification deems safe, or a documented case where a system passes all behavior-level competencies yet still exhibits an AI-induced hazardous event that no ODD or behavior artifact could have captured.","supporting_citations":[{"cited_title":"Addressing Uncertainty in the Safety Assurance of Machine-Learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the uncertainty model that separates open-context uncertainty from AI-specific uncertainty, which the paper uses to argue that only the inner rings are truly AI-specific."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ISO/PAS 8800 defines AI system and AI model in the narrow sense the paper adopts, and explicitly demands traceability between system-level and AI-specific analyses without providing methodical guidance."},{"cited_title":"Methods and Tools for the Engineering and Assurance of Safe Autonomous Systems,","cited_arxiv_id":null,"evidence_quote":"Troubitsyna et al. emphasize that safety is a system property and call for relating AI performance metrics to system-level safety metrics, the direct research gap the paper addresses."},{"cited_title":"L145 Challenges in Autonomous Vehicle Safety Assessment,","cited_arxiv_id":null,"evidence_quote":"Koopman's claim that 'Machine Learning breaks the Vee' is the position the paper challenges, framing the prior-building challenge that motivates the behavior-based approach."},{"cited_title":"A VSC Best Practice for Evaluation of Behavioral Competencies for Automated Driving System Dedicated Vehicles (ADS-DVs),","cited_arxiv_id":null,"evidence_quote":"The AVSC best practice supplies the behavioral competency concept used as the intermediate artifact for decomposing system-level indicators into pass-fail criteria."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ISO 21448 (SOTIF) is used in the case study to identify functional and specification insufficiencies, showing how behavior-level analysis connects to an established safety domain."}],"review_version":1}