REVIEW 2 major objections 6 minor 44 references
What's Really Different with AI? -- A Behavior-based Perspective on System Safety for Automated Driving Systems
T0 review · 2 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper asserts that traditional engineering rigor—captured in ODD, behavior specification, and behavioral competencies—is a necessary condition for safe AI-based automated driving, and that this behavior-based layer enables traceable…
desk verdict A useful, well-grounded position paper on separating AI-specific from open-context risk in ADS safety, but its central claim about decomposing behavior specifications into AI performance metrics is asserted, not demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the behavior-based safety analysis chain: stakeholder needs, use cases and abstract scenarios, ODD and system context, behavior specification, and behavioral competencies. Each step stays solution- and technology-neutral, so hazard analysis can be performed at the level of maneuvers and required capabilities before any AI-specific or implementation-specific analysis begins. The paper maps these artifacts onto systems-engineering concepts such as the operational concept and capabilities, and argues that this chain lets safety engineers trace a system-level safety indicator, such as 'collisions with vulnerable road users must be prevented', down to an AI performance metric, such as 'the dataset must contain labeled occluded areas', with behavioral competencies serving as the intermediate pass-fail criteria.
What would settle it
A concrete counterexample would be an AI-specific risk that cannot be expressed as a behavioral competency at the system level, such as a nondeterministic failure mode of a perception model that produces hazardous behavior in situations the behavior specification deems safe, or a documented case where a system passes all behavior-level competencies yet still exhibits an AI-induced hazardous event that no ODD or behavior artifact could have captured.
Extended reading notes
Core claim
The authors' central claim is that the risks introduced by AI-based components in automated driving are real but narrow: they are performance insufficiencies of AI models, captured in the inner rings of Burton and Herd's uncertainty model. Everything else that complicates assurance, such as unpredictability of the open world, incomplete knowledge, and system complexity, is shared by any ADS, whether or not it uses AI. The paper therefore argues that engineering rigor, understood as diligent problem-space analysis through operational concepts, ODD, behavior specification, and behavioral competencies, is a necessary condition for building safe AI-based systems, and that this behavior-based foundation provides the missing link that ISO/PAS 8800 calls for: traceable decomposition of system-level safety metrics into AI component performance metrics. The illustrative occlusion-pedestrian case study shows how a SOTIF functional insufficiency (inability to predict occluded pedestrians) translates first into a behavior-level competency and then into a concrete dataset requirement (labeled occluded areas).
Load-bearing premise
The load-bearing assumption is that a behavior specification expressed through abstract scenarios and behavioral competencies can be decomposed without loss into AI-component-level requirements and performance metrics, so that system-level safety remains traceable down to AI-specific indicators; the paper itself acknowledges that the comprehensive metamodel for this is still to be established.
Editorial extensions
If this is right
- Safety analyses can be performed once at the behavior level and reused for functional safety, SOTIF, and AI safety analyses, reducing duplicated effort across standards.
- Standards with broad AI definitions can be scoped out of early development stages, because problem-space analyses are technology-neutral.
- AI component requirements, such as dataset labels and coverage criteria, become derived artifacts from behavior-level safety goals rather than ad hoc additions.
- Behavioral competencies provide a natural basis for pass-fail criteria in verification and validation, linking system-level safety indicators to test outcomes.
- The identified gap in traceability between specification, test results, and field monitoring motivates further work on ontologies and metamodels for behavior specification.
Reading between the lines
- If the behavior-based bridge works, regulators could focus on ODD, behavior, and competencies as the assurance backbone, treating AI-specific metrics as implementation details to be checked against that backbone.
- The same decomposition logic could generalize beyond driving to other open-context AI systems, such as robots or drones, where an abstract behavior specification can anchor downstream AI assurance.
- A testable extension would be to build the missing metamodel and run it on a set of known incidents to see whether every AI-contributed factor maps back to a behavior-level competency gap; any that do not would signal the decomposition needs revision.
- The position paper implies that investment in formalizing ODD and behavior reasoning may yield higher safety leverage than investment in AI-specific verification tools alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that safety assurance for AI-based automated driving systems is hampered by imprecise AI definitions and by an overemphasis on AI-specific challenges. The authors propose a behavior-based perspective in which system-level safety analyses—using ODD, behavior specifications, and behavioral competencies—can provide grounds for connecting system-level requirements to AI-related performance metrics. The paper surveys definitions in the EU AI Act and ISO standards, discusses sources of uncertainty following Burton and Herd, and presents a short illustrative case study on occluded pedestrians. It concludes with recommendations for standards navigation and future research.
Significance. If substantiated, the proposed behavior-based framework would be a valuable integration of SOTIF, functional safety, and AI safety analyses, and would offer practical guidance for the application of ISO/PAS 8800. The paper is a useful and mostly well-grounded position statement: it correctly observes that many assurance challenges attributed to AI actually stem from open-context uncertainty, and it leverages established systems engineering concepts. The authors are transparent about the need for future work (Section V-D). Its main weakness is that the central decomposition claim is asserted rather than demonstrated; the illustrative case study stops at a dataset requirement and does not show a traceable performance metric.
major comments (2)
- [§IV-B4 and §V-C1] The claim that behavioral competencies "can provide guidance for a traceable definition of pass-fail criteria from system-level safety indicators, as well as the definition of meaningful performance criteria for AI components" is not demonstrated. The case study ends with a dataset requirement (labeled occluded areas) rather than an AI performance metric, and no argument is provided that satisfying such a requirement guarantees, or even measurably advances, the system-level safety goal of preventing collisions with pedestrians. This is load-bearing because the paper's answer to "What's really different with AI?" rests on this decomposition. Please either extend the example to a concrete performance metric with a traceability argument, or reframe the claim as a research hypothesis that requires the metamodel identified in Section V-D.
- [§V-D and §VI] The admission that "a comprehensive approach that has been fully connected to AI-specific needs is still yet to be established" is in tension with the strength of the conclusion that behavior-based analyses "can provide solid grounds" for the decomposition. Given the conceded gap, the conclusion should be conditioned on the future development of the metamodel, or the authors should specify conditions under which the decomposition is expected to be lossless—particularly with respect to unmodeled interactions among perception, prediction, planning, and control components. Without such qualification, the central contribution is an interesting assertion rather than a supported result.
minor comments (6)
- [Abstract and §III-A] There is a typo in the quoted definition: "abscence" should be "absence".
- [§II] The discussion of differing AI definitions would benefit from a table comparing the EU AI Act, ISO/IEC 22989, ISO/IEC TR 24028, and ISO/PAS 8800 to improve readability and scannability.
- [§III-A] The relationship between "risk" and "assurance uncertainty" in Fig. 1a is stated but not formally defined; a brief mapping would help readers follow the rephrasing.
- [§III-C] The reference to Koopman's "Machine Learning breaks the Vee" and the Bayesian analogy is underdeveloped; a sentence on how the prior is updated by evidence in the assurance setting would clarify the point.
- [§IV-B2] The sentence "The elements of the system context are, e.g., sources of requirements for must-have class labels" is terse; specify how ODD elements map to input-space requirements for ML models.
- [References] References [27] and [42] contain repeated author/consortium names; please clean them up.
Circularity Check
No significant circularity: the paper's behavior-based assurance argument is asserted with an acknowledged gap, not reduced to its own inputs.
full rationale
This is a position paper whose central claims are recommendations and conceptual arguments, not quantitative predictions derived from fitted parameters. The key conceptual separation of risk sources is grounded in external work by Burton and Herd [2] and in standards such as ISO/PAS 8800 and ISO 21448. The proposed behavior-based decomposition in Sections IV-B3, IV-B4 and VI is explicitly offered as a 'starting point' rather than a derived result, and Section V-D concedes that 'a comprehensive approach that has been fully connected to AI-specific needs is still yet to be established.' That is an acknowledged support gap, not circularity: the paper does not fit a parameter and rename it a prediction, nor does it define the conclusion in terms of the premise. The self-citations ([3], [32], [34], [42]) are used as pointers to prior systems-engineering artifacts and illustrative case-study work; they are not invoked as an authority that makes the conclusion true by construction, and there is no exhibited equation or derivation where an output reduces to an input. Accordingly, the derivation chain is not circular, just—by the authors' own admission—incomplete.
Assumptions & free parameters
assumptions (5)
- domain assumption Safety is the absence of unreasonable risk, and risk is the effect of uncertainty on objectives.
- domain assumption Open-world uncertainty is irreducible and can be mitigated but not fully eliminated.
- domain assumption Adopting ISO/PAS 8800's definition of an AI system (uses an AI model not completely defined by human knowledge) is appropriate for scoping the discussion.
- domain assumption Behavior-based artifacts (ODD, behavior specification, behavioral competencies) can support traceability from system-level safety indicators to AI-specific metrics.
- domain assumption Stakeholder needs can be systematically captured from normative sources and represented in traceable form.
Cite this review
Pith. "Pith review of What's Really Different with AI? -- A Behavior-based Perspective on System Safety for Automated Driving Systems." pith.science (2026). https://pith.science/paper/6HJUISWV
@misc{pith2026250720685,
author = {Pith},
title = {Pith review of: What's Really Different with AI? -- A Behavior-based Perspective on System Safety for Automated Driving Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/6HJUISWV}},
note = {Machine review of arXiv:2507.20685}
}
read the original abstract
Assuring safety for ``AI-based'' systems is one of the current challenges in safety engineering. For automated driving systems, in particular, further assurance challenges result from the open context that the systems need to operate in after deployment. The current standardization and regulation landscape for ``AI-based'' systems is becoming ever more complex, as standards and regulations are being released at high frequencies. This position paper seeks to provide guidance for making qualified arguments which standards should meaningfully be applied to (``AI-based'') automated driving systems. Furthermore, we argue for clearly differentiating sources of risk between AI-specific and general uncertainties related to the open context. In our view, a clear conceptual separation can help to exploit commonalities that can close the gap between system-level and AI-specific safety analyses, while ensuring the required rigor for engineering safe ``AI-based'' systems.
Figures
Reference graph
Works this paper leans on
-
[1]
S. Burton and J. A. McDermid, “Closing the Gaps: Complexity and Uncertainty in the Safety Assurance and Regulation of Automated Driving,” Technical Rep. 2023
work page 2023
-
[2]
Addressing Uncertainty in the Safety Assurance of Machine-Learning,
S. Burton and B. Herd, “Addressing Uncertainty in the Safety Assurance of Machine-Learning,” Frontiers Comput. Sci., vol. 5, Apr. 6,
-
[3]
M. Nolte, “Werte- und fähigkeitsbasierte Bewegungsplanung für autonome Straßenfahrzeuge – Ein systemischer Ansatz,” (in German), Ph.D. dissertation, TU Braunschweig, 2025
work page 2025
-
[4]
Road Vehicles — Safety and Artificial Intelligence , ISO Publ. Avail. Spec. 8800:2024
work page 2024
-
[5]
Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges,
R. Ashmore, R. Calinescu, and C. Paterson, “Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges,” ACM Comput. Surv., vol. 54, no. 5, 111:1–111:39, May 25, 2021. DOI: 10.1145/3453444
doi:10.1145/3453444 2021
-
[6]
R. Schnitzer, L. Kilian, S. Roessner, K. Theodorou, and S. Zillner, Landscape of AI Safety Concerns – A Methodology to Support Safety Assurance for AI-Based Autonomous Systems , Dec. 18, 2024. DOI: 10.48550/arXiv.2412.14020. arXiv: 2412.14020[cs]
work page Pith review arXiv doi:10.48550/arxiv.2412.14020 2024
- [7]
-
[8]
Information Technology — Artificial Intelligence — Artificial Intelli- gence Concepts and Terminology , ISO/IEC Standard 22989:2020
work page 2020
Show all 44 references
-
[9]
Information Technology — Artificial Intelligence — Al System Life Cycle Processes, ISO/IEC Standard 5338:2023
2023
-
[10]
Information Technology — Artificial Intelligence — Guidance on Risk Management, ISO/IEC Standard 23894:2023
2023
-
[11]
Information Technology — Artificial Intelligence — Overview of Trust- worthiness in Artificial Intelligence , ISO/IEC Tech. Rep. 24028:2020
2020
-
[12]
Software and Systems Engineering — Software Testing , ISO/IEC Tech. Rep. 29119:2020
2020
-
[13]
Systems and Software Engineering — Systems and Software Assurance — Part 1: Concepts and Vocabulary , ISO/IEC/IEEE Standard 15026- 1:2023
2023
-
[14]
Road Vehicles — Functional Safety , ISO Standard 26262:2018
2018
-
[15]
Risk Management — Guidelines , ISO Standard 31000:2018
2018
-
[16]
Safety and Risk – Why Their Definitions Matter,
N. F. Salem, S. Le Page, J. Millar, P. Junietz, M. Nolte, R. Graubohm, and M. Maurer, “Safety and Risk – Why Their Definitions Matter,” in Handbook Assisted Automated Driving , H. Winner, K. Dietmayer, L. Eckstein, M. Jipp, M. Maurer, and C. Stiller, Eds., 4th ed., Heidelberg:...
2025
-
[17]
Redefining Safety for Autonomous Vehicles,
P. Koopman and W. Widen, “Redefining Safety for Autonomous Vehicles,” in 2024 Int. Conf. Comput. Saf., Rel., Secur. (SAFECOMP) , A. Ceccarelli, M. Trapp, A. Bondavalli, and F. Bitsch, Eds., ser. Lecture Notes Comput. Sci. V ol. 14988, Florence, Italy: Springer, Cham, pp. 300–3...
2024 doi
-
[18]
Koopman, How Safe Is Safe Enough? Measuring and Predicting Autonomous Vehicle Safety, 1st ed
P. Koopman, How Safe Is Safe Enough? Measuring and Predicting Autonomous Vehicle Safety, 1st ed. Pittsburgh, PA: Carnegie Mellon Univ., 2022, 352 pp
2022
-
[19]
Towards a Framework to Manage Perceptual Uncertainty for Safe Automated Driving,
K. Czarnecki and R. Salay, “Towards a Framework to Manage Perceptual Uncertainty for Safe Automated Driving,” in 2018 Int. Conf. Comput. Saf., Rel., Secur. (SAFECOMP) , B. Gallina, A. Skavhaug, E. Schoitsch, and F. Bitsch, Eds., ser. Lecture Notes Comput. Sci. Västerås, Sweden...
2018 doi
-
[20]
Methods and Tools for the Engineering and Assurance of Safe Autonomous Systems,
E. Troubitsyna, I. J. Alvarez, P. Koopman, and M. Trapp, “Methods and Tools for the Engineering and Assurance of Safe Autonomous Systems,” Schloss Dagstuhl – Leibniz-Zentrum für Informatik, 2024, ISSN: 2192-5283 Issue: 4 V olume: 14, pp. 23–41. DOI: 10.4230/ DAGREP.14.4.23
2024
-
[21]
Trustworthiness assurance assessment for high-risk AI-based systems,
G. Stettinger, P. Weissensteiner, and S. Khastgir, “Trustworthiness assurance assessment for high-risk AI-based systems,” IEEE Access, vol. 12, pp. 22 718–22 745, 2024. DOI: 10 . 1109 / ACCESS . 2024 . 3364387
2024
-
[22]
Multilayer Feedforward Networks Are Universal Approximators,
K. Hornik, M. Stinchcombe, and H. White, “Multilayer Feedforward Networks Are Universal Approximators,” Neural Netw.s, vol. 2, no. 5, pp. 359–366, Jan. 1989. DOI: 10.1016/0893-6080(89)90020-8
1989 doi
- [23]
-
[24]
L145 Challenges in Autonomous Vehicle Safety Assessment,
P. Koopman, “L145 Challenges in Autonomous Vehicle Safety Assessment,” US DOT Workshop, US DOT Workshop, online, Mar. 2024
2024
-
[25]
Criticality Analysis for the Verification and Validation of Automated Vehicles,
C. Neurohr, L. Westhofen, M. Butz, M. H. Bollmann, U. Eberle, and R. Galbas, “Criticality Analysis for the Verification and Validation of Automated Vehicles,” IEEE Access, vol. 9, pp. 18 016–18 041, 2021. DOI: 10.1109/ACCESS.2021.3053159
2021
- [26]
-
[27]
VV Methods Safety Assurance Position Paper,
R. Galbas, M. Nolte, U. Eberle, H. Hungar, H. Mosebach, N. F. Salem, H. Schittenhelm, J. Reich, T. Kirschbaum, L. Westhofen, PEGASUS VVM Consortium, R. Galbas, M. Nolte, U. Eberle, H. Hungar, H. Mosebach, N. F. Salem, H. Schittenhelm, J. Reich, T. Kirschbaum, and L. Westhofen,...
2024 doi
- [28]
-
[29]
A VSC Best Practice for Evaluation of Behavioral Competencies for Automated Driving System Dedicated Vehicles (ADS-DVs),
Automated Vehicle Safety Consortium (A VSC), “A VSC Best Practice for Evaluation of Behavioral Competencies for Automated Driving System Dedicated Vehicles (ADS-DVs),” A VSC00008202111, Nov. 2021
2021
-
[30]
Behaviour Taxonomy for Automated Driving System (ads) Applications – Specification, BSI Standard BSI Flex 1891 v1.0:Jan. 2025
2025
-
[31]
A Behavioural Safety Centric Approach for E2E ADS,
G. Price, “A Behavioural Safety Centric Approach for E2E ADS,” Presentation, The26262Club, online, Mar. 2025
2025
-
[32]
An Ontology-Based Approach Toward Traceable Behavior Specifications in Automated Driving,
N. F. Salem, M. Nolte, V . Haber, T. Menzel, H. Steege, R. Graubohm, and M. Maurer, “An Ontology-Based Approach Toward Traceable Behavior Specifications in Automated Driving,” IEEE Access, vol. 12, pp. 165 203–165 226, 2024. DOI: 10.1109/ACCESS.2024.3494036
2024
-
[33]
Opera- tional Design Domain-Driven Coverage for the Safety Argumentation of Automated Vehicles,
P. Weissensteiner, G. Stettinger, S. Khastgir, and D. Watzenig, “Opera- tional Design Domain-Driven Coverage for the Safety Argumentation of Automated Vehicles,” IEEE Access , vol. 11, pp. 12 263–12 284,
-
[34]
Towards a Skill- and Ability-Based Development Process for Self-Aware Automated Road Vehicles,
M. Nolte, G. Bagschik, I. Jatzkowski, T. Stolte, A. Reschka, and M. Maurer, “Towards a Skill- and Ability-Based Development Process for Self-Aware Automated Road Vehicles,” in 2017 IEEE Int. Conf. Intell. Transp. Syst. (ITSC) , Yokohama, Japan: IEEE, pp. 739–744. DOI: 10.1109/...
2017
-
[35]
DOI: 10.1109/ACCESS.2023.3242127
2023
-
[36]
Defining and Substantiating the Terms Scene, Situation and Scenario for Automated Driving,
S. Ulbrich, A. Reschka, T. Menzel, F. Schuldt, and M. Maurer, “Defining and Substantiating the Terms Scene, Situation and Scenario for Automated Driving,” in 2015 18th IEEE Int. Annu. Conf. Intell. Transp. Syst. (ITSC) , Las Palmas, Spain: IEEE, pp. 982–988
2015
-
[37]
D. D. Walden and International Council on Systems Engineering, Eds., INCOSE Systems Engineering Handbook , 5th ed., Hoboken, NJ: John Wiley Sons Ltd, 2023
2023
-
[38]
Risk management core – towards an explicit representation of risk in automated driving,
N. F. Salem, T. Kirschbaum, M. Nolte, C. Lalitsch-Schneider, R. Graubohm, J. Reich, and M. Maurer, “Risk management core – towards an explicit representation of risk in automated driving,” IEEE Access, vol. 12, pp. 33 200–33 217, 2024, tex.publisher: IEEE. DOI: 10.1109/ ACCESS...
2024
-
[39]
Road Vehicles — Safety of the Intended Functionality , ISO Standard 21448:2022
2022
-
[40]
On Assumptions with Respect to Occlusions in Urban Environments for Automated Vehicle Speed Decisions,
R. Graubohm, N. F. Salem, M. Nolte, and M. Maurer, “On Assumptions with Respect to Occlusions in Urban Environments for Automated Vehicle Speed Decisions,” in 2023 IEEE 26th Int. Conf. Intell. Transp. Syst. (ITSC), citation-key: graubohm2023, Bilbao, Spain: IEEE, pp. 738–745. ...
2023
-
[41]
C. S. Wasson, System Engineering Analysis, Design, and Development: Concepts, Principles, and Practices . Hoboken, NJ: John Wiley Sons Inc, 2005
2005
-
[42]
Towards Closing the Gap between Model-Based Systems Engineering and Automated Vehicle Assurance: Tailoring Generic Methods by Integrating Domain Knowledge,
M. Nolte and M. Maurer, “Towards Closing the Gap between Model-Based Systems Engineering and Automated Vehicle Assurance: Tailoring Generic Methods by Integrating Domain Knowledge,” presented at the 16. Uni-DAS e.V . Workshop Fahrerassistenz und automatisiertes Fahren, Irsee: ...
2025
-
[43]
Using Ontologies for the Formalization and Recognition of Criticality for Automated Driving,
L. Westhofen, C. Neurohr, M. Butz, M. Scholtes, and M. Schuldes, “Using Ontologies for the Formalization and Recognition of Criticality for Automated Driving,” IEEE Open J. Intell. Transp. Syst. , vol. 3, pp. 519–538, 2022. DOI: 10.1109/OJITS.2022.3187247
2022
-
[2023]
DOI: 10.3389/fcomp.2023.1132580
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.