REVIEW 1 major objections 1 minor 22 references
Operationalization of Scenario-Based Safety Assessment of Automated Driving Systems
T0 review · 1 major / 1 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proposes the Safety Assessment Framework (SAF), a concrete process that operationalizes the UNECE NATM multi-pillar approach by federating heterogeneous scenario databases into a single queryable dataspace for generating…
desk verdict A coherent, honest synthesis of NATM scenario-based safety assessment, but the central coverage-metric claim is asserted, not shown — valuable as a roadmap, not yet an operationalization. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Safety Assessment Framework (SAF), a process whose load-bearing component is the federation layer: a common interface through which a single query retrieves scenarios from multiple heterogeneous databases, including knowledge-based scenarios, accident databases, and data-driven real-world scenarios. The federation layer, developed in the SUNRISE project and extended in SYNERGIES with traceability and quality metrics, is what makes the scenario set for a given ODD sufficiently broad and defensible. The framework then converts those scenarios into test descriptions, allocates them to virtual, proving-ground, or real-world execution, and feeds results into a safety case organized as test matrices per requirement, argumentation for general requirements, and quantified risk estimates. The metric of residual safety risk, computed from scenario exposure and crash probability, is the quantitative anchor for the safety case.
What would settle it
Take a defined set of operating conditions (ODD), build the federated scenario dataspace as described, and compare the set of test scenarios selected through the federation layer against a ground-truth set of scenarios extracted from real-world field data and accident reports for that ODD; if the selected set misses a scenario that produced a known collision or a known unsafe interaction, the coverage claim fails.
Extended reading notes
Core claim
Following the NATM multi-pillar approach, the paper argues that scenario-based safety assessment of an automated driving system is feasible in practice if the scenario databases that hold knowledge-based, accident, and data-driven real-world scenarios are connected through a federation layer. The Safety Assessment Framework (SAF) operationalizes the NATM by defining processes to generate relevant test scenarios from the ODD and system requirements, query the federated scenario databases, allocate test scenarios to virtual, proving-ground, and real-world testing, analyse test results for coverage and system safety, and assemble the evidence into a safety case. The central insight is that the federation layer plus quality metrics for scenario representativeness and coverage turns the collection of heterogeneous databases into one EU-wide scenario dataspace, making the scenario selection process traceable and repeatable for manufacturers and authorities.
Load-bearing premise
The whole safety case rests on the assumption that a federated set of scenario databases, measured for coverage and representativeness, can actually supply enough scenarios to cover the full set of conditions the automated driving system is designed to handle, including rare edge cases; the paper describes this challenge but does not demonstrate that the metrics deliver such coverage.
Editorial extensions
If this is right
- Manufacturers can build a safety case for type approval by following the SAF process instead of inventing an ad hoc assessment procedure.
- Authorities can audit the safety case against the documented test matrix, scenario coverage metrics, and validation reports rather than relying on unspecified internal methods.
- Scenario database owners can contribute to a shared EU-wide dataspace without ceding control of their data, since the federation layer provides a single access point and search documentation.
- The same scenario dataspace can support risk quantification: exposure values from continuous vehicle data streams, combined with simulated crash probabilities, yield a residual safety risk estimate for positive risk balance or similar acceptance decisions.
- Traceability of scenario selection improves with metadata that allows any search through the dataspace to be reconstructed for later reference or spot checks.
Reading between the lines
- If the coverage metrics do not actually guarantee that rare but critical ODD scenarios are present, the whole SAF safety case inherits that gap; a testable extension is to benchmark the metrics against known accident databases.
- The federation-layer architecture could be extended from pre-deployment testing to continuous in-service monitoring, using the same dataspace to detect unknown scenarios after deployment; the paper treats in-service monitoring as a separate pillar.
- Converting the regulatory phrase 'competent and carefully driven manual vehicle' into quantitative acceptance criteria will require the assertion-based scaling-up that the paper cites; the SAF itself does not supply those criteria.
- The governance of the federated dataspace after the SYNERGIES project ends is a non-technical precondition for the framework to work at scale; the paper flags this as a recommendation but does not solve it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes a Safety Assessment Framework (SAF, Fig. 2) intended to operationalize the UNECE WP.29 New Assessment/Test Method (NATM) multi-pillar approach for safety assessment of Automated Driving Systems (ADS). The framework comprises scenario database searching via a federation layer, test scenario generation, allocation of tests across virtual, physical, and real-world pillars, and analysis of test results to support a safety case. The authors draw on concepts from the Horizon Europe projects SUNRISE and SYNERGIES, particularly a federated scenario dataspace and quality metrics for coverage and representativeness, and they discuss related topics such as Acceptable Means of Compliance, regulatory requirements, risk quantification, and validation of test environments. The paper explicitly acknowledges that several important aspects (e.g., acceptance criteria, in-service monitoring) are not fully addressed.
Significance. If the framework were fully specified and validated, it could provide a valuable common structure for manufacturers and authorities to plan and evaluate the safety evidence of an ADS, and the federated scenario dataspace could be a practical contribution to collaborative safety assessment. The paper is useful as a synthesis of ongoing EU project directions and as a high-level architectural proposal. However, its significance as a scientific contribution is limited because the central operationalization claim rests on components that are not defined or validated in the manuscript: the federation layer interface, the coverage and representativeness metrics, and the inference from scenario coverage to a safety conclusion. The paper is therefore best read as a position or roadmap paper rather than as a completed operationalization.
major comments (1)
- [Section V and Fig. 2] The SAF diagram includes blocks for 'coverage analysis, system analysis' and 'In-service monitoring and reporting', but the paper explicitly states that acceptance criteria are not discussed and that ISMR is left to future work in the CERTAIN project. This creates a discrepancy between the claim of an 'operationalized' framework and the actual completeness of the presented components. Please either provide initial specifications for these blocks or narrow the paper's claim to the pre-deployment testing pillars and clearly mark the other blocks as future extensions.
minor comments (1)
- [Section II.A] The five NATM pillars are listed, but the relationship between the 'audit' pillar and the 'safety management system' is not explained; consider clarifying how the audit process relates to the four test-related components in Fig. 1.
Circularity Check
No circularity found: the SAF is a proposed process architecture, not a derived result, and the self-citations are supporting references rather than inputs that reappear as outputs.
full rationale
The paper contains no fitted parameters, no equations, and no prediction whose value is forced by construction. The central claim is that the Safety Assessment Framework (SAF, Fig. 2) operationalizes the UNECE NATM multi-pillar approach; this is a process proposal, and its mapping to the NATM pillars is anchored to the external NATM master document [1] and Fig. 1 adapted from [2]. The authors' self-citations ([7], [13], [14], [17]) and references to their own EU projects (SUNRISE, SYNERGIES, CERTAIN) point to published methods or project plans; for example, '[13]' is cited for a published statistical method and '[17]' for risk-estimation procedures. These are deferrals to external works, not definitions of the SAF's conclusion. One citation is anomalous: 'as proposed by UNECE [7]' cites the authors' own TNO StreetWise report where an external UNECE source would be expected; however, the UNECE document is also cited as [1], and the SAF is not derived from [7] by construction. The paper's own limitations—validation of methods/tools is outside scope (Section IV-D) and the considerations are 'not complete' (Section V)—are completeness gaps, not circularity: the paper does not claim to derive coverage from the SYNERGIES metrics or to validate ODD coverage. There is therefore no circular step to report.
Assumptions & free parameters
assumptions (4)
- domain assumption Scenario-based testing with sufficient ODD coverage is a valid basis for ADS safety assurance.
- domain assumption The federated approach accessing heterogeneous scenario databases can provide coverage and representativeness metrics that are meaningful for ODD coverage.
- domain assumption Risk estimates based on scenario statistics and simulation can be used as acceptance evidence in a safety case.
- domain assumption Quantitative acceptance criteria can be derived from competent human driving references.
Cite this review
Pith. "Pith review of Operationalization of Scenario-Based Safety Assessment of Automated Driving Systems." pith.science (2026). https://pith.science/paper/OTP6U73E
@misc{pith2026250722433,
author = {Pith},
title = {Pith review of: Operationalization of Scenario-Based Safety Assessment of Automated Driving Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/OTP6U73E}},
note = {Machine review of arXiv:2507.22433}
}
read the original abstract
Before introducing an Automated Driving System (ADS) on the road at scale, the manufacturer must conduct some sort of safety assurance. To structure and harmonize the safety assurance process, the UNECE WP.29 Working Party on Automated/Autonomous and Connected Vehicles (GRVA) is developing the New Assessment/Test Method (NATM) that indicates what steps need to be taken for safety assessment of an ADS. In this paper, we will show how to practically conduct safety assessment making use of a scenario database, and what additional steps must be taken to fully operationalize the NATM. In addition, we will elaborate on how the use of scenario databases fits with methods developed in the Horizon Europe projects that focus on safety assessment following the NATM approach.
Reference graph
Works this paper leans on
-
[1]
Ne Assessment/ est Method for Automated Driving (NATM) - Master Document,
UNECE WP29 G A, “Ne Assessment/ est Method for Automated Driving (NATM) - Master Document,” UNECE, Geneva, Switzerland, 2021
work page 2021
-
[2]
R. Donà, B. Ciuffo, A. Tsakalidis, L. D. Cesare, C. Sollima, M. Sangiogi and M. C. Galassi, “ ecent Advancements in Automated Vehicle Certification: How the Experience from the Nuclear Sector Contributed to Ma ing hem a eality,” Energies, vol. 15, 2022
work page 2022
-
[3]
Harmonised European Solutions for esting Automated road ransport,
HEADS A , “Harmonised European Solutions for esting Automated road ransport,” EU Horizon 2020 project, 2021. [Online]. Available: https://www.headstart-project.eu/
work page 2020
-
[4]
Safety Assurance Frame or for Connected am Automated Mobility Systems,
SUN SE, “Safety Assurance Frame or for Connected am Automated Mobility Systems,” EU Horizon Europe project, 2023. [Online]. Available: https://ccam-sunrise-project.eu/
work page 2023
-
[5]
SYNE G ES, “ eal and synthetic scenarios generated for the development, training, virtual testing and validation of CCAM systems,” EU Horizon Europe project, 2024. [Online]. Available: https://synergies-ccam.eu/
work page 2024
-
[6]
“SafetyPool,” deepen & WMG, 2025. [Online]. Available: https://www.safetypool.ai/database
work page 2025
-
[7]
NO StreetWise, Scenario- Based Safety Assessment of Automated Driving Systems, TNO 2024 10983,
E. de Gelder, O. Op den Camp, J. Broos, J. -P. Paardekooper, S. van Montfort, S. Kalisvaart and H. Goossens, “ NO StreetWise, Scenario- Based Safety Assessment of Automated Driving Systems, TNO 2024 10983,” 28 May 2024. [Online]. Available: https:// .tno. nl/ en/newsroom/papers/scenario-based-safety-assessment/
work page 2024
-
[8]
L. Guyonvarcha, T. Hermitte, R. Kroger, C. Chauvel, E. Arnoux and S. Geronimi, “ADSCENE scenarios data base: Focus on accident data support for validation of Automated Driving Functions,” ScienceDirect, Transportation Research Procedia, vol. 72, pp. 9 -16, 2023
work page 2023
Show all 22 references
-
[9]
scenario.center: Methods from Real- orld Data to a Scenario Database,
M. Schuldes, C. Glasmacher and L. Ec stein, “scenario.center: Methods from Real- orld Data to a Scenario Database,” in 35th IEEE Intelligent Vehicles Symposium, Korea, 2024
2024
-
[10]
[Online]
European Union Aviation Safety Agency (EASA), „Acceptable Means of Compliance (AMC) and Alternative Means of Compliance (AltMoC),” 2025. [Online]. Available: https:// .easa.europa.eu/ en/document-library/
2025
-
[11]
UNECE, „UN egulation No. 157 on uniform provisions concerning the approval of vehicles with regards to Automated Lane Keeping System,” UN Economic Commission for Europe, nland ransport Committee, World Forum for Harmonization of Vehicle Regulations, Geneva, 2021
2021
-
[12]
157 (ALKS),” UN Economic Commission for Europe, nland Transport Committee, World Forum for Harmonization of Vehicle Regulations, Geneva, 2022
UNECE, „Proposal for the 01 series of amendments to UN egulation No. 157 (ALKS),” UN Economic Commission for Europe, nland Transport Committee, World Forum for Harmonization of Vehicle Regulations, Geneva, 2022
2022
-
[13]
A uantitative method to determine what collisions are reasonably foreseeable and preventable,
E. de Gelder and O. Op den Camp, “A uantitative method to determine what collisions are reasonably foreseeable and preventable,” Safety Science, vol. 167, p. 106233, 2023
2023
-
[14]
Procedure for the Safety Assessment of an Autonomous Vehicle using Real-World Scenarios,
E. de Gelder and O. Op den Camp, “Procedure for the Safety Assessment of an Autonomous Vehicle using Real-World Scenarios,” in FISITA 2020 World Congress, Prague, 2020
2020
-
[15]
Positive ris balance: a comprehensive frame or to ensure vehicle safety,
F. Kauffmann, F. Fahren rog, L. Drees and F. aisch, “Positive ris balance: a comprehensive frame or to ensure vehicle safety,” Ethics and Information Technology, vol. 24, no. 1, 2022
2022
-
[16]
Koopman, How Safe is Safe Enough? Measuring and Predicting Autonomous Vehicle Safety, Pittsburgh PA: Carnegie Mellon University, 2022
P. Koopman, How Safe is Safe Enough? Measuring and Predicting Autonomous Vehicle Safety, Pittsburgh PA: Carnegie Mellon University, 2022
2022
-
[17]
Ho certain are e that our automated driving system is safe?,
E. de Gelder and O. Op den Camp, “Ho certain are e that our automated driving system is safe?,” Traffic Injury Prevention, Special issue for the 27th International Technical Conference andf the Enhanced Safety of Vehicles (ESV), vol. 24, no. sup 1, pp. S131-S140, 2023
2023
-
[18]
Data Sources for Baseline Generation - Overview, Grading, and ecommendations,
U. Sander, P. Ek, D. Sander, S. Breunig, J. Bärgman, T. Menzel, C. Glasmacher, M. Urban, H. Chajmowicz, R. Davidse, G. Schermers, J. Lorente Mallada, O. Op den Camp, E. Charoniti, M. Meocci, F. La orre and J. Hay, “Data Sources for Baseline Generation - Overview, Grading, and ...
2024
-
[19]
Fahrenkrog, F., Das, A., Sander, D., Bärgman, J., Urban, M., Pohl, M., Glasmacher, C., et al., „Prospective Safety Assessment Frame or - Instruction, Deliverable D2.1 of the Horizon Europe project 4SAFE Y,” 2024
2024
-
[20]
Being good (at driving): Characterizing behavioral expectations on automated and human driven vehicles.,
L. Fraade-Blanar, F. Favarò, J. Engstrom, M. Cefkin, R. Best, J. Lee and . ictor, “Being good (at driving): Characterizing behavioral expectations on automated and human driven vehicles.,” 2025. [Online]. Available: https://doi.org/10.48550/arXiv.2502.08121
-
[21]
European Commission, „Commission mplementing egulation (EU) 2022/1426 for the Application of Regulation (EU) 2019/2144 as Regards Uniform Procedures and Technical Specifications for the Type-Approval of the Automated Driving System (ADS) of Fully Automated ehicles,” EC - Europ...
2022
-
[22]
o ards a Practical Methodology for Defining Competent Driving,
A. Tejada, M. Legius, A. Kalose, P. Oliveira, E. van Dam and J. Hogema, “ o ards a Practical Methodology for Defining Competent Driving,” in 2023 IEEE 26th International Conference on Intelligent Transportation Systems (ITSC), Bilbao, Bizkaia, Spain, 2023
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.