REVIEW 2 major objections 3 minor 36 references
Using Formal Models, Safety Shields and Certified Control to Validate AI-Based Train Systems
T0 review · 2 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Coupled runs of a formal B-model, a real YOLO perception system, and a runtime certificate checker inside ProB/SimB can expose safety-critical interactions—in particular, false rejections that defeat the safety shield—and yield…
desk verdict A modest, honest workshop paper that demonstrates a useful integration of a real YOLO detector and certificate checker into a formal B model; the surprising false-rejection result is the takeaway. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the linked demonstrator: a formal B model of the shunting yard (environment, steering system, and perception events), executed by the ProB animator and SimB simulator, with the real YOLOv8 detector and a deterministic certificate checker connected through SimB's external-simulation interface. The formal model doubles as a safety shield that disables train movement when a signal is expected but not detected, while the certificate checker rejects false positive detections by inspecting cropped bounding-box image features with computer vision. The key mechanism is that both monitors operate on the same formal state: the shield checks the AI's detections against known signal positions, and the checker validates the AI's image-level output. The interaction between these two monitors is what the simulation measures, and it is exactly where the paper finds the safety-critical failure.
What would settle it
Run the identical 500-run Monte Carlo protocol with images generated from a state-responsive simulator or real onboard footage synchronized to the formal state instead of sampled video frames. If the combined shield-plus-checker configuration becomes as safe as or safer than the shield alone, for example if certified control no longer lowers correct detections from 9155 to 2094, then the reported safety degradation is an artifact of the sampling method rather than an inherent property of the two-layer monitor.
Extended reading notes
Core claim
The central claim is that replacing hand-coded detection probabilities with the real AI and the real certificate checker inside a formal B-model simulation makes hidden interactions visible. In the combined configuration with both the safety shield and certified control, the train travelled only 63.0% of the safe distance and reached a safe outcome in 63.0% of runs, which is worse than the shield alone (82.4% distance, 82.8% safe). The cause identified with ProB is that the certificate checker falsely rejects correct permission-signal detections, so the shield does not know a signal is present when that signal later changes to stop. The paper's point is not that this particular system is safe; it is that the setup surfaces such vulnerabilities systematically and gives statistical evidence through the formal properties.
Load-bearing premise
The results hinge on the assumption that randomly sampling a video frame by train position and signal state gives the perception system inputs that faithfully correspond to the formal model's current state, so that measured detection errors reflect what would happen in real operation.
Editorial extensions
If this is right
- When the real AI drives the formal simulation, the measured safety properties can be reported as statistics rather than assumed from fixed probabilities; for example, 82.8% of runs are safe with the safety shield alone versus 63.0% when the certificate checker is added.
- The certificate checker removes all false-positive detections in the tested simple environment, but its false rejections are numerous enough to reduce correct signal detections from 9155 to 2094, and a falsely rejected permission signal can become an undetected stop signal.
- The methodology identifies concrete weaknesses: the fine-tuned YOLO model produces many false positive stop detections that block progress without monitors, and particular stop-sign images are hard for the checker to certify.
- Linking the formal model, the real AI, and the checker inside ProB/SimB gives a repeatable validation loop that can be rerun after improving the AI or the certificate checker.
- The Monte Carlo results provide a way to compare configurations, such as with and without the shield or with and without certified control, using the same formal safety properties.
Reading between the lines
- Beyond the paper, the shield and the checker are logically independent subsystems, and the case study shows they need a shared protocol: a detection that is 'rejected' by the checker should still inform the formal model, otherwise the shield treats an unseen permission signal as absent. This suggests a combined monitor with three-valued output (detected, rejected, absent) as a natural extension.
- The quantitative results should be read as a demonstration of the method rather than as deployment estimates, because the paper itself flags that images are sampled from static video collections instead of a state-responsive simulation; a configurable railway co-simulation would be needed before the percentages transfer to real operation.
- The same harness could be used to tune the checker's acceptance threshold against the shield's tolerance, for example by varying the probability that a permission aspect falls back to stop and measuring how false rejections affect the safety statistics.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper describes a demonstrator that couples a formal B model of an AI-controlled train in a shunting yard with a real YOLO-based perception system and a runtime certificate checker, using ProB and SimB for closed-loop simulation. The B model acts as a safety shield, and the certificate checker is intended to filter false positive detections. The authors run 500 Monte Carlo simulations with and without each component, report distance travelled, percentage of safe runs, and false/correct detection counts, and identify an interaction in which false rejections by the certificate checker can defeat the safety shield when a permission signal falls back to stop. They position the work as a method for runtime monitoring, runtime verification, and statistical validation of formal safety properties, and for surfacing AI and checker weaknesses.
Significance. The main contribution is a concrete integration of a formal model with real AI components, replacing hand-coded error probabilities with actual detector and checker behavior. This is a useful step toward validating AI-based railway systems, and the paper reports a non-obvious, credible failure mode: certified control's false rejections can create safety-critical situations despite the safety shield. The experimental setup is transparent about the random sampling of images and environment changes, and the use of 500 runs is reasonable for a demonstrator. The paper is honest about its main limitation (no real train rides) and about the fact that the components come from prior work. If the reported effects are confirmed with the additional statistical detail requested below, the work will be a solid contribution.
major comments (2)
- [§3.2, Table 1] The statement 'With certified control, it can be obtained that all false detections are correctly rejected' is not supported by the metric defined in the table caption. The caption says False/Correct Det. count 'activated operations for false/correct detections', so a detection rejected by the certificate checker produces no operation and is not counted. The zero entries in the certified-control columns therefore show that no false operation was executed, not that every false positive was rejected by the checker. Please report the raw YOLO detections split by ground-truth class and the checker's accept/reject decisions, so that the reader can verify the rejection claim.
- [Abstract, §3.2, §5] The 'statistical validation' claim needs uncertainty quantification. Table 1 reports point estimates from 500 runs, but for the safety percentages (binomial proportions) and for the distance values no confidence intervals or standard errors are given, and the text says nothing about random seeds or the number of distinct images sampled. Without this, the reader cannot assess whether differences such as Safe 82.8% vs 80.4% (Safety Shield, No vs NoStop) are meaningful, and the phrase 'statistical validation of formal safety properties' in the abstract and conclusion overstates what Table 1 shows. Please add at least binomial confidence intervals for the safety percentages and state the sampling details.
minor comments (3)
- [§3.1] The simulation procedure says 'if no signal has been detected: ignore and do not execute any operation in the B model', but §3.2 states that the safety shield stops the train when no signal is detected at an expected position. Please clarify how the absence of a detection is communicated to the B model, presumably through the subsequent movement or environment-change event that consults the shield's condition.
- [§3.3] The image-sampling representativeness limitation is acknowledged in §3.3, but the abstract and §5 still claim 'statistical validation of formal safety properties' without qualification. Please rephrase these statements to make explicit that the statistics are conditional on the synthetic image-selection distribution and do not directly transfer to real operation.
- [Table 1] The '-' entries for Correct Det. in the no-safety-shield columns should be replaced with 0 or explained, since the text says the train never reaches a position where a correct detection is possible; as printed, the table invites confusion about missing data.
Circularity Check
No significant circularity: the central failure-mode finding is an emergent result of running a real YOLO model and certificate checker against a formal B model, not a consequence forced by definitions or self-citations.
full rationale
The paper makes no first-principles prediction that reduces to its inputs. Its load-bearing claim is that coupling the previously developed B model [14], the previously developed certificate checker [31], and ProB/SimB's external-simulation interface [34] lets the authors surface failure modes in Monte Carlo runs. The central reported result—that false rejections by the certificate checker can defeat the safety shield when a correctly detected permission signal later falls back to stop (Section 3.2, Table 1)—is an emergent interaction observed in runs, not a property of the model's event definitions. The certificate checker's false-rejection behavior is measured on real YOLO outputs and real images rather than assumed; the false-detection counts and safety-critical percentages in Table 1 come from simulation statistics, not from a fitted parameter renamed as a prediction. The only self-citations are to prior artifacts used as components; no uniqueness theorem is imported and no alternative is forbidden by citation. The explicitly acknowledged simplification that images are randomly sampled from video collections rather than from a real train ride (Sections 3.1 and 3.3: 'we have not yet simulated real train rides') concerns external validity of the measured statistics, not circularity of the derivation. Consequently, no circular step can be exhibited.
Assumptions & free parameters
free parameters (4)
- signal_visibility_distance =
10 distance units (freely selectable)
- environment_change_probability =
25% environment change vs. 75% train movement
- detection_placement_distance =
fixed distance in front of the train
- simulation_termination_condition =
stop at safety-critical situation or when train can no longer proceed
assumptions (5)
- domain assumption The formal B model is a faithful abstraction of the real shunting yard, steering logic, and signal configuration.
- ad hoc to paper Randomly sampled images from video collections are representative of the formal model's current state.
- domain assumption The safety properties SAF1-5 correctly encode what counts as a safety-critical situation.
- domain assumption The certificate checker is a deterministic trusted monitor whose accept/reject decisions reflect genuine perceptual properties.
- ad hoc to paper Detections can be mapped to formal events without position information.
Cite this review
Pith. "Pith review of Using Formal Models, Safety Shields and Certified Control to Validate AI-Based Train Systems." pith.science (2026). https://pith.science/paper/QNA4J767
@misc{pith2026241114374,
author = {Pith},
title = {Pith review of: Using Formal Models, Safety Shields and Certified Control to Validate AI-Based Train Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/QNA4J767}},
note = {Machine review of arXiv:2411.14374}
}
read the original abstract
The certification of autonomous systems is an important concern in science and industry. The KI-LOK project explores new methods for certifying and safely integrating AI components into autonomous trains. We pursued a two-layered approach: (1) ensuring the safety of the steering system by formal analysis using the B method, and (2) improving the reliability of the perception system with a runtime certificate checker. This work links both strategies within a demonstrator that runs simulations on the formal model, controlled by the real AI output and the real certificate checker. The demonstrator is integrated into the validation tool ProB. This enables runtime monitoring, runtime verification, and statistical validation of formal safety properties using a formal B model. Consequently, one can detect and analyse potential vulnerabilities and weaknesses of the AI and the certificate checker. We apply these techniques to a signal detection case study and present our findings.
Figures
Reference graph
Works this paper leans on
-
[1]
Hoare (2005): The B-Book: Assigning Programs to Meanings
Jean-Raymond Abrial & A. Hoare (2005): The B-Book: Assigning Programs to Meanings . Cambridge University Press, doi:10.1017/CBO9780511624162
-
[2]
In: Proceedings FMICS, LNCS 12863, Springer, pp
Jens Bendisposto, David Geleßus, Yumiko Jansing, Michael Leuschel, Antonia Pütz, Fabian Vu & Michelle Werth (2021): ProB2-UI: A Java-Based User Interface for ProB . In: Proceedings FMICS, LNCS 12863, Springer, pp. 193–201, doi:10.1007/978-3-030-85248-1_12
-
[3]
Michael J. Butler, Philipp Körner, Sebastian Krings, Thierry Lecomte, Michael Leuschel, Luis-Fernando Mejia & Laurent V oisin (2020):The First Twenty-Five Years of Industrial Use of the B-Method. In: Proceed- ings FMICS, LNCS 12327, pp. 189–209, doi:10.1007/978-3-030-58298-2_8
-
[4]
Technical Report EN50128, European Standard
CENELEC (2011): Railway Applications – Communication, signalling and processing systems – Software for railway control and protection systems. Technical Report EN50128, European Standard
work page 2011
-
[5]
Mathieu Comptier, David Déharbe, Julien Molinero Perez, Louis Mussat, Pierre Thibaut & Denis Sabatier (2017): Safety Analysis of a CBTC System: A Rigorous Approach with Event-B . In: Proceedings RSSRail, pp. 148–159, doi:10.1007/978-3-319-68499-4_10
-
[6]
Mathieu Comptier, Michael Leuschel, Luis-Fernando Mejia, Julien Molinero Perez & Mareike Mutz (2019): Property-Based Modelling and Validation of a CBTC Zone Controller in Event-B. In: Proceedings RSSRail, pp. 202–212, doi:10.1007/978-3-030-18744-6_13. 158 Using Formal Models, Safety Shields and Certified Control to Validate AI-Based Train Systems
-
[7]
L’expérience de Siemens Transportation Systems
Daniel Dollé, Didier Essamé & Jérôme Falampin (2003): B dans le transport ferroviaire. L’expérience de Siemens Transportation Systems. Technique et Science Informatiques 22(1), pp. 11–32, doi:10.3166/tsi.22.11-32
-
[8]
López & Vladlen Koltun (2017): CARLA: An Open Urban Driving Simulator
Alexey Dosovitskiy, Germán Ros, Felipe Codevilla, Antonio M. López & Vladlen Koltun (2017): CARLA: An Open Urban Driving Simulator . CoRR abs/1711.03938, doi:10.48550/arXiv.1711.03938. arXiv:1711.03938
Show all 36 references
-
[9]
IEEE Transactions on Intelligent Transportation Systems 24(12), pp
Gianluca D’Amico, Mauro Marinoni, Federico Nesti, Giulio Rossolini, Giorgio Buttazzo, Salvatore Sabina & Gianluigi Lauro (2023): TrainSim: A Railway Simulation Framework for LiDAR and Camera Dataset Generation. IEEE Transactions on Intelligent Transportation Systems 24(12), pp...
2023
-
[10]
In: Proceedings B (B2007), LNCS 4355, Springer, Besancon, France, pp
Didier Essamé & Daniel Dollé (2007): B in Large Scale Projects: The Canarsie Line CBTC Experience. In: Proceedings B (B2007), LNCS 4355, Springer, Besancon, France, pp. 252–254, doi:10.1007/11955757_21
2007 doi
-
[11]
In: 2018 IEEE symposium on security and privacy (SP), IEEE, pp
Timon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov, Swarat Chaudhuri & Martin Vechev (2018): Ai2: Safety and robustness certification of neural networks with abstract interpretation . In: 2018 IEEE symposium on security and privacy (SP), IEEE, pp. 3–18, doi:10.110...
2018
-
[12]
P˘as˘areanu & Clark Barrett (2018):DeepSafe: A Data-Driven Approach for Assessing Robustness of Neural Networks
Divya Gopinath, Guy Katz, Corina S. P˘as˘areanu & Clark Barrett (2018):DeepSafe: A Data-Driven Approach for Assessing Robustness of Neural Networks . In: Proceedings ATV A, LNCS 11138, Springer, pp. 3–19, doi:10.1007/978-3-030-01090-4_1
2018 doi
-
[14]
In: Proceedings RSSRail, LNCS 14198, Springer, pp
Jan Gruteser, David Geleßus, Michael Leuschel, Jan Roßbach & Fabian Vu (2023): A Formal Model of Train Control with AI-based Obstacle Detection. In: Proceedings RSSRail, LNCS 14198, Springer, pp. 128–145, doi:10.1007/978-3-031-43366-5_8
2023 doi
-
[15]
In: Proceedings ICECCS 2024, LNCS 14784, Springer, pp
Jan Gruteser & Michael Leuschel (2024): Validation of RailML Using ProB. In: Proceedings ICECCS 2024, LNCS 14784, Springer, pp. 245–256, doi:10.1007/978-3-031-66456-4_13
2024 doi
-
[16]
Dominik Hansen, Michael Leuschel, Philipp Körner, Sebastian Krings, Thomas Naulin, Nader Nayeri, David Schneider & Frank Skowron (2020):Validation and real-life demonstration of ETCS hybrid level 3 principles using a formal B model. Int. J. Softw. Tools Technol. Transf. 22(3),...
2020 doi
-
[17]
STTT 24(4), pp
Christian Hensel, Sebastian Junges, Joost-Pieter Katoen, Tim Quatmann & Matthias V olk (2022): The prob- abilistic model checker Storm. STTT 24(4), pp. 589–610, doi:10.1007/s10009-021-00633-z
2022 doi
-
[18]
In: Proceedings CA V, LNCS 10426, Springer, pp
Xiaowei Huang, Marta Kwiatkowska, Sen Wang & Min Wu (2017): Safety Verification of Deep Neural Networks. In: Proceedings CA V, LNCS 10426, Springer, pp. 3–29, doi:10.1007/978-3-319-63387-9_1
2017 doi
-
[19]
CoRR abs/2104.06178, doi:10.48550/arXiv.2104.06178
Daniel Jackson, Valerie Richmond, Mike Wang, Jeff Chow, Uriel Guajardo, Soonho Kong, Sergio Cam- pos, Geoffrey Litt & Nikos Aréchiga (2021): Certified Control: An Architecture for Verifiable Safety of Autonomous Vehicles. CoRR abs/2104.06178, doi:10.48550/arXiv.2104.06178. arX...
-
[20]
P ˘as˘areanu & Huafeng Yu (2022): Case Study: Analysis of Autonomous Center Line Tracking Neural Networks
Ismet Burak Kadron, Divya Gopinath, Corina S. P ˘as˘areanu & Huafeng Yu (2022): Case Study: Analysis of Autonomous Center Line Tracking Neural Networks. In: Proceedings VSTTE 2021, LNCS 13124, Springer, pp. 104–121, doi:10.1007/978-3-030-95561-8_7
2022 doi
-
[21]
In: Proceedings CA V, LNCS 10426, Springer, pp
Guy Katz, Clark Barrett, David L Dill, Kyle Julian & Mykel J Kochenderfer (2017): Reluplex: An efficient SMT solver for verifying deep neural networks . In: Proceedings CA V, LNCS 10426, Springer, pp. 97–117, doi:10.1007/978-3-319-63387-9_5
2017 doi
-
[22]
In: Proceedings CA V, LNCS 6806, Springer, pp
Marta Kwiatkowska, Gethin Norman & David Parker (2011): PRISM 4.0: Verification of probabilistic real- time systems. In: Proceedings CA V, LNCS 6806, Springer, pp. 585–591, doi:10.1007/978-3-642-22110-1_- 47. J. Gruteser, J. Roßbach, F. Vu & M. Leuschel 159
2011 doi
-
[23]
STTT 10(2), pp
Michael Leuschel & Michael Butler (2008): ProB: an automated analysis toolset for the B method . STTT 10(2), pp. 185–203, doi:10.1007/s10009-007-0063-9
2008 doi
-
[24]
In: Proceedings RSSRail , LNCS 14198, Springer, pp
Michael Leuschel & Nader Nayeri (2023): Modelling, Visualisation and Proof of an ETCS Level 3 Moving Block System . In: Proceedings RSSRail , LNCS 14198, Springer, pp. 193–210, doi:10.1007/978-3-031- 43366-5_12
2023 doi
-
[25]
Nurminen (2021): Sys- tematic literature review of validation methods for AI systems
Lalli Myllyaho, Mikko Raatikainen, Tomi Männistö, Tommi Mikkonen & Jukka K. Nurminen (2021): Sys- tematic literature review of validation methods for AI systems . Journal of Systems and Software 181, p. 111050, doi:10.1016/j.jss.2021.111050
2021
-
[26]
Springer Science & Business Media, doi:10.1007/978-4-431-53856-1
Kenzo Nonami, Farid Kendoul, Satoshi Suzuki, Wei Wang & Daisuke Nakazawa (2010): Autonomous fly- ing robots: unmanned aerial vehicles and micro aerial vehicles . Springer Science & Business Media, doi:10.1007/978-4-431-53856-1
2010 doi
-
[27]
P ˘as˘areanu, Ravi Mangal, Divya Gopinath, Sinem Getir Yaman, Calum Imrie, Radu Calinescu & Huafeng Yu (2023): Closed-Loop Analysis of Vision-Based Autonomous Systems: A Case Study
Corina S. P ˘as˘areanu, Ravi Mangal, Divya Gopinath, Sinem Getir Yaman, Calum Imrie, Radu Calinescu & Huafeng Yu (2023): Closed-Loop Analysis of Vision-Based Autonomous Systems: A Case Study . In: Proceedings CA V, LNCS 13964, Springer, pp. 289–303, doi:10.1007/978-3-031-37706-8_15
2023 doi
-
[28]
In: Proceedings ISoLA, LNCS 13704, Springer, pp
Jan Peleska, Anne E Haxthausen & Thierry Lecomte (2022): Standardisation considerations for autonomous train control. In: Proceedings ISoLA, LNCS 13704, Springer, pp. 286–307, doi:10.1007/978-3-031-19762- 8_22
2022 doi
-
[29]
Girshick & Ali Farhadi (2016): You Only Look Once: Uni- fied, Real-Time Object Detection
Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick & Ali Farhadi (2016): You Only Look Once: Uni- fied, Real-Time Object Detection. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA, pp. 779–788, doi:10...
2016 doi
-
[30]
In: Proceedings KI 2024 , LNAI 14992, Springer, pp
Jan Roßbach, Oliver De Candido, Ahmed Hamman & Michael Leuschel (2024): Evaluating AI-based Com- ponents for Autonomous Railway System . In: Proceedings KI 2024 , LNAI 14992, Springer, pp. 190–203, doi:10.1007/978-3-031-70893-0_14
2024 doi
-
[31]
EPTCS 395, pp
Jan Roßbach & Michael Leuschel (2023): Certified Control for Train Sign Classification. EPTCS 395, pp. 69–76, doi:10.4204/eptcs.395.5
2023 doi
-
[32]
In: Proceedings IJCAI, pp
Wenjie Ruan, Xiaowei Huang & Marta Kwiatkowska (2018):Reachability Analysis of Deep Neural Networks with Provable Guarantees. In: Proceedings IJCAI, pp. 2651–2659, doi:10.24963/ijcai.2018/368
2018 doi
-
[33]
(2020): Scalability in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine et al. (2020): Scalability in perception for autonomous driving: Waymo open dataset. In: Proceedings of the IEEE/CVF conference on com...
-
[34]
In: NASA Formal Methods Symposium , LNCS 14627, Springer, pp
Fabian Vu, Jannik Dunkelau & Michael Leuschel (2024): Validation of Reinforcement Learning Agents and Safety Shields with ProB . In: NASA Formal Methods Symposium , LNCS 14627, Springer, pp. 279–297, doi:10.1007/978-3-031-60698-4_16
2024 doi
-
[35]
In: Proceedings ABZ, LNCS 12709, Springer, pp
Fabian Vu, Michael Leuschel & Atif Mashkoor (2021): Validation of Formal Models by Timed Probabilistic Simulation. In: Proceedings ABZ, LNCS 12709, Springer, pp. 81–96, doi:10.1007/978-3-030-77543-8_6
2021 doi
-
[36]
In: Proceedings ABZ, LNCS 12071, Springer, pp
Michelle Werth & Michael Leuschel (2020): VisB: A Lightweight Tool to Visualize Formal Models with SVG Graphics. In: Proceedings ABZ, LNCS 12071, Springer, pp. 260–265, doi:10.1007/978-3-030-48077-6_21
2020 doi
-
[37]
In: Proceedings RSSRail, LNCS 14198, Springer, pp
Michael Wild, Jan Steffen Becker, Günter Ehmen & Eike Möhlmann (2023): Towards Scenario-Based Certi- fication of Highly Automated Railway Systems. In: Proceedings RSSRail, LNCS 14198, Springer, pp. 78–97, doi:10.1007/978-3-031-43366-5_5
2023 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.