REVIEW 3 major objections 6 minor 24 references
A Systematic Mapping Study on Software Architecture for AI-based Mobility Systems
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A systematic mapping study of 1,639 papers argues that the architecture of AI-based mobility systems is still immature, with most proposed solutions unvalidated and research concentrated in automotive and aircraft domains.
desk verdict A competent mapping study in a real gap, but the missing list of included studies and selection criteria make its central claims unverifiable; fixable with major revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the systematic mapping study procedure itself, following the cited guidelines for systematic literature reviews: keyword-based search over three digital libraries, duplicate removal, three iterative voting rounds with an inter-rater agreement measure, and structured data extraction against a classification scheme. Each selected study is coded along axes such as architectural style, safety pattern, architectural framework and notation, application area, validation method, and type of AI. The maturity assessment uses a six-facet taxonomy, namely evaluation, validation, solution proposal, philosophical, opinion, and experience papers, so that an area counts as mature only when studies address diverse facets. These coding axes are the machinery that turns 38 papers into the quantitative maps and gap claims of the results.
What would settle it
Retrieve the 141 studies that survived the full-text voting round, apply a written criterion for whether each study addresses the research questions, and publish the resulting list; if applying that criterion yields a set substantially different from the 38 studies, or if the re-derived counts for architectures, safety patterns, and validation methods differ materially from the figures in the results, then the mapping's picture of the field would be called into question.
Extended reading notes
Core claim
The paper's central discovery is a descriptive map of the field rather than a new architecture. It finds that most published architectures are layered or microservice styles; that redundancy, in homogeneous, heterogeneous, or on-demand forms, and runtime monitoring are the most used safety patterns; that a small set of automotive-oriented frameworks recurs across the selected studies; and that informal box-and-arrow notation is used by more than half of the studies, with formal or semi-formal notation appearing only in a minority. Applying a standard six-facet maturity classification, the authors find a strong concentration of solution-proposal and validation papers and very few evaluation, philosophical, or opinion papers, and they report that over half of the approaches are validated by no method at all. From this they conclude that the area is relatively immature and that the missing evaluation on real, industrial systems is the most significant gap.
Load-bearing premise
The load-bearing premise is that the final selection step, dropping from 141 full-text-reviewed studies to 38, was systematic and unbiased, even though the criteria for those exclusions are not listed, the excluded studies are not identified, and the 38 selected papers are not listed in this version.
Editorial extensions
If this is right
- Because most solutions are solution proposals or validation studies, the paper implies that industry-oriented evaluation on real systems is the next prerequisite for safe deployment.
- The dominance of layered and microservice styles suggests that separating AI components from the rest of the system is the current default strategy for containing AI uncertainty.
- The concentration of research in automotive and aircraft, with sparse coverage in railway and spacecraft, points to the least explored application areas.
- The paper's maturity maps imply that a researcher entering the field should expect to find few reusable, validated reference architectures for AI-based safety-critical mobility systems.
- The finding that formal and semi-formal notations are rare implies a communication gap across disciplines that the field will need to close as systems grow more complex.
Reading between the lines
- Inference: The reported underuse of reinforcement learning and end-to-end learning may partly be a classification artifact, because many primary studies are imprecise and may label such systems simply as AI or ML; a finer re-coding of the same 38 studies could change the counts for the AI-type questions.
- Inference: The missing list of the 38 selected papers and the undocumented reasons for excluding 103 full-text-reviewed studies mean the maps' stability cannot be checked; re-running the selection with explicit exclusion criteria and publishing the list would test the reproducibility of the picture.
- Inference: The maturity classification treats evaluation in a real-world setting as the sign of maturity, so the paper implicitly predicts that the most cited, industry-facing architectures are the ones most likely to persist as reference points.
- Inference: A testable extension is to run the same coding scheme on a second, independently drawn sample obtained by forward snowballing from the 38 studies and compare the architecture-style and validation-method distributions; if the distributions diverge sharply, the selection was not representative.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a systematic mapping study (following Kitchenham's guidelines) of software architectures for AI-based mobility systems. The authors searched ACM DL, IEEE Xplore, and DBLP, screened 1,520 unique studies down to 38 primary studies, and classified them according to research questions covering architectural styles, safety patterns, frameworks, notations, application areas, maturity, validation methods, and AI types. The main findings are that layered and microservice architectures dominate, redundancy is the most common safety pattern, informal notations predominate, and validation in industry settings is scarce. The paper then discusses research gaps, particularly in spacecraft and railway domains and in end-to-end and reinforcement learning.
Significance. If the mapping is reliable, the paper would be a useful resource for researchers and practitioners, consolidating scattered evidence on architecture-level safety for AI mobility systems. The study's strengths include adherence to Kitchenham's guidelines, explicit search strings, inter-rater agreement reporting (Cohen's kappa), and the use of the Wieringa et al. classification framework to assess maturity. However, its current value is contingent on the transparency of study selection, which is the main weakness assessed below.
major comments (3)
- [III-B, Footnote 1, Fig. 1] The reduction from 141 full-text-reviewed studies to 38 is not transparent. The text states only that 'additional scrutiny during the full-text reviews led to their exclusion if they were found not to address the research questions adequately,' without operationalizing 'adequately,' listing the excluded studies, or listing the included studies (Footnote 1 says the 38 papers 'will be mentioned' only upon acceptance). Because all quantitative results in Section IV (Figs. 2–10) and the gap analysis in Section V are derived from these 38 studies, every central claim in the paper is untraceable and the study cannot be replicated. This is a load-bearing violation of the transparency expected in a systematic mapping study.
- [Abstract vs. Section I vs. Section III-B] The reported number of primary studies is inconsistent: the abstract says 1,639, Section I says 1,693, and Section III-B says 1,639, of which 119 duplicates were removed to yield 1,520. Since 1,639−119=1,520 but 1,693−119≠1,520, the 1,693 figure is not a simple typo in the derived count. The authors must correct this inconsistency and reconcile the search pipeline numbers before the selection process can be audited.
- [III-C and Table III] The classification scheme is not defined precisely enough to support the reported results. Table III lists the data items to be extracted (e.g., architecture type, research facet, validation method) but does not define the category sets used in Figs. 2, 3, 6, 7, 8, and 9, nor explain whether these categories were predefined or derived during coding. Inter-rater reliability is reported only for the selection rounds, not for the extraction/classification phase, so the figures' category assignments cannot be checked for consistency or bias.
minor comments (6)
- [Section IV, Fig. 9] The caption of Fig. 9 reads 'The correlation between architectures and AI types,' but the figure plots validation/evaluation methods (Experiments, Simulation, etc.) against AI types; the caption should be corrected to match the content.
- [Section II] The text refers to 'Guassi et al. [8]' but the reference list entry is 'Guessi et al.'; please correct the spelling.
- [Abstract] The aim phrase 'all existing architectures' overstates the scope, which is limited to AI-based mobility systems; recommend changing to 'all existing architectures in AI-based mobility systems' or an equivalent formulation.
- [Section III-B] The phrase 'each author either voted for or against the paper' is ambiguous; rephrase as 'each author either voted for or against the inclusion of each paper.'
- [Table I] The DBLP search string includes 'Software Design' as an alternative to 'Software Architecture,' but this term is not listed in the keyword enumeration in Section III-A; please clarify whether the search relies only on the listed strings.
- [Section V-A] The authors state that data collection dates are documented for replication, but no search date or retrieval date is given in the paper; please add this information.
Circularity Check
No significant circularity: the mapping's outputs are descriptive summaries of an external corpus, not restatements of its inputs.
full rationale
The paper reports a systematic mapping study rather than a derived mathematical or predictive result. Its central outputs (architecture counts, maturity classifications, validation-method tallies, and gap statements) are descriptive summaries of the 38 selected external studies, and they are classified using an external framework (Wieringa et al. [23]) and standard agreement statistics (Cohen's kappa [7]), not using the authors' own definitions or prior results. The inclusion and exclusion criteria in Tab. II are stated independently of the findings, and no conclusion is forced by construction: the classification facets and the counts in Figs. 2-10 are empirical observations about the selected literature. The only self-citation, [21] by Shafaei et al. including author Kugele, appears in the introduction as motivational context ('a more rigorous consideration of uncertainties inherent to AI components is crucial... (e. g., [21], [24])') and is not used to justify any classification, architecture count, or gap claim. The paper's genuine weaknesses are reproducibility and verifiability, not circularity: the final screening step from 141 full-text-reviewed studies to 38 is described in Sect. III-B only as excluding studies that 'were found not to address the research questions adequately,' the 103 excluded studies are not listed, and footnote 1 states the 38 included papers will be named only 'in case of acceptance.' Similarly, the abstract's 1,639 primary studies vs. the introduction's 1,693 is an internal inconsistency. These are threats to validity and auditability—legitimate correctness concerns—but they do not make any claimed result equivalent to its own input. No circular step is exhibited, so the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Kitchenham and Charters' guidelines for systematic literature reviews are a valid basis for this mapping study.
- domain assumption Wieringa et al.'s six-facet classification accurately captures research maturity.
- domain assumption The ACM DL, IEEE Xplore, and DBLP, together with the constructed search strings, provide adequate coverage of the relevant literature.
- ad hoc to paper The final 38-study selection is representative of the 141 full-text-reviewed studies.
Cite this review
Pith. "Pith review of A Systematic Mapping Study on Software Architecture for AI-based Mobility Systems." pith.science (2026). https://pith.science/paper/CDTQK2O3
@misc{pith2026250601595,
author = {Pith},
title = {Pith review of: A Systematic Mapping Study on Software Architecture for AI-based Mobility Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/CDTQK2O3}},
note = {Machine review of arXiv:2506.01595}
}
read the original abstract
Background: Due to their diversity, complexity, and above all importance, safety-critical and dependable systems must be developed with special diligence. Criticality increases as these systems likely contain artificial intelligence (AI) components known for their uncertainty. As software and reference architectures form the backbone of any successful system, including safety-critical dependable systems with learning-enabled components, choosing the suitable architecture that guarantees safety despite uncertainties is of great eminence. Aim: We aim to provide the missing overview of all existing architectures, their contribution to safety, and their level of maturity in AI-based safety-critical systems. Method: To achieve this aim, we report a systematic mapping study. From a set of 1,639 primary studies, we selected 38 relevant studies dealing with safety assurance through software architecture in AI-based safety-critical systems. The selected studies were then examined using various criteria to answer the research questions and identify gaps in this area of research. Results: Our findings showed which architectures have been proposed and to what extent they have been implemented. Furthermore, we identified gaps in different application areas of those systems and explained these gaps with various arguments. Conclusion: As the AI trend continues to grow, the system complexity will inevitably increase, too. To ensure the lasting safety of the systems, we provide an overview of the state of the art, intending to identify best practices and research gaps and direct future research more focused.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Experience, results and lessons learned from automated driving on germany’s highways,
M. Aeberhard, S. Rauch, M. Bahram, G. Tanzmeister, J. Thomas, Y . Pilat, F. Homm, W. Huber, and N. Kaempchen, “Experience, results and lessons learned from automated driving on germany’s highways,” IEEE Intelligent transportation systems magazine, vol. 7, no. 1, pp. 42– 57, 2015
work page 2015
-
[2]
Security at software architecture level: A systematic mapping study,
A. Arshad and M. Usman, “Security at software architecture level: A systematic mapping study,” in15th Annual Conference on Evaluation & Assessment in Software Engineering (EASE 2011), 2011, pp. 164–168
work page 2011
-
[3]
A. Biondi, D. Casini, G. Cicero, N. Borgioli, G. Buttazzo, G. Patti, L. Leonardi, L. L. Bello, M. Solieri, P. Burgio,et al., “Sphere: A multi- soc architecture for next-generation cyber-physical systems based on heterogeneous platforms,”IEEE Access, vol. 9, pp. 75 446–75 459, 2021
work page 2021
-
[4]
A safe, secure, and predictable software architecture for deep learning in safety- critical systems,
A. Biondi, F. Nesti, G. Cicero, D. Casini, and G. Buttazzo, “A safe, secure, and predictable software architecture for deep learning in safety- critical systems,”IEEE Embedded Systems Letters, vol. 12, no. 3, pp. 78–82, 2019
work page 2019
-
[5]
The Ethics of Safety-Critical Systems,
J. Bowen, “The Ethics of Safety-Critical Systems,”Commun. ACM, vol. 43, no. 4, p. 91–97, apr 2000. [Online]. Available: https://doi.org/10.1145/332051.332078
-
[6]
Technical architectures for automotive systems,
A. Bucaioni and P. Pelliccione, “Technical architectures for automotive systems,” in2020 IEEE International Conference on Software Architec- ture (ICSA). IEEE, 2020, pp. 46–57
work page 2020
-
[7]
A coefficient of agreement for nominal scales,
J. Cohen, “A coefficient of agreement for nominal scales,”Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960. [Online]. Available: https://doi.org/10.1177/001316446002000104
-
[8]
A systematic literature review on the description of software architectures for systems of systems,
M. Guessi, V . V . G. Neto, T. Bianchi, K. R. Felizardo, F. Oquendo, and E. Y . Nakagawa, “A systematic literature review on the description of software architectures for systems of systems,” in30th Annual ACM Symposiums on Applied Computing. ACM, 2015, pp. 1433–1440. [Online]. Available: https://doi.org/10.1145/2695664.2695795
Show all 24 references
-
[9]
Architectural description of embedded sys- tems: A systematic review,
Guessi, Milena and Nakagawa, Elisa Yumi and Oquendo, Flavio and Maldonado, Jos ´e Carlos, “Architectural description of embedded sys- tems: A systematic review,” in3rd International ACM SIGSOFT Sym- posium on Architecting Critical Systems. ACM, 2012, p. 31–40
2012
-
[10]
Selene: Self-monitored dependable platform for high-performance safety-critical systems,
C. Hernandez, J. Flieh, R. Paredes, C.-A. Lefebvre, I. Allende, J. Abella, D. Trillin, M. Matschnig, B. Fischer, K. Schwarz,et al., “Selene: Self-monitored dependable platform for high-performance safety-critical systems,” in2020 23rd euromicro conference on digital system des...
2020
-
[11]
Welcome aadl resource pages
J. Hugues, “Welcome aadl resource pages.” [Online]. Available: http://www.openaadl.org/
-
[12]
Ml-based fault injection for autonomous vehicles: A case for bayesian fault injection,
S. Jha, S. S. Banerjee, T. Tsai, S. K. S. Hari, M. B. Sullivan, Z. T. Kalbarczyk, S. W. Keckler, and R. K. Iyer, “Ml-based fault injection for autonomous vehicles: A case for bayesian fault injection,” in49th Annual IEEE/IFIP International Conference on Dependable Systems and ...
2019
-
[13]
Guidelines for performing systematic literature reviews in software engineering,
B. A. Kitchenham and S. Charters, “Guidelines for performing systematic literature reviews in software engineering,” Keele University and Durham University Joint Report, Tech. Rep. EBSE 2007-001, 07 2007. [Online]. Available: https://www.elsevier.com/ data/promis misc/525444sy...
2007
-
[14]
A systematic review of system-of-systems architecture research,
J. Klein and H. van Vliet, “A systematic review of system-of-systems architecture research,” in9th international ACM SIGSOFT conference on Quality of Software Architectures, QoSA. ACM, 2013, pp. 13–22
2013
-
[15]
Safety critical systems: challenges and directions,
J. C. Knight, “Safety critical systems: challenges and directions,” in24th International Conference on Software Engineering, ICSE 2002. ACM, 2002, pp. 547–550
2002
-
[16]
Microservice archi- tectures for advanced driver assistance systems: A case-study,
J. Lotz, A. V ogelsang, O. Benderius, and C. Berger, “Microservice archi- tectures for advanced driver assistance systems: A case-study,” in2019 IEEE International Conference on Software Architecture Companion (ICSA-C). IEEE, 2019, pp. 45–52
2019
-
[17]
Software engineering for ai-based systems: A survey,
S. Mart ´ınez-Fern´andez, J. Bogner, X. Franch, M. Oriol, J. Siebert, A. Trendowicz, A. M. V ollmer, and S. Wagner, “Software engineering for ai-based systems: A survey,”ACM Trans. Softw. Eng. Methodol., vol. 31, no. 2, pp. 37e:1–37e:59, 2022
2022
-
[18]
Systems for safety and autonomous behavior in cars: The darpa grand challenge experience,
U. Ozguner, C. Stiller, and K. Redmill, “Systems for safety and autonomous behavior in cars: The darpa grand challenge experience,” Proceedings of the IEEE, vol. 95, no. 2, pp. 397–412, 2007
2007
-
[19]
Micro-safe: Microservices- and deep learning-based safety-as-a-service architecture for 6g-enabled intelligent transportation system,
C. Roy, R. Saha, S. Misra, and K. Dev, “Micro-safe: Microservices- and deep learning-based safety-as-a-service architecture for 6g-enabled intelligent transportation system,”IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 7, pp. 9765–9774, 2021
2021
-
[20]
Designing safety critical software systems to manage inherent uncertainty,
A. C. Serban, “Designing safety critical software systems to manage inherent uncertainty,” inIEEE International Conference on Software Architecture Companion (ICSA-C). IEEE, 2019, pp. 246–249
2019
-
[21]
Uncertainty in machine learning: A safety perspective on autonomous driving,
S. Shafaei, S. Kugele, M. H. Osman, and A. C. Knoll, “Uncertainty in machine learning: A safety perspective on autonomous driving,” in Computer Safety, Reliability, and Security - SAFECOMP Workshops, ser. LNCS, vol. 11094. Springer, 2018, pp. 458–464
2018
-
[22]
Autonomous driving architectures, perception and data fusion: A review,
G. Velasco-Hernandez, J. Barry, J. Walsh,et al., “Autonomous driving architectures, perception and data fusion: A review,” in2020 IEEE 16th International Conference on Intelligent Computer Communication and Processing (ICCP). IEEE, 2020, pp. 315–321
2020
-
[23]
Requirements engineering paper classification and evaluation criteria: A proposal and a discussion,
R. Wieringa, N. Maiden, N. Mead, and C. Rolland, “Requirements engineering paper classification and evaluation criteria: A proposal and a discussion,”Requir. Eng., vol. 11, pp. 102–107, 03 2006
2006
-
[24]
Un- certainties in onboard algorithms for autonomous vehicles: Challenges, mitigation, and perspectives,
K. Yang, X. Tang, J. Li, H. Wang, G. Zhong, J. Chen, and D. Cao, “Un- certainties in onboard algorithms for autonomous vehicles: Challenges, mitigation, and perspectives,”IEEE Trans. Intell. Transp. Syst., vol. 24, no. 9, pp. 8963–8987, 2023
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.