Pith. sign in

REVIEW 3 major objections 6 minor 24 references

A Systematic Mapping Study on Software Architecture for AI-based Mobility Systems

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A systematic mapping study of 1,639 papers argues that the architecture of AI-based mobility systems is still immature, with most proposed solutions unvalidated and research concentrated in automotive and aircraft domains.

desk verdict A competent mapping study in a real gap, but the missing list of included studies and selection criteria make its central claims unverifiable; fixable with major revision. read the letter →

arxiv 2506.01595 v1 pith:CDTQK2O3 submitted 2025-06-02 cs.SE

classification cs.SE
keywords systematicmappingstudysoftwarearchitectureAI-basedmobilitysystemssafety-criticalautonomousdrivingsafetypatternsmaturityresearchgaps
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a systematic mapping study that tries to establish a comprehensive overview of software architectures for AI-based mobility systems and of how those architectures contribute to safety. Starting from 1,639 primary studies, the authors narrow the set through duplicate removal and three voting rounds to 38 studies, which are then coded against a structured classification scheme. The paper's core claim is that the field is still immature: layered and microservice styles dominate, safety is pursued mostly through redundancy and monitoring, over half of the proposed solutions are never validated or evaluated, and research is concentrated in automotive and aircraft while railway and spacecraft are neglected. The authors present this map as a basis for identifying best practices and for steering future research toward validated, industry-ready architectures.

What carries the argument

The carrying mechanism is the systematic mapping study procedure itself, following the cited guidelines for systematic literature reviews: keyword-based search over three digital libraries, duplicate removal, three iterative voting rounds with an inter-rater agreement measure, and structured data extraction against a classification scheme. Each selected study is coded along axes such as architectural style, safety pattern, architectural framework and notation, application area, validation method, and type of AI. The maturity assessment uses a six-facet taxonomy, namely evaluation, validation, solution proposal, philosophical, opinion, and experience papers, so that an area counts as mature only when studies address diverse facets. These coding axes are the machinery that turns 38 papers into the quantitative maps and gap claims of the results.

What would settle it

Retrieve the 141 studies that survived the full-text voting round, apply a written criterion for whether each study addresses the research questions, and publish the resulting list; if applying that criterion yields a set substantially different from the 38 studies, or if the re-derived counts for architectures, safety patterns, and validation methods differ materially from the figures in the results, then the mapping's picture of the field would be called into question.

Watch

Extended reading notes

Core claim

The paper's central discovery is a descriptive map of the field rather than a new architecture. It finds that most published architectures are layered or microservice styles; that redundancy, in homogeneous, heterogeneous, or on-demand forms, and runtime monitoring are the most used safety patterns; that a small set of automotive-oriented frameworks recurs across the selected studies; and that informal box-and-arrow notation is used by more than half of the studies, with formal or semi-formal notation appearing only in a minority. Applying a standard six-facet maturity classification, the authors find a strong concentration of solution-proposal and validation papers and very few evaluation, philosophical, or opinion papers, and they report that over half of the approaches are validated by no method at all. From this they conclude that the area is relatively immature and that the missing evaluation on real, industrial systems is the most significant gap.

Load-bearing premise

The load-bearing premise is that the final selection step, dropping from 141 full-text-reviewed studies to 38, was systematic and unbiased, even though the criteria for those exclusions are not listed, the excluded studies are not identified, and the 38 selected papers are not listed in this version.

Editorial extensions

If this is right

  • Because most solutions are solution proposals or validation studies, the paper implies that industry-oriented evaluation on real systems is the next prerequisite for safe deployment.
  • The dominance of layered and microservice styles suggests that separating AI components from the rest of the system is the current default strategy for containing AI uncertainty.
  • The concentration of research in automotive and aircraft, with sparse coverage in railway and spacecraft, points to the least explored application areas.
  • The paper's maturity maps imply that a researcher entering the field should expect to find few reusable, validated reference architectures for AI-based safety-critical mobility systems.
  • The finding that formal and semi-formal notations are rare implies a communication gap across disciplines that the field will need to close as systems grow more complex.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: The reported underuse of reinforcement learning and end-to-end learning may partly be a classification artifact, because many primary studies are imprecise and may label such systems simply as AI or ML; a finer re-coding of the same 38 studies could change the counts for the AI-type questions.
  • Inference: The missing list of the 38 selected papers and the undocumented reasons for excluding 103 full-text-reviewed studies mean the maps' stability cannot be checked; re-running the selection with explicit exclusion criteria and publishing the list would test the reproducibility of the picture.
  • Inference: The maturity classification treats evaluation in a real-world setting as the sign of maturity, so the paper implicitly predicts that the most cited, industry-facing architectures are the ones most likely to persist as reference points.
  • Inference: A testable extension is to run the same coding scheme on a second, independently drawn sample obtained by forward snowballing from the 38 studies and compare the architecture-style and validation-method distributions; if the distributions diverge sharply, the selection was not representative.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper reports a systematic mapping study (following Kitchenham's guidelines) of software architectures for AI-based mobility systems. The authors searched ACM DL, IEEE Xplore, and DBLP, screened 1,520 unique studies down to 38 primary studies, and classified them according to research questions covering architectural styles, safety patterns, frameworks, notations, application areas, maturity, validation methods, and AI types. The main findings are that layered and microservice architectures dominate, redundancy is the most common safety pattern, informal notations predominate, and validation in industry settings is scarce. The paper then discusses research gaps, particularly in spacecraft and railway domains and in end-to-end and reinforcement learning.

Significance. If the mapping is reliable, the paper would be a useful resource for researchers and practitioners, consolidating scattered evidence on architecture-level safety for AI mobility systems. The study's strengths include adherence to Kitchenham's guidelines, explicit search strings, inter-rater agreement reporting (Cohen's kappa), and the use of the Wieringa et al. classification framework to assess maturity. However, its current value is contingent on the transparency of study selection, which is the main weakness assessed below.

major comments (3)
  1. [III-B, Footnote 1, Fig. 1] The reduction from 141 full-text-reviewed studies to 38 is not transparent. The text states only that 'additional scrutiny during the full-text reviews led to their exclusion if they were found not to address the research questions adequately,' without operationalizing 'adequately,' listing the excluded studies, or listing the included studies (Footnote 1 says the 38 papers 'will be mentioned' only upon acceptance). Because all quantitative results in Section IV (Figs. 2–10) and the gap analysis in Section V are derived from these 38 studies, every central claim in the paper is untraceable and the study cannot be replicated. This is a load-bearing violation of the transparency expected in a systematic mapping study.
  2. [Abstract vs. Section I vs. Section III-B] The reported number of primary studies is inconsistent: the abstract says 1,639, Section I says 1,693, and Section III-B says 1,639, of which 119 duplicates were removed to yield 1,520. Since 1,639−119=1,520 but 1,693−119≠1,520, the 1,693 figure is not a simple typo in the derived count. The authors must correct this inconsistency and reconcile the search pipeline numbers before the selection process can be audited.
  3. [III-C and Table III] The classification scheme is not defined precisely enough to support the reported results. Table III lists the data items to be extracted (e.g., architecture type, research facet, validation method) but does not define the category sets used in Figs. 2, 3, 6, 7, 8, and 9, nor explain whether these categories were predefined or derived during coding. Inter-rater reliability is reported only for the selection rounds, not for the extraction/classification phase, so the figures' category assignments cannot be checked for consistency or bias.
minor comments (6)
  1. [Section IV, Fig. 9] The caption of Fig. 9 reads 'The correlation between architectures and AI types,' but the figure plots validation/evaluation methods (Experiments, Simulation, etc.) against AI types; the caption should be corrected to match the content.
  2. [Section II] The text refers to 'Guassi et al. [8]' but the reference list entry is 'Guessi et al.'; please correct the spelling.
  3. [Abstract] The aim phrase 'all existing architectures' overstates the scope, which is limited to AI-based mobility systems; recommend changing to 'all existing architectures in AI-based mobility systems' or an equivalent formulation.
  4. [Section III-B] The phrase 'each author either voted for or against the paper' is ambiguous; rephrase as 'each author either voted for or against the inclusion of each paper.'
  5. [Table I] The DBLP search string includes 'Software Design' as an alternative to 'Software Architecture,' but this term is not listed in the keyword enumeration in Section III-A; please clarify whether the search relies only on the listed strings.
  6. [Section V-A] The authors state that data collection dates are documented for replication, but no search date or retrieval date is given in the paper; please add this information.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the mapping's outputs are descriptive summaries of an external corpus, not restatements of its inputs.

full rationale

The paper reports a systematic mapping study rather than a derived mathematical or predictive result. Its central outputs (architecture counts, maturity classifications, validation-method tallies, and gap statements) are descriptive summaries of the 38 selected external studies, and they are classified using an external framework (Wieringa et al. [23]) and standard agreement statistics (Cohen's kappa [7]), not using the authors' own definitions or prior results. The inclusion and exclusion criteria in Tab. II are stated independently of the findings, and no conclusion is forced by construction: the classification facets and the counts in Figs. 2-10 are empirical observations about the selected literature. The only self-citation, [21] by Shafaei et al. including author Kugele, appears in the introduction as motivational context ('a more rigorous consideration of uncertainties inherent to AI components is crucial... (e. g., [21], [24])') and is not used to justify any classification, architecture count, or gap claim. The paper's genuine weaknesses are reproducibility and verifiability, not circularity: the final screening step from 141 full-text-reviewed studies to 38 is described in Sect. III-B only as excluding studies that 'were found not to address the research questions adequately,' the 103 excluded studies are not listed, and footnote 1 states the 38 included papers will be named only 'in case of acceptance.' Similarly, the abstract's 1,639 primary studies vs. the introduction's 1,693 is an internal inconsistency. These are threats to validity and auditability—legitimate correctness concerns—but they do not make any claimed result equivalent to its own input. No circular step is exhibited, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No fitted parameters or invented entities appear because this is a review. The four axioms above are the load-bearing premises; the last one is currently unsupported by the manuscript.

assumptions (4)
  • domain assumption Kitchenham and Charters' guidelines for systematic literature reviews are a valid basis for this mapping study.
    The entire study design in Section III follows these guidelines; if the guidelines are insufficient or misapplied, the selection and classification could be invalid.
  • domain assumption Wieringa et al.'s six-facet classification accurately captures research maturity.
    Used in Section IV (RQ2.1/RQ2.2) to judge maturity per application area; a flawed framework would distort the reported maturity gaps.
  • domain assumption The ACM DL, IEEE Xplore, and DBLP, together with the constructed search strings, provide adequate coverage of the relevant literature.
    Section III-A uses these three sources; the authors acknowledge in Section V-A that some studies might have been missed, so coverage is an unverified premise.
  • ad hoc to paper The final 38-study selection is representative of the 141 full-text-reviewed studies.
    The reduction from 141 to 38 is described only as additional scrutiny excluding studies not addressing the RQs, with no criteria or list; the entire results section depends on this premise.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Systematic Mapping Study on Software Architecture for AI-based Mobility Systems." pith.science (2026). https://pith.science/paper/CDTQK2O3

@misc{pith2026250601595,
  author       = {Pith},
  title        = {Pith review of: A Systematic Mapping Study on Software Architecture for AI-based Mobility Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CDTQK2O3}},
  note         = {Machine review of arXiv:2506.01595}
}
read the original abstract

Background: Due to their diversity, complexity, and above all importance, safety-critical and dependable systems must be developed with special diligence. Criticality increases as these systems likely contain artificial intelligence (AI) components known for their uncertainty. As software and reference architectures form the backbone of any successful system, including safety-critical dependable systems with learning-enabled components, choosing the suitable architecture that guarantees safety despite uncertainties is of great eminence. Aim: We aim to provide the missing overview of all existing architectures, their contribution to safety, and their level of maturity in AI-based safety-critical systems. Method: To achieve this aim, we report a systematic mapping study. From a set of 1,639 primary studies, we selected 38 relevant studies dealing with safety assurance through software architecture in AI-based safety-critical systems. The selected studies were then examined using various criteria to answer the research questions and identify gaps in this area of research. Results: Our findings showed which architectures have been proposed and to what extent they have been implemented. Furthermore, we identified gaps in different application areas of those systems and explained these gaps with various arguments. Conclusion: As the AI trend continues to grow, the system complexity will inevitably increase, too. To ensure the lasting safety of the systems, we provide an overview of the state of the art, intending to identify best practices and research gaps and direct future research more focused.

Figures

Figures reproduced from arXiv: 2506.01595 by the authors.

Figure 1
Figure 1. Summary of the selection procedure TABLE III: Data Extraction RQ Information Description RQ1.1 Architecture type Proposed architecture or a design pattern RQ1.2 Architectural framework Architectural framework used RQ1.3 Architecture notation Notation language used for the graphical or textual representation RQ2.1 Research facet The classification of a paper according to [23] RQ2.2 Validation method The method the pr… view at source ↗
Figure 2
Figure 2. Software architectures [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 4
Figure 4. Architectural Frameworks Inf. notation MDP AADL UML FOCUS Diff. equations PCM None 0 10 20 30 27 2 1 1 1 1 1 6 [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: Used architectural notations elements such as boxes and arrows. The second group includes semi-formal notations, which facilitate the description of the architecture, design, and implementation of complex software systems. One notable example of semi-formal notations i…
Figure 6
Figure 6. Figure 6: Application areas and their maturity Experiments Simulation Use case Prototyping Case study Survey SWOT None 0 10 20 7 6 5 2 1 1 1 23 [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 7
Figure 7. Figure 7: Validation/Evaluation methods such as simulation and experimentation, do not constitute validation of the actual system. Therefore, a notable gap exists in evaluating the approach and its outcomes on real systems. RQ3.1: Analysis of the primary studies revealed the dis…
Figure 8
Figure 8. Figure 8: The correlation between software architectures and the contained AI types. [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 10
Figure 10. Figure 10: Publication venues trend over the years gap in areas like reinforcement learning or end-to-end learning. Limited Validation and Evaluation of Proposed Solu￾tions. Except for a few instances involving DNN and some AI and ML applications, most proposed solutions lack th…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

24 extracted references · 23 canonical work pages

  1. [1]

    Experience, results and lessons learned from automated driving on germany’s highways,

    M. Aeberhard, S. Rauch, M. Bahram, G. Tanzmeister, J. Thomas, Y . Pilat, F. Homm, W. Huber, and N. Kaempchen, “Experience, results and lessons learned from automated driving on germany’s highways,” IEEE Intelligent transportation systems magazine, vol. 7, no. 1, pp. 42– 57, 2015

  2. [2]

    Security at software architecture level: A systematic mapping study,

    A. Arshad and M. Usman, “Security at software architecture level: A systematic mapping study,” in15th Annual Conference on Evaluation & Assessment in Software Engineering (EASE 2011), 2011, pp. 164–168

  3. [3]

    Sphere: A multi- soc architecture for next-generation cyber-physical systems based on heterogeneous platforms,

    A. Biondi, D. Casini, G. Cicero, N. Borgioli, G. Buttazzo, G. Patti, L. Leonardi, L. L. Bello, M. Solieri, P. Burgio,et al., “Sphere: A multi- soc architecture for next-generation cyber-physical systems based on heterogeneous platforms,”IEEE Access, vol. 9, pp. 75 446–75 459, 2021

  4. [4]

    A safe, secure, and predictable software architecture for deep learning in safety- critical systems,

    A. Biondi, F. Nesti, G. Cicero, D. Casini, and G. Buttazzo, “A safe, secure, and predictable software architecture for deep learning in safety- critical systems,”IEEE Embedded Systems Letters, vol. 12, no. 3, pp. 78–82, 2019

  5. [5]

    The Ethics of Safety-Critical Systems,

    J. Bowen, “The Ethics of Safety-Critical Systems,”Commun. ACM, vol. 43, no. 4, p. 91–97, apr 2000. [Online]. Available: https://doi.org/10.1145/332051.332078

  6. [6]

    Technical architectures for automotive systems,

    A. Bucaioni and P. Pelliccione, “Technical architectures for automotive systems,” in2020 IEEE International Conference on Software Architec- ture (ICSA). IEEE, 2020, pp. 46–57

  7. [7]

    A coefficient of agreement for nominal scales,

    J. Cohen, “A coefficient of agreement for nominal scales,”Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960. [Online]. Available: https://doi.org/10.1177/001316446002000104

  8. [8]

    A systematic literature review on the description of software architectures for systems of systems,

    M. Guessi, V . V . G. Neto, T. Bianchi, K. R. Felizardo, F. Oquendo, and E. Y . Nakagawa, “A systematic literature review on the description of software architectures for systems of systems,” in30th Annual ACM Symposiums on Applied Computing. ACM, 2015, pp. 1433–1440. [Online]. Available: https://doi.org/10.1145/2695664.2695795

Show all 24 references
  1. [9]

    Architectural description of embedded sys- tems: A systematic review,

    Guessi, Milena and Nakagawa, Elisa Yumi and Oquendo, Flavio and Maldonado, Jos ´e Carlos, “Architectural description of embedded sys- tems: A systematic review,” in3rd International ACM SIGSOFT Sym- posium on Architecting Critical Systems. ACM, 2012, p. 31–40

  2. [10]

    Selene: Self-monitored dependable platform for high-performance safety-critical systems,

    C. Hernandez, J. Flieh, R. Paredes, C.-A. Lefebvre, I. Allende, J. Abella, D. Trillin, M. Matschnig, B. Fischer, K. Schwarz,et al., “Selene: Self-monitored dependable platform for high-performance safety-critical systems,” in2020 23rd euromicro conference on digital system des...

  3. [11]

    Welcome aadl resource pages

    J. Hugues, “Welcome aadl resource pages.” [Online]. Available: http://www.openaadl.org/

  4. [12]

    Ml-based fault injection for autonomous vehicles: A case for bayesian fault injection,

    S. Jha, S. S. Banerjee, T. Tsai, S. K. S. Hari, M. B. Sullivan, Z. T. Kalbarczyk, S. W. Keckler, and R. K. Iyer, “Ml-based fault injection for autonomous vehicles: A case for bayesian fault injection,” in49th Annual IEEE/IFIP International Conference on Dependable Systems and ...

  5. [13]

    Guidelines for performing systematic literature reviews in software engineering,

    B. A. Kitchenham and S. Charters, “Guidelines for performing systematic literature reviews in software engineering,” Keele University and Durham University Joint Report, Tech. Rep. EBSE 2007-001, 07 2007. [Online]. Available: https://www.elsevier.com/ data/promis misc/525444sy...

  6. [14]

    A systematic review of system-of-systems architecture research,

    J. Klein and H. van Vliet, “A systematic review of system-of-systems architecture research,” in9th international ACM SIGSOFT conference on Quality of Software Architectures, QoSA. ACM, 2013, pp. 13–22

  7. [15]

    Safety critical systems: challenges and directions,

    J. C. Knight, “Safety critical systems: challenges and directions,” in24th International Conference on Software Engineering, ICSE 2002. ACM, 2002, pp. 547–550

  8. [16]

    Microservice archi- tectures for advanced driver assistance systems: A case-study,

    J. Lotz, A. V ogelsang, O. Benderius, and C. Berger, “Microservice archi- tectures for advanced driver assistance systems: A case-study,” in2019 IEEE International Conference on Software Architecture Companion (ICSA-C). IEEE, 2019, pp. 45–52

  9. [17]

    Software engineering for ai-based systems: A survey,

    S. Mart ´ınez-Fern´andez, J. Bogner, X. Franch, M. Oriol, J. Siebert, A. Trendowicz, A. M. V ollmer, and S. Wagner, “Software engineering for ai-based systems: A survey,”ACM Trans. Softw. Eng. Methodol., vol. 31, no. 2, pp. 37e:1–37e:59, 2022

  10. [18]

    Systems for safety and autonomous behavior in cars: The darpa grand challenge experience,

    U. Ozguner, C. Stiller, and K. Redmill, “Systems for safety and autonomous behavior in cars: The darpa grand challenge experience,” Proceedings of the IEEE, vol. 95, no. 2, pp. 397–412, 2007

  11. [19]

    Micro-safe: Microservices- and deep learning-based safety-as-a-service architecture for 6g-enabled intelligent transportation system,

    C. Roy, R. Saha, S. Misra, and K. Dev, “Micro-safe: Microservices- and deep learning-based safety-as-a-service architecture for 6g-enabled intelligent transportation system,”IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 7, pp. 9765–9774, 2021

  12. [20]

    Designing safety critical software systems to manage inherent uncertainty,

    A. C. Serban, “Designing safety critical software systems to manage inherent uncertainty,” inIEEE International Conference on Software Architecture Companion (ICSA-C). IEEE, 2019, pp. 246–249

  13. [21]

    Uncertainty in machine learning: A safety perspective on autonomous driving,

    S. Shafaei, S. Kugele, M. H. Osman, and A. C. Knoll, “Uncertainty in machine learning: A safety perspective on autonomous driving,” in Computer Safety, Reliability, and Security - SAFECOMP Workshops, ser. LNCS, vol. 11094. Springer, 2018, pp. 458–464

  14. [22]

    Autonomous driving architectures, perception and data fusion: A review,

    G. Velasco-Hernandez, J. Barry, J. Walsh,et al., “Autonomous driving architectures, perception and data fusion: A review,” in2020 IEEE 16th International Conference on Intelligent Computer Communication and Processing (ICCP). IEEE, 2020, pp. 315–321

  15. [23]

    Requirements engineering paper classification and evaluation criteria: A proposal and a discussion,

    R. Wieringa, N. Maiden, N. Mead, and C. Rolland, “Requirements engineering paper classification and evaluation criteria: A proposal and a discussion,”Requir. Eng., vol. 11, pp. 102–107, 03 2006

  16. [24]

    Un- certainties in onboard algorithms for autonomous vehicles: Challenges, mitigation, and perspectives,

    K. Yang, X. Tang, J. Li, H. Wang, G. Zhong, J. Chen, and D. Cao, “Un- certainties in onboard algorithms for autonomous vehicles: Challenges, mitigation, and perspectives,”IEEE Trans. Intell. Transp. Syst., vol. 24, no. 9, pp. 8963–8987, 2023

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.