Pith. sign in

REVIEW 3 major objections 5 minor 60 references

Industry practice in autonomous-driving testing centers on scenario-based, X-in-the-loop methods, and its central unresolved problems are scenario realism, coverage, and acceptance criteria.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 22:12 UTC pith:KQN73OVY

load-bearing objection An honest, methodologically standard exploratory interview study of ADS testing practice; the abstract overreaches a bit, but the descriptive core and the synthesized framework deserve referee time. the 3 major comments →

arxiv 2607.15820 v1 pith:KQN73OVY submitted 2026-07-17 cs.SE

In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing

classification cs.SE
keywords autonomous driving systemsADS testingscenario-based testingX-in-the-loopindustry interviewstesting practicesacceptance criteriaclosed-loop testing framework
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish what autonomous driving system testing actually looks like inside companies, rather than what textbooks or academic surveys describe. Based on nine interviews with practitioners from nine companies in six countries, it claims that industry practice is dominated by scenario-based testing executed through an X-in-the-loop pipeline (model-, software-, hardware-, and vehicle-in-the-loop), and that the field's persistent pain points are scenario realism, scenario coverage, simulation fidelity, and the absence of agreed acceptance criteria. If true, this matters because it means ADS testing is hindered less by a shortage of tools than by a shortage of shared standards and definitions of 'safe enough.' The paper also synthesizes the findings into a six-stage evidence-centered closed-loop framework intended as actionable guidance for practitioners and researchers.

Core claim

On the paper's own terms, the central discovery is a cross-company picture: despite different strategies and system types, the interviewed experts converge on scenario-based, X-in-the-loop testing as the dominant process, with metrics such as scenario coverage, collision rate, success rate, and fault rate; and they converge on four unresolved challenges—scenario realism, scenario coverage, simulation fidelity, and acceptance criteria. From that picture the paper derives an evidence-centered closed-loop testing framework that starts from safety claims rather than scenarios, builds a curated scenario portfolio, routes scenarios to appropriate environments (simulation, proving ground, real road

What carries the argument

The load-bearing synthesis is the six-stage evidence-centered closed-loop framework: (1) define safety claims and testing intent, (2) construct and maintain a scenario portfolio, (3) plan evidence production and route test environments, (4) apply evidence gates and acceptance criteria, (5) build a safety argument and decide progression or release, and (6) learn closed-loop from operation. It unifies the interview themes and converts scattered practices into a single process in which scenario selection, environment choice, and acceptance are explicitly linked to evidence. The X-in-the-loop pipeline—moving from models to software, then hardware, then full vehicles—is the concrete testing mecha

Load-bearing premise

The account depends on nine self-selected interviewees, mostly from large motor-vehicle manufacturers, giving truthful and sufficiently complete descriptions of their companies' testing, with data saturation after nine interviews taken as evidence that additional voices would not change the synthesized picture.

What would settle it

A broader, representative survey of ADS developers—covering robotaxi operators, Tier-1 suppliers, simulation vendors, and small startups—that finds a majority do not organize testing around scenarios and the X-in-the-loop pipeline, or that reports widely shared, well-defined acceptance criteria, would directly contradict the paper's central picture.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the framework is adopted, ADS testing shifts from 'run many scenarios' to 'decide what claim each scenario evidences.'
  • The open challenge list gives researchers a clear target: metrics for scenario realism and coverage, plus operational acceptance criteria.
  • Because current practice is scenario-based and X-in-the-loop, new techniques that ignore these anchors will struggle to transfer into industry.
  • Operational feedback—logs, disengagements, and non-reproducible failures—is treated not as a bonus but as a required loop for improving coverage.
  • The framework can serve as a gap-analysis checklist for companies designing or refining their own testing processes.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If acceptance criteria remain undefined, the framework's 'evidence gates' must ultimately be stocked with quantitative thresholds that the paper does not supply; a testable next step is deriving gate thresholds from operational disengagement and incident data.
  • The reported gap between official automation level (L2/L2+) and actual capability (L4) suggests that safety standards keyed to level labels may evaluate the wrong target; a follow-up study could quantify how widespread this mislabeling is.
  • If world models and end-to-end architectures reduce the sim-to-real gap as participants expect, scenario generation may shift from enumerating cases to learned generation, which the framework's scenario-portfolio stage anticipates but does not specify.
  • Because the sample is small and self-selected, the framework is best read as a synthesizing hypothesis about industry practice rather than a verified standard; a quantitative survey across a larger, more diverse sample would test whether the six stages actually appear in practice.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a semi-structured interview study of nine practitioners from nine companies across six countries, analyzed thematically, to map current ADS testing practices, challenges, solutions, and future outlooks. The paper's findings are organized as thematic models: testing processes and activities center on X-in-the-loop (model-, software-, hardware-, vehicle-in-the-loop), complemented by simulation, proving-ground, and on-road testing; approaches include scenario-based and data-driven testing with metrics such as coverage, success rate, collision, fault, and comfort. Challenges include sim-to-real gaps, incomplete scenario coverage, unclear acceptance criteria, resource constraints, and safety-argumentation gaps. The authors also propose a six-stage evidence-centered closed-loop framework linking safety claims, scenario portfolio, test-environment routing, evidence gates, safety arguments, and operational feedback. The abstract claims that current practices primarily focus on scenario-based and X-in-the-loop testing and that participants highlighted scenario realism, coverage, simulation fidelity, and acceptance criteria as major challenges.

Significance. If the empirical claims are properly calibrated, the paper offers a useful practitioner perspective on ADS testing and complements prior interview-based studies (Beringhoff et al.; Song et al.; Lou et al.) by broadening scope to practices, challenges, solutions, and outlooks. Strengths include a semi-structured protocol with member-checked transcripts, a publicly available thematic model and supplementary material on Zenodo, and a framework that is a plausible synthesis of the interview themes. The six-stage evidence-centered framework is not validated beyond the interview synthesis, but the authors present it as a reference model, which is appropriate for an exploratory study. The main threat is over-generalization from a small, self-selected sample; the contribution can be made sound by systematically reporting coding densities, saturation evidence, and limitations, and by aligning the abstract and discussion claims with the reported counts.

major comments (3)
  1. [§3.1, §3.4, §9.1, abstract] The external-validity claim is load-bearing but unsupported. Nine of 21 invited companies participated; sampling was convenience/snowball/social; Table 2 shows five large motor-vehicle manufacturers plus one large automation-machinery company, one small automotive company, one software company, and one tech/internet company, with no robotaxi operator or AI-first autonomy provider. Section 3.4 asserts saturation only impressionistically ('we observed... few new codes and themes') with no cumulative-code counts or participant-order analysis. Yet the abstract promises a 'comprehensive industry-grounded overview' and §9.1 states 'the industry has reached a relatively high consensus.' Those claims exceed the evidence. I recommend either adding a quantitative saturation analysis and participant-by-participant code coverage, or rephrasing to 'the sampled companies/participants.'
  2. [§5.2.1, Figure 3, §6.13, Figure 4, abstract] The paper's own coding contradicts the headline emphasis. Figure 3 codes 'Scenario-Based' for only P3–P6 (4/9), and §5.2.1 text concedes 'only a few participants explicitly described them as such'; the additional phrase 'all participants referred to the use of scenarios' is weaker and not shown in the coding. Similarly, Figure 4 codes 'Unclear Acceptance Criteria' for P1 only, and §6.13 says it was 'discussed by P1,' yet the abstract lists acceptance criteria as a participant-highlighted major challenge. The same applies to 'Simulation Fidelity' (P2, P8, P9) and 'Incomplete Coverage' (P4, P5, P7, P8) in Figure 4. Please report systematic participant counts for each theme and either support aggregate claims with those counts or attribute findings to specific participants.
  3. [§3.4, §5.2.1] The thematic analysis does not provide inter-coder agreement or codebook counts, and the text alternates between 'only a few participants explicitly described' and 'all participants referred to scenarios.' Since the central claim is about the prevalence of scenario-based and X-in-the-loop practice, the reader needs to know how many participants described each practice and how the coding handled implicit references. Without this, the synthesis cannot be independently checked. Adding a table of themes with participant counts and an explicit rule for when a practice is counted (explicit mention vs. inferred) would resolve the ambiguity.
minor comments (5)
  1. [§5.2.1 / Figure 3] The text says 'all participants referred to the use of scenarios,' but Figure 3 lists Scenario-Based for P3–P6 only. Please reconcile this inconsistency or clarify that Figure 3 shows explicit scenario-based testing, not scenario references.
  2. [§3.4] The saturation statement is impressionistic. Please describe the exact procedure used to assess saturation (e.g., code emergence by interview order) or acknowledge that saturation was not formally measured.
  3. [§6.13] The paper frames unclear acceptance criteria as a major challenge but only one participant (P1) explicitly discussed it. The discussion in §9.2 should be adjusted to reflect the participant-level evidence, or additional supporting data should be provided.
  4. [Title page / References] The ACM reference format and received date (2007/2009) appear to be template artifacts; the publication year and venue should be corrected.
  5. [Figure 2 and Figure 3] Figure labels mix participant IDs and theme names in a way that makes it hard to read counts; consider using a table with explicit counts alongside the figures.

Circularity Check

0 steps flagged

No circular derivation: the findings are inductive interview summaries and the framework is a synthesis of those data, not a verification of them.

full rationale

This paper contains no derivation chain of the kind that can be circular: it reports qualitative findings from nine semi-structured interviews and synthesizes them into a framework. The abstract's claims about scenario-based and X-in-the-loop practices, and about challenges such as realism, coverage, fidelity, and acceptance criteria, are presented as summaries of participant reports, not as predictions derived from a fitted model or from prior results. The proposed six-stage framework is explicitly introduced as a synthesis of the same interview findings (Section 8: 'Based on the cross-company findings presented in Section 4–7, we synthesize an evidence-centered closed-loop framework'), so it does not independently verify or generate those findings. Self-citations, such as [38] for the notion of scenario-based testing and [41] for the research gap, are contextual and not load-bearing; the new interview data are independent of those prior works. The broad coding of 'scenario-based' in Section 5.2.1 and the impressionistic saturation claim in Section 3.4 are validity and representativeness limitations, not circular reductions: no result is true by construction, no fitted parameter is relabeled as a prediction, and no uniqueness theorem from the authors' prior work is invoked to force the conclusions. Therefore the paper exhibits no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 1 invented entities

No free parameters are fitted because the study is qualitative. The central claim rests on four domain assumptions about self-report validity, saturation, sample diversity, and coding stability. One conceptual entity is introduced — the six-stage evidence-centered framework — with no independent evidence of utility.

axioms (4)
  • domain assumption Participant self-reports accurately and sufficiently describe the actual testing practices of their companies.
    All findings about 'industry practice' derive from nine interviewees' testimony (Section 3.1, Tables 1-2); no direct observation or artifact triangulation is provided.
  • domain assumption Nine interviews reached thematic saturation, so additional participants would not materially change the themes.
    Section 3.4 asserts 'data saturation had been reached' without saturation counts, inter-rater metrics, or stopping-rule criteria.
  • domain assumption The recruited companies are diverse enough to support a 'comprehensive industry-grounded overview' across countries and company sizes.
    Sampling was convenience/purposive/snowball/social (Section 3.1); Table 2 shows 7 of 9 companies with 1,001+ employees and 6 of 9 motor-vehicle manufacturers, limiting small-company and supplier perspectives.
  • domain assumption Thematic coding was applied consistently and the resulting themes are stable.
    Section 3.3 follows Cruzes et al. guidelines, but no inter-coder agreement or audit trail is reported.
invented entities (1)
  • Evidence-centered closed-loop ADS testing framework no independent evidence
    purpose: Six-stage process (safety claims, scenario portfolio, environment routing, evidence gates, safety argument, operational feedback) claimed to provide structured guidance for ADS testing.
    Proposed in Section 8 and synthesized from the same nine interviews; no independent application, case study, or benchmark validates its utility.

pith-pipeline@v1.3.0-alltime-deepseek · 27488 in / 9908 out tokens · 90639 ms · 2026-08-01T22:12:32.777298+00:00 · methodology

0 comments
read the original abstract

Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This creates growing demands for effective testing to ensure system functionality and safety. However, ADS testing remains complex and lacks well-established standards for scenario selection, performance evaluation, and acceptance criteria. To better understand current ADS testing practices and challenges, we conducted an interview study with experts working on ADS development and testing in nine companies from six different countries. Through thematic analysis, we synthesized industrial testing practices, challenges, potential solutions, future trends, and proposed an evidence-centered closed-loop testing framework for ADS testing. Our findings show that current practices primarily focus on scenario-based and X-in-the-loop testing approaches, supported by diverse tools, metrics, benchmarks, and testing strategies. The participants highlighted major challenges related to scenario realism, scenario coverage, simulation fidelity, and acceptance criteria, while also discussing potential solutions such as the use of AI, world models, and end-to-end approaches. Furthermore, participants envisioned future ADS testing to become more automated, data-driven, and transparent across the industry. Overall, this study provides a comprehensive industry-grounded overview of ADS testing, proposes an evidence-centered closed-loop testing framework to provide actionable guidance for ADS testing, and outlines important directions for future research and practice.

Figures

Figures reproduced from arXiv: 2607.15820 by Dietmar Pfahl, Federica Sarro, Johannes Betz, Mohammad Reza Mousavi, Qunying Song, Yuan Gao.

Figure 1
Figure 1. Figure 1: Characteristics of the ADS worked on by the participants, including their functionalities, operational [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: A thematic model of testing processes, including overall strategies, pipelines, activities, and the flows [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: A thematic model of testing approaches, metrics, benchmarks, and tools. [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: A thematic model of challenges for testing ADS. [PITH_FULL_IMAGE:figures/full_fig_p019_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: A thematic model of outlook and future trends for testing ADS. [PITH_FULL_IMAGE:figures/full_fig_p024_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: An evidence-centered closed-loop testing framework for ADS synthesized from the interview findings, [PITH_FULL_IMAGE:figures/full_fig_p027_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

60 extracted references · 1 canonical work pages

  1. [1]

    Sebastian Baltes and Paul Ralph. 2022. Sampling in software engineering research: A critical review and guidelines. Empirical Software Engineering27, 4 (2022), 94. doi:10.1007/s10664-021-10072-8

  2. [2]

    Felix Beringhoff, Joel Greenyer, Christian Roesener, and Matthias Tichy. 2022. Thirty-one challenges in testing automated vehicles: Interviews with experts from industry and research. In2022 IEEE Intelligent Vehicles Symposium (IV). IEEE, 360–366

  3. [3]

    Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. 2020. nuscenes: A multimodal dataset for autonomous driving. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11621–11631

  4. [4]

    Jinkang Cai, Weiwen Deng, Haoran Guang, Ying Wang, Jiangkun Li, and Juan Ding. 2022. A survey on data-driven scenario generation for automated vehicle testing.Machines10, 11 (2022), 1101

  5. [5]

    Li Chen, Penghao Wu, Kashyap Chitta, Bernhard Jaeger, Andreas Geiger, and Hongyang Li. 2024. End-to-End Autonomous Driving: Challenges and Frontiers.IEEE Transactions on Pattern Analysis and Machine Intelligence46, 12 (2024), 10164–10183. doi:10.1109/TPAMI.2024.3435937

  6. [6]

    Daniela S Cruzes and Tore Dybå. 2011. Recommended steps for thematic synthesis in software engineering. In2011 international symposium on empirical software engineering and measurement. IEEE, 275–284. doi:10.1109/ESEM.2011.36

  7. [7]

    Rafael Maiani De Mello, Pedro Correa Da Silva, and Guilherme Horta Travassos. 2014. Sampling improvement in software engineering surveys. InProceedings of the 8th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement. 1–4

  8. [8]

    Rafael Maiani de Mello, Pedro Corrêa Da Silva, and Guilherme Horta Travassos. 2015. Investigating probabilistic sampling approaches for large-scale surveys in software engineering.Journal of Software Engineering Research and Development3, 1 (2015), 8

  9. [9]

    Joanna F DeFranco and Phillip A Laplante. 2017. A content analysis process for qualitative software engineering research.Innovations in Systems and Software Engineering13, 2 (2017), 129–141

  10. [10]

    Wenhao Ding, Chejian Xu, Mansur Arief, Haohong Lin, Bo Li, and Ding Zhao. 2023. A survey on safety-critical driving scenario generation—a methodological perspective.IEEE Transactions on Intelligent Transportation Systems24, 7 (2023), 6971–6988

  11. [11]

    Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. 2017. CARLA: An open urban driving simulator. InConference on robot learning. PMLR, 1–16

  12. [12]

    Ilker Etikan, Sulaiman Abubakar Musa, and Rukayya Sunusi Alkassim. 2016. Comparison of convenience sampling and purposive sampling.American journal of theoretical and applied statistics5, 1 (2016), 1–4. doi:10.11648/j.ajtas.20160501.11

  13. [13]

    Yuan Gao, Mattia Piccinini, Yuchen Zhang, Dingrui Wang, Korbinian Moller, Roberto Brusnicki, Baha Zarrouki, Alessio Gambi, Jan Frederik Totz, Kai Storms, et al. 2026. Foundation models in autonomous driving: A survey on scenario generation and scenario analysis.IEEE Open Journal of Intelligent Transportation Systems(2026)

  14. [14]

    Ahmad Nauman Ghazi, Kai Petersen, Sri Sai Vijay Raj Reddy, and Harini Nekkanti. 2019. Survey research in software engineering: Problems and mitigation strategies.IEEE Access7 (2019), 24703–24718. doi:10.1109/ACCESS.2018.2881041

  15. [15]

    Yanchen Guan, Haicheng Liao, Zhenning Li, Jia Hu, Runze Yuan, Guohui Zhang, and Chengzhong Xu. 2024. World models for autonomous driving: An initial survey.IEEE Transactions on Intelligent Vehicles(2024)

  16. [16]

    Huxin Luo. 2025. Chinese Car Media Organizes ADAS Test, Triggering Safety Debate and Industry Re- sponse. https://www.ichongqing.info/2025/07/31/chinese-car-media-organizes-adas-test-triggering-safety-debate- and-industry-response/?utm_source=chatgpt.com (last accessed: July 13 2026)

  17. [17]

    Standard

    ISO 21448:2022(en) 2022.Road vehicles — Safety of the intended functionality. Standard. International Organization for Standardization. https://www.iso.org/obp/ui/#iso:std:77490:en , Vol. 1, No. 1, Article . Publication date: July 2018. In the Driver’s Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing 33

  18. [18]

    Standard

    J3016_202104 2021.Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles. Standard. SAE International. https://www.sae.org/standards/content/j3016_202104/

  19. [19]

    Pengliang Ji, Ruan Li, Yunzhi Xue, Qian Dong, Limin Xiao, and Rui Xue. 2021. Perspective, survey and trends: Public driving datasets and toolsets for autonomous driving virtual test. In2021 IEEE International Intelligent Transportation Systems Conference (ITSC). IEEE, 264–269

  20. [20]

    Jiri Opletal. 2025. China’s massive ADAS test: 36 cars, 15 hazard scenarios, 216 crashes. https://carnewschina.com/ 2025/07/24/chinas-massive-adas-test-36-cars-15-hazard-scenarios-216-crashes/ (last accessed: July 13 2026)

  21. [21]

    Yue Kang, Hang Yin, and Christian Berger. 2019. Test your self-driving algorithm: An overview of publicly available driving datasets and virtual testing environments.IEEE Transactions on Intelligent Vehicles4, 2 (2019), 171–185

  22. [22]

    Fauzia Khan, Mariana Falco, Hina Anwar, and Dietmar Pfahl. 2023. Safety Testing of Automated Driving Systems: A Literature Review.IEEE Access11 (2023), 120049–120072. doi:10.1109/ACCESS.2023.3327918

  23. [23]

    Patricia Lago, Per Runeson, Qunying Song, and Roberto Verdecchia. 2024. Threats to validity in software engineering– hypocritical paper section or essential analysis?. InProceedings of the 18th ACM/IEEE International symposium on empirical software engineering and measurement. 314–324

  24. [24]

    Yueyuan Li, Wei Yuan, Songan Zhang, Weihao Yan, Qiyuan Shen, Chunxiang Wang, and Ming Yang. 2024. Choose your simulator wisely: A review on open-source simulators for autonomous driving.IEEE Transactions on Intelligent Vehicles9, 5 (2024), 4861–4876

  25. [25]

    Yihan Liao, Jingyu Zhang, Jacky Keung, Yan Xiao, and Yurou Dai. 2025. Advancing autonomous driving system testing: Demands, challenges, and future directions.Information and Software Technology(2025), 107859

  26. [26]

    Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, et al. 2024. Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv:2402.17177(2024)

  27. [27]

    Guannan Lou, Yao Deng, Xi Zheng, Mengshi Zhang, and Tianyi Zhang. 2022. Testing of autonomous driving systems: where are we and where should we go?. InProceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 31–43

  28. [28]

    Jing Ma, Xiaobo Che, Yanqiang Li, and Edmund M-K Lai. 2021. Traffic scenarios for automated vehicle testing: A review of description languages and systems.Machines9, 12 (2021), 342

  29. [29]

    Stefan Riedmaier, Thomas Ponn, Dieter Ludwig, Bernhard Schick, and Frank Diermeyer. 2020. Survey on scenario-based safety assessment of automated vehicles.IEEE access8 (2020), 87456–87477

  30. [30]

    Francisca Rosique, Pedro J Navarro, Carlos Fernández, and Antonio Padilla. 2019. A systematic review of perception system and simulators for autonomous vehicles research.Sensors19, 3 (2019), 648

  31. [31]

    Jennifer Rowley. 2012. Conducting research interviews.Management research review35, 3/4 (2012), 260–271. doi:10. 1108/01409171211210154

  32. [32]

    Per Runeson and Martin Höst. 2009. Guidelines for conducting and reporting case study research in software engineering.Empirical software engineering14 (2009), 131–164. doi:10.1007/s10664-008-9102-8

  33. [33]

    Janet Siegmund, Norbert Siegmund, and Sven Apel. 2015. Views on internal and external validity in empirical software engineering. In2015 IEEE/ACM 37th IEEE International Conference on Software Engineering, Vol. 1. IEEE, 9–19

  34. [34]

    Amolika Sinha, Sai Chand, Vincent Vu, Huang Chen, and Vinayak Dixit. 2021. Crash and disengagement data of autonomous vehicles on public roads in California.Scientific data8, 1 (2021), 298

  35. [35]

    Dag IK Sjøberg and Gunnar Rye Bergersen. 2022. Construct validity in software engineering.IEEE Transactions on Software Engineering49, 3 (2022), 1374–1396

  36. [36]

    Qunying Song, Avner Bensoussan, and Mohammad Reza Mousavi. 2025. Synthetic versus real: an analysis of critical scenarios for autonomous vehicle testing.Automated Software Engineering32, 2 (2025), 37

  37. [37]

    Qunying Song, Emelie Engström, and Per Runeson. 2021. Concepts in testing of autonomous systems: Academic literature and industry practice. In2021 IEEE/ACM 1st Workshop on AI Engineering-Software Engineering for AI (W AIN). IEEE, 74–81

  38. [38]

    Qunying Song, Emelie Engström, and Per Runeson. 2024. An empirically grounded path forward for scenario-based testing of autonomous driving systems. InCompanion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering. 232–243

  39. [39]

    Qunying Song, Emelie Engström, and Per Runeson. 2024. Industry practices for challenging autonomous driving systems with critical scenarios.ACM Transactions on Software Engineering and Methodology33, 4 (2024), 1–35

  40. [40]

    Qunying Song, Yuan Gao, Johannes Betz, Dietmar Pfahl, Mohammad Reza Mousavi, and Federica Sarro. 2026. Supple- mentary Material for an Interview Study on ADS Testing in Industry. doi:10.5281/zenodo.21408632

  41. [41]

    Qunying Song, Ali Nouri, Håkan Sivencrona, and Federica Sarro. 2026. From Research to Practice: An Interactive Rapid Review of Autonomous Driving System Testing in Industry.arXiv preprint arXiv:2605.00531(2026)

  42. [42]

    Qunying Song and Per Runeson. 2023. Industry-academia collaboration for realism in software engineering research: Insights and recommendations.Information and Software Technology156 (2023), 107135. , Vol. 1, No. 1, Article . Publication date: July 2018. 34 Song et al

  43. [43]

    Qunying Song, He Ye, Mark Harman, and Federica Sarro. 2025. Generative AI for Testing of Autonomous Driving Systems: A Survey.arXiv preprint arXiv:2508.19882(2025)

  44. [44]

    Jack Stilgoe. 2019. Who killed elaine herzberg? InWho’s driving innovation? New technologies and the collaborative state. Springer, 1–6

  45. [45]

    Andrea Stocco, Brian Pulfer, and Paolo Tonella. 2023. Mind the Gap! A Study on the Transferability of Virtual Versus Physical-World Testing of Autonomous Driving Systems.IEEE Transactions on Software Engineering49, 4 (2023), 1928–1940. doi:10.1109/TSE.2022.3202311

  46. [46]

    Andrea Stocco, Brian Pulfer, and Paolo Tonella. 2023. Model vs system level testing of autonomous driving systems: a replication and extension study.Empirical Software Engineering28, 3 (2023), 73. doi:10.1007/s10664-023-10306-x

  47. [47]

    Jian Sun, He Zhang, Huajun Zhou, Rongjie Yu, and Ye Tian. 2021. Scenario-based test automation for highly automated vehicles: A review and paving the way for systematic safety assurance.IEEE transactions on intelligent transportation systems23, 9 (2021), 14088–14103

  48. [48]

    Shuncheng Tang, Zhenya Zhang, Yi Zhang, Jixiang Zhou, Yan Guo, Shuang Liu, Shengjian Guo, Yan-Fu Li, Lei Ma, Yinxing Xue, et al. 2023. A survey on automated driving system testing: Landscapes and trends.ACM Transactions on Software Engineering and Methodology32, 5 (2023), 1–62

  49. [49]

    Hanlin Tian, Kethan Reddy, Yuxiang Feng, Mohammed Quddus, Yiannis Demiris, and Panagiotis Angeloudis. 2025. Large (vision) language models for autonomous vehicles: Current trends and future directions.IEEE Transactions on Intelligent Transportation Systems27, 1 (2025), 187–210

  50. [50]

    TÜV SÜD AG. 2025. TÜV SÜD. https://www.tuvsud.com/en (last accessed: May 19 2026)

  51. [51]

    Roberto Verdecchia, Emelie Engström, Patricia Lago, Per Runeson, and Qunying Song. 2023. Threats to validity in software engineering research: A critical reflection.Information and Software Technology164 (2023), 107329

  52. [52]

    Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, Jiagang Zhu, and Jiwen Lu. 2024. Drivedreamer: Towards real-world-drive world models for autonomous driving. InEuropean conference on computer vision. Springer, 55–72

  53. [53]

    Yuqi Wang, Jiawei He, Lue Fan, Hongxin Li, Yuntao Chen, and Zhaoxiang Zhang. 2024. Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14749–14759

  54. [54]

    Yan Wang, Wenjie Luo, Junjie Bai, Yulong Cao, Tong Che, Ke Chen, Yuxiao Chen, Jenna Diamond, Yifan Ding, Wenhao Ding, et al. 2025. Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail.arXiv preprint arXiv:2511.00088(2025)

  55. [55]

    Waymo LLC. 2019. Waymo Open Dataset. https://github.com/waymo-research/waymo-open-dataset (last accessed: May 19 2026)

  56. [56]

    Xiongfei Wu, Mingfei Cheng, Xiaoning Ren, Qiang Hu, Jianlang Chen, Yuheng Huang, Maxime Cordy, Yao Zhang, Xiaofei Xie, Lei Ma, et al . 2026. Foundation Models for Autonomous Driving Systems: An Initial Roadmap.ACM Transactions on Software Engineering and Methodology(2026)

  57. [57]

    Xinhai Zhang, Jianbo Tao, Kaige Tan, Martin Törngren, José Manuel Gaspar Sánchez, Muhammad Rusyadi Ramli, Xin Tao, Magnus Gyllenhammar, Franz Wotawa, Naveen Mohan, et al. 2022. Finding critical scenarios for automated driving systems: A systematic mapping study.IEEE Transactions on Software Engineering49, 3 (2022), 991–1026

  58. [58]

    Jingyuan Zhao, Wenyi Zhao, Bo Deng, Zhenghong Wang, Feng Zhang, Wenxiang Zheng, Wanke Cao, Jinrui Nan, Yubo Lian, and Andrew F Burke. 2024. Autonomous driving system: A comprehensive survey.Expert Systems with Applications242 (2024), 122836

  59. [59]

    Yongqi Zhao, Ji Zhou, Dong Bi, Tomislav Mihalj, Jia Hu, and Arno Eichberger. 2026. A survey on the application of large language models in scenario-based testing of automated driving systems.IEEE Transactions on Intelligent Transportation Systems(2026)

  60. [60]

    Xiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. 2024. Drivinggaussian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 21634–21643. Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009...