Pith. sign in

REVIEW 4 major objections 5 minor 66 references

SimADFuzz: Simulation-Feedback Fuzz Testing for Autonomous Driving Systems

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read SimADFuzz claims that feeding simulation feedback into both scenario selection and mutation lets a fuzzer find more safety-critical scenarios for autonomous driving systems, reporting 35 unique violations in 6 hours.

desk verdict Worth a serious referee, but the superiority claim rests on a single-run comparison with unverifiable baselines and a metric the method directly optimizes. read the letter →

arxiv 2412.13802 v1 pith:HD6R5XZK submitted 2024-12-18 cs.SE cs.RO

classification cs.SEcs.RO
keywords AutonomousdrivingsystemsFuzztestingSimulation-basedViolationpredictionmodelDistance-guidedmutationTransformerencoderCARLAInterFuser
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

SimADFuzz is a fuzz-testing framework for autonomous driving systems that tries to turn simulator data—vehicle positions, speeds, accelerations, and headings—into better test scenarios. Most prior fuzzers score scenarios with simple aggregates like minimum distance or counts of hard brakes, ignoring the sequence of interactions over time, and they mutate scenarios randomly. SimADFuzz instead learns a violation-prediction model from the temporal feedback and uses distance-guided mutation that removes vehicles unlikely to interact with the ego vehicle. In a 6-hour CARLA campaign against InterFuser, it reports 35 unique violations, including 4 reproducible collisions, compared with 26 for TM-Fuzzer, 18 for DriveFuzz, and 3 for AV-Fuzzer. If the comparisons hold, the method would give ADS testers a way to expose more safety failures per simulation hour.

What carries the argument

The load-bearing object is the violation prediction model (VPM), a Transformer encoder over a $T \times N_{\text{info}}$ tensor of per-vehicle coordinates and physical states that emits a scalar violation probability used as the primary fitness score. It is paired with a distance-guided mutation procedure (Algorithm 1) that computes an inter-vehicle Euclidean distance matrix $m_{\text{dis}}(t, v_1, v_2)$, flags vehicles with route length below threshold $w$ as stuck and vehicles whose cumulative distance to the ego vehicle never decreases over sliding window $u$ as leaving, removes them, and spawns new NPC vehicles with routes that cross the ego vehicle's path. NSGA-2 selects scenarios on the Pareto frontier of VPM probability, minimum distance, unique violations, and SDC-Scissor static road features.

What would settle it

Re-run the three baselines from their official implementations with their originally reported settings and the same seed scenarios for 6 hours in CARLA Town03; if AV-Fuzzer's unique-violation count rises to the levels reported in its own paper, the claimed 32-violation advantage would be an artifact of baseline configuration rather than evidence for SimADFuzz.

Watch

Extended reading notes

Core claim

The paper's central claim is that simulation feedback should drive both halves of the genetic algorithm, not just selection. SimADFuzz embeds each scenario as a sequence of scenes and feeds the ego and NPC vehicle states into a Transformer encoder; a violation prediction layer outputs a probability that the scenario triggers a violation. That probability is combined with minimum distance, number of unique violations, and SDC-Scissor road-attribute scores under NSGA-2 to select Pareto-optimal parent scenarios. For mutation, the algorithm removes NPC vehicles that are stuck or persistently moving away from the ego vehicle and replaces them with vehicles whose routes intersect the ego trajectory, increasing interaction likelihood. In the evaluation, the full pipeline detected 35 unique violations in 6 hours—4 collisions, 20 lane invasions, 11 stuck violations—and achieved 61.25% map trajectory coverage, and the authors state that all detected scenarios can be replayed to reproduce the violations.

Load-bearing premise

The results assume that AV-Fuzzer, DriveFuzz, and TM-Fuzzer were faithfully reproduced and fairly configured, so the reported gap in unique violations reflects SimADFuzz's design rather than weak or mis-tuned baselines.

Editorial extensions

If this is right

  • If the 6-hour comparison holds, SimADFuzz finds more unique safety violations than TM-Fuzzer, DriveFuzz, and AV-Fuzzer under the same simulation budget.
  • Each added component contributes: the full selection-and-mutation combination finds 24 unique violations in 3 hours versus 14 for random selection and random mutation.
  • Distance-guided mutation increases nearby NPC vehicles from 23 to 35 over 3 hours, supporting the claim that proximity drives interaction and violation discovery.
  • SimADFuzz finds its first collision within 17 minutes, so the method provides early safety-critical signal during a campaign.
  • The generated scenarios cover 61.25% of Town03 waypoints, versus 3.04% for AV-Fuzzer, 13.85% for DriveFuzz, and 24.88% for TM-Fuzzer, so diversity is a measured side effect.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: because only InterFuser on Town03 is tested, the method's generality to other ADS stacks and maps remains open; the feedback features are generic, so retraining the VPM per domain is the natural extension.
  • Editorial inference: the reported 35-to-3 gap over AV-Fuzzer may partly reflect baseline configuration rather than method quality; re-running the baselines with their original tuned settings would separate those effects.
  • Editorial inference: the 1,000-scenario training set needed for the violation prediction model is a real adoption cost that the 6-hour comparison excludes; teams would either reuse the released model or generate their own labeled scenarios.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents SimADFuzz, a fuzz testing framework for autonomous driving systems (ADS) in simulation. It augments a genetic algorithm with a Transformer-based violation prediction model (VPM) for scenario fitness evaluation and a distance-guided mutation strategy that removes stuck or departing NPC vehicles. The method is evaluated on InterFuser in CARLA Town03 against three baselines (AV-Fuzzer, DriveFuzz, TM-Fuzzer), reporting 35 unique violations (UVs) in 6 hours, including 4 reproducible collisions, versus 26, 18, and 3 for the baselines. An ablation study over selection/mutation variants supports the contribution of each component.

Significance. If the quantitative results are trustworthy, the work makes a useful contribution to simulation-based ADS testing: the model-based fitness evaluation addresses a known weakness of aggregate fitness functions, and the distance-guided mutation is a simple, well-motivated mechanism for increasing interaction density. The paper's ablation structure (V/S/R × D/R) is clear, Algorithm 1 is concrete, and the four collision case studies with controller displays are valuable qualitative evidence. However, the significance is conditional on the evaluation being statistically sound and the baselines being faithfully implemented; in its current form, the headline 32-violation advantage may be an artifact of a single run and a possibly weak AV-Fuzzer baseline.

major comments (4)
  1. [§5.2.2, Figure 8] The central comparative claim rests on a single execution of each fuzzing configuration: no error bars, confidence intervals, or statistical tests are reported. With stochastic genetic search, the 35 vs 26 UV gap over TM-Fuzzer could easily arise from run-to-run variance. Please run each configuration (including the ablation variants in Figure 7) multiple times (at least 3–5 independent runs), report the distribution (e.g., median with IQR) and apply an appropriate non-parametric test or effect-size measure on the UV counts.
  2. [§5.1.3 (Baselines), Figure 8] The manuscript gives no implementation details, configuration parameters, or adaptation notes for AV-Fuzzer, DriveFuzz, and TM-Fuzzer. In particular, AV-Fuzzer's result of 3 UVs in 6 hours is far below the numbers reported in the original AV-Fuzzer paper, so the reader cannot rule out a degenerate or unfairly configured baseline. Please provide the exact versions, parameters, and porting protocol, and ideally release the baseline implementations alongside the artifact so the comparison can be independently reproduced.
  3. [Introduction and §5.2.2 (Answer to RQ2)] The claimed differences are internally inconsistent: the introduction and the RQ2 answer state that SimADFuzz finds 32, 27, and 9 more violations than AV-Fuzzer, DriveFuzz, and TM-Fuzzer, respectively, but Figure 8 reports final UV counts of 35 (SimADFuzz), 3 (AV-Fuzzer), 18 (DriveFuzz), and 26 (TM-Fuzzer), which give differences of 32, 17, and 9. The number 27 appears to be a typo, but it affects the headline result and must be corrected consistently.
  4. [§5.3.1 (Internal Validity)] The paper states that "the source code of SimADFuzz is publicly available," but no repository URL, DOI, or artifact identifier appears anywhere in the manuscript. Without an accessible artifact, the reproducibility claim and the promised mitigation of implementation threats cannot be verified. Please provide a stable link or artifact ID.
minor comments (5)
  1. [§5.1.2 (Fuzzing Configurations)] Please clarify how the VPM input handles variable-length scenarios: a scenario can run up to 10 minutes at 20 Hz (12,000 frames), so the Transformer's T dimension must be subsampled or padded; the paper does not state which method is used.
  2. [§4.2.1 (Model-based Fitness Evaluation)] The VPM is trained on 1,000 simple 2-minute scenarios but used to score longer, more complex fuzzing scenarios; the paper does not assess the VPM's prediction accuracy or the distribution shift. Please report the VPM's validation performance and, if possible, its correlation with actual violations on the fuzzing distribution.
  3. [Abstract and Introduction] The abstract's phrasing "identifying 32 more unique violations" is ambiguous: it could be read as a total advantage over all baselines combined, whereas the introduction clarifies it is the advantage over AV-Fuzzer alone. Please rephrase for clarity.
  4. [§5.2.1, Figure 7] The ablation figure shows single trajectories; a note that these are single runs would help, and ideally the same multiple-run protocol as RQ2 should be used.
  5. [§4.2.3, Algorithm 1] The definition of Δdis sums distance increases over a sliding window; the description "consistently moving away" matches the condition Δdis ≥ 0, but the pseudocode's loop structure (checking all t and breaking at the first negative window) could be made clearer with an explicit all() quantification.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's central claims are empirical and are not forced by construction or by a self-citation chain.

full rationale

SimADFuzz does not present a mathematical derivation chain in which an output equals an input by definition. The only learned component, the violation prediction model (VPM), is a supervised surrogate trained on 1,000 labeled short scenarios (Section 5.1.2) and used only as a fitness heuristic inside the genetic search. The reported outcome, unique violations, is measured afterward by CARLA's built-in sensors and the paper's own violation detectors (Section 4.3), not by the VPM. Thus the VPM's predictions are not renamed as results. The dual use of 'unique violations' as one of the four NSGA-II fitness objectives (Section 4.2.1) and as the evaluation metric (Section 5.1.3) is an evaluate-what-you-optimize alignment, but it is not a definitional reduction: the fitness objective counts distinct violations within an individual parent scenario, while the reported result is the cumulative number of distinct violations across the entire 6-hour fuzzing campaign. Those quantities are related but not identical by construction. The baselines (AV-Fuzzer, DriveFuzz, TM-Fuzzer) are external independent tools; no load-bearing self-citation is present, and the absence of baseline source code or configuration details (Section 5.3.1) is a reproducibility and fairness threat, not a circularity threat. Likewise, the unusually low 3-UV count for AV-Fuzzer in Figure 8 is a validity concern about the baseline port, not evidence that SimADFuzz's advantage is constructed from its own assumptions. The reproducible collision case studies (Section 5.2.2) provide independent grounding that the fuzzer finds real failures. Overall, the derivation is self-contained and no circular step can be exhibited from the paper's equations or self-citations.

Assumptions & free parameters 10 free parameters · 5 assumptions · 0 invented entities

No new physical or conceptual entities beyond a neural network model and heuristic mutation rule; these are methods, not postulated entities. The free parameters are mostly standard hyperparameters, but several (UV thresholds, stuck/leaving thresholds) directly shape the reported results.

free parameters (10)
  • VPM embedding dimension = 128
    Chosen by hand for the Transformer encoder; no sensitivity analysis reported.
  • VPM head number = 3
    Chosen by hand; no sensitivity analysis reported.
  • VPM encoder layer number = 3
    Chosen by hand; no sensitivity analysis reported.
  • Stuck vehicle threshold w = 10 meters
    Threshold in Algorithm 1 for marking a vehicle as stuck based on total route length; directly affects which vehicles are mutated away.
  • Leaving vehicle time window u = 10 seconds
    Sliding window size in Algorithm 1 for detecting vehicles moving away from the ego vehicle.
  • Max pedestrian-ego distance = 20 meters
    Pedestrians are spawned within this distance to increase interaction.
  • Unique violation temporal threshold = ±10 seconds
    Used to decide whether two violations are distinct; directly controls the UV count metric.
  • Unique violation spatial threshold = ±30 meters
    Used to decide whether two violations are distinct; directly controls the UV count metric.
  • Population size / crossover / mutation probabilities = 20 / 0.5 / 0.5
    Standard genetic algorithm settings chosen without sensitivity analysis.
  • Training scenario count and duration = 1,000 scenarios of 2 minutes
    Used to train VPM and SDC-Scissor; no analysis of how this choice affects downstream performance.
assumptions (5)
  • domain assumption CARLA simulator's built-in sensors (collision, lane invasion) and the paper's rule-based detectors provide correct ground truth for the five violation types.
    The entire violation count depends on these oracles; no validation against real-world data or alternative oracles is provided (Section 4.3).
  • domain assumption InterFuser is representative of production-grade ADS; results generalize beyond this single agent.
    Only InterFuser is tested; authors acknowledge this in threats to validity (Section 5.3.2).
  • domain assumption The unique violation definition (temporal ±10s, spatial ±30m) yields a meaningful and fair comparison metric across fuzzers.
    The central comparison uses UV counts; thresholds are set by hand and no robustness check is reported (Section 5.1.3).
  • ad hoc to paper The violation prediction model trained on 1,000 simple 2-minute scenarios transfers to the longer, more complex fuzzing scenarios.
    The VPM is central to scenario selection; its training distribution differs from the test-time scenarios, and no generalization analysis is given (Section 5.1.2).
  • domain assumption Baseline fuzzers (AV-Fuzzer, DriveFuzz, TM-Fuzzer) are faithfully reimplemented and configured with equivalent budgets.
    The comparison claims superiority, but no baseline source code, configuration, or verification is provided; AV-Fuzzer's very low count (3 UVs) raises concern (Section 5.2.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of SimADFuzz: Simulation-Feedback Fuzz Testing for Autonomous Driving Systems." pith.science (2026). https://pith.science/paper/HD6R5XZK

@misc{pith2026241213802,
  author       = {Pith},
  title        = {Pith review of: SimADFuzz: Simulation-Feedback Fuzz Testing for Autonomous Driving Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HD6R5XZK}},
  note         = {Machine review of arXiv:2412.13802}
}
read the original abstract

Autonomous driving systems (ADS) have achieved remarkable progress in recent years. However, ensuring their safety and reliability remains a critical challenge due to the complexity and uncertainty of driving scenarios. In this paper, we focus on simulation testing for ADS, where generating diverse and effective testing scenarios is a central task. Existing fuzz testing methods face limitations, such as overlooking the temporal and spatial dynamics of scenarios and failing to leverage simulation feedback (e.g., speed, acceleration and heading) to guide scenario selection and mutation. To address these issues, we propose SimADFuzz, a novel framework designed to generate high-quality scenarios that reveal violations in ADS behavior. Specifically, SimADFuzz employs violation prediction models, which evaluate the likelihood of ADS violations, to optimize scenario selection. Moreover, SimADFuzz proposes distance-guided mutation strategies to enhance interactions among vehicles in offspring scenarios, thereby triggering more edge-case behaviors of vehicles. Comprehensive experiments demonstrate that SimADFuzz outperforms state-of-the-art fuzzers by identifying 32 more unique violations, including 4 reproducible cases of vehicle-vehicle and vehicle-pedestrian collisions. These results demonstrate SimADFuzz's effectiveness in enhancing the robustness and safety of autonomous driving systems.

Figures

Figures reproduced from arXiv: 2412.13802 by the authors.

Figure 1
Figure 1. Simulation-based Fuzz Testing Framework for ADS [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Multi-Vehicle Interaction at a T-junction [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. presents an overview of SimADFuzz , which comprises three modules: the simulation test engine, the simulation-feedback genetic algorithm, and the violation detector. ① Simulation Test Engine ③ Violation Detector Simulation Execctor Violation Analysis Violation Report Vehicle States Seed Scenarios ② Simulation-Feedback Genetic Algorithm Model-based Fitness Evaluation Distance-guided Mutation Strategy 1 1 1 1 t1 Tempo… view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Process of the Simulation-Feedback Genetic Algorithm [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: illustrates an example of distance-guided mutation. In the parent scenario (left), the red vehicle is identified as a stuck vehicle, as it remains stationary at a long red traffic light, while the yellow one is identified as a leaving vehicle, as its trajectory indicat…
Figure 6
Figure 6. Figure 6: Scenarios Generated During the Fuzz Testing [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: The number of UVs detected by variants of [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: The number of UVs detected by SimADFuzz and baselines To further analyze the results, we categorized the violations discovered by SimADFuzz according to their types [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: Pedestrian Collision under Low-Light Conditions (#1) [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 10
Figure 10. Figure 10: Vehicle Collision during Lane Change at the Roundabout (#2) [PITH_FULL_IMAGE:figures/full_fig_p019_10.png]
Figure 11
Figure 11. Figure 11: Vehicle Collision during Lane Change at the Intersection (#3) [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]
Figure 12
Figure 12. Figure 12: Vehicle Collision during Turn Right (#4) [PITH_FULL_IMAGE:figures/full_fig_p020_12.png]
Figure 13
Figure 13. Figure 13: Trajectory Coverage of Fuzzers in Town03 [PITH_FULL_IMAGE:figures/full_fig_p021_13.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

66 extracted references · 30 canonical work pages

  1. [1]

    Bushra Alhijawi and Arafat Awajan. 2023. Genetic algorithms: Theory, genetic operators, solutions, and applications. Evolutionary Intelligence (2023), 1–12

  2. [2]

    Baidu Apollo. 2023. Dreamview. https://developer.apollo.auto/platform/simulation.html

  3. [3]

    Christian Birchler, Nicolas Ganz, Sajad Khatiri, Alessio Gambi, and Sebastiano Panichella. 2022. Cost-effective simulation-based test selection in self-driving cars software with SDC-Scissor. In 2022 IEEE international conference on software analysis, evolution and reengineering (SANER) . IEEE, 164–168

  4. [4]

    Christian Birchler, Sajad Khatiri, Bill Bosshard, Alessio Gambi, and Sebastiano Panichella. 2023. Machine learning-based test selection for simulation-based testing of self-driving cars software. Empirical Software Engineering 28, 3 (2023), 71. https://doi.org/10.1007/s10664-023-10286-y

  5. [5]

    Marcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, and Abhik Roychoudhury. 2017. Directed Greybox Fuzzing. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security (Dallas, Texas, USA) (CCS ’17). Association for Computing Machinery, New York, NY, USA, 2329–2344. https://doi.org/10.1145/3133956.3134020

  6. [6]

    Ezequiel Castellano, Ahmet Cetinkaya, and Paolo Arcaini. 2021. Analysis of Road Representations in Search-Based Testing of Autonomous Driving Systems. In 2021 IEEE 21st International Conference on Software Quality, Reliability and Security (QRS). 167–178. https://doi.org/10.1109/QRS54544.2021.00028

  7. [7]

    Ching-Yao Chan. 2017. Advancements, prospects, and impacts of automated driving systems. International journal of transportation science and technology 6, 3 (2017), 208–216

  8. [8]

    Chongpu Chen, Xinbo Chen, Chong Guo, and Peng Hang. 2023. Trajectory Prediction for Autonomous Driving Based on Structural Informer Method. IEEE Transactions on Automation Science and Engineering (2023), 1–12. https: //doi.org/10.1109/TASE.2023.3342978

Show all 66 references
  1. [9]

    Chenyi Chen, Ari Seff, Alain Kornhauser, and Jianxiong Xiao. 2015. DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)

  2. [10]

    Mingfei Cheng, Yuan Zhou, and Xiaofei Xie. 2023. Behavexplor: Behavior diversity guided testing for autonomous driving systems. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis . 488–500

  3. [11]

    Jiarun Dai, Bufan Gao, Mingyuan Luo, Zongan Huang, Zhongrui Li, Yuan Zhang, and Min Yang. 2024. SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems. In Proceedings of the 46th IEEE/ACM International Conference on Software ...

  4. [12]

    Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. 2017. CARLA: An open urban driving simulator. In Conference on robot learning . PMLR, 1–16

  5. [13]

    Vinicius H. S. Durelli, Rafael S. Durelli, Simone S. Borges, Andre T. Endo, Marcelo M. Eler, Diego R. C. Dias, and Marcelo P. Guimarães. 2019. Machine Learning Applied to Software Testing: A Systematic Mapping Study. IEEE Transactions on Reliability 68, 3 (2019), 1189–1212. ht...

  6. [14]

    Alessio Gambi, Tri Huynh, and Gordon Fraser. 2019. Generating effective test cases for self-driving cars from police reports. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering ...

  7. [15]

    Alessio Gambi, Marc Müller, and Gordon Fraser. 2019. Asfault: Testing self-driving car software using search-based procedural content generation. In 2019 IEEE/ACM 41st International Conference on Software Engineering: Companion Proceedings (ICSE-Companion). IEEE, 27–30

  8. [16]

    Hongbo Gao, Hang Su, Yingfeng Cai, Renfei Wu, Zhengyuan Hao, Yongneng Xu, Wei Wu, Jianqing Wang, Zhijun Li, and Zhen Kan. 2021. Trajectory prediction of cyclist based on dynamic Bayesian network and long short-term memory model at unsignalized intersections. Science China Info...

  9. [17]

    Zhenhai Gao, Mingxi Bao, Fei Gao, and Minghong Tang. 2023. Probabilistic multi-modal expected trajectory prediction based on LSTM for autonomous driving. Proceedings of the Institution of Mechanical Engineers, Part D: Journal of Automobile Engineering (2023), 09544070231167906

  10. [18]

    Joshua Garcia, Yang Feng, Junjie Shen, Sumaya Almanee, Yuan Xia, and Qi Alfred Chen. 2020. A comprehensive study of autonomous vehicle bugs. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering (Seoul, South Korea) (ICSE ’20). Association for Co...

  11. [19]

    Maosi Geng, Junyi Li, Yingji Xia, and Xiqun Michael Chen. 2023. A physics-informed Transformer model for vehicle trajectory prediction on highways. Transportation research part C: emerging technologies 154 (2023), 104272

  12. [20]

    Francesco Giuliari, Irtiza Hasan, Marco Cristani, and Fabio Galasso. 2021. Transformer Networks for Trajectory Forecasting. In 2020 25th International Conference on Pattern Recognition (ICPR) . 10335–10342. https://doi.org/10.1109/ ICPR48806.2021.9412190

  13. [21]

    Fitash Ul Haq, Donghwan Shin, and Lionel C. Briand. 2022. Efficient Online Testing for DNN-Enabled Systems using Surrogate-Assisted and Many-Objective Optimization.. In ICSE. ACM, 811–822. http://dblp.uni-trier.de/db/conf/icse/ icse2022.html#Haq0B22

  14. [22]

    Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation 9, 8 (1997), 1735–1780

  15. [23]

    Zhisheng Hu, Shengjian Guo, Zhenyu Zhong, and Kang Li. 2021. Coverage-based scene fuzzing for virtual autonomous driving testing. arXiv preprint arXiv:2106.00873 (2021)

  16. [24]

    Yuqi Huai, Sumaya Almanee, Yuntianyi Chen, Xiafa Wu, Qi Alfred Chen, and Joshua Garcia. 2023. scenoRITA: Generating Diverse, Fully Mutable, Test Scenarios for Autonomous Vehicle Planning. IEEE Transactions on Software Engineering 49, 10 (2023), 4656–4676. https://doi.org/10.11...

  17. [25]

    Yuqi Huai, Yuntianyi Chen, Sumaya Almanee, Tuan Ngo, Xiang Liao, Ziwen Wan, Qi Alfred Chen, and Joshua Garcia

  18. [26]

    Seulbae Kim, Major Liu, Junghwan "John" Rhee, Yuseok Jeon, Yonghwi Kwon, and Chung Hwan Kim. 2022. DriveFuzz: Discovering Autonomous Driving Bugs through Driving Quality-Guided Fuzzing. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security (L...

  19. [27]

    Guanpeng Li, Yiran Li, Saurabh Jha, Timothy Tsai, Michael Sullivan, Siva Kumar Sastry Hari, Zbigniew Kalbarczyk, and Ravishankar Iyer. 2020. AV-FUZZER: Finding Safety Violations in Autonomous Driving Systems. In 2020 IEEE 31st International Symposium on Software Reliability En...

  20. [28]

    Yuekang Li, Yinxing Xue, Hongxu Chen, Xiuheng Wu, Cen Zhang, Xiaofei Xie, Haijun Wang, and Yang Liu. 2019. Cerebro: context-aware adaptive fuzzing for effective vulnerability detection. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conferen...

  21. [29]

    Hongliang Liang, Xiaoxiao Pei, Xiaodong Jia, Wuwei Shen, and Jian Zhang. 2018. Fuzzing: State of the art. IEEE Transactions on Reliability 67, 3 (2018), 1199–1218

  22. [30]

    Shenghao Lin, Fansong Chen, Laile Xi, Gaosheng Wang, Rongrong Xi, Yuyan Sun, and Hongsong Zhu. 2024. TM-fuzzer: fuzzing autonomous driving systems through traffic management. Automated Software Engineering 31, 2 (2024), 61

  23. [32]

    Chengjie Lu, Yize Shi, Huihui Zhang, Man Zhang, Tiexin Wang, Tao Yue, and Shaukat Ali. 2023. Learning Configurations of Operating Environment of Autonomous Vehicles to Maximize their Collisions. IEEE Transactions on Software Engineering 49, 1 (2023), 384–402. https://doi.org/1...

  24. [33]

    Xiaolei Ma, Zhimin Tao, Yinhai Wang, Haiyang Yu, and Yunpeng Wang. 2015. Long short-term memory neural network for traffic speed prediction using remote microwave sensor data. Transportation Research Part C: Emerging Technologies 54 (2015), 187–197

  25. [34]

    Oded Maler and Dejan Nickovic. 2004. Monitoring temporal properties of continuous signals. In International Symposium on Formal Techniques in Real-Time and Fault-Tolerant Systems . Springer, 152–166

  26. [35]

    Manès, HyungSeok Han, Choongwoo Han, Sang Kil Cha, Manuel Egele, Edward J

    Valentin J.M. Manès, HyungSeok Han, Choongwoo Han, Sang Kil Cha, Manuel Egele, Edward J. Schwartz, and Maverick Woo. 2021. The Art, Science, and Engineering of Fuzzing: A Survey. IEEE Transactions on Software Engineering 47, 11 (2021), 2312–2331. https://doi.org/10.1109/TSE.20...

  27. [36]

    Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015. Human-level control through deep reinforcement learning. nature 518, 7540 (2015), 529–533

  28. [37]

    Tobias Moers, Lennart Vater, Robert Krajewski, Julian Bock, Adrian Zlocki, and Lutz Eckstein. 2022. The exiD Dataset: A Real-World Trajectory Dataset of Highly Interactive Highway Scenarios in Germany. In2022 IEEE Intelligent Vehicles Symposium (IV). 958–964. https://doi.org/1...

  29. [38]

    Demin Nalic, Tomislav Mihalj, Maximilian Bäumler, Matthias Lehmann, Arno Eichberger, and Stefan Bernsteiner. 2020. Scenario based testing of automated driving systems: A literature survey. In FISITA web Congress, Vol. 10. 1

  30. [39]

    Van-Thuan Pham, Marcel Böhme, Andrew E Santosa, Alexandru Răzvan Căciulescu, and Abhik Roychoudhury. 2019. Smart greybox fuzzing. IEEE Transactions on Software Engineering 47, 9 (2019), 1980–1997

  31. [40]

    Guodong Rong, Byung Hyun Shin, Hadi Tabatabaee, Qiang Lu, Steve Lemke, M¯artin, š Možeiko, Eric Boise, Geehoon Uhm, Mark Gerow, Shalin Mehta, et al. 2020. Lgsvl simulator: A high fidelity simulator for autonomous driving. In 2020 IEEE 23rd International conference on intellige...

  32. [41]

    Barbara Schütt, Markus Steimle, Birte Kramer, Danny Behnecke, and Eric Sax. 2022. A Taxonomy for Quality in Simulation-Based Development and Testing of Automated Driving Systems. IEEE Access 10 (2022), 18631–18644. https://doi.org/10.1109/ACCESS.2022.3149542

  33. [42]

    Hao Shao, Letian Wang, Ruobing Chen, Hongsheng Li, and Yu Liu. 2023. Safety-enhanced autonomous driving using interpretable sensor fusion transformer. In Conference on Robot Learning . PMLR, 726–737

  34. [43]

    Dongdong She, Rahul Krishna, Lu Yan, Suman Jana, and Baishakhi Ray. 2020. MTFuzz: fuzzing with a multi-task neural network. In Proceedings of the 28th ACM joint meeting on European software engineering conference and symposium on the foundations of software engineering . 737–749

  35. [44]

    Dongdong She, Kexin Pei, Dave Epstein, Junfeng Yang, Baishakhi Ray, and Suman Jana. 2019. Neuzz: Efficient fuzzing with neural program smoothing. In 2019 IEEE Symposium on Security and Privacy (SP) . IEEE, 803–817

  36. [45]

    Dongdong She, Abhishek Shah, and Suman Jana. 2022. Effective seed scheduling for fuzzing with graph centrality analysis. In 2022 IEEE Symposium on Security and Privacy (SP) . IEEE, 2194–2211

  37. [46]

    Sunbeom So, Seongjoon Hong, and Hakjoo Oh. 2021. SmarTest: Effectively Hunting Vulnerable Transaction Sequences in Smart Contracts through Language Model-Guided Symbolic Execution. In30th USENIX Security Symposium (USENIX Security 21). USENIX Association, 1361–1378. https://ww...

  38. [47]

    Yang Sun, Christopher M Poskitt, Jun Sun, Yuqi Chen, and Zijiang Yang. 2022. LawBreaker: An approach for specifying traffic laws and fuzzing autonomous vehicles. InProceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering. 1–12

  39. [48]

    Haoxiang Tian, Yan Jiang, Guoquan Wu, Jiren Yan, Jun Wei, Wei Chen, Shuo Li, and Dan Ye. 2022. MOSAT: finding safety violations of autonomous driving systems using multi-objective genetic algorithm. In Proceedings of the 30th ACM Joint European Software Engineering Conference ...

  40. [49]

    Haoxiang Tian, Guoquan Wu, Jiren Yan, Yan Jiang, Jun Wei, Wei Chen, Shuo Li, and Dan Ye. 2023. Generating Critical Test Scenarios for Autonomous Driving Systems via Influential Behavior Patterns. In Proceedings of the 37th IEEE/ACM International Conference on Automated Softwar...

  41. [50]

    tier4. 2023. AWSIM. https://github.com/tier4/AWSIM

  42. [51]

    Simon Ulbrich, Till Menzel, Andreas Reschka, Fabian Schuldt, and Markus Maurer. 2015. Defining and Substantiating the Terms Scene, Situation, and Scenario for Automated Driving. In2015 IEEE 18th International Conference on Intelligent Transportation Systems. 982–988. https://d...

  43. [52]

    A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017)

  44. [53]

    Jingkang Wang, Ava Pun, James Tu, Sivabalan Manivasagam, Abbas Sadat, Sergio Casas, Mengye Ren, and Raquel Urtasun. 2021. AdvSim: Generating Safety-Critical Scenarios for Self-Driving Vehicles. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 990...

  45. [54]

    Sen Wang, Zhuheng Sheng, Jingwei Xu, Taolue Chen, Junjun Zhu, Shuhui Zhang, Yuan Yao, and Xiaoxing Ma

  46. [55]

    Tong Wang, Taotao Gu, Huan Deng, Hu Li, Xiaohui Kuang, and Gang Zhao. 2024. Dance of the ADS: Orchestrating Failures through Historically-Informed Scenario Fuzzing. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis . 1086–1098

  47. [56]

    Cheng Wen, Haijun Wang, Yuekang Li, Shengchao Qin, Yang Liu, Zhiwu Xu, Hongxu Chen, Xiaofei Xie, Geguang Pu, and Ting Liu. 2020. Memlock: Memory usage guided fuzzing. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering . 765–777

  48. [57]

    Mingyuan Wu, Ling Jiang, Jiahong Xiang, Yuqun Zhang, Guowei Yang, Huixin Ma, Sen Nie, Shi Wu, Heming Cui, and Lingming Zhang. 2022. Evaluating and improving neural program-smoothing-based fuzzing. In Proceedings of the 44th International Conference on Software Engineering . 847–858

  49. [58]

    Yufei Xu, Yu Wang, and Srinivas Peeta. 2023. Leveraging transformer model to predict vehicle trajectories in congested urban traffic. Transportation research record 2677, 2 (2023), 898–909. , Vol. 1, No. 1, Article . Publication date: December 2024. SimADFuzz: Simulation-Feedb...

  50. [59]

    Chi Zhang, Yuehu Liu, Danchen Zhao, and Yuanqi Su. 2014. RoadView: A traffic scene simulator for autonomous vehicle simulation testing. In 17th International IEEE Conference on Intelligent Transportation Systems (ITSC) . 1160–1165. https://doi.org/10.1109/ITSC.2014.6957844

  51. [60]

    Kunpeng Zhang, Xi Xiao, Xiaogang Zhu, Ruoxi Sun, Minhui Xue, and Sheng Wen. 2022. Path transitions tell more: Optimizing fuzzing schedules via runtime program states. InProceedings of the 44th International Conference on Software Engineering. 1658–1668

  52. [61]

    Xiaodong Zhang, Wei Zhao, Yang Sun, Jun Sun, Yulong Shen, Xuewen Dong, and Zijiang Yang. 2023. Testing automated driving systems by breaking many laws efficiently. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis. 942–953

  53. [62]

    Ziyuan Zhong, Gail Kaiser, and Baishakhi Ray. 2022. Neural network guided evolutionary fuzzing for finding traffic violations of autonomous vehicles. IEEE Transactions on Software Engineering (2022)

  54. [63]

    Husheng Zhou, Wei Li, Zelun Kong, Junfeng Guo, Yuqun Zhang, Bei Yu, Lingming Zhang, and Cong Liu. 2020. DeepBillboard: systematic physical-world testing of autonomous driving systems. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering (Seoul, ...

  55. [64]

    Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, Vol. 35. 11106–11115

  56. [65]

    Eckart Zitzler and Lothar Thiele. 1999. Multiobjective evolutionary algorithms: a comparative case study and the strength Pareto approach. IEEE transactions on Evolutionary Computation 3, 4 (1999), 257–271. , Vol. 1, No. 1, Article . Publication date: December 2024

  57. [2022]

    In 37th IEEE/ACM International Conference on Automated Software Engineering, ASE 2022, Rochester, MI, USA, October 10-14, 2022

    ADEPT: A Testing Platform for Simulated Autonomous Driving. In 37th IEEE/ACM International Conference on Automated Software Engineering, ASE 2022, Rochester, MI, USA, October 10-14, 2022 . ACM, 150:1–150:4. https: //doi.org/10.1145/3551349.3559528

  58. [2023]

    In Proceedings of the 45th International Conference on Software Engineering (Melbourne, Victoria, Australia) (ICSE ’23)

    Doppelgänger Test Generation for Revealing Bugs in Autonomous Driving Software. In Proceedings of the 45th International Conference on Software Engineering (Melbourne, Victoria, Australia) (ICSE ’23). IEEE Press, 2591–2603. https://doi.org/10.1109/ICSE48619.2023.00216

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.