REVIEW 4 major objections 6 minor 97 references
Curiosity-Driven Testing for Sequential Decision-Making Process
T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that a curiosity signal—neural prediction error—lets black-box fuzzing find more, and more varied, crashes in deep-learning decision-makers.
desk verdict Solid RND-based fuzzing for SDMs with real efficiency gains, but the unqualified superiority claim doesn't survive contact with Coop Navi; the repair experiment is also circular. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two interacting mechanisms. First, a curiosity module: a fixed randomly initialized target network and a trainable predictor network with identical MLP architecture; the mean squared error between their outputs on a state sequence is the intrinsic reward, updated online as the fuzzer sees new states. High prediction error means the scenario is unfamiliar, and seeds that produce it are favored. Second, a multi-objective seed energy score: E(s) = e^(-alpha*r) + e^(beta*i) + gamma*r', where r is cumulative reward (low reward raises energy), i is the intrinsic curiosity reward, and r' is robustness, the Euclidean distance between final states of the original and slightly perturbed runs. Energy-p
What would settle it
Replace the trained predictor network in the curiosity module with a second frozen random network, so the 'curiosity' score is pure noise, while keeping seed selection, mutation, budget, and environments unchanged. If CureFuzz still finds as many distinct crash-triggering scenarios as reported, the prediction-error signal is not what drives the gain; if performance falls back to MDPFuzz levels, the causal role of curiosity is confirmed.
Extended reading notes
Core claim
The paper's central claim is that curiosity—measured as the prediction error between a frozen random target network and a learned predictor network trained to mimic it—is a practical novelty measure for fuzzing sequential decision-makers. CureFuzz is a black-box fuzzer built on this signal: it randomly seeds an initial corpus of legal environment states, mutates them with small perturbations, executes the SDM to obtain state sequences, and assigns each candidate an energy score combining the mean prediction-error intrinsic reward, an exponential low-cumulative-reward term, and a robustness term that measures how far the final state moves under perturbation. Seeds are selected proportionally
Load-bearing premise
CureFuzz assumes the tester can sample and mutate valid starting states inside a known legitimate state space, with a simulator to reject illegal states; for a closed SDM whose initial states cannot be controlled or validated, the approach does not apply.
Editorial extensions
If this is right
- Crash discovery without white-box access: any SDM that can be exercised by setting a legal initial state and observing state sequences can be fuzzed with CureFuzz, even if its weights and gradients are hidden.
- A 12-hour CureFuzz run detects more distinct failure modes, not just more crashes: distinct crash-type counts are up to 200% higher than MDPFuzz at 100-bin discretization, reducing duplicate debugging effort.
- Per-iteration novelty analysis costs about 0.005–0.011 seconds versus 0.033–0.985 seconds for MDPFuzz, so the diversity signal scales to high-dimensional continuous state spaces where density-based novelty is expensive.
- Discovered scenarios can be used to repair SDMs: after fine-tuning ACAS Xu on CureFuzz crashes, the number of faults found on re-testing drops by 73%.
- CureFuzz works across policy types—DNN, DRL, MARL, and IL—suggesting the method depends on the environment interface rather than on a particular learning algorithm.
Reading between the lines
- The Coop Navi result hints at a boundary-vs-interior trade-off: the generative baseline found more crashes by concentrating agents near the state-space boundary, while CureFuzz spread crashes across the interior. A hybrid that adds boundary-aware robustness to curiosity scoring might dominate both methods—a testable direction the paper does not pursue.
- Because the novelty signal is an online-trained predictor, CureFuzz's curriculum is order-dependent; early random seeds shape what counts as 'curious' later. Scheduling the corpus and updating the predictor on non-crash sequences could change the diversity of found crashes as much as the energy weights do.
- The 73% repair result is shown for one DNN policy; an untested extension is to feed CureFuzz crashes back into reinforcement-learning training as negative demonstrations or safety constraints, which would test whether the same diversity signal improves policies during learning, not just after fine-tuning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CureFuzz, a curiosity-driven black-box fuzz testing approach for sequential decision-makers (SDMs). CureFuzz uses prediction error between a fixed random target network and a learnable predictor network as an intrinsic novelty signal, combined with a multi-objective seed energy function (cumulative reward, robustness, intrinsic reward) to select seeds for mutation. The approach is evaluated on five SDMs across CARLA (RL and IL), ACAS Xu, Coop Navi (MARL), and BipedalWalker, against MDPFuzz and a generative-model baseline (G-Model). The paper reports that CureFuzz finds more total crashes and more distinct crash types than the baselines in most settings, with lower per-iteration analysis time than MDPFuzz, and that crashes found by CureFuzz can be used to fine-tune ACAS Xu, reducing detected faults by 73%.
Significance. If the results hold, CureFuzz would be a useful addition to the SDM testing toolbox: its RND-style curiosity signal is computationally cheap, the comparison against two external baselines is meaningful, and the authors report five repetitions, confidence intervals, Mann-Whitney U tests, and effect sizes. The Coop Navi exception is disclosed in the text, and the ablation study (RQ2) supports the usefulness of the curiosity mechanism. The replication package and reuse of official baseline implementations are also positives. However, the headline claim of universal superiority over the state of the art is contradicted by the paper's own Table 2/Table 3 for Coop Navi, and the RQ3 repair result is evaluated on the same crash distribution used for repair, so the practical-utility claim needs additional support. The contribution is plausible but currently over-stated.
major comments (4)
- [Abstract; §5 RQ1] The abstract claims CureFuzz 'outperforms the state-of-the-art method by a substantial margin in the total number of faults and distinct types of crash-triggering scenarios' without qualification. This is contradicted by Table 2 (Coop Navi: G-Model 185.4 vs CureFuzz 85) and Table 3 (100 bins: G-Model 184.6 vs CureFuzz 85.3). The RQ1 answer ('CureFuzz statistically significantly and substantially outperforms both baselines in these cases') likewise overstates the evidence, since the text later acknowledges G-Model performs better on Coop Navi. The abstract, RQ1 conclusion, and general claims should be revised to state explicitly that CureFuzz outperforms MDPFuzz on all five SDMs and outperforms G-Model on four of five, with Coop Navi as a notable exception.
- [§5 RQ3] The repair experiment is internally circular. The same CureFuzz-discovered crash scenarios are used to fine-tune ACAS Xu, and then CureFuzz is re-run on the repaired model to count remaining faults. A 73% reduction in this setting largely measures how well the model memorizes or fits the specific crash distribution that CureFuzz already explored; it does not demonstrate generalization to other crash-triggering scenarios. To support the claim that CureFuzz's findings 'can repair SDMs', the evaluation should include held-out crashes (e.g., crashes found by MDPFuzz or G-Model, or crashes from an independent seed/run) and report the reduction on those scenarios. Without this, the RQ3 conclusion is over-stated.
- [§3.4, Eq. (2); §4.3] The paper does not report the values of α, β, γ in Eq. (2), the isInteresting threshold (Algorithm 2, line 23), the curiosity network architecture and training hyperparameters, or the seed mutation magnitude (only stated as reused from Pang et al.). These parameters are load-bearing for the reported improvements, since fuzzing performance is often sensitive to such choices. Please provide a configuration table and, ideally, a sensitivity analysis or at least a statement of the ranges explored. Without this, the empirical comparison is not fully reproducible.
- [§3.4 and §4.2] The 'black-box' framing is weakened by the assumption in Section 3.4 that 'we are aware of the legitimate state space of the environment' and by the reliance on environment-specific validity oracles (e.g., 'We use the CARLA simulator itself to check for the validity of the mutated state' in Section 4.2). In a genuinely black-box or closed deployment where initial states cannot be freely sampled and legality cannot be checked, CureFuzz is not directly applicable. This limitation should be stated in the assumptions or in the threat-to-validity section, and the abstract/introduction should not imply that the method applies to arbitrary black-box SDMs without such state-space access.
minor comments (6)
- [Algorithm 1] The function initCuriosity returns phi_target twice; the second return value should be phi_pred.
- [Algorithm 2] Typos: 'runnign time' in Lines 6 and 16 should be 'running time'.
- [§1] 'remains a to be an ongoing challenge' is ungrammatical; should read 'remains an ongoing challenge'.
- [§5, Table 2 text] 'CureFuzz also archives an improvement of 94.2%' should be 'achieves'.
- [§8.4] The sentence 'Gong et al. [?]' has a missing reference; either cite the work or remove the placeholder.
- [§4.3] The text says curiosity module is implemented with 'Pytorch Library' and uses ReLU; consider adding the predictor's learning rate, batch size, and number of training steps per state sequence for reproducibility.
Circularity Check
No significant circularity: core claims are evaluated against external baselines and the curiosity mechanism is an external RND adaptation.
full rationale
The paper's central derivation chain is not circular. CureFuzz's curiosity mechanism is explicitly based on Random Network Distillation (Burda et al., [11]), an external prior work, and the fuzzing loop is a standard mutation-based process with a novelty-guided selection function. The performance claims are established by comparing against two external baselines (MDPFuzz and G-Model) using their official replication packages, so the central comparison is not self-referential. The evaluation metrics (total crashes and distinct crash types) are defined from the environment state space and are not constructed from CureFuzz's internal parameters. The RQ3 repair experiment uses crashes found by CureFuzz to fine-tune the model and then re-tests with CureFuzz again; this is an empirical measurement of improvement, not a definitional equivalence. Although the repair and evaluation share the same fuzzer, the fuzzer is re-initialized with a fresh predictor network and random seeds, so the 73% reduction is not forced by construction. No parameter is fitted to the target result and no 'prediction' is renamed from an input. Self-citations appear only in related-work discussions, not as load-bearing justifications for the approach. The only notable issue is an overclaim in the abstract regarding universal superiority, since G-Model outperforms CureFuzz on Coop Navi (MARL), but this is a correctness/consistency concern, not circularity.
Assumptions & free parameters
free parameters (7)
- alpha (a) in Eq. (2)
- beta (b) in Eq. (2)
- gamma (g) in Eq. (2)
- isInteresting threshold
- curiosity network hyperparameters
- robustness perturbation size
- seed mutation magnitude
assumptions (5)
- domain assumption Environments satisfy the Markov Property (MDP framework).
- domain assumption The legitimate state space of each environment is known, sampleable, and controllable by the tester, with an oracle for legal mutations.
- domain assumption The SDM policy is fixed during fuzzing and is queried only through environment interaction.
- domain assumption Random Network Distillation prediction error is a valid novelty proxy for diverse crash discovery.
- domain assumption Catastrophic failure (crash) is the testing oracle, defined per environment by the authors.
Cite this review
Pith. "Pith review of Curiosity-Driven Testing for Sequential Decision-Making Process." pith.science (2026). https://pith.science/paper/MFRF2WM7
@misc{pith2026250902025,
author = {Pith},
title = {Pith review of: Curiosity-Driven Testing for Sequential Decision-Making Process},
year = {2026},
howpublished = {\url{https://pith.science/paper/MFRF2WM7}},
note = {Machine review of arXiv:2509.02025}
}
read the original abstract
Sequential decision-making processes (SDPs) are fundamental for complex real-world challenges, such as autonomous driving, robotic control, and traffic management. While recent advances in Deep Learning (DL) have led to mature solutions for solving these complex problems, SDMs remain vulnerable to learning unsafe behaviors, posing significant risks in safety-critical applications. However, developing a testing framework for SDMs that can identify a diverse set of crash-triggering scenarios remains an open challenge. To address this, we propose CureFuzz, a novel curiosity-driven black-box fuzz testing approach for SDMs. CureFuzz proposes a curiosity mechanism that allows a fuzzer to effectively explore novel and diverse scenarios, leading to improved detection of crashtriggering scenarios. Additionally, we introduce a multi-objective seed selection technique to balance the exploration of novel scenarios and the generation of crash-triggering scenarios, thereby optimizing the fuzzing process. We evaluate CureFuzz on various SDMs and experimental results demonstrate that CureFuzz outperforms the state-of-the-art method by a substantial margin in the total number of faults and distinct types of crash-triggering scenarios. We also demonstrate that the crash-triggering scenarios found by CureFuzz can repair SDMs, highlighting CureFuzz as a valuable tool for testing SDMs and optimizing their performance.
Figures
Reference graph
Works this paper leans on
-
[1]
Abien Fred Agarap. 2018. Deep learning using rectified linear units (relu). arXiv preprint arXiv:1803.08375 (2018)
arXiv 2018
-
[2]
Briand, Ramesh S, and Mojtaba Bagherzadeh
Zohreh Aghababaeyan, Manel Abdellatif, Lionel C. Briand, Ramesh S, and Mojtaba Bagherzadeh. 2023. Black-Box Testing of Deep Neural Networks through Test ICSE ’24, April 14–20, 2024, Lisbon, Portugal Junda He, Zhou Yang, Jieke Shi, Chengran Yang, Kisub Kim, Bowen Xu, Xin Zhou, and David Lo Case Diversity. IEEE Trans. Software Eng. 49, 5 (2023), 3182–3204. ...
-
[3]
Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, and Anil Anthony Bharath. 2017. Deep reinforcement learning: A brief survey.IEEE Signal Processing Magazine 34, 6 (2017), 26–38
2017
-
[4]
Muhammad Hilmi Asyrofi, Zhou Yang, and David Lo. 2021. CrossASR++: A Modular Differential Testing Framework for Automatic Speech Recognition. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Athens, Greece) (ESEC/FSE 2021). Association for Computing Machinery...
-
[5]
Muhammad Hilmi Asyrofi, Zhou Yang, Imam Nur Bani Yusuf, Hong Jin Kang, Ferdian Thung, and David Lo. 2021. Biasfinder: Metamorphic test generation to uncover bias for sentiment analysis systems. IEEE Transactions on Software Engineering 48, 12 (2021), 5087–5101
2021
-
[6]
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Sax- ton, and Remi Munos. 2016. Unifying count-based exploration and intrinsic motivation. Advances in neural information processing systems 29 (2016)
2016
-
[7]
Robert Binder. 2000. Testing object-oriented systems: models, patterns, and tools . Addison-Wesley Professional
2000
-
[8]
Marcel Böhme. 2018. STADS: Software testing as species discovery. ACM Trans- actions on Software Engineering and Methodology (TOSEM) 27, 2 (2018), 1–52
2018
Show all 97 references
-
[9]
Marcel Böhme, Danushka Liyanage, and Valentin Wüstholz. 2021. Estimating residual risk in greybox fuzzing. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 230–241
2021
-
[10]
Marcel Böhme, Valentin J. M. Manès, and Sang Kil Cha. 2020. Boosting fuzzer efficiency: an information theoretic perspective. In ESEC/FSE ’20: 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Virtual Event, USA, ...
2020
-
[11]
Storkey, and Oleg Klimov
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov. 2019. Explo- ration by random network distillation. In7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net. https://openreview.net/forum?id=H1lJJnR5Ym
2019
-
[12]
Marcel Böhme, Van-Thuan Pham, and Abhik Roychoudhury. 2019. Coverage- Based Greybox Fuzzing as Markov Chain. IEEE Transactions on Software Engi- neering 45, 5 (2019), 489–506. https://doi.org/10.1109/TSE.2017.2785841
2019
-
[13]
Emanuela G Cartaxo, Patrícia DL Machado, and Francisco G Oliveira Neto. 2011. On the use of a similarity function for test case selection in the context of model- based testing. Software Testing, Verification and Reliability 21, 2 (2011), 75–100
2011
-
[14]
Sang Kil Cha, Maverick Woo, and David Brumley. 2015. Program-adaptive mutational fuzzing. In 2015 IEEE Symposium on Security and Privacy . IEEE, 725– 741
2015
-
[15]
Dian Chen, Brady Zhou, Vladlen Koltun, and Philipp Krähenbühl. 2020. Learning by Cheating. In Proceedings of the Conference on Robot Learning (Proceedings of Machine Learning Research, Vol. 100), Leslie Pack Kaelbling, Danica Kragic, and Komei Sugiura (Eds.). PMLR, 66–75. http...
2020
-
[16]
López, and Vladlen Koltun
Alexey Dosovitskiy, Germán Ros, Felipe Codevilla, Antonio M. López, and Vladlen Koltun. 2017. CARLA: An Open Urban Driving Simulator. In1st Annual Conference on Robot Learning, CoRL 2017, Mountain View, California, USA, November 13-15, 2017, Proceedings (Proceedings of Machine...
2017
-
[17]
Maxim Egorov. 2016. Multi-agent deep reinforcement learning. CS231n: convolu- tional neural networks for visual recognition (2016), 1–8
2016
-
[18]
Jan Eisenhut, Alvaro Torralba, Maria Christakis, and Jörg Hoffmann. 2023. Auto- matic Metamorphic Test Oracles for Action-Policy Testing. (2023)
2023
-
[19]
Hasan Ferit Eniser, Timo P Gros, Valentin Wüstholz, Jörg Hoffmann, and Maria Christakis. 2022. Metamorphic relations via relaxations: An approach to ob- tain oracles for action-policy testing. In Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing a...
2022
-
[20]
Florian Fuchs, Yunlong Song, Elia Kaufmann, Davide Scaramuzza, and Peter Dürr
- [21]
-
[22]
Robert J Grissom and John J Kim. 2005. Effect sizes for research: A broad practical approach. Lawrence Erlbaum Associates Publishers
2005
-
[23]
Simon Haykin. 1994. Neural networks: a comprehensive foundation . Prentice Hall PTR
1994
-
[24]
Junda He, Bowen Xu, Zhou Yang, DongGyun Han, Chengran Yang, and David Lo
-
[25]
Junda He, Xin Zhou, Bowen Xu, Ting Zhang, Kisub Kim, Zhou Yang, Ferdian Thung, Ivana Clairine Irsan, and David Lo. 2023. Representation Learning for Stack Overflow Posts: How Far Are We? ACM Trans. Softw. Eng. Methodol. (dec 2023). https://doi.org/10.1145/3635711 Just Accepted
2023 doi
-
[26]
Hadi Hemmati, Andrea Arcuri, and Lionel C. Briand. 2013. Achieving scalable model-based testing through test case diversity. ACM Trans. Softw. Eng. Methodol. 22, 1 (2013), 6:1–6:42. https://doi.org/10.1145/2430536.2430540
2013
-
[27]
Hadi Hemmati, Zhihan Fang, and Mika V Mantyla. 2015. Prioritizing manual test cases in traditional and rapid release environments. In2015 IEEE 8th international conference on software testing, verification and validation (ICST) . IEEE, 1–10
2015
-
[28]
Eduard Hofer, Martina Kloos, Bernard Krzykacz-Hausmann, Jörg Peschke, and Martin Woltereck. 2002. An approximate epistemic uncertainty analysis approach in the presence of epistemic and aleatory uncertainties. Reliability Engineering & System Safety 77, 3 (2002), 229–238
2002
-
[29]
Ionel-Alexandru Hosu and Traian Rebedea. 2016. Playing Atari Games with Deep Reinforcement Learning and Human Checkpoint Replay. CoRR abs/1607.05077 (2016). arXiv:1607.05077 http://arxiv.org/abs/1607.05077
2016 arXiv
-
[30]
Fan Hu, Shisong Qin, Zheyu Ma, Bodong Zhao, Tingting Yin, and Chao Zhang
-
[31]
Yuqi Huai, Sumaya Almanee, Yuntianyi Chen, Xiafa Wu, Qi Alfred Chen, and Joshua Garcia. 2023. scenoRITA: Generating Diverse, Fully Mutable, Test Scenar- ios for Autonomous Vehicle Planning. IEEE Transactions on Software Engineering 49, 10 (2023), 4656–4676. https://doi.org/10....
2023
-
[32]
Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne. 2017. Imitation learning: A survey of learning methods. ACM Computing Surveys (CSUR) 50, 2 (2017), 1–35
2017
-
[33]
David Isele, Reza Rahimi, Akansel Cosgun, Kaushik Subramanian, and Kikuo Fu- jimura. 2018. Navigating occluded intersections with autonomous vehicles using deep reinforcement learning. In 2018 IEEE international conference on robotics and automation (ICRA). IEEE, 2034–2039
2018
-
[34]
Larson, Magnus Almgren, and Vin- cenzo Gulisano
William Johansson, Martin Svensson, Ulf E. Larson, Magnus Almgren, and Vin- cenzo Gulisano. 2014. T-Fuzz: Model-Based Fuzzing for Robustness Testing of Telecommunication Protocols. In 2014 IEEE Seventh International Conference on Software Testing, Verification and Validation. ...
2014
-
[35]
Julian, Mykel J
Kyle D. Julian, Mykel J. Kochenderfer, and Michael P. Owen. 2018. Deep Neu- ral Network Compression for Aircraft Collision Avoidance Systems. CoRR abs/1810.04240 (2018). arXiv:1810.04240 http://arxiv.org/abs/1810.04240
2018 arXiv
-
[36]
George Klees, Andrew Ruef, Benji Cooper, Shiyi Wei, and Michael Hicks. 2018. Evaluating Fuzz Testing. In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security, CCS 2018, Toronto, ON, Canada, October 15-19, 2018, David Lie, Mohammad Mannan, Micha...
2018
-
[37]
W Bradley Knox, Adam Bradley Setapen, and Peter Stone. 2011. Reinforcement Learning with Human Feedback in Mountain Car.. In AAAI Spring Symposium: Help Me Help You: Bridging the Gaps in Human-Agent Collaboration
2011
-
[38]
Jens Kober, J Andrew Bagnell, and Jan Peters. 2013. Reinforcement learning in robotics: A survey. The International Journal of Robotics Research 32, 11 (2013), 1238–1274
2013
-
[39]
Mario Köppen. 2000. The curse of dimensionality. In 5th online world conference on soft computing in industrial applications (WSC5) , Vol. 1. 4–8
2000
-
[40]
Nishanth Kumar. 2020. The Past and Present of Imitation Learning: A Citation Chain Study. CoRR abs/2001.02328 (2020). arXiv:2001.02328 http://arxiv.org/abs/ 2001.02328
2020 arXiv
-
[41]
Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, and Dmitry Vetrov
-
[42]
Yuxi Li. 2017. Deep reinforcement learning: An overview. arXiv preprint arXiv:1701.07274 (2017)
2017 arXiv
-
[43]
David Lo. 2023. Trustworthy and Synergistic Artificial Intelligence for Software Engineering: Vision and Roadmaps. arXiv preprint arXiv:2309.04142 (2023)
2023 arXiv
-
[44]
Karol Lina López, Christian Gagné, and Marc-André Gardner. 2018. Demand-side management using deep learning for smart charging of electric vehicles. IEEE Transactions on Smart Grid 10, 3 (2018), 2683–2691
2018
-
[45]
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. In Advances in Neural Information Processing Systems 30: Annual Conference on Neu- ral Information Processing Systems 2017,...
2017
-
[46]
Yuteng Lu, Weidi Sun, and Meng Sun. 2022. Towards mutation testing of rein- forcement learning systems. Journal of Systems Architecture 131 (2022), 102701. Curiosity-Driven Testing for Sequential Decision-Making Process ICSE ’24, April 14–20, 2024, Lisbon, Portugal
2022
-
[47]
Lei Ma, Felix Juefei-Xu, Minhui Xue, Bo Li, Li Li, Yang Liu, and Jianjun Zhao
-
[48]
Lei Ma, Felix Juefei-Xu, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Chun- yang Chen, Ting Su, Li Li, Yang Liu, et al. 2018. Deepgauge: Multi-granularity testing criteria for deep learning systems. In Proceedings of the 33rd ACM/IEEE international conference on automated soft...
2018
-
[49]
Mike Marston and Gabe Baca. 2015. ACAS-Xu initial self-separation flight tests . Technical Report
2015
-
[50]
Ruijie Meng, Martin Mirchev, Marcel Böhme, and Abhik Roychoudhury. 2024. Large language model guided protocol fuzzing. In Proceedings of the 31st Annual Network and Distributed System Security Symposium (NDSS)
2024
-
[51]
Rusu, Joel Veness, Marc G
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, ...
2015 doi
-
[52]
Grégoire Montavon, Wojciech Samek, and Klaus-Robert Müller. 2018. Methods for interpreting and understanding deep neural networks. Digital signal processing 73 (2018), 1–15
2018
-
[53]
Jean-Baptiste Mouret and Jeff Clune. 2015. Illuminating search spaces by mapping elites. arXiv preprint arXiv:1504.04909 (2015)
2015 arXiv
-
[54]
Nadim Nachar et al. 2008. The Mann-Whitney U: A test for assessing whether two independent samples come from the same distribution. Tutorials in quantitative Methods for Psychology 4, 1 (2008), 13–20
2008
-
[55]
Roberto Natella. 2022. StateAFL: Greybox fuzzing for stateful network servers. Empir. Softw. Eng. 27, 7 (2022), 191. https://doi.org/10.1007/S10664-022-10233-3
2022 doi
-
[56]
Qi Pang, Yuanyuan Yuan, and Shuai Wang. 2022. MDPFuzz: testing models solving Markov decision processes. In ISSTA ’22: 31st ACM SIGSOFT International Symposium on Software Testing and Analysis, Virtual Event, South Korea, July 18 - 22, 2022 , Sukyoung Ryu and Yannis Smaragdaki...
2022
-
[57]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Des- maison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, L...
2019
-
[58]
Ketan Patil and Aditya Kanade. 2018. Greybox fuzzing as a contextual bandits problem. arXiv preprint arXiv:1806.03806 (2018)
2018 arXiv
-
[59]
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana. 2017. Deepxplore: Au- tomated whitebox testing of deep learning systems. In proceedings of the 26th Symposium on Operating Systems Principles . 1–18
2017
-
[60]
Van-Thuan Pham, Marcel Böhme, and Abhik Roychoudhury. 2020. AFLNET: A Greybox Fuzzer for Network Protocols. In 13th IEEE International Conference on Software Testing, Validation and Verification, ICST 2020, Porto, Portugal, October 24-28, 2020. IEEE, 460–465. https://doi.org/1...
2020
-
[61]
Martin L Puterman. 1990. Markov decision processes. Handbooks in operations research and management science 2 (1990), 331–434
1990
-
[62]
Jürgen Schmidhuber. 2015. Deep learning in neural networks: An overview. Neural networks 61 (2015), 85–117
2015
-
[63]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov
-
[64]
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Pan- neershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicra...
2016 doi
-
[65]
Marcel Steinmetz, Daniel Fišer, Hasan Ferit Eniser, Patrick Ferber, Timo P Gros, Philippe Heim, Daniel Höller, Xandra Schuler, Valentin Wüstholz, Maria Chris- takis, et al. 2022. Debugging a Policy: Automatic Action-Policy Testing in AI Planning. In Proceedings of the Internat...
2022
-
[66]
Andrea Stocco, Brian Pulfer, and Paolo Tonella. 2022. Mind the gap! a study on the transferability of virtual vs physical-world testing of autonomous driving systems. IEEE Transactions on Software Engineering (2022)
2022
-
[67]
Ari Takanen, Jared D Demott, Charles Miller, and Atte Kettunen. 2018. Fuzzing for software security testing and quality assurance . Artech House
2018
-
[68]
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel. 2017. #Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning. In Advances in Neu- ral Information Processing Systems 30: Annual...
2017
-
[69]
Aichernig, and Bettina Könighofer
Martin Tappler, Filip Cano Córdoba, Bernhard K. Aichernig, and Bettina Könighofer. 2022. Search-Based Testing of Reinforcement Learning. In Pro- ceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI 2022, Vienna, Austria, 23-29 July 2022...
2022 doi
-
[70]
CARLA Team. 2021. CARLA Challenge. https://carlachallenge.org/. Accessed on May 6, 2023
2021
-
[71]
OpenAI Team. 2021. rlbaselines3-zoo. https://github.com/DLR-RM/rlbaselines3- zoo. Accessed on May 6, 2023
2021
-
[72]
Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray. 2018. Deeptest: Automated testing of deep-neural-network-driven autonomous cars. In Proceedings of the 40th international conference on software engineering . 303–314
2018
-
[73]
Miller Trujillo, Mario Linares-Vásquez, Camilo Escobar-Velásquez, Ivana Dus- paric, and Nicolás Cardozo. 2020. Does Neuron Coverage Matter for Deep Reinforcement Learning?: A Preliminary Study. In ICSE ’20: 42nd International Conference on Software Engineering, Workshops, Seou...
2020
-
[74]
Shiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang, and Suman Jana. 2018. Formal Security Analysis of Neural Networks using Symbolic Intervals. In 27th USENIX Security Symposium, USENIX Security 2018, Baltimore, MD, USA, August 15-17, 2018, William Enck and Adrienne Porter...
2018
-
[75]
Cathy Wu, Aboudy Kreidieh, Kanaad Parvate, Eugene Vinitsky, and Alexandre M Bayen. 2017. Flow: Architecture and benchmarking for reinforcement learning in traffic control. arXiv preprint arXiv:1710.05465 10 (2017)
2017 arXiv
-
[76]
Xiaofei Xie, Lei Ma, Felix Juefei-Xu, Minhui Xue, Hongxu Chen, Yang Liu, Jianjun Zhao, Bo Li, Jianxiong Yin, and Simon See. 2019. Deephunter: a coverage-guided fuzz testing framework for deep neural networks. In Proceedings of the 28th ACM SIGSOFT International Symposium on So...
2019
-
[77]
Wen Xu, Hyungon Moon, Sanidhya Kashyap, Po-Ning Tseng, and Taesoo Kim
-
[78]
Zhou Yang, Jieke Shi, Muhammad Hilmi Asyrofi, and David Lo. 2022. Revisit- ing Neuron Coverage Metrics and Quality of Deep Neural Networks. In 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). 408–419. https://doi.org/10.1109/SANER53...
2022
- [79]
-
[80]
Zhou Yang, Jieke Shi, Junda He, and David Lo. 2022. Natural Attack for Pre- Trained Models of Code. In Proceedings of the 44th International Conference on Software Engineering (Pittsburgh, Pennsylvania) (ICSE ’22). Association for Com- puting Machinery, New York, NY, USA, 1482...
2022
-
[81]
Deheng Ye, Guibin Chen, Wen Zhang, Sheng Chen, Bo Yuan, Bo Liu, Jia Chen, Zhao Liu, Fuhao Qiu, Hongsheng Yu, et al. 2020. Towards playing full moba games with deep reinforcement learning. Advances in Neural Information Processing Systems 33 (2020), 621–632
2020
-
[82]
Mengshi Zhang, Yuqun Zhang, Lingming Zhang, Cong Liu, and Sarfraz Khurshid
-
[83]
Zhenya Zhang, Deyun Lyu, Paolo Arcaini, Lei Ma, Ichiro Hasuo, and Jianjun Zhao
-
[84]
In 2019 IEEE Symposium on Security and Privacy (SP)
Fuzzing file systems via two-dimensional input space exploration. In 2019 IEEE Symposium on Security and Privacy (SP) . IEEE, 818–834
2019
-
[85]
Xin Zhou, Kisub Kim, Bowen Xu, DongGyun Han, Junda He, and David Lo. 2023. Generation-based Code Review Automation: How Far Are We? arXiv preprint arXiv:2303.07221 (2023)
2023 arXiv
-
[86]
Li ZHUO, Xiongfei WU, Derui ZHU, Mingfei CHENG, Siyuan CHEN, Fuyuan ZHANG, Xiaofei XIE, Lei MA, and Jianjun ZHAO. 2023. Generative model-based testing on decision-making policies. ASE
2023
-
[87]
Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, and Paolo Tonella. 2021. DeepHyperion: Exploring the Feature Space of Deep Learning-Based Systems through Illumination Search. InProceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis (Vi...
2021
-
[88]
Briand, Mojtaba Bagherzadeh, and Ramesh S
Amirhossein Zolfagharian, Manel Abdellatif, Lionel C. Briand, Mojtaba Bagherzadeh, and Ramesh S. 2023. A Search-Based Testing Approach for Deep Re- inforcement Learning Agents. IEEE Transactions on Software Engineering (2023), 1–22. https://doi.org/10.1109/TSE.2023.3269804
2023
-
[92]
IEEE Trans
FalsifAI: Falsification of AI-Enabled Hybrid Control Systems Guided by Time-Aware Coverage Criteria. IEEE Trans. Software Eng. 49, 4 (2023), 1842–1859. https://doi.org/10.1109/TSE.2022.3194640
2023
-
[93]
Ziyuan Zhong, Gail Kaiser, and Baishakhi Ray. 2022. Neural network guided evolutionary fuzzing for finding traffic violations of autonomous vehicles. IEEE Transactions on Software Engineering (2022)
2022
-
[2017]
arXiv preprint arXiv:1707.06347 (2017)
Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017)
2017 arXiv
-
[2018]
In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering (Montpellier, France) (ASE ’18)
DeepRoad: GAN-Based Metamorphic Testing and Input Validation Frame- work for Autonomous Driving Systems. In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering (Montpellier, France) (ASE ’18). Association for Computing Machinery, New Yor...
-
[2019]
In 2019 IEEE 26th International Conference on Software Analysis, Evolution and Reengineering (SANER)
Deepct: Tomographic combinatorial testing for deep learning systems. In 2019 IEEE 26th International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 614–618
2019
-
[2020]
In International Conference on Machine Learning
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics. In International Conference on Machine Learning . PMLR, 5556–5566
-
[2021]
IEEE Robotics and Automation Letters 6, 3 (2021), 4257–4264
Super-human performance in gran turismo sport using deep reinforcement learning. IEEE Robotics and Automation Letters 6, 3 (2021), 4257–4264
2021
-
[2022]
In Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension (Virtual Event) (ICPC ’22)
PTM4Tag: Sharpening Tag Recommendation of Stack Overflow Posts with Pre-Trained Models. In Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension (Virtual Event) (ICPC ’22). Association for Computing Machinery, New York, NY, USA, 1–11. https://doi.o...
-
[2023]
ACM Trans
NSFuzz: Towards Efficient and State-Aware Network Service Fuzzing - RCR Report. ACM Trans. Softw. Eng. Methodol. 32, 6 (2023), 161:1–161:8. https: //doi.org/10.1145/3580599
2023 doi
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.