Pith. sign in

REVIEW 4 major objections 6 minor 97 references

Curiosity-Driven Testing for Sequential Decision-Making Process

T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper claims that a curiosity signal—neural prediction error—lets black-box fuzzing find more, and more varied, crashes in deep-learning decision-makers.

desk verdict Solid RND-based fuzzing for SDMs with real efficiency gains, but the unqualified superiority claim doesn't survive contact with Coop Navi; the repair experiment is also circular. read the letter →

arxiv 2509.02025 v1 pith:MFRF2WM7 submitted 2025-09-02 cs.SE

classification cs.SE
keywords fuzztestingsequentialdecision-makingdeepreinforcementlearningcuriosity-drivenexplorationrandomnetworkdistillationcrash-triggeringscenariosautonomousdrivingblack-box
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Sequential decision-makers—deep-learning policies that drive cars, pilot aircraft, or control robots—are trained to maximize performance and can fail in rare states. The paper claims that a black-box fuzzer can find those failures more often and in greater variety if it treats 'unpredictability to a small neural network' as the signal of novelty, an idea borrowed from reinforcement-learning exploration. The proposed CureFuzz mutates legal starting states, runs the policy to collect state sequences, and scores each seed by a combination of prediction-error curiosity, low cumulative reward, and behavioral sensitivity to perturbation; seeds with higher scores are mutated preferentially. Evaluated over 12-hour runs on CARLA driving, ACAS Xu collision avoidance, Coop Navi, and BipedalWalker, CureFuzz reports more total crash-triggering scenarios than the MDPFuzz baseline in all five settings, with a 422% gain for BipedalWalker, and more distinct crash types at every discretization level except Coop Navi, where a generative-model baseline finds more crashes near the boundary. The paper also reports that fine-tuning an ACAS Xu network on CureFuzz-discovered crashes reduces re-found faults by 73%. If correct, the contribution is a computationally cheap, model-agnostic diversity signal for safety testing of deep-learning decision-makers.

What carries the argument

Two interacting mechanisms. First, a curiosity module: a fixed randomly initialized target network and a trainable predictor network with identical MLP architecture; the mean squared error between their outputs on a state sequence is the intrinsic reward, updated online as the fuzzer sees new states. High prediction error means the scenario is unfamiliar, and seeds that produce it are favored. Second, a multi-objective seed energy score: E(s) = e^(-alpha*r) + e^(beta*i) + gamma*r', where r is cumulative reward (low reward raises energy), i is the intrinsic curiosity reward, and r' is robustness, the Euclidean distance between final states of the original and slightly perturbed runs. Energy-p

What would settle it

Replace the trained predictor network in the curiosity module with a second frozen random network, so the 'curiosity' score is pure noise, while keeping seed selection, mutation, budget, and environments unchanged. If CureFuzz still finds as many distinct crash-triggering scenarios as reported, the prediction-error signal is not what drives the gain; if performance falls back to MDPFuzz levels, the causal role of curiosity is confirmed.

Watch

Extended reading notes

Core claim

The paper's central claim is that curiosity—measured as the prediction error between a frozen random target network and a learned predictor network trained to mimic it—is a practical novelty measure for fuzzing sequential decision-makers. CureFuzz is a black-box fuzzer built on this signal: it randomly seeds an initial corpus of legal environment states, mutates them with small perturbations, executes the SDM to obtain state sequences, and assigns each candidate an energy score combining the mean prediction-error intrinsic reward, an exponential low-cumulative-reward term, and a robustness term that measures how far the final state moves under perturbation. Seeds are selected proportionally

Load-bearing premise

CureFuzz assumes the tester can sample and mutate valid starting states inside a known legitimate state space, with a simulator to reject illegal states; for a closed SDM whose initial states cannot be controlled or validated, the approach does not apply.

Editorial extensions

If this is right

  • Crash discovery without white-box access: any SDM that can be exercised by setting a legal initial state and observing state sequences can be fuzzed with CureFuzz, even if its weights and gradients are hidden.
  • A 12-hour CureFuzz run detects more distinct failure modes, not just more crashes: distinct crash-type counts are up to 200% higher than MDPFuzz at 100-bin discretization, reducing duplicate debugging effort.
  • Per-iteration novelty analysis costs about 0.005–0.011 seconds versus 0.033–0.985 seconds for MDPFuzz, so the diversity signal scales to high-dimensional continuous state spaces where density-based novelty is expensive.
  • Discovered scenarios can be used to repair SDMs: after fine-tuning ACAS Xu on CureFuzz crashes, the number of faults found on re-testing drops by 73%.
  • CureFuzz works across policy types—DNN, DRL, MARL, and IL—suggesting the method depends on the environment interface rather than on a particular learning algorithm.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The Coop Navi result hints at a boundary-vs-interior trade-off: the generative baseline found more crashes by concentrating agents near the state-space boundary, while CureFuzz spread crashes across the interior. A hybrid that adds boundary-aware robustness to curiosity scoring might dominate both methods—a testable direction the paper does not pursue.
  • Because the novelty signal is an online-trained predictor, CureFuzz's curriculum is order-dependent; early random seeds shape what counts as 'curious' later. Scheduling the corpus and updating the predictor on non-crash sequences could change the diversity of found crashes as much as the energy weights do.
  • The 73% repair result is shown for one DNN policy; an untested extension is to feed CureFuzz crashes back into reinforcement-learning training as negative demonstrations or safety constraints, which would test whether the same diversity signal improves policies during learning, not just after fine-tuning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes CureFuzz, a curiosity-driven black-box fuzz testing approach for sequential decision-makers (SDMs). CureFuzz uses prediction error between a fixed random target network and a learnable predictor network as an intrinsic novelty signal, combined with a multi-objective seed energy function (cumulative reward, robustness, intrinsic reward) to select seeds for mutation. The approach is evaluated on five SDMs across CARLA (RL and IL), ACAS Xu, Coop Navi (MARL), and BipedalWalker, against MDPFuzz and a generative-model baseline (G-Model). The paper reports that CureFuzz finds more total crashes and more distinct crash types than the baselines in most settings, with lower per-iteration analysis time than MDPFuzz, and that crashes found by CureFuzz can be used to fine-tune ACAS Xu, reducing detected faults by 73%.

Significance. If the results hold, CureFuzz would be a useful addition to the SDM testing toolbox: its RND-style curiosity signal is computationally cheap, the comparison against two external baselines is meaningful, and the authors report five repetitions, confidence intervals, Mann-Whitney U tests, and effect sizes. The Coop Navi exception is disclosed in the text, and the ablation study (RQ2) supports the usefulness of the curiosity mechanism. The replication package and reuse of official baseline implementations are also positives. However, the headline claim of universal superiority over the state of the art is contradicted by the paper's own Table 2/Table 3 for Coop Navi, and the RQ3 repair result is evaluated on the same crash distribution used for repair, so the practical-utility claim needs additional support. The contribution is plausible but currently over-stated.

major comments (4)
  1. [Abstract; §5 RQ1] The abstract claims CureFuzz 'outperforms the state-of-the-art method by a substantial margin in the total number of faults and distinct types of crash-triggering scenarios' without qualification. This is contradicted by Table 2 (Coop Navi: G-Model 185.4 vs CureFuzz 85) and Table 3 (100 bins: G-Model 184.6 vs CureFuzz 85.3). The RQ1 answer ('CureFuzz statistically significantly and substantially outperforms both baselines in these cases') likewise overstates the evidence, since the text later acknowledges G-Model performs better on Coop Navi. The abstract, RQ1 conclusion, and general claims should be revised to state explicitly that CureFuzz outperforms MDPFuzz on all five SDMs and outperforms G-Model on four of five, with Coop Navi as a notable exception.
  2. [§5 RQ3] The repair experiment is internally circular. The same CureFuzz-discovered crash scenarios are used to fine-tune ACAS Xu, and then CureFuzz is re-run on the repaired model to count remaining faults. A 73% reduction in this setting largely measures how well the model memorizes or fits the specific crash distribution that CureFuzz already explored; it does not demonstrate generalization to other crash-triggering scenarios. To support the claim that CureFuzz's findings 'can repair SDMs', the evaluation should include held-out crashes (e.g., crashes found by MDPFuzz or G-Model, or crashes from an independent seed/run) and report the reduction on those scenarios. Without this, the RQ3 conclusion is over-stated.
  3. [§3.4, Eq. (2); §4.3] The paper does not report the values of α, β, γ in Eq. (2), the isInteresting threshold (Algorithm 2, line 23), the curiosity network architecture and training hyperparameters, or the seed mutation magnitude (only stated as reused from Pang et al.). These parameters are load-bearing for the reported improvements, since fuzzing performance is often sensitive to such choices. Please provide a configuration table and, ideally, a sensitivity analysis or at least a statement of the ranges explored. Without this, the empirical comparison is not fully reproducible.
  4. [§3.4 and §4.2] The 'black-box' framing is weakened by the assumption in Section 3.4 that 'we are aware of the legitimate state space of the environment' and by the reliance on environment-specific validity oracles (e.g., 'We use the CARLA simulator itself to check for the validity of the mutated state' in Section 4.2). In a genuinely black-box or closed deployment where initial states cannot be freely sampled and legality cannot be checked, CureFuzz is not directly applicable. This limitation should be stated in the assumptions or in the threat-to-validity section, and the abstract/introduction should not imply that the method applies to arbitrary black-box SDMs without such state-space access.
minor comments (6)
  1. [Algorithm 1] The function initCuriosity returns phi_target twice; the second return value should be phi_pred.
  2. [Algorithm 2] Typos: 'runnign time' in Lines 6 and 16 should be 'running time'.
  3. [§1] 'remains a to be an ongoing challenge' is ungrammatical; should read 'remains an ongoing challenge'.
  4. [§5, Table 2 text] 'CureFuzz also archives an improvement of 94.2%' should be 'achieves'.
  5. [§8.4] The sentence 'Gong et al. [?]' has a missing reference; either cite the work or remove the placeholder.
  6. [§4.3] The text says curiosity module is implemented with 'Pytorch Library' and uses ReLU; consider adding the predictor's learning rate, batch size, and number of training steps per state sequence for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: core claims are evaluated against external baselines and the curiosity mechanism is an external RND adaptation.

full rationale

The paper's central derivation chain is not circular. CureFuzz's curiosity mechanism is explicitly based on Random Network Distillation (Burda et al., [11]), an external prior work, and the fuzzing loop is a standard mutation-based process with a novelty-guided selection function. The performance claims are established by comparing against two external baselines (MDPFuzz and G-Model) using their official replication packages, so the central comparison is not self-referential. The evaluation metrics (total crashes and distinct crash types) are defined from the environment state space and are not constructed from CureFuzz's internal parameters. The RQ3 repair experiment uses crashes found by CureFuzz to fine-tune the model and then re-tests with CureFuzz again; this is an empirical measurement of improvement, not a definitional equivalence. Although the repair and evaluation share the same fuzzer, the fuzzer is re-initialized with a fresh predictor network and random seeds, so the 73% reduction is not forced by construction. No parameter is fitted to the target result and no 'prediction' is renamed from an input. Self-citations appear only in related-work discussions, not as load-bearing justifications for the approach. The only notable issue is an overclaim in the abstract regarding universal superiority, since G-Model outperforms CureFuzz on Coop Navi (MARL), but this is a correctness/consistency concern, not circularity.

Assumptions & free parameters 7 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the RND curiosity heuristic and on the MDP-plus-controllable-state assumptions. The free parameters that most affect the results are the three coefficients in the seed energy function and the threshold for admitting seeds to the corpus; none are reported. The whole empirical comparison is thus under-specified at the level needed for exact replication. No new physical or conceptual entities are postulated beyond the curiosity module, which is a standard RND pair of networks.

free parameters (7)
  • alpha (a) in Eq. (2)
    Exponent coefficient for cumulative reward in seed energy; value never reported or justified.
  • beta (b) in Eq. (2)
    Exponent coefficient for intrinsic reward in seed energy; value never reported or justified.
  • gamma (g) in Eq. (2)
    Coefficient for robustness term in seed energy; value never reported or justified.
  • isInteresting threshold
    Seed evaluation adds a seed to the corpus only if intrinsic reward surpasses a 'pre-set threshold' (Section 3.4); threshold value and selection method are absent.
  • curiosity network hyperparameters
    MLP architecture depth/width, learning rate, L2 regularization strength, and optimizer for the predictor network are not reported; only ReLU activation is stated (Section 4.3).
  • robustness perturbation size
    Algorithm 3 perturbs the seed by 'a tiny random perturbation'; the magnitude is not specified.
  • seed mutation magnitude
    The random perturbation for mutation is reused from Pang et al. [56], so its value is inherited and not reported here.
assumptions (5)
  • domain assumption Environments satisfy the Markov Property (MDP framework).
    Section 3.1: 'We focus on environments where the transition dynamics satisfy the Markov Property'. This restricts the approach to MDP-based SDPs.
  • domain assumption The legitimate state space of each environment is known, sampleable, and controllable by the tester, with an oracle for legal mutations.
    Section 3.4: 'we are aware of the legitimate state space of the environment'; Section 4.2 uses the CARLA simulator itself to check validity. This limits the claimed black-box applicability.
  • domain assumption The SDM policy is fixed during fuzzing and is queried only through environment interaction.
    Section 3.1: 'the SDM's policy remains fixed and will not be updated during the fuzzing process'.
  • domain assumption Random Network Distillation prediction error is a valid novelty proxy for diverse crash discovery.
    Section 3.3 borrows RND from Burda et al. [11] as the curiosity signal; its suitability for testing is demonstrated empirically via RQ2, not derived.
  • domain assumption Catastrophic failure (crash) is the testing oracle, defined per environment by the authors.
    Section 3.1 defines crash as catastrophic failure, with environment-specific definitions; results depend on these hand-set oracles.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Curiosity-Driven Testing for Sequential Decision-Making Process." pith.science (2026). https://pith.science/paper/MFRF2WM7

@misc{pith2026250902025,
  author       = {Pith},
  title        = {Pith review of: Curiosity-Driven Testing for Sequential Decision-Making Process},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MFRF2WM7}},
  note         = {Machine review of arXiv:2509.02025}
}
read the original abstract

Sequential decision-making processes (SDPs) are fundamental for complex real-world challenges, such as autonomous driving, robotic control, and traffic management. While recent advances in Deep Learning (DL) have led to mature solutions for solving these complex problems, SDMs remain vulnerable to learning unsafe behaviors, posing significant risks in safety-critical applications. However, developing a testing framework for SDMs that can identify a diverse set of crash-triggering scenarios remains an open challenge. To address this, we propose CureFuzz, a novel curiosity-driven black-box fuzz testing approach for SDMs. CureFuzz proposes a curiosity mechanism that allows a fuzzer to effectively explore novel and diverse scenarios, leading to improved detection of crashtriggering scenarios. Additionally, we introduce a multi-objective seed selection technique to balance the exploration of novel scenarios and the generation of crash-triggering scenarios, thereby optimizing the fuzzing process. We evaluate CureFuzz on various SDMs and experimental results demonstrate that CureFuzz outperforms the state-of-the-art method by a substantial margin in the total number of faults and distinct types of crash-triggering scenarios. We also demonstrate that the crash-triggering scenarios found by CureFuzz can repair SDMs, highlighting CureFuzz as a valuable tool for testing SDMs and optimizing their performance.

Figures

Figures reproduced from arXiv: 2509.02025 by the authors.

Figure 1
Figure 1. Illustration of the overall workflow of the proposed ap [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Visualization of Agent Positions in Crash-Triggering Sce [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Results for RQ2: Comparison between CureFuzz with and without the curiosity mechanism. The x-axis represents the time passed in hours and the y-axis represents the number of found crashes. The complete CureFuzz is represented with green color and triangles, the ablated CureFuzz is represented with blue color and circles [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

97 extracted references · 61 canonical work pages

  1. [1]

    Abien Fred Agarap. 2018. Deep learning using rectified linear units (relu). arXiv preprint arXiv:1803.08375 (2018)

  2. [2]

    Briand, Ramesh S, and Mojtaba Bagherzadeh

    Zohreh Aghababaeyan, Manel Abdellatif, Lionel C. Briand, Ramesh S, and Mojtaba Bagherzadeh. 2023. Black-Box Testing of Deep Neural Networks through Test ICSE ’24, April 14–20, 2024, Lisbon, Portugal Junda He, Zhou Yang, Jieke Shi, Chengran Yang, Kisub Kim, Bowen Xu, Xin Zhou, and David Lo Case Diversity. IEEE Trans. Software Eng. 49, 5 (2023), 3182–3204. ...

  3. [3]

    Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, and Anil Anthony Bharath. 2017. Deep reinforcement learning: A brief survey.IEEE Signal Processing Magazine 34, 6 (2017), 26–38

  4. [4]

    Muhammad Hilmi Asyrofi, Zhou Yang, and David Lo. 2021. CrossASR++: A Modular Differential Testing Framework for Automatic Speech Recognition. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Athens, Greece) (ESEC/FSE 2021). Association for Computing Machinery...

  5. [5]

    Muhammad Hilmi Asyrofi, Zhou Yang, Imam Nur Bani Yusuf, Hong Jin Kang, Ferdian Thung, and David Lo. 2021. Biasfinder: Metamorphic test generation to uncover bias for sentiment analysis systems. IEEE Transactions on Software Engineering 48, 12 (2021), 5087–5101

  6. [6]

    Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Sax- ton, and Remi Munos. 2016. Unifying count-based exploration and intrinsic motivation. Advances in neural information processing systems 29 (2016)

  7. [7]

    Robert Binder. 2000. Testing object-oriented systems: models, patterns, and tools . Addison-Wesley Professional

  8. [8]

    Marcel Böhme. 2018. STADS: Software testing as species discovery. ACM Trans- actions on Software Engineering and Methodology (TOSEM) 27, 2 (2018), 1–52

Show all 97 references
  1. [9]

    Marcel Böhme, Danushka Liyanage, and Valentin Wüstholz. 2021. Estimating residual risk in greybox fuzzing. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 230–241

  2. [10]

    Marcel Böhme, Valentin J. M. Manès, and Sang Kil Cha. 2020. Boosting fuzzer efficiency: an information theoretic perspective. In ESEC/FSE ’20: 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Virtual Event, USA, ...

  3. [11]

    Storkey, and Oleg Klimov

    Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov. 2019. Explo- ration by random network distillation. In7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net. https://openreview.net/forum?id=H1lJJnR5Ym

  4. [12]

    Marcel Böhme, Van-Thuan Pham, and Abhik Roychoudhury. 2019. Coverage- Based Greybox Fuzzing as Markov Chain. IEEE Transactions on Software Engi- neering 45, 5 (2019), 489–506. https://doi.org/10.1109/TSE.2017.2785841

  5. [13]

    Emanuela G Cartaxo, Patrícia DL Machado, and Francisco G Oliveira Neto. 2011. On the use of a similarity function for test case selection in the context of model- based testing. Software Testing, Verification and Reliability 21, 2 (2011), 75–100

  6. [14]

    Sang Kil Cha, Maverick Woo, and David Brumley. 2015. Program-adaptive mutational fuzzing. In 2015 IEEE Symposium on Security and Privacy . IEEE, 725– 741

  7. [15]

    Dian Chen, Brady Zhou, Vladlen Koltun, and Philipp Krähenbühl. 2020. Learning by Cheating. In Proceedings of the Conference on Robot Learning (Proceedings of Machine Learning Research, Vol. 100), Leslie Pack Kaelbling, Danica Kragic, and Komei Sugiura (Eds.). PMLR, 66–75. http...

  8. [16]

    López, and Vladlen Koltun

    Alexey Dosovitskiy, Germán Ros, Felipe Codevilla, Antonio M. López, and Vladlen Koltun. 2017. CARLA: An Open Urban Driving Simulator. In1st Annual Conference on Robot Learning, CoRL 2017, Mountain View, California, USA, November 13-15, 2017, Proceedings (Proceedings of Machine...

  9. [17]

    Maxim Egorov. 2016. Multi-agent deep reinforcement learning. CS231n: convolu- tional neural networks for visual recognition (2016), 1–8

  10. [18]

    Jan Eisenhut, Alvaro Torralba, Maria Christakis, and Jörg Hoffmann. 2023. Auto- matic Metamorphic Test Oracles for Action-Policy Testing. (2023)

  11. [19]

    Hasan Ferit Eniser, Timo P Gros, Valentin Wüstholz, Jörg Hoffmann, and Maria Christakis. 2022. Metamorphic relations via relaxations: An approach to ob- tain oracles for action-policy testing. In Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing a...

  12. [20]

    Florian Fuchs, Yunlong Song, Elia Kaufmann, Davide Scaramuzza, and Peter Dürr

  13. [21]

    Chen Gong, Zhou Yang, Yunpeng Bai, Junda He, Jieke Shi, Arunesh Sinha, Bowen Xu, Xinwen Hou, Guoliang Fan, and David Lo. 2022. Mind Your Data! Hiding Backdoors in Offline Reinforcement Learning Datasets. https://doi.org/10.48550/ ARXIV.2210.04688

  14. [22]

    Robert J Grissom and John J Kim. 2005. Effect sizes for research: A broad practical approach. Lawrence Erlbaum Associates Publishers

  15. [23]

    Simon Haykin. 1994. Neural networks: a comprehensive foundation . Prentice Hall PTR

  16. [24]

    Junda He, Bowen Xu, Zhou Yang, DongGyun Han, Chengran Yang, and David Lo

  17. [25]

    Junda He, Xin Zhou, Bowen Xu, Ting Zhang, Kisub Kim, Zhou Yang, Ferdian Thung, Ivana Clairine Irsan, and David Lo. 2023. Representation Learning for Stack Overflow Posts: How Far Are We? ACM Trans. Softw. Eng. Methodol. (dec 2023). https://doi.org/10.1145/3635711 Just Accepted

  18. [26]

    Hadi Hemmati, Andrea Arcuri, and Lionel C. Briand. 2013. Achieving scalable model-based testing through test case diversity. ACM Trans. Softw. Eng. Methodol. 22, 1 (2013), 6:1–6:42. https://doi.org/10.1145/2430536.2430540

  19. [27]

    Hadi Hemmati, Zhihan Fang, and Mika V Mantyla. 2015. Prioritizing manual test cases in traditional and rapid release environments. In2015 IEEE 8th international conference on software testing, verification and validation (ICST) . IEEE, 1–10

  20. [28]

    Eduard Hofer, Martina Kloos, Bernard Krzykacz-Hausmann, Jörg Peschke, and Martin Woltereck. 2002. An approximate epistemic uncertainty analysis approach in the presence of epistemic and aleatory uncertainties. Reliability Engineering & System Safety 77, 3 (2002), 229–238

  21. [29]

    Ionel-Alexandru Hosu and Traian Rebedea. 2016. Playing Atari Games with Deep Reinforcement Learning and Human Checkpoint Replay. CoRR abs/1607.05077 (2016). arXiv:1607.05077 http://arxiv.org/abs/1607.05077

  22. [30]

    Fan Hu, Shisong Qin, Zheyu Ma, Bodong Zhao, Tingting Yin, and Chao Zhang

  23. [31]

    Yuqi Huai, Sumaya Almanee, Yuntianyi Chen, Xiafa Wu, Qi Alfred Chen, and Joshua Garcia. 2023. scenoRITA: Generating Diverse, Fully Mutable, Test Scenar- ios for Autonomous Vehicle Planning. IEEE Transactions on Software Engineering 49, 10 (2023), 4656–4676. https://doi.org/10....

  24. [32]

    Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne. 2017. Imitation learning: A survey of learning methods. ACM Computing Surveys (CSUR) 50, 2 (2017), 1–35

  25. [33]

    David Isele, Reza Rahimi, Akansel Cosgun, Kaushik Subramanian, and Kikuo Fu- jimura. 2018. Navigating occluded intersections with autonomous vehicles using deep reinforcement learning. In 2018 IEEE international conference on robotics and automation (ICRA). IEEE, 2034–2039

  26. [34]

    Larson, Magnus Almgren, and Vin- cenzo Gulisano

    William Johansson, Martin Svensson, Ulf E. Larson, Magnus Almgren, and Vin- cenzo Gulisano. 2014. T-Fuzz: Model-Based Fuzzing for Robustness Testing of Telecommunication Protocols. In 2014 IEEE Seventh International Conference on Software Testing, Verification and Validation. ...

  27. [35]

    Julian, Mykel J

    Kyle D. Julian, Mykel J. Kochenderfer, and Michael P. Owen. 2018. Deep Neu- ral Network Compression for Aircraft Collision Avoidance Systems. CoRR abs/1810.04240 (2018). arXiv:1810.04240 http://arxiv.org/abs/1810.04240

  28. [36]

    George Klees, Andrew Ruef, Benji Cooper, Shiyi Wei, and Michael Hicks. 2018. Evaluating Fuzz Testing. In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security, CCS 2018, Toronto, ON, Canada, October 15-19, 2018, David Lie, Mohammad Mannan, Micha...

  29. [37]

    W Bradley Knox, Adam Bradley Setapen, and Peter Stone. 2011. Reinforcement Learning with Human Feedback in Mountain Car.. In AAAI Spring Symposium: Help Me Help You: Bridging the Gaps in Human-Agent Collaboration

  30. [38]

    Jens Kober, J Andrew Bagnell, and Jan Peters. 2013. Reinforcement learning in robotics: A survey. The International Journal of Robotics Research 32, 11 (2013), 1238–1274

  31. [39]

    Mario Köppen. 2000. The curse of dimensionality. In 5th online world conference on soft computing in industrial applications (WSC5) , Vol. 1. 4–8

  32. [40]

    Nishanth Kumar. 2020. The Past and Present of Imitation Learning: A Citation Chain Study. CoRR abs/2001.02328 (2020). arXiv:2001.02328 http://arxiv.org/abs/ 2001.02328

  33. [41]

    Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, and Dmitry Vetrov

  34. [42]

    Yuxi Li. 2017. Deep reinforcement learning: An overview. arXiv preprint arXiv:1701.07274 (2017)

  35. [43]

    David Lo. 2023. Trustworthy and Synergistic Artificial Intelligence for Software Engineering: Vision and Roadmaps. arXiv preprint arXiv:2309.04142 (2023)

  36. [44]

    Karol Lina López, Christian Gagné, and Marc-André Gardner. 2018. Demand-side management using deep learning for smart charging of electric vehicles. IEEE Transactions on Smart Grid 10, 3 (2018), 2683–2691

  37. [45]

    Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. In Advances in Neural Information Processing Systems 30: Annual Conference on Neu- ral Information Processing Systems 2017,...

  38. [46]

    Yuteng Lu, Weidi Sun, and Meng Sun. 2022. Towards mutation testing of rein- forcement learning systems. Journal of Systems Architecture 131 (2022), 102701. Curiosity-Driven Testing for Sequential Decision-Making Process ICSE ’24, April 14–20, 2024, Lisbon, Portugal

  39. [47]

    Lei Ma, Felix Juefei-Xu, Minhui Xue, Bo Li, Li Li, Yang Liu, and Jianjun Zhao

  40. [48]

    Lei Ma, Felix Juefei-Xu, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Chun- yang Chen, Ting Su, Li Li, Yang Liu, et al. 2018. Deepgauge: Multi-granularity testing criteria for deep learning systems. In Proceedings of the 33rd ACM/IEEE international conference on automated soft...

  41. [49]

    Mike Marston and Gabe Baca. 2015. ACAS-Xu initial self-separation flight tests . Technical Report

  42. [50]

    Ruijie Meng, Martin Mirchev, Marcel Böhme, and Abhik Roychoudhury. 2024. Large language model guided protocol fuzzing. In Proceedings of the 31st Annual Network and Distributed System Security Symposium (NDSS)

  43. [51]

    Rusu, Joel Veness, Marc G

    Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, ...

  44. [52]

    Grégoire Montavon, Wojciech Samek, and Klaus-Robert Müller. 2018. Methods for interpreting and understanding deep neural networks. Digital signal processing 73 (2018), 1–15

  45. [53]

    Jean-Baptiste Mouret and Jeff Clune. 2015. Illuminating search spaces by mapping elites. arXiv preprint arXiv:1504.04909 (2015)

  46. [54]

    Nadim Nachar et al. 2008. The Mann-Whitney U: A test for assessing whether two independent samples come from the same distribution. Tutorials in quantitative Methods for Psychology 4, 1 (2008), 13–20

  47. [55]

    Roberto Natella. 2022. StateAFL: Greybox fuzzing for stateful network servers. Empir. Softw. Eng. 27, 7 (2022), 191. https://doi.org/10.1007/S10664-022-10233-3

  48. [56]

    Qi Pang, Yuanyuan Yuan, and Shuai Wang. 2022. MDPFuzz: testing models solving Markov decision processes. In ISSTA ’22: 31st ACM SIGSOFT International Symposium on Software Testing and Analysis, Virtual Event, South Korea, July 18 - 22, 2022 , Sukyoung Ryu and Yannis Smaragdaki...

  49. [57]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Des- maison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, L...

  50. [58]

    Ketan Patil and Aditya Kanade. 2018. Greybox fuzzing as a contextual bandits problem. arXiv preprint arXiv:1806.03806 (2018)

  51. [59]

    Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana. 2017. Deepxplore: Au- tomated whitebox testing of deep learning systems. In proceedings of the 26th Symposium on Operating Systems Principles . 1–18

  52. [60]

    Van-Thuan Pham, Marcel Böhme, and Abhik Roychoudhury. 2020. AFLNET: A Greybox Fuzzer for Network Protocols. In 13th IEEE International Conference on Software Testing, Validation and Verification, ICST 2020, Porto, Portugal, October 24-28, 2020. IEEE, 460–465. https://doi.org/1...

  53. [61]

    Martin L Puterman. 1990. Markov decision processes. Handbooks in operations research and management science 2 (1990), 331–434

  54. [62]

    Jürgen Schmidhuber. 2015. Deep learning in neural networks: An overview. Neural networks 61 (2015), 85–117

  55. [63]

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov

  56. [64]

    David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Pan- neershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicra...

  57. [65]

    Marcel Steinmetz, Daniel Fišer, Hasan Ferit Eniser, Patrick Ferber, Timo P Gros, Philippe Heim, Daniel Höller, Xandra Schuler, Valentin Wüstholz, Maria Chris- takis, et al. 2022. Debugging a Policy: Automatic Action-Policy Testing in AI Planning. In Proceedings of the Internat...

  58. [66]

    Andrea Stocco, Brian Pulfer, and Paolo Tonella. 2022. Mind the gap! a study on the transferability of virtual vs physical-world testing of autonomous driving systems. IEEE Transactions on Software Engineering (2022)

  59. [67]

    Ari Takanen, Jared D Demott, Charles Miller, and Atte Kettunen. 2018. Fuzzing for software security testing and quality assurance . Artech House

  60. [68]

    Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel. 2017. #Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning. In Advances in Neu- ral Information Processing Systems 30: Annual...

  61. [69]

    Aichernig, and Bettina Könighofer

    Martin Tappler, Filip Cano Córdoba, Bernhard K. Aichernig, and Bettina Könighofer. 2022. Search-Based Testing of Reinforcement Learning. In Pro- ceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI 2022, Vienna, Austria, 23-29 July 2022...

  62. [70]

    CARLA Team. 2021. CARLA Challenge. https://carlachallenge.org/. Accessed on May 6, 2023

  63. [71]

    OpenAI Team. 2021. rlbaselines3-zoo. https://github.com/DLR-RM/rlbaselines3- zoo. Accessed on May 6, 2023

  64. [72]

    Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray. 2018. Deeptest: Automated testing of deep-neural-network-driven autonomous cars. In Proceedings of the 40th international conference on software engineering . 303–314

  65. [73]

    Miller Trujillo, Mario Linares-Vásquez, Camilo Escobar-Velásquez, Ivana Dus- paric, and Nicolás Cardozo. 2020. Does Neuron Coverage Matter for Deep Reinforcement Learning?: A Preliminary Study. In ICSE ’20: 42nd International Conference on Software Engineering, Workshops, Seou...

  66. [74]

    Shiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang, and Suman Jana. 2018. Formal Security Analysis of Neural Networks using Symbolic Intervals. In 27th USENIX Security Symposium, USENIX Security 2018, Baltimore, MD, USA, August 15-17, 2018, William Enck and Adrienne Porter...

  67. [75]

    Cathy Wu, Aboudy Kreidieh, Kanaad Parvate, Eugene Vinitsky, and Alexandre M Bayen. 2017. Flow: Architecture and benchmarking for reinforcement learning in traffic control. arXiv preprint arXiv:1710.05465 10 (2017)

  68. [76]

    Xiaofei Xie, Lei Ma, Felix Juefei-Xu, Minhui Xue, Hongxu Chen, Yang Liu, Jianjun Zhao, Bo Li, Jianxiong Yin, and Simon See. 2019. Deephunter: a coverage-guided fuzz testing framework for deep neural networks. In Proceedings of the 28th ACM SIGSOFT International Symposium on So...

  69. [77]

    Wen Xu, Hyungon Moon, Sanidhya Kashyap, Po-Ning Tseng, and Taesoo Kim

  70. [78]

    Zhou Yang, Jieke Shi, Muhammad Hilmi Asyrofi, and David Lo. 2022. Revisit- ing Neuron Coverage Metrics and Quality of Deep Neural Networks. In 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). 408–419. https://doi.org/10.1109/SANER53...

  71. [79]

    Zhou Yang, Jieke Shi, Muhammad Hilmi Asyrofi, Bowen Xu, Xin Zhou, DongGyun Han, and David Lo. 2023. Prioritizing Speech Test Cases. https://doi.org/10. 48550/ARXIV.2302.00330

  72. [80]

    Zhou Yang, Jieke Shi, Junda He, and David Lo. 2022. Natural Attack for Pre- Trained Models of Code. In Proceedings of the 44th International Conference on Software Engineering (Pittsburgh, Pennsylvania) (ICSE ’22). Association for Com- puting Machinery, New York, NY, USA, 1482...

  73. [81]

    Deheng Ye, Guibin Chen, Wen Zhang, Sheng Chen, Bo Yuan, Bo Liu, Jia Chen, Zhao Liu, Fuhao Qiu, Hongsheng Yu, et al. 2020. Towards playing full moba games with deep reinforcement learning. Advances in Neural Information Processing Systems 33 (2020), 621–632

  74. [82]

    Mengshi Zhang, Yuqun Zhang, Lingming Zhang, Cong Liu, and Sarfraz Khurshid

  75. [83]

    Zhenya Zhang, Deyun Lyu, Paolo Arcaini, Lei Ma, Ichiro Hasuo, and Jianjun Zhao

  76. [84]

    In 2019 IEEE Symposium on Security and Privacy (SP)

    Fuzzing file systems via two-dimensional input space exploration. In 2019 IEEE Symposium on Security and Privacy (SP) . IEEE, 818–834

  77. [85]

    Xin Zhou, Kisub Kim, Bowen Xu, DongGyun Han, Junda He, and David Lo. 2023. Generation-based Code Review Automation: How Far Are We? arXiv preprint arXiv:2303.07221 (2023)

  78. [86]

    Li ZHUO, Xiongfei WU, Derui ZHU, Mingfei CHENG, Siyuan CHEN, Fuyuan ZHANG, Xiaofei XIE, Lei MA, and Jianjun ZHAO. 2023. Generative model-based testing on decision-making policies. ASE

  79. [87]

    Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, and Paolo Tonella. 2021. DeepHyperion: Exploring the Feature Space of Deep Learning-Based Systems through Illumination Search. InProceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis (Vi...

  80. [88]

    Briand, Mojtaba Bagherzadeh, and Ramesh S

    Amirhossein Zolfagharian, Manel Abdellatif, Lionel C. Briand, Mojtaba Bagherzadeh, and Ramesh S. 2023. A Search-Based Testing Approach for Deep Re- inforcement Learning Agents. IEEE Transactions on Software Engineering (2023), 1–22. https://doi.org/10.1109/TSE.2023.3269804

  81. [92]

    IEEE Trans

    FalsifAI: Falsification of AI-Enabled Hybrid Control Systems Guided by Time-Aware Coverage Criteria. IEEE Trans. Software Eng. 49, 4 (2023), 1842–1859. https://doi.org/10.1109/TSE.2022.3194640

  82. [93]

    Ziyuan Zhong, Gail Kaiser, and Baishakhi Ray. 2022. Neural network guided evolutionary fuzzing for finding traffic violations of autonomous vehicles. IEEE Transactions on Software Engineering (2022)

  83. [2017]

    arXiv preprint arXiv:1707.06347 (2017)

    Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017)

  84. [2018]

    In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering (Montpellier, France) (ASE ’18)

    DeepRoad: GAN-Based Metamorphic Testing and Input Validation Frame- work for Autonomous Driving Systems. In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering (Montpellier, France) (ASE ’18). Association for Computing Machinery, New Yor...

  85. [2019]

    In 2019 IEEE 26th International Conference on Software Analysis, Evolution and Reengineering (SANER)

    Deepct: Tomographic combinatorial testing for deep learning systems. In 2019 IEEE 26th International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 614–618

  86. [2020]

    In International Conference on Machine Learning

    Controlling overestimation bias with truncated mixture of continuous distributional quantile critics. In International Conference on Machine Learning . PMLR, 5556–5566

  87. [2021]

    IEEE Robotics and Automation Letters 6, 3 (2021), 4257–4264

    Super-human performance in gran turismo sport using deep reinforcement learning. IEEE Robotics and Automation Letters 6, 3 (2021), 4257–4264

  88. [2022]

    In Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension (Virtual Event) (ICPC ’22)

    PTM4Tag: Sharpening Tag Recommendation of Stack Overflow Posts with Pre-Trained Models. In Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension (Virtual Event) (ICPC ’22). Association for Computing Machinery, New York, NY, USA, 1–11. https://doi.o...

  89. [2023]

    ACM Trans

    NSFuzz: Towards Efficient and State-Aware Network Service Fuzzing - RCR Report. ACM Trans. Softw. Eng. Methodol. 32, 6 (2023), 161:1–161:8. https: //doi.org/10.1145/3580599

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.