Pith. sign in

REVIEW 3 major objections 5 minor 80 references

SoK: Where to Fuzz? Assessing Target Selection Methods in Directed Fuzzing

T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper shows that choosing where to fuzz is an under-studied bottleneck, and that a simple program metric, Leopard-V, outperforms every advanced alternative at picking crash-relevant functions.

desk verdict Useful SoK-with-artifact: clean isolation of target selection as an IR task, solid benchmark, honest limitations; but the 'only viable candidate' conclusion overreaches what NDCG can support without an end-to-end fuzzing check. read the letter →

arxiv 2502.08341 v1 pith:FQLSPRJZ submitted 2025-02-12 cs.SE cs.CR

classification cs.SEcs.CR
keywords directedfuzzingtargetselectionsoftwaremetricsinformationretrievalOSS-Fuzzcrashreproductionvulnerabilitypredictionlargelanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks a question fuzzing research mostly skips: once a fuzzer is directed, who decides where to point it? It argues that target selection is an independent, under-studied component of directed fuzzing, and that it can be measured in isolation by treating selection as an information-retrieval task. Using more than 1,600 reproduced crashes from 97 real-world projects as ground truth, it compares metric-based, pattern-based, static-analysis, and machine-learning selection methods. The paper's central finding is that simple software metrics—above all the vulnerability metric Leopard-V—significantly outperform every other tested method, with the top-ranked Leopard-V function landing in the crash stack trace 13% of the time. A sympathetic reader should take away that 'where to fuzz' is itself a decisive performance factor, and that current default heuristics like recently changed code or sanitizer locations are not the best choices.

What carries the argument

The machinery is a formal analogy between target selection and information retrieval. A selection method is treated as a scoring function $\rho: \mathcal{F} \to \mathbb{R}^+$ that ranks functions in a project; the top-$k$ functions form a retrieval, and quality is measured by NDCG, a ranking-aware metric that rewards relevant functions placed high. Relevance is assigned by an oracle $\mathcal{O}$ that labels a function relevant if it appears in the reproduced crash's stack trace, with ubiquitous frames like main and sanitizer helpers zeroed out. To absorb label noise, the paper uses two matching policies: NDCG$^-$ treats every stack-trace function as relevant and NDCG$^+$ only the first retrieved one. The evaluation corpus comes from reproducing OSS-Fuzz crashes at a pinned commit and extracting post-preprocessor functions, giving 1,621 labelled crashes across 97 C/C++ projects.

What would settle it

Run a controlled head-to-head where a directed fuzzer is given targets from Leopard-V and, separately, an equal-sized set drawn from random functions on the same OSS-Fuzz projects, measuring time-to-first-crash over many repetitions; if the Leopard-V targets do not hit crashes faster than random targets, the paper's central ranking claim would not translate into fuzzing performance.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that target selection for directed fuzzing can be evaluated as ranked retrieval, and that when it is, the ranking is dominated by a classic program-metric method. The authors distill 25 directed-fuzzing papers into scoring mechanisms, then test representative methods in isolation: Leopard's complexity and vulnerability metrics, sanitizer-instrumentation counts, recency of code change, two SAST tools, three deep-learning vulnerability predictors, and random scoring. Across 1,621 reproduced OSS-Fuzz crashes, Leopard-V—a function-level score built from pointer use, control-flow nesting, and dependency-related metrics—significantly outperforms all competitors for nearly every retrieval size and crash class, and Leopard-C is close behind. The paper states that Leopard-V's highest-ranked function matches a crash-stack-trace function in 13% of cases across the whole corpus, making it 'the most natural and really only viable candidate' for directed fuzzers that need a discrete target set. It also reports that a fine-tuned code language model, CodeT5+, approaches the metric-based methods' performance, identifying learned models as the most promising direction for improving selection.

Load-bearing premise

The load-bearing premise is that a function appearing in a crash's stack trace is a good target for a directed fuzzer, so the whole ranking evaluation inherits that labelling; if the real crash-triggering code is missing from the trace, the measured rankings mislabel what good target selection means.

Editorial extensions

If this is right

  • For directed fuzzers that require a discrete target set, defaulting to Leopard-V's top-ranked functions should outperform the field's common heuristics.
  • Continuous target selection can afford to use methods that only shine at larger retrieval sizes; discrete selection should be chosen only from methods strong at small $k$.
  • Sanitizer-instrumentation counts and recently modified code, two widely used heuristics, are measurably worse than simple metrics and should not be assumed safe defaults.
  • Machine-learned vulnerability predictors, particularly CodeT5+, are close enough to metrics to justify further work on learned target selection.
  • Target-selection quality varies strongly by crash type, so a single selection method may need to be tuned per bug class.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's retrieval results, if Leopard-V's top-1 hit rate carries over to a real fuzzing loop, then simply trying the top handful of metric-ranked functions in order is a discrete selection strategy that should beat current heuristics end-to-end; the paper measures ranking, not end-to-end fuzzing, so this is my extrapolation.
  • The stack-trace oracle labels functions implicated in crashes, not functions whose targeting shortens time-to-crash; feeding Leopard-V targets into an existing directed fuzzer and measuring time-to-exposure would test the transfer.
  • The same retrieval framing could be reused for ranking suspicious functions in patch review or static-analysis triage, where the bottleneck is also 'where to look first'.
  • A practical yardstick falls out of the corpus: future target-selection methods can be compared against Leopard-V's top-1 hit rate on the released crash set.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This SoK paper presents the first systematic study of target selection methods for directed fuzzing. The authors review 25 papers from top venues, distill target selection into four characteristics (information source, scoring type, granularity, scoring mechanism), and model target selection as an information retrieval problem. They assemble a corpus of 1,621 reproducible OSS-Fuzz crashes across 97 C/C++ projects, label functions as relevant if they appear in the crash stack trace, and compare ten selection methods (Leopard-C, Leopard-V, sanitizer instrumentation, recent-code-change heuristic, Rats, Cppcheck, ReVeal, Linevul, CodeT5+, and random) using NDCG with optimistic and pessimistic matching policies. The central finding is that simple software metrics, especially Leopard-V, significantly outperform all other methods, with the strongest claim being that Leopard-V is 'the most natural and really only viable candidate' for fuzzing approaches requiring a discrete selection method. The paper also reports breakdowns by sanitizer and crash type, and provides public artifacts.

Significance. If the central claim holds, the paper would provide a practically important and somewhat surprising result: decades of research on sophisticated target selection in directed fuzzing has not surpassed simple code metrics. The work is also valuable as a reproducible benchmark: a corpus of over 1,600 real-world crashes, an explicit information-retrieval formulation, two matching policies that bracket label noise, a random baseline that anchors significance, and statistical tests across retrieval sizes. The systematic review itself (Table 1) is a useful contribution for future work. The paper is careful in many design choices, including using post-preprocessor source code to avoid preprocessor-directive confounds and reporting results broken down by sanitizer and crash type. However, the headline practical claim is currently supported only by a retrieval-based evaluation, not by end-to-end fuzzing experiments, which creates a gap between the measured NDCG performance and the paper's conclusions about fuzzing practice.

major comments (3)
  1. [Section 4.1 and Section 6] The relevance oracle O in Section 4.1 labels a function as relevant if and only if it appears in the crash's stack trace, and Section 6 'False positives' concedes that the root cause may be absent from the traceback. The cited prior work [7, 15, 25] shows that traceback-based targets beat undirected fuzzing, but it does not establish that every trace frame is a good fuzzing target or that the relative ordering of selection methods transfers to time-to-crash. Because Section 5.2 uses this retrieval measure to conclude that Leopard-V is 'the most natural and really only viable candidate' for discrete selection, the paper's central practical claim goes beyond what the measurement actually supports.
  2. [Section 5.1, Leopard-V] Leopard-V's vulnerability metrics include the number of parameters of a function and the number of parameters to its callees. These are call-graph connectivity features, and hub-like functions are over-represented in stack traces independently of whether they are good fuzzing targets. The reported 13% top-1 NDCG− result may therefore partly reward ranking trace-frequent functions rather than ranking functions whose targeting would help a directed fuzzer find the crash. I ask for a control: compare Leopard-V against a simple baseline that ranks functions by call-graph degree or by historical trace frequency, or restrict the oracle to the deepest stack frame that is not a sanitizer or main function.
  3. [Section 4.1, matching policies] The optimistic and pessimistic matching policies bound the label noise that arises from uncertainty about which trace functions are defective, but they do not bound the semantic gap identified in Section 6, where the relevant function may be missing from the trace entirely. Consequently, NDCG− and NDCG+ are not lower and upper bounds on true fuzzing utility; they are bounds only relative to the trace-membership labeling. This should be stated explicitly in Section 4.1 so that readers do not interpret the reported intervals as bracketing end-to-end fuzzing performance.
minor comments (5)
  1. [Abstract] The first two sentences of the abstract are repeated verbatim, which appears to be a formatting error.
  2. [Figure 3] The step numbering in Figure 3 lists '1, 2, 44, 43'; the last two labels should likely be '3' and '4'.
  3. [Section 5.2] There is a duplicated word in 'for for k > 1 retrieved functions'; please fix the typo.
  4. [Figure 8] The legend in Figure 8 groups methods by class using color; adding distinct line styles or markers would improve readability for color-blind readers and in grayscale printouts.
  5. [Section 3] The literature-filtering heuristic (selecting papers containing the word 'directed' at least three times) is coarse; a brief discussion of how many papers were discarded and whether any known directed-fuzzing papers were missed would strengthen confidence in the systematization.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the NDCG evaluation is a self-contained retrieval comparison against external OSS-Fuzz crash labels, with no target-selection method fit to those labels.

full rationale

The paper's central comparison treats target selection as an information-retrieval problem and evaluates pre-existing scoring functions (Leopard-C/V, sanitizer callbacks, recency, SAST tools, ML models, random) against a relevance oracle derived from externally reproduced OSS-Fuzz crash stack traces. No method's scoring parameters are fitted to these labels: Leopard metrics, Linevul, CodeT5+, ReVeal, Cppcheck, Rats, and the sanitizer/recency heuristics are all taken as-is from prior work or standard tooling, so the measured NDCG values are not predictions that reduce by construction to their inputs. The only author-self-citation ([36]) is used as methodological justification for normalizing Linevul inputs and as a post-hoc explanation for Linevul's baseline-level performance; it does not feed back into the ranking computation or into the central Leopard-V result. The stack-trace oracle concern (root causes may be absent from traces, and call-graph hubs may be over-represented) is a measurement-validity limitation that the paper itself acknowledges in Section 6 under 'False positives', but it does not make the derivation circular because the labels are external ground truth rather than a function of the scorings being compared. The comparison is self-contained against an external benchmark, so no circularity score above 0 is warranted.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters are introduced by the paper; the target selection methods are pre-existing and the evaluation settings (k, NDCG variants) are analysis choices rather than fitted constants. No new entities or mediators are postulated. The 'oracle O' is a labeling function derived from stack traces, not an invented entity.

assumptions (4)
  • domain assumption Functions appearing in a crash's stack trace are valid targets for directed fuzzing; functions not appearing in any reproduced crash's stack trace are irrelevant.
    This is the labeling oracle in Section 4.1. The paper acknowledges false positives (root cause absent from traceback) and false negatives (undiscovered bugs) in Section 6.
  • domain assumption The reproduced OSS-Fuzz crash corpus (1,621 crashes, 97 C/C++ projects) is representative of realistic defect locations for fuzzing targets.
    The corpus is built from OSS-Fuzz historical crashes, biased toward ASan/MSan/UBSan-detectable bugs and projects that OSS-Fuzz continuously fuzzes. Discussed in Section 4.2 and Section 6.
  • domain assumption Function-level granularity is sufficient for comparing target selection methods.
    The authors adopt a meet-in-the-middle approach in Section 6 to allow all methods to be compared, accepting a loss of line-level precision.
  • domain assumption NDCG@k computed on stack-trace labels measures how useful a target ranking is for directed fuzzing.
    The information retrieval framing in Section 4.1 assumes ranking quality correlates with fuzzing efficacy; the paper does not validate this end-to-end with real fuzzing campaigns.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SoK: Where to Fuzz? Assessing Target Selection Methods in Directed Fuzzing." pith.science (2026). https://pith.science/paper/FQLSPRJZ

@misc{pith2026250208341,
  author       = {Pith},
  title        = {Pith review of: SoK: Where to Fuzz? Assessing Target Selection Methods in Directed Fuzzing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FQLSPRJZ}},
  note         = {Machine review of arXiv:2502.08341}
}
read the original abstract

A common paradigm for improving fuzzing performance is to focus on selected regions of a program rather than its entirety. While previous work has largely explored how these locations can be reached, their selection, that is, the where, has received little attention so far. A common paradigm for improving fuzzing performance is to focus on selected regions of a program rather than its entirety. While previous work has largely explored how these locations can be reached, their selection, that is, the where, has received little attention so far. In this paper, we fill this gap and present the first comprehensive analysis of target selection methods for fuzzing. To this end, we examine papers from leading security and software engineering conferences, identifying prevalent methods for choosing targets. By modeling these methods as general scoring functions, we are able to compare and measure their efficacy on a corpus of more than 1,600 crashes from the OSS-Fuzz project. Our analysis provides new insights for target selection in practice: First, we find that simple software metrics significantly outperform other methods, including common heuristics used in directed fuzzing, such as recently modified code or locations with sanitizer instrumentation. Next to this, we identify language models as a promising choice for target selection. In summary, our work offers a new perspective on directed fuzzing, emphasizing the role of target selection as an orthogonal dimension to improve performance.

Figures

Figures reproduced from arXiv: 2502.08341 by the authors.

Figure 1
Figure 1. Target selection. The initial step in target selection is the extraction of code locations from the SUT (Step ❶). The granularity of this extraction depends on the specifics of the selection method and could, for example, be on a function-level or basic-block level. After extraction, the locations and (optionally) the SUT or external information (e.g., code change timestamps) are forwarded to the target selection me… view at source ↗
Figure 2
Figure 2. Overview of a retrieval. We compute a ranking using target selection method 𝜌 which assigns each function f ∈ F a rele￾vance score 𝑟ˆf . To measure the quality of target selection method, we compute the 𝑁𝐷𝐶𝐺𝑘 for a retrieval 𝐹 with cardinality 𝑘. As ground truth, we use the oracle O to assign relevance scores 𝑟f to each location f ∈ F. Requirements. Therefore, we identify two requirements for con￾ducting a thorough … view at source ↗
Figure 3
Figure 3. Dataset generation process. As basis for our analysis, we collect 1, 621 reproducible crashes. We crawl OSS-Fuzz, which yields a crashing input and a fuzzing configuration for a project (❶). To reproduce the crash, we search for a commit of the project which crashes (❷) when executed under the input from OSS-Fuzz (❸). Once we reproduce a crash, we extract functions from the project’s code and label them according to… view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: Stack trace lengths. We show the frequency of stack trace lengths from reproduced crashes. We observe the dominant peak at 7 functions per stack trace. code would thus pose a disadvantage to target selection methods working on the source code level: The source code may…
Figure 6
Figure 6. Figure 6: Number of crashes per sanitizer. We show the number of crashes in the corpus broken down by sanitizers. 5 COMPARISON OF TARGET SELECTIONS With our analysis framework at hand, we can now dive into the comparison of target selection methods. Before we get started, howeve…
Figure 7
Figure 7. Figure 7: Retrieval and evaluation process. Our crash corpus consists of various projects with functions labelled as relevant (9) or irrelevant. A target selection method is used to create a ranking of the functions with the 𝑘 highest ranks forming the retrieval. For each retrie…
Figure 8
Figure 8. Figure 8: Mean retrieval scores of the target selection methods. [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Retrieval performance of Leopard-V across different [PITH_FULL_IMAGE:figures/full_fig_p011_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

80 extracted references · 64 canonical work pages

  1. [1]

    Andrei Arusoaie, Stefan Ciobâca, Vlad Craciun, Dragos Gavrilut, and Dorel Lucanu. 2017. A Comparison of Open-Source Static Analysis Tools for Vul- nerability Detection in C/C++ Code. In 2017 19th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC) . 161–168. https://doi.org/10.1109/SYNASC.2017.00035

  2. [3]

    Cornelius Aschermann, Sergej Schumilo, Tim Blazytko, Robert Gawlik, and Thorsten Holz. 2019. REDQUEEN: Fuzzing with Input-to-State Correspondence.. In Proc. of the Network and Distributed System Security Symposium (NDSS)

  3. [4]

    Nils Bars, Moritz Schloegel, Tobias Scharnowski, Nico Schiller, and Thorsten Holz

  4. [5]

    Tim Blazytko, Cornelius Aschermann, Moritz Schlögel, Ali Abbasi, Sergej Schu- milo, Simon Wörner, and Thorsten Holz. 2019. GRIMOIRE: Synthesizing Struc- ture while Fuzzing.. In Proc. of the USENIX Security Symposium . 1985–2002

  5. [6]

    Tim Blazytko, Moritz Schlögel, Cornelius Aschermann, Ali Abbasi, Joel Frank, Simon Wörner, and Thorsten Holz. 2020. AURORA: Statistical Crash Analysis for Automated Root Cause Explanation.. In Proc. of the USENIX Security Symposium . 235–252

  6. [7]

    Marcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, and Abhik Roychoud- hury. 2017. Directed Greybox Fuzzing. InProc. of the ACM Conference on Computer and Communications Security (CCS) . 2329–2344

  7. [8]

    Marcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, and Abhik Roychoud- hury. 2017. Directed Greybox Fuzzing.. In Proc. of the ACM Conference on Com- puter and Communications Security (CCS) . 2329–2344. https://doi.org/10.1145/ 3133956.3134020

  8. [9]

    Marcel Böhme, Van-Thuan Pham, and Abhik Roychoudhury. 2016. Coverage- based Greybox Fuzzing as Markov Chain.. In Proc. of the ACM Conference on Computer and Communications Security (CCS) . 1032–1043. https://doi.org/10. 1145/2976749.2978428

Show all 80 references
  1. [10]

    Marcel Böhme, László Szekeres, and Jonathan Metzman. 2022. On the Reliability of Coverage-Based Fuzzer Benchmarking.. InProc. of the International Conference on Software Engineering. 1621–1633. https://doi.org/10.1145/3510003.3510230

  2. [11]

    Cristiano Calcagno, Dino Distefano, Jeremy Dubreil, Dominik Gabi, Pieter Hooimeijer, Martino Luca, Peter O’Hearn, Irene Papakonstantinou, Jim Purbrick, and Dulma Rodriguez. 2015. Moving Fast with Software Verification. Technical Report. Facebook Inc. (Meta)

  3. [12]

    Sadullah Canakci, Nikolay Matyunin, Kalman Graffi, Ajay Joshi, and Manuel Egele. 2022. Targetfuzz: Using darts to guide directed greybox fuzzers. In Pro- ceedings of the 2022 ACM on Asia conference on computer and communications security. 561–573

  4. [13]

    Sicong Cao, Biao He, Xiaobing Sun, Yu Ouyang, Chao Zhang, Xiaoxue Wu, Ting Su, Lili Bo, Bin Li, Chuanlei Ma, Jiajia Li, and Tao Wei. 2023. ODDFuzz: Discovering Java Deserialization Vulnerabilities via Structure-Aware Directed Greybox Fuzzing.. In Proc. of the IEEE Symposium on...

  5. [14]

    Saikat Chakraborty, Rahul Krishna, Yangruibo Ding, and Baishakhi Ray. 2022. Deep learning based vulnerability detection: Are we there yet. IEEE Transactions on Software Engineering (2022)

  6. [16]

    Jiongyi Chen, Wenrui Diao, Qingchuan Zhao, Chaoshun Zuo, Zhiqiang Lin, XiaoFeng Wang, Wing Cheong Lau, Menghan Sun, Ronghai Yang, and Kehuan Zhang. 2018. IoTFuzzer: Discovering Memory Corruptions in IoT Through App- based Fuzzing.. In Proc. of the Network and Distributed Syste...

  7. [17]

    Peng Chen and Hao Chen. 2018. Angora: Efficient Fuzzing by Principled Search.. In Proc. of the IEEE Symposium on Security and Privacy . 711–725. https://doi.org/ 10.1109/SP.2018.00046

  8. [18]

    Yizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen, and David Wagner

  9. [19]

    Yaohui Chen, Peng Li, Jun Xu, Shengjian Guo, Rundong Zhou, Yulong Zhang, Tao Wei, and Long Lu. 2020. SAVIOR: Towards Bug-Driven Hybrid Testing.. In Proc. of the IEEE Symposium on Security and Privacy . 1580–1596. https: //doi.org/10.1109/SP40000.2020.00002

  10. [20]

    Timothy Clem and Patrick Thomson. 2021. Static Analysis at GitHub: An Expe- rience Report. Queue 19, 4 (sep 2021), 42–67. https://doi.org/10.1145/3487019. 3487022

  11. [21]

    In Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses

    Diversevul: A new vulnerable source code dataset for deep learning based vulnerability detection. In Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses . 654–668

  12. [22]

    José D’Abruzzo Pereira and Marco Vieira. 2020. On the Use of Open-Source C/C++ Static Analysis Tools in Large Projects. In 2020 16th European Dependable Computing Conference (EDCC). https://doi.org/10.1109/EDCC51268.2020.00025

  13. [24]

    Ali Babar

    Roland Croft, Dominic Newlands, Ziyu Chen, and M. Ali Babar. 2021. An Empir- ical Study of Rule-Based and Learning-Based Approaches for Static Application Security Testing. In Proceedings of the 15th ACM / IEEE International Sympo- sium on Empirical Software Engineering and Me...

  14. [25]

    Zhengjie Du, Yuekang Li, Yang Liu, and Bing Mao. 2022. Windranger: A Directed Greybox Fuzzer driven by Deviation Basic Blocks.. In Proc. of the International Conference on Software Engineering . 2440–2451. https://doi.org/10.1145/3510003. 3510197

  15. [26]

    Max Eisele, Marcello Maugeri, Rachna Shriwas, Christopher Huth, and Gi- ampaolo Bella. 2022. Embedded fuzzing: a review of challenges, tools, and solutions. Cybersecurity 5, 1 (2022), 18

  16. [27]

    Xiaoning Du, Bihuan Chen, Yuekang Li, Jianmin Guo, Yaqin Zhou, Yang Liu, and Yu Jiang. 2019. Leopard: identifying vulnerable code for vulnerability assessment through program metrics.. In Proc. of the International Conference on Software Engineering. 60–71. https://doi.org/10....

  17. [28]

    Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP ...

  18. [29]

    Fortify. 2013. RATS on Google Code. https://code.google.com/archive/p/rough- auditing-tool-for-security/issues

  19. [30]

    Jiahao Fan, Yi Li, Shaohua Wang, and Tien N. Nguyen. 2020. A C/C++ Code Vulnerability Dataset with Code Changes and CVE Summaries. In Proceedings of the 17th International Conference on Mining Software Repositories (MSR) (Seoul, Republic of Korea). Association for Computing Ma...

  20. [31]

    Shuitao Gan, Chao Zhang, Xiaojun Qin, Xuwen Tu, Kang Li, Zhongyu Pei, and Zuoning Chen. 2018. CollAFL: Path Sensitive Fuzzing.. In Proc. of the IEEE Symposium on Security and Privacy . 679–696. https://doi.org/10.1109/SP.2018. 00040

  21. [32]

    István Haller, Asia Slowinska, Matthias Neugschwandtner, and Herbert Bos. 2013. Dowsing for Overflows: A Guided Fuzzer to Find Buffer Boundary Violations.. In Proc. of the USENIX Security Symposium . 49–64

  22. [33]

    Michael Fu and Chakkrit Tantithamthavorn. 2022. LineVul: A Transformer- based Line-Level Vulnerability Prediction. In International Conference on Mining Software Repositories (MSR)

  23. [35]

    Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. CoRR abs/1909.09436 (2019). arXiv:1909.09436 http://arxiv.org/ abs/1909.09436

  24. [36]

    Adrian Herrera, Hendra Gunadi, Shane Magrath, Michael Norrish, Mathias Payer, and Antony L Hosking. 2021. Seed selection for successful fuzzing. InProceedings of the 30th ACM SIGSOFT international symposium on software testing and analysis. 230–243

  25. [37]

    Zu-Ming Jiang, Jia-Ju Bai, Kangjie Lu, and Shi-Min Hu. 2022. Context-Sensitive and Directional Concurrency Fuzzing for Data-Race Detection.. In Proc. of the Network and Distributed System Security Symposium (NDSS)

  26. [38]

    Arvinder Kaur and Ruchikaa Nayyar. 2020. A Comparative Study of Static Code Analysis tools for Vulnerability Detection in C/C++ and JAVA Source Code. Procedia Computer Science 171 (2020), 2023–2029. https://doi.org/10.1016/ j.procs.2020.04.217 Third International Conference on...

  27. [39]

    Erik Imgrund, Tom Ganz, Martin Härterich, Lukas Pirch, Niklas Risse, and Konrad Rieck. 2023. Broken Promises: Measuring Confounding Effects in Learning-based Vulnerability Discovery. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec), Maura...

  28. [40]

    Tae Eun Kim, Jaeseung Choi, Seongjae Im, Kihong Heo, and Sang Kil Cha. 2024. Evaluating Directed Fuzzers: Are We Heading in the Right Direction?. In Pro- ceedings of the 32st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engine...

  29. [41]

    George Klees, Andrew Ruef, Benji Cooper, Shiyi Wei, and Michael Hicks. 2018. Evaluating Fuzz Testing.. In Proc. of the ACM Conference on Computer and Com- munications Security (CCS). 2123–2138. https://doi.org/10.1145/3243734.3243804

  30. [42]

    Tae Eun Kim, Jaeseung Choi, Kihong Heo, and Sang Kil Cha. 2023. DAFL: Directed Grey-box Fuzzing guided by Data Dependency.. In Proc. of the USENIX Security Symposium

  31. [43]

    Gwangmu Lee, Woochul Shim, and Byoungyoung Lee. 2021. Constraint-guided Directed Greybox Fuzzing.. In Proc. of the USENIX Security Symposium . 3559– 3576

  32. [44]

    Hongliang Liang, Xiaoxiao Pei, Xiaodong Jia, Wuwei Shen, and Jian Zhang. 2018. Fuzzing: State of the art. IEEE Transactions on Reliability 67, 3 (2018), 1199–1218

  33. [45]

    Johannes Krupp, Ilya Grishchenko, and Christian Rossow. 2022. AmpFuzz: Fuzzing for Amplification DDoS Vulnerabilities.. In Proc. of the USENIX Security Symposium. 1043–1060

  34. [46]

    Changhua Luo, Wei Meng, and Penghui Li. 2023. SelectFuzz: Efficient Directed Fuzzing with Selective Path Exploration.. In Proc. of the IEEE Symposium on Security and Privacy. 2693–2707. https://doi.org/10.1109/SP46215.2023.10179296

  35. [47]

    Rahma Mahmood and Qusay H. Mahmoud. 2018. Evaluation of Static Anal- ysis Tools for Finding Vulnerabilities in Java and C/C++ Source Code. CoRR abs/1805.09040 (2018). arXiv:1805.09040 http://arxiv.org/abs/1805.09040

  36. [49]

    Paul Dan Marinescu and Cristian Cadar. 2013. KATCH: high-coverage test- ing of software patches. In Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on the Foundations of Software Engineering, (ESEC/FSE)

  37. [50]

    Daniel Marjamäki. 2007. Cppcheck: Tool for static C/C++ code analysis. https: //github.com/danmar/cppcheck

  38. [51]

    Manès, HyungSeok Han, Choongwoo Han, Sang Kil Cha, Manuel Egele, Edward J

    Valentin J.M. Manès, HyungSeok Han, Choongwoo Han, Sang Kil Cha, Manuel Egele, Edward J. Schwartz, and Maverick Woo. 2021. The Art, Science, and Engineering of Fuzzing: A Survey. IEEE Transactions on Software Engineering 47, 11 (2021), 2312–2331

  39. [52]

    Marius Muench, Jan Stijohann, Frank Kargl, Aurélien Francillon, and Davide Balzarotti. 2018. What You Corrupt Is Not What You Crash: Challenges in Fuzzing Embedded Devices.. In Proc. of the Network and Distributed System Security Symposium (NDSS)

  40. [53]

    Manh-Dung Nguyen, Sébastien Bardin, Richard Bonichon, Roland Groz, and Matthieu Lemerre. 2020. Binary-level directed fuzzing for {use-after-free} vulnerabilities. In 23rd International Symposium on Research in Attacks, Intrusions and Defenses (RAID 2020) . 47–62

  41. [54]

    Ruijie Meng, Zhen Dong, Jialin Li, Ivan Beschastnikh, and Abhik Roychoudhury

  42. [55]

    Jiaqi Peng, Feng Li, Bingchang Liu, Lili Xu, Binghong Liu, Kai Chen, and Wei Huo. 2019. 1dvul: Discovering 1-day vulnerabilities through binary patches. In 2019 49th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE, 605–616

  43. [56]

    Sanjay Rawat, Vivek Jain, Ashish Kumar, Lucian Cojocar, Cristiano Giuffrida, and Herbert Bos. 2017. VUzzer: Application-aware Evolutionary Fuzzing.. In Proc. of the Network and Distributed System Security Symposium (NDSS)

  44. [57]

    Gary J Saavedra, Kathryn N Rodhouse, Daniel M Dunlavy, and Philip W Kegelmeyer. 2019. A Review of Machine Learning Applications in Fuzzing. arXiv:1906.11133 [cs.CR]

  45. [58]

    Lin Padgham, Young Lee, Shazia Sadiq, Michael Winikoff, Alan Fekete, Stephen MacDonell, Dali Kaafar, and Stefanie Zollmann. [n. d.]. CORE Rankings. https: //www.core.edu.au/conference-portal

  46. [59]

    Nico Schiller, Merlin Chlosta, Moritz Schloegel, Nils Bars, Thorsten Eisenhofer, Tobias Scharnowski, Felix Domke, Lea Schönherr, and Thorsten Holz. 2023. Drone Security and the Mysterious Case of DJI’s DroneID.. InProc. of the Network and Distributed System Security Symposium (NDSS)

  47. [60]

    Moritz Schloegel, Nils Bars, Nico Schiller, Lukas Bernhard, Tobias Scharnowski, Addison Crump, Arash Ale-Ebrahim, Nicolai Bissantz, Marius Muench, and Thorsten Holz. 2024. SoK: Prudent Evaluation Practices for Fuzzing. In 2024 IEEE Symposium on Security and Privacy (SP) . IEEE...

  48. [61]

    Lukas Seidel, Dominik Christian Maier, and Marius Muench. 2023. Forming Faster Firmware Fuzzers.. In Proc. of the USENIX Security Symposium

  49. [62]

    Caitlin Sadowski, Edward Aftandilian, Alex Eagle, Liam Miller-Cushon, and Ciera Jaspan. 2018. Lessons from Building Static Analysis Tools at Google. Communications of the ACM (CACM) 61 Issue 4 (2018), 58–66. https://dl.acm. org/citation.cfm?id=3188720

  50. [63]

    Abhishek Shah, Dongdong She, Samanway Sadhu, Krish Singal, Peter Coffman, and Suman Jana. 2022. MC2: Rigorous and Efficient Directed Greybox Fuzzing.. In Proc. of the ACM Conference on Computer and Communications Security (CCS) . 2595–2609. https://doi.org/10.1145/3548606.3560648

  51. [64]

    Dongdong She, Kexin Pei, Dave Epstein, Junfeng Yang, Baishakhi Ray, and Suman Jana. 2019. NEUZZ: Efficient Fuzzing with Neural Program Smoothing.. In Proc. of the IEEE Symposium on Security and Privacy . 803–817. https://doi.org/10.1109/ SP.2019.00052

  52. [66]

    2017.{OSS-Fuzz}-Google’s continuous fuzzing service for open source software

    Kostya Serebryany. 2017.{OSS-Fuzz}-Google’s continuous fuzzing service for open source software. (2017)

  53. [67]

    VulDetProject. 2021. VulDetProject Implementation of ReVeal. https://github. com/VulDetProject/ReVeal

  54. [68]

    Pengfei Wang, Xu Zhou, Tai Yue, Peihong Lin, Yingying Liu, and Kai Lu. [n. d.]. The progress, challenges, and perspectives of directed greybox fuzzing. Softw. Test. Verification Reliab. 34, 2 ([n. d.])

  55. [69]

    Yan Wang, Peng Jia, Luping Liu, Cheng Huang, and Zhonglin Liu. 2020. A systematic review of fuzzing based on machine learning techniques. PLOS ONE 15, 8 (2020), e0237749

  56. [70]

    Sutton, A

    M. Sutton, A. Greene, and P. Amini. 2007. Fuzzing: Brute Force Vulnerability Discovery. Addison-Wesley Professional

  57. [71]

    Bui, Junnan Li, and Steven C

    Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi D.Q. Bui, Junnan Li, and Steven C. H. Hoi. 2023. CodeT5+: Open Code Large Language Models for Code Understanding and Generation. arXiv preprint (2023)

  58. [72]

    Yining Wang, Liwei Wang, Yuanzhi Li, Di He, and Tie-Yan Liu. 2013. Conference on Learning Theory (COLT)

  59. [73]

    Mingyuan Wu, Ling Jiang, Jiahong Xiang, Yanwei Huang, Heming Cui, Lingming Zhang, and Yuqun Zhang. 2022. One Fuzzing Strategy to Rule Them All.. In Proc. of the International Conference on Software Engineering . 1634–1645. https: //doi.org/10.1145/3510003.3510174

  60. [74]

    Yanhao Wang, Xiangkun Jia, Yuwei Liu, Kyle Zeng, Tiffany Bao, Dinghao Wu, and Purui Su. 2020. Not All Coverage Measurements Are Equal: Fuzzing by Coverage Accounting for Input Prioritization.. In Proc. of the Network and Distributed System Security Symposium (NDSS)

  61. [75]

    Fabian Yamaguchi, Nico Golde, Daniel Arp, and Konrad Rieck. 2014. Modeling and Discovering Vulnerabilities with Code Property Graphs.. In Proc. of the IEEE Symposium on Security and Privacy . 590–604. https://doi.org/10.1109/SP.2014.44

  62. [76]

    Ming Yuan, Bodong Zhao, Penghui Li, Jiashuo Liang, Xinhui Han, Xiapu Luo, and Chao Zhang. 2023. DDRace: Finding Concurrency UAF Vulnerabilities in Linux Drivers with Directed Fuzzing.. In Proc. of the USENIX Security Symposium

  63. [77]

    Lei Zhang, Keke Lian, Haoyu Xiao, Zhibo Zhang, Peng Liu, Yuan Zhang, Min Yang, and Haixin Duan. 2022. Exploit the Last Straw That Breaks Android Systems.. In Proc. of the IEEE Symposium on Security and Privacy . 2230–2247. https://doi.org/10.1109/SP46214.2022.9833563

  64. [78]

    Valentin Wüstholz and Maria Christakis. 2020. Targeted greybox fuzzing with static lookahead analysis.. In Proc. of the International Conference on Software Engineering. 789–800. https://doi.org/10.1145/3377811.3380388

  65. [79]

    Han Zheng, Jiayuan Zhang, Yuhang Huang, Zezhong Ren, He Wang, Chun- jie Cao, Yuqing Zhang, Flavio Toffalini, and Mathias Payer. 2023. FISHFUZZ: Catch Deeper Bugs by Throwing Larger Nets.. In Proc. of the USENIX Security Symposium

  66. [80]

    Yaqin Zhou, Shangqing Liu, Jing Kai Siow, Xiaoning Du, and Yang Liu. 2019. Devign: Effective Vulnerability Identification by Learning Comprehensive Pro- gram Semantics via Graph Neural Networks.. In Proc. of the Conference on Neural Information Processing Systems (NeurIPS) . 1...

  67. [81]

    Xiaogang Zhu and Marcel Böhme. 2021. Regression Greybox Fuzzing.. In Proc. of the ACM Conference on Computer and Communications Security (CCS). 2169–2182. https://doi.org/10.1145/3460120.3484596

  68. [82]

    Yujian Zhang, Yaokun Liu, Jinyu Xu, and Yanhao Wang. 2023. Predecessor-aware Directed Greybox Fuzzing. In 2024 IEEE Symposium on Security and Privacy (SP) . IEEE Computer Society, 40–40

  69. [83]

    Sebastian Österlund, Kaveh Razavi, Herbert Bos, and Cristiano Giuffrida. 2020. ParmeSan: Sanitizer-guided Greybox Fuzzing.. In Proc. of the USENIX Security Symposium. 2289–2306. Assessing Target Selection Methods in Directed Fuzzing A ADDITIONAL EXPERIMENTS Alongside the impac...

  70. [86]

    Peiyuan Zong, Tao Lv, Dawei Wang, Zizhuang Deng, Ruigang Liang, and Kai Chen. 2020. FuzzGuard: Filtering out Unreachable Inputs in Directed Grey-box Fuzzing through Deep Learning.. In Proc. of the USENIX Security Symposium . 2255–2269

  71. [2022]

    Linear-time Temporal Logic guided Greybox Fuzzing.. In Proc. of the International Conference on Software Engineering . 1343–1355. https://doi.org/10. 1145/3510003.3510082

  72. [2023]

    Fuzztruction: Using Fault Injection-based Fuzzing to Leverage Implicit Domain Knowledge.. In Proc. of the USENIX Security Symposium

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.