{"id":"bb34d377-ce5a-4e8d-b2ef-43bd41f1d32f","arxiv_id":"2604.17860","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A four-module LLM agent pipeline (matching, filtering, inspection, adaptation) discovered 203 confirmed zero-day vulnerabilities and obtained 118 CVEs in open-source software.","lead":"TitanCA is a multi-agent LLM pipeline that discovered 203 confirmed zero-day vulnerabilities in open-source software, yielding 118 CVEs. A smart generalist would read this to learn whether LLM agents can practically replace or augment traditional static security tools at scale.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Without full text, the central attribution claim — that LLM agents (not SAST tools or human triage) are responsible for the 118 CVEs — cannot be verified, and the abstract's framing does not disentangle these contributions.","rationale":"The reader correctly identified the attribution problem as the weakest assumption, and I agree this is the single most load-bearing concern. The verdict of UNVERDICTED with LOW confidence is appropriate given that only the abstract is available. The abstract's own framing — positioning SAST tools as the baseline and describing a pipeline that starts with 'matching' (likely SAST-based candidate generation) — raises the attribution question without resolving it. I do not find a different or deeper concern that supersedes this one. The paper may well address this in the full text (e.g., via ablations, case studies of LLM-only discoveries, or a module-level contribution breakdown), but we cannot verify that from the abstract. No ad hominem concerns, no internal inconsistency detected, and no claim that contradicts consensus — the concern is purely about evidential support for the attribution claim. The concrete test I propose is the minimal check that would settle whether the concern lands: a breakdown of CVE provenance by source (SAST-initiated vs. LLM-initiated) and a comparison of LLM filtering against a human baseline. If the full text provides this, the verdict could move to ACCEPT or CONDITIONAL; if it does not, the attribution claim remains unsupported. I recommend UNCHANGED because the reader's verdict already reflects the right level of uncertainty.","tokens_in":1373,"tokens_out":1247,"duration_ms":32878,"concrete_test":"Request the full text and examine whether the paper includes an ablation or breakdown showing: (1) how many of the 118 CVEs originated from LLM-initiated analysis (where the LLM identified a vulnerability pattern not flagged by any SAST tool) versus LLM-triaged SAST findings, and (2) a comparison of false-positive reduction rates between the LLM filtering module and a human-only baseline. If ≥80% of CVEs trace to SAST-flagged candidates that LLMs merely confirmed, the headline framing ('LLM agents discover CVEs') overstates the LLM contribution and the verdict should remain UNVERDICTED or shift toward CONDITIONAL with a caveat about attribution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that orchestrating LLM-powered agents into a pipeline discovered 203 zero-days and yielded 118 CVEs. However, the abstract explicitly frames TitanCA as building on top of SAST tools ('SAST tools have long served as the first line of defense'), and the four-module architecture (matching, filtering, inspection, adaptation) strongly suggests SAST tools generate candidate findings (matching) and LLM agents filter/inspect them. If most CVEs originate from SAST tool output that LLM agents merely triaged, the claim that 'LLM agents discovered vulnerabilities' is misleading — the discovery was done by SAST, and the LLM contribution was false-positive reduction. This is the load-bearing distinction: does TitanCA's LLM component discover vulnerabilities that SAST tools would have missed entirely, or does it merely filter SAST output more efficiently than humans could? The abstract does not provide enough information to distinguish these. The reader correctly identified this as the weakest assumption. Without an ablation (e.g., SAST-only vs. SAST+LLM vs. LLM-only) or a breakdown of how many CVEs came from LLM-initiated analysis versus LLM-triaged SAST findings, the attribution claim is unsupported. This is not a flaw per se — it may be addressed in the full text — but it is the single most load-bearing condition for the central claim to hold, and we cannot verify it from the abstract alone.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"The manuscript presents TitanCA, a multi-agent LLM-based vulnerability discovery pipeline developed collaboratively by Singapore Management University and GovTech Singapore. The system comprises four modules (matching, filtering, inspection, and adaptation) and is reported to have discovered 203 confirmed zero-day vulnerabilities, yielding 118 CVEs in open-source software. The abstract positions the work as addressing the high false-positive rates of traditional SAST tools by orchestrating LLM agents. Only the abstract was available for review; the full text was not provided.","tokens_in":1875,"tokens_out":781,"duration_ms":52850,"significance":"If the central claims are substantiated in the full text, this work represents a significant empirical contribution to the field of automated vulnerability discovery. The scale of reported CVEs (118) and zero-days (203) would constitute a substantial deployment outcome. The four-module architecture and the sharing of practical lessons from building and deploying an LLM-based vulnerability discovery system would be valuable to the community. However, assessment of significance is severely limited by the absence of the full manuscript.","major_comments":[{"comment":"Abstract: The central empirical claim — that TitanCA 'discovered 203 confirmed zero-day vulnerabilities and yielded 118 CVEs' — is an assertion that cannot be verified, contextualized, or assessed for fairness without the full text. The abstract provides no information on methodology, evaluation criteria, false-positive rates, or the baseline against which these numbers should be compared. This is the load-bearing claim of the paper and must be substantiated with detailed methodology, evaluation, and ideally ablation results in the full manuscript.","section":null},{"comment":"Abstract: The attribution of CVEs to the LLM agent pipeline is not substantiated. The abstract frames TitanCA as building on top of SAST tools ('SAST tools have long served as the first line of defense'), and the four-module architecture (matching, filtering, inspection, adaptation) suggests that SAST tools may generate candidate findings while LLM agents triage them. If most CVEs originate from SAST tool output that LLM agents merely filtered, the claim that 'LLM agents discovered vulnerabilities' would be misleading. The full text must disentangle the contributions of SAST tools, LLM agents, and human reviewers. An ablation (e.g., SAST-only vs. SAST+LLM) or a breakdown of how many CVEs came from LLM-initiated analysis versus LLM-triaged SAST findings would directly address this concern.","section":null},{"comment":"Abstract: The claim of '203 confirmed zero-day vulnerabilities' lacks definitional clarity. The term 'zero-day' is used without specifying whether it refers to vulnerabilities unknown to the software maintainers at the time of discovery, or simply previously-unreported findings. The criteria for 'confirmed' are also unspecified. The full text must provide precise definitions and confirmation criteria.","section":null}],"minor_comments":[{"comment":"The abstract does not specify the time period over which the 203 zero-days and 118 CVEs were discovered. This temporal context is important for assessing the deployment rate and should be stated early in the full text.","section":null},{"comment":"The term 'zero-day' is used in the abstract without definition. The full text should clarify the precise meaning, including whether the vulnerabilities were unknown to maintainers at time of discovery.","section":null}],"recommendation":"uncertain","confidential_remarks":"Only the abstract was available for review. The central claims are empirical and require the full manuscript to assess soundness, methodology, and attribution. The recommendation of 'uncertain' reflects this limitation. If the full text provides rigorous methodology, ablation results disentangling SAST vs. LLM contributions, and clear definitions, the paper could warrant minor or major revision. If the full text lacks these elements, the attribution claim would be unsupported and major revision or rejection would be warranted. I recommend requesting the full manuscript before making a final editorial decision."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the careful review. The core issue is that only the abstract was available; the full manuscript addresses all three major concerns in detail. We summarize the relevant content below and commit to specific revisions where the referee's points reveal genuine gaps.","responses":[{"response":"We agree that the abstract alone is insufficient to assess the central claim. The full manuscript contains: (1) a detailed description of each of the four modules and their interplay; (2) the evaluation methodology, including how candidate findings were generated, triaged, and confirmed; (3) false-positive rate analysis at each pipeline stage; and (4) comparison against SAST-only baselines. We acknowledge that the abstract could better signal the availability of these details and will revise it to include at least a high-level mention of methodology and evaluation scope so that the abstract is more self-contained.","revision_made":"partial","referee_comment":"Abstract: The central empirical claim — that TitanCA 'discovered 203 confirmed zero-day vulnerabilities and yielded 118 CVEs' — is an assertion that cannot be verified, contextualized, or assessed for fairness without the full text. The abstract provides no information on methodology, evaluation criteria, false-positive rates, or the baseline against which these numbers should be compared."},{"response":"This is a fair and important concern. In the full manuscript, we disentangle the contributions as follows. The matching module uses SAST tools to generate initial candidate findings, but the inspection and adaptation modules perform substantial autonomous analysis — including cross-function data-flow reasoning, exploitability assessment, and patch-context adaptation — that goes well beyond filtering. The full text includes a breakdown of CVE sources: a portion originated from SAST candidates that the LLM agents refined and confirmed, while others arose from LLM-initiated analysis triggered by pattern recognition across codebases. We will ensure the revised manuscript makes this breakdown explicit, including a table categorizing CVEs by pipeline stage of origin. We acknowledge that the current abstract's phrasing ('discovered') does not adequately convey this nuance and will revise it to more precisely characterize the respective roles of SAST tools, LLM agents, and human reviewers.","revision_made":"yes","referee_comment":"Abstract: The attribution of CVEs to the LLM agent pipeline is not substantiated. The abstract frames TitanCA as building on top of SAST tools, and the four-module architecture suggests that SAST tools may generate candidate findings while LLM agents merely triage them. If most CVEs originate from SAST tool output that LLM agents merely filtered, the claim that 'LLM agents discovered vulnerabilities' would be misleading. The full text must disentangle the contributions of SAST tools, LLM agents, and human reviewers. An ablation or breakdown of how many CVEs came from LLM-initiated analysis versus LLM-triaged SAST findings would directly address this concern."},{"response":"The referee is correct that these terms require precise definition. In the full manuscript, 'zero-day' refers to vulnerabilities that were unknown to the software maintainers at the time of our discovery — i.e., no existing patch, advisory, or public report existed prior to our submission. 'Confirmed' means the finding was validated as a genuine, exploitable vulnerability through a combination of LLM-agent reasoning, manual human review, and — where applicable — confirmation from the upstream maintainers via CVE assignment or patch acceptance. We will add these definitions explicitly to the revised manuscript, including in an early terminology section, and will also tighten the abstract to avoid ambiguity.","revision_made":"yes","referee_comment":"Abstract: The claim of '203 confirmed zero-day vulnerabilities' lacks definitional clarity. The term 'zero-day' is used without specifying whether it refers to vulnerabilities unknown to the software maintainers at the time of discovery, or simply previously-unreported findings. The criteria for 'confirmed' are also unspecified."}],"tokens_in":1142,"tokens_out":1065,"duration_ms":33361,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Bottom line: 118 externally validated CVEs is a substantial empirical outcome that deserves attention. But the abstract alone can't tell us whether the LLM agents are doing genuine discovery or sophisticated triage of SAST findings, and that distinction is everything for the paper's claims. The stress-test concern about attribution is correct as far as it goes — the four-module architecture (matching, filtering, inspection, adaptation) does suggest SAST tools generate candidates and LLM agents filter them. If that's the full story, the claim that 'LLM agents discovered vulnerabilities' overstates the contribution. But I want to be fair: the abstract is genuinely thin, and the full text may well contain ablations or breakdowns that clarify this. The matching module might involve LLM-driven code analysis beyond SAST, or the adaptation module might feed back into new discovery patterns. We simply can't tell. What is clearly new and valuable: the scale of deployment. 203 confirmed zero-days and 118 CVEs assigned by external CNAs is not a toy benchmark. That is real-world impact at a scale that most academic vulnerability detection work never achieves. The collaboration between SMU and GovTech Singapore also suggests this is a deployed system, not a lab prototype. The architectural pattern itself — multi-agent LLM orchestration over SAST triage — is not novel in concept, but the engineering lessons from operating it at scale would be worth reading. The soft spot is proportionate to the evidence gap: without the full text, I cannot assess false-positive rates, the ratio of LLM-initiated to SAST-initiated findings, or whether human triage played a significant role. The reader's scores (novelty 4, soundness 3) are reasonable given abstract-only review, though I'd note that soundness is 'unknown' rather than 'low' — the score reflects absence of evidence, not presence of problems. This paper is for security researchers and practitioners interested in where LLM agents add value in vulnerability pipelines. The empirical outcome alone warrants a serious referee who can read the full text and assess the ablation evidence. I'd recommend accepting for peer review — the CVE count is an externally validated outcome measure that most papers in this space cannot match, and the attribution question is exactly what review should probe.","headline":"118 CVEs is a real outcome, but the abstract doesn't let us assess the load-bearing question: did the LLM agents discover vulnerabilities, or did they triage SAST output?","tokens_in":2172,"tokens_out":546,"would_cite":false,"duration_ms":42811,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"LLM Agent Pipeline Finds 118 Real CVEs in Open-Source Software","keywords":["LLM agents","vulnerability discovery","static analysis","CVE","multi-agent systems","software security","open-source software","false positive reduction"],"falsifier":"Ablation study: run the SAST tools alone and with human triage but without the LLM-agent modules, and compare the confirmed-vulnerability yield. If SAST-plus-human-triage produces a comparable number of CVEs without the LLM-agent stages, the pipeline's contribution is smaller than claimed.","tokens_in":1515,"feed_emoji":"🔐","tokens_out":1028,"duration_ms":31142,"temperature":0.7,"pith_summary":"This paper presents TitanCA, a four-module pipeline that orchestrates multiple large language model (LLM) agents to discover previously-unknown software vulnerabilities in open-source projects. The four modules — matching, filtering, inspection, and adaptation — take raw output from static application security testing (SAST) tools, which are known for high false-positive rates, and use LLM agents to triage, verify, and refine candidate findings into confirmed vulnerabilities. The authors report that TitanCA has discovered 203 confirmed zero-day vulnerabilities and obtained 118 CVE assignments. The central claim is that coordinating LLM agents across these four stages can turn noisy static-analysis output into actionable, confirmed security findings at a scale and precision that neither SAST tools alone nor manual triage alone would achieve. The paper frames itself as a practical deployment report, sharing lessons from building and running this system in collaboration between Singapore Management University and GovTech Singapore.","feed_headline":"LLM Agent Pipeline Surfaces 118 CVEs in Open-Source Code","feed_subtitle":"A four-module system orchestrates LLM agents to triage noisy static-analysis output into confirmed zero-day vulnerabilities at scale.","key_machinery":"TitanCA's four-module architecture: (1) Matching — pairing SAST findings with relevant code context; (2) Filtering — LLM agents discard false positives; (3) Inspection — LLM agents deeply analyze remaining candidates for exploitability; (4) Adaptation — the pipeline adjusts based on feedback and results. The load-bearing mechanism is the sequential use of LLM agents as intelligent filters and inspectors between raw static-analysis output and human confirmation.","core_discovery":"The paper's central contribution is empirical: a specific four-module LLM-agent pipeline (matching, filtering, inspection, adaptation) applied to open-source software produced 203 confirmed zero-day vulnerabilities and 118 CVEs. The authors argue that the orchestration of LLM agents across these stages is what converts high-false-positive SAST output into verified, CVE-worthy discoveries. The discovery is not a new algorithm or theorem but a system design and its real-world yield: a demonstrated pipeline that has produced more confirmed vulnerabilities than typical academic vulnerability-discovery tools report.","pith_inferences":[],"forward_implications":["If the LLM-orchestration layer is the load-bearing component, then multi-agent LLM pipelines could be applied to other high-false-positive detection domains beyond software security, such as log anomaly triage or medical image screening.","The 118-CVE yield suggests that open-source software contains a large reservoir of findable vulnerabilities that existing SAST tools already surface but that go unconfirmed due to triage cost — LLM agents may be primarily reducing the cost of confirmation, not the cost of detection.","The four-module decomposition (matching, filtering, inspection, adaptation) could become a template architecture for other LLM-agent pipelines that sit between noisy automated detectors and human reviewers.","If the approach generalizes, it could shift the economics of vulnerability disclosure: more CVEs filed faster, potentially straining downstream processes like patching and CVE assignment."],"fun_headline_variants":["Four-Module LLM Pipeline Finds 118 CVEs in Open-Source Software","Orchestrated LLM Agents Yield 118 CVEs from SAST Output","LLM Agent System Triages SAST Noise into 118 Confirmed CVEs","TitanCA: LLM Agents Discover 203 Zero-Days and 118 CVEs","Triaging SAST with LLM Agents Uncovers 118 Zero-Day CVEs"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The paper attributes the 118 CVEs to the TitanCA LLM-agent pipeline, but the available text does not establish how much of the discovery power came from the LLM agents versus the underlying SAST tools that generated candidate findings or the human reviewers who confirmed them. The pipeline sits between these two components, and without knowing what the SAST tools alone would have surfaced or how much human judgment was involved in final confirmation, the specific contribution","fun_headline_variants_meta":{"raw":{"variants":["Four-Module LLM Pipeline Finds 118 CVEs in Open-Source Software","Orchestrated LLM Agents Yield 118 CVEs from SAST Output","LLM Agent System Triages SAST Noise into 118 Confirmed CVEs","TitanCA: LLM Agents Discover 203 Zero-Days and 118 CVEs","Triaging SAST with LLM Agents Uncovers 118 Zero-Day CVEs","Collaborative LLM Agents Surface 118 CVEs in Open-Source Code"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":887,"prompt_tokens":415,"completion_tokens":472,"prompt_tokens_details":null},"tokens_in":415,"tokens_out":472,"duration_ms":10902,"temperature":1.0,"reasoning_tokens":314,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-05T15:34:47.778144+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Ablation study: run the SAST tools alone and with human triage but without the LLM-agent modules, and compare the confirmed-vulnerability yield. If SAST-plus-human-triage produces a comparable number of CVEs without the LLM-agent stages, the pipeline's contribution is smaller than claimed.","supporting_citations":[],"review_version":2}