{"id":"b22ffd16-b5b7-42a7-97e9-c0528d40f891","arxiv_id":"2607.13085","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Plain coding CLI agents, especially Codex with newer GPT models, solve 70-96 of 104 XBOW benchmark tasks, matching or exceeding some published security-harness scores under model-matched comparisons.","lead":"This paper tests whether ordinary coding assistants, without any security-specific add-ons, can already solve a large share of an automated hacking benchmark. It finds plain agents match or beat some purpose-built pentesting systems when the underlying AI model is the same, so future claims should compare against these baselines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RQ3's cross-infrastructure residuals and the 'exceeds 91%' claim depend on published scores from a different benchmark fork; the baseline-first message survives, but the quantitative architecture gap is not yet established.","rationale":"The reader's weakest assumption—that RQ3 assumes direct comparability of published MAPTA and PentestGPT V2 results to the paper's repaired-fork runs—is exactly the load-bearing concern I identify. It is load-bearing because the paper's most striking quantitative statements (the +9.6pp and +5.2pp residuals, and the 95.2% plain-agent union exceeding PentestGPT V2's 91% headline) depend on it. The internal Phase 1 and Phase 2 results are independently valuable: they show a same-model plain coding agent solving 70/104 GPT-5 tasks, which already demonstrates the attribution problem the paper is making. But the paper goes beyond claiming 'baselines should be reported' and asserts specific residual sizes and an 'exceeds' ordering. Those specific numbers are not yet settled because the comparison systems were not rerun under the same infrastructure, budget, and benchmark fork. The paper's own threat model admits this, and the artifact provides the repair patch, so the concern is checkable rather than fatal. I agree with the reader's conditional posture: the central methodological message stands, but the quantitative architecture-residual claims should be either reproduced in a matched setting or softened to explicitly descriptive status. No change to the CONDITIONAL verdict is needed because the reader already captured this conditionality; my read reinforces it rather than moving it.","tokens_in":11413,"tokens_out":5314,"duration_ms":58326,"concrete_test":"Use the released artifact to identify the ~40 repaired targets and re-run the plain Codex GPT-5 and GPT-5.2 rows only on the 64 byte-identical challenges; re-express the MAPTA and PentestGPT V2 published scores on the same subset if per-challenge results are available, and cap the plain-agent budget at the published median per-task cost of each system. If the residuals and the 95.2% vs 91% ordering persist on the matched subset under matched budget, the concern is resolved. If the plain-agent advantage concentrates in the repaired targets or the residuals shrink below the paper's own interpretation thresholds, then RQ3 and the 'exceeds' headline should be reframed as descriptive cross-study comparisons pending direct reruns.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central recommendation—report model-matched plain-agent baselines—is well supported by the internal RQ1/RQ2/RQ4 runs. However, the specific quantitative claims that give the paper its headline force are cross-infrastructure comparisons: the RQ3 architecture residuals (MAPTA +9.6pp, PentestGPT V2 +5.2pp) and the claim that a plain Codex GPT-5.5 two-pass union (95.2%) exceeds PentestGPT V2's 91% headline. These use published scores from MAPTA and PentestGPT V2 that were obtained on a different benchmark checkout. Section VII-D acknowledges that roughly 40 of 104 targets required local repairs, making the fork 'not byte-identical to cross-paper benchmark checkouts.' Since 40/104 is 38% of the benchmark, if the repaired targets differ systematically in difficulty, both the plain-agent absolute scores and the residuals shift. The paper also uses different cost caps: plain Codex GPT-5.2 had a $1.50 cap and an average cost of $0.486/task, while PentestGPT V2 reports a median cost of $0.18/task; the GPT-5 comparison has a $0.75 cap against MAPTA's reported $0.206 average. Inference settings (temperature, reasoning effort, retries) for the published systems are unknown. The paper is transparent about these limitations, but it still presents the residuals and the 'exceeds' comparison as results rather than purely illustrative cross-study observations. The load-bearing issue is not that the plain-agent baseline is weak—it is strong internally—but that the headline quantitative comparison to published architectures is not a controlled measurement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the attribution of XBOW benchmark performance to security-specific architecture versus backbone model capability. It runs three default coding CLIs (Codex, OpenCode, Pi) under a fixed GPT-5 model on the 104-task XBOW benchmark with fresh flags, matched interfaces, and a fixed cost cap; finds Codex the strongest baseline; tests whether security-specific prompt variants help (they do not); compares the default Codex scaffold against published MAPTA and PentestGPT V2 results under the closest claimed model matches; and scales the same scaffold to GPT-5.2 and GPT-5.5. The central recommendation is that future penetration-testing-agent evaluations should report model-matched plain-agent baselines before attributing gains to architecture. The paper also releases a reproducibility artifact.","tokens_in":11835,"tokens_out":4156,"duration_ms":44585,"significance":"The baseline-first message is timely and valuable: it directly addresses a real confound in the recent autonomous-pentesting literature, and the controlled internal comparisons (RQ1/RQ2/RQ4) are designed carefully, with fresh flags, a fixed target interface, and budget-censored analyses. The released artifact with per-attempt tables is a genuine strength. If the paper limited its quantitative claims to its own runs, the contribution would be solid. However, the headline architecture-residual and 'exceeds 91%' claims depend on published scores from a different benchmark checkout, and the residual estimates are reported without uncertainty intervals. The methodological recommendation survives, but the specific quantitative architecture gap is not yet established.","major_comments":[{"comment":"The architecture residuals are computed as published architecture score minus plain-agent score, using MAPTA and PentestGPT V2 results that were not rerun in this infrastructure. Section VII-D states that roughly 40 of 104 targets required local image/build repairs, so the fork is not byte-identical to cross-paper benchmark checkouts; Table VIII shows different cost caps ($0.75 vs $0.206 average per task for GPT-5; $1.50 vs $0.18 median for GPT-5.2) and unknown inference settings. These differences shift both the absolute plain-agent scores and the residuals. The paper is transparent about this, but it still presents +9.6pp and +5.2pp as results rather than as illustrative cross-study observations. Please either rerun the published systems in the same harness, restrict the comparison to the subset of unchanged targets, or explicitly label these residuals as non-quantitative and move them","section":"V-C, Eq. (1), Table VII, VII-D"},{"comment":"Only two passes per condition are reported, and the paper itself notes 22 challenges flip between solved and unsolved across the two GPT-5 passes. Yet no confidence intervals or per-challenge variance estimates are provided for the key pass@1 means (67.3%, 79.8%, 92.3%) or for the residuals in Table VII. For a 104-task binary outcome, the single-pass binomial standard error at 67.3% is roughly 4.6 pp, so the 9.6 pp MAPTA residual and the 5.2 pp PentestGPT V2 residual are not clearly separated from noise. Please report bootstrap intervals, per-task regression estimates, or at least explicitly state that the residuals are not statistically distinguishable from zero under the current trial count. This is load-bearing because the quantitative force of RQ3 depends on these gaps.","section":"VI-B, Tables V and IX"},{"comment":"The claim that plain Codex with GPT-5.5 'exceeds PentestGPT V2's 91% headline' compares a different, newer model (GPT-5.5) against Opus 4.5 and compares a two-pass union (95.2%) or pass@1 mean (92.3%) to a single-run published score. This is a model-scaling observation, not an architecture-attribution result, and presenting it in the conclusion as 'exceeding that headline' overstates the comparison. Reframe all such statements as illustrative of why current plain baselines matter, and separate them from the internal RQ1/RQ2/RQ4 results. This is especially important because the abstract and conclusion lean on the strong plain-agent performance to motivate the methodology.","section":"V-D, VII-C, IX"}],"minor_comments":[{"comment":"'OW ASP' should be 'OWASP' in the text and references.","section":"II-A, II-B, References [8],[10]"},{"comment":"'A WE' should be 'AWE'.","section":"II-C"},{"comment":"The sentence before Eq. (1), 'For each comparison systems, let P_s be ...' has a grammar error; also define P_s and C_s explicitly as percentages or fractions to avoid ambiguity.","section":"VI-A, Eq. (1)"},{"comment":"The 'Run $' cell for PentestGPT V2 is blank; consider writing 'not reported' rather than leaving it empty.","section":"Table VIII"},{"comment":"'The MAPTA residual is about ten percentage points' is imprecise; Table VII reports +9.6 pp. Use the exact number.","section":"V-C"},{"comment":"The caption states 'Plain P@2 is a two-pass union, not the residual baseline.' This is helpful, but consider adding a visual marker to distinguish union from single-run averages in the figure itself.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper's central methodological message is important and likely correct, and the internal RQ1/RQ2/RQ4 runs are a useful controlled contribution. The main risk is overstatement: the RQ3 residuals and the GPT-5.5 'exceeds 91%' comparisons rest on cross-infrastructure published numbers and lack statistical uncertainty. If the authors reframe those parts as illustrative and add proper uncertainty quantification for the internal comparisons, the paper could be a solid contribution to the evaluation methodology of LLM-based pentesting agents. I would not reject, but the current version needs a substantial revision of the quantitative claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the paper's main claim is right. If you're comparing autonomous pentesting architectures, you need a model-matched plain coding-agent baseline, because the model and scaffold are confounded otherwise. The internal experiments—Codex vs OpenCode vs Pi on the same GPT-5 with same budget and scoring, the negative prompt results, and the model scaling within the same Codex scaffold—are carefully done and internally consistent. The artifact looks genuinely reproducible: harness, prompts, configs, repair patch, per-attempt 1,456-row table, analysis scripts. That's real evidence and worth credit.\n\nThe soft spot is the part that gives the paper its headline force. RQ3's residuals—MAPTA +9.6pp, PentestGPT V2 +5.2pp—and the claim that a plain Codex GPT-5.5 two-pass union (95.2%) exceeds PentestGPT V2's 91% headline are not controlled measurements. They rely on published scores from a different benchmark checkout. The paper says about 40 of 104 targets needed local repairs; that's 38% of the benchmark, and if repaired targets differ in difficulty, the absolute numbers shift. Cost caps differ: plain GPT-5.2 had a $1.50 cap, PentestGPT V2 reports median $0.18 per task; MAPTA $0.206 average vs the plain GPT-5's $0.266. Inference settings for the published systems are unknown. The paper acknowledges all of this in Section VII-D, and I believe the acknowledgments are honest—but then it still presents the residuals and the 'exceeds' comparison as results, not as descriptive cross-study observations. That's a framing issue, not a fatal flaw.\n\nAlso minor: only two passes per condition, no confidence intervals, and 22 challenges flip between passes. RQ4 changes model and cost cap together; the budget-censored analysis helps but can't fully reproduce a prospective lower-cap run. These are minor because the direction of the findings is robust.\n\nWho is this for? Anyone doing LLM-agent evaluation on security benchmarks, and anyone citing MAPTA/PentestGPT V2 numbers. The paper is an evaluation study, not a new system, but the methodological point is timely and the internal comparisons are solid. I'd send it to peer review. I'd ask the authors to (1) add more trials and report intervals, (2) either rerun the comparison systems in their harness or explicitly demote RQ3 to descriptive status, and (3) separate model and cap in RQ4 if feasible. The core argument doesn't depend on the exact residuals, and it holds up.","headline":"The baseline-first message is right and well supported by the internal runs; the cross-study residuals and 'beats 91%' claim are not controlled measurements, but the paper is honest about that and deserves a serious referee.","tokens_in":12251,"tokens_out":2260,"would_cite":true,"duration_ms":21367,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Default coding agents may match security-specific pentesting systems once models are matched","keywords":["autonomous penetration testing","XBOW benchmark","coding agents","LLM agent baselines","model scaling","architecture attribution","prompt engineering","security testing"],"falsifier":"Rerun MAPTA and PentestGPT V2 inside the paper's harness with identical model, budget cap, and target images; if the residuals vanish or reverse, the architecture contribution is smaller than reported. Conversely, run plain Codex on the original unrepaired benchmark; if scores drop substantially, the local repairs rather than agent ability explain the baseline.","tokens_in":11356,"feed_emoji":"🛡️","tokens_out":3196,"duration_ms":26290,"temperature":0.7,"pith_summary":"The paper argues that the field of autonomous penetration testing cannot attribute benchmark gains to security-specific architecture unless it first measures a strong, model-matched plain-agent baseline. Running default coding CLI agents on the 104-task XBOW benchmark, it finds that Codex with GPT-5 already solves 67.3% of tasks on average, and with GPT-5.5 reaches 92.3% pass@1 and 95.2% two-pass coverage, exceeding the 91% headline of a purpose-built system. Security-specific prompt variants did not improve over the default prompt. The paper concludes that architecture can add a few points on single-attempt efficiency, but model progress alone closes most of the gap.","feed_headline":"Plain coding agent outruns specialized pentesting systems","feed_subtitle":"Unmodified Codex hits 92.3% pass@1 with GPT-5.5, topping PentestGPT V2's 91% headline.","key_machinery":"The central object is the 'architecture residual' — the difference between a published purpose-built system's solve rate and the closest model-matched plain coding-agent solve rate — computed under a controlled harness that varies only one factor at a time (agent CLI, prompt, model) while holding model, budget, target interface, and scoring rule fixed. The plain-agent baseline is a default coding CLI (Codex, OpenCode, Pi) with no security-specific modifications, whose general loop of write-code, execute, observe, and revise is argued to already instantiate the iterative probing mechanics of web exploitation.","core_discovery":"On the 104-task XBOW benchmark, a default, unmodified Codex CLI agent is a strong baseline for autonomous penetration testing. Across two full passes, Codex with GPT-5 averages 67.3% pass@1, GPT-5.2 averages 79.8%, and GPT-5.5 averages 92.3%, with two-pass union coverage of 77.9%, 88.5%, and 95.2% respectively. The GPT-5.5 plain-agent pass@1 exceeds the 91% headline reported for PentestGPT V2 with Opus 4.5, and the two-pass union exceeds it by 4.2 points. Against model-matched published results, MAPTA retains a 9.6-point residual over GPT-5 Codex, and PentestGPT V2 retains a 5.2-point residual over GPT-5.2 Codex, but both residuals are smaller than headline system-to-system gaps suggest, and","pith_inferences":["The results suggest that some reported advantages of multi-agent security harnesses may shrink or invert when re-tested with the same model generation and a strong generic agent loop, a hypothesis testable on the same benchmark.","The Codex advantage over OpenCode and Pi may be partly explained by OpenAI-model integration maturity; a replication with non-GPT models could reveal whether the ranking is model-agnostic.","The benchmark repair of about 40 targets introduces a possible difficulty shift; an ablation on the original unrepaired targets would clarify whether plain agents benefit from the repairs.","The methodology transfers to other agentic security tasks, such as CTF-style challenges or bug bounty reproduction, where the same model/architecture confound exists."],"forward_implications":["Future autonomous-pentesting papers should report a model-matched plain coding-agent baseline before attributing benchmark gains to architecture.","Repeated plain-agent runs (pass@2) can match or exceed some published architecture scores, so single-pass comparisons overstate harness value.","Backbone model progress alone substantially lifts a fixed minimal scaffold, meaning headline scores conflate model and architecture gains.","Security-specific prompt text did not help in this study; prompt design should be compared against the agent's shipped default.","Architecture can still deliver value on single-attempt efficiency and cost per solved challenge, even when union coverage is matched."],"fun_headline_variants":["Plain Codex with GPT-5.5 hits 92.3%, beats PentestGPT V2","Unmodified coding agent outruns specialized pentesting tools","Default Codex surpasses PentestGPT V2 without any harness","Coding agent alone reaches 92.3% pass@1 on XBOW, topping V2"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That the published MAPTA and PentestGPT V2 scores are directly comparable to this paper's runs, despite the paper's fork repairing about 40 benchmark targets, different cost caps, and unknown inference settings, and despite those systems not being rerun in the same harness.","fun_headline_variants_meta":{"raw":{"variants":["Plain Codex with GPT-5.5 hits 92.3%, beats PentestGPT V2","Unmodified coding agent outruns specialized pentesting tools","Default Codex surpasses PentestGPT V2 without any harness","Coding agent alone reaches 92.3% pass@1 on XBOW, topping V2"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1367,"prompt_tokens":833,"completion_tokens":534,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":445}},"tokens_in":577,"tokens_out":534,"duration_ms":5304,"temperature":1.0,"reasoning_tokens":445,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:50:12.077502+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun MAPTA and PentestGPT V2 inside the paper's harness with identical model, budget cap, and target images; if the residuals vanish or reverse, the architecture contribution is smaller than reported. Conversely, run plain Codex on the original unrepaired benchmark; if scores drop substantially, the local repairs rather than agent ability explain the baseline.","supporting_citations":[],"review_version":1}